Can code similarity prove student AI use?
An accepted computing-education study compared 29,970 student submissions with 90,000 attempts from three frontier models. Similarity revealed population-level convergence and useful assignment-design signals, but the authors say it cannot attribute AI use to an individual without direct process evidence.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
What generated-reference code matching can reveal about population-level changes in introductory programming submissions, and why it cannot by itself establish individual AI use

At a glance
- 1The study analysed 29,970 final Python submissions from ten weekly labs in 2021, 2023 and 2025, then generated 90,000 retrospective attempts with three 2026 frontier models under a consistent protocol.
- 2Across Labs 1–9, 95.48%–96.08% of cross-model solution pairs matched after docstrings were removed; constrained tasks therefore produced similar code even across different model families.
- 3Student-to-model matching rose in later cohorts, but the study had no verified submission-level AI-use labels and could not attribute the change—or any individual file—to AI assistance.
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Living evidence record
Impact record IAI-1919R3I
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
3 October 2026
Source trail
3 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
The study asks what a code match actually means
Universities increasingly face a difficult evidential problem: code written with an AI assistant can resemble other model output, but short programming exercises can also funnel independent learners toward the same natural solution. This accepted conference paper examines that ambiguity at a scale rarely available in computing-education research. Its central conclusion is deliberately narrow. Generated-reference matching can describe changes across groups and expose assignments that invite convergence; it does not prove how one student's submission was produced.
The manuscript was posted to arXiv on 1 October 2026 and is listed for Koli Calling 2026 with an ACM DOI. Researchers from the University of Toronto, Aalto University, McMaster University, Maranatha Christian University in Indonesia and the University of Auckland studied ten Python labs from a large research-intensive university. The project received research-ethics approval, used de-identified submissions and reports only aggregate results. The authors explicitly say no similarity result was used to make or imply a misconduct determination.[1][2][3]
The denominator is 29,970 student files and 90,000 model attempts
The researchers used the final submitted file for each student and lab from three course offerings: 9,735 graded submissions in 2021, 10,987 in 2023 and 9,248 in 2025. The same lead instructor, delivery mode and core materials linked the cohorts, although other instructors, teaching assistants and some checker details varied. The later two offerings permitted AI for learning and coding support but required individually produced assessed work. Two labs also offered an institution-hosted tutor that was supposed to guide without revealing direct solutions.
For the retrospective comparison, GPT-5.5, Gemini 3.1 Pro and Claude Opus 4.8 each made 1,000 attempts per lab and cohort: 30,000 attempts per model and 90,000 overall. Every attempt received the relevant year's handout and starter file, could revise at most twice after the student-visible checker, and was then graded with hidden instructor tests never shown in the prompt. The run consumed 491.8 million tokens and cost about US$4,284. Applying the same 2026 model set to every cohort made the protocol consistent, but it did not recreate the tools students could actually access in 2021, 2023 or 2025.[2]
Model code was usually correct—and often looked alike
Across model-and-year combinations, mean hidden-test scores ranged from 98.93% to 99.99%, and full-pass rates ranged from 86.97% to 99.95%. Those figures show that the reference bank was usually functionally plausible. They also expose why a visible checker is not enough: on one 2021 lab, two models almost always passed the public checker on the first prompt, yet their full-pass rates on the hidden suite were only 0.4% and 0%. A code detector built from unchecked model answers could therefore treat recurring mistakes as evidence without first knowing whether the references solve the task.
After removing starter code and, for the model-to-model comparison, docstrings, MOSS reported matches for 87.79%–88.14% of cross-model pairs when all ten labs were pooled. Across Labs 1–9, the cross-model range was 95.48%–96.08%; the more open tenth lab was markedly lower. The researchers also counted exact abstract-syntax-tree forms for five functions. Constrained functions collapsed toward a few forms, while a more open-ended string task remained diverse. That is the essential warning: similarity partly measures the shape of the assignment, not just the origin of the answer.[2]
Later student cohorts moved closer to the reference bank
By 2025, student-to-model match incidence was 1.76 to 2.12 times its 2021 level, depending on the model. The pattern persisted when analysis was restricted to submissions that passed every hidden test. Four constrained functions also appeared in fewer exact forms in later cohorts, including after the researchers equalised the number of correct submissions being compared. These are meaningful population-level signals: the submitted code became more alike and more similar to contemporary model solutions.
But timing is not causation. The 2023 cohort showed only a small increase over 2021 even though general-purpose chat assistants were already available, while the larger increase arrived in 2025. Course-team composition, student mix, permitted support, shared resources and small changes to specifications could all affect the pattern. The dataset contains no verified record of which students used AI, how they used it, or whether their final file reflected generation, debugging, explanation, peer help or independent work. The study therefore cannot estimate AI-use prevalence or calculate detector accuracy.[2]
A match can start a review; it cannot finish one
MOSS retrieves overlapping snippets for human inspection; it does not determine authorship. Most overlaps in this study were short, and changing the minimum accepted match length materially changed how many student files were retained. Starter-code exclusion, docstring treatment and the number and diversity of generated references also shape the outcome. A school that converts one threshold into a misconduct score would hide these choices behind a number that looks more objective than it is.
For an individual case, the paper says attribution needs evidence about process: prompts, revisions, intermediate code, student explanations and protected disclosures, gathered with consent and kept separate from premature disciplinary judgement. Even then, an institution would need task-specific false-positive and false-negative estimates from held-out work with verified histories. A correct, conventional solution to a constrained exercise should never become suspicious merely because modern models also produce the conventional solution.[2]
The safest use is before students receive the task
The most constructive application is assignment design. An instructor can generate a modest reference bank from the exact handout and starter code, test every output with the hidden suite, and inspect whether valid solutions collapse into one form. If they do, the assignment is poorly suited to authorship inference and may not expose the reasoning the course wants to assess. The instructor can redesign it to require tests, explanations, intermediate artefacts, comparisons between approaches or a later unaided modification.
This shifts similarity from policing to quality assurance. The paper estimates that 100 attempts would have cost about US$5 at the batch rates used, although it does not establish that 100 is sufficient. Institutions must also consider privacy: the study ran MOSS locally so student code stayed on university infrastructure. Sending identifiable student work to external model or detection services would create separate data-protection, contractual and fairness questions that the reported performance cannot answer.[2]
Funding, model use and the evidence still missing
The authors acknowledge support from the University of Toronto's Learning & Education Advancement Fund, an NSERC Discovery Grant and two Research Council of Finland grants. They also disclose using Codex and Claude Code to assist with analysis and visualisation scripts, followed later by language editing and claim cross-checking; the researchers say they manually reviewed the work and take responsibility for the results. That disclosure is especially relevant in a paper about model-generated artefacts.
Confidence would rise with preregistered reference banks tested on held-out submissions whose production histories are independently verified; replication across institutions, languages, courses and less constrained tasks; and direct measurement of learning and later unaided performance. Until then, the study supports cohort research and pre-release assessment review. It is evidence against treating a similarity hit as a verdict on a student.[2]
What this means for people
- Students should not face misconduct findings based solely on a similarity score from code that a constrained task naturally makes alike.
- Educators can use generated reference banks to redesign weak assignments before release and ask for evidence of understanding rather than stylistic uniqueness.
- Universities need privacy-preserving review processes, transparent thresholds and a meaningful human appeal route wherever similarity tools are used.
Global context
The author team spans Canada, Finland, Indonesia and New Zealand, but the student corpus comes from one Canadian university course. The practical warning is global: any institution using generated-reference matching must validate it against its own tasks, languages, policies and student population.
What the evidence does not yet show
- The three cohorts were observational rather than randomised, and instructor mix, teaching assistants, student composition, permitted support and some task details varied over time.
- No submission had a verified AI-use label or contemporaneous process history, so detector accuracy, individual attribution and AI-use prevalence were outside the study's scope.
- The reference bank used three frontier models available in 2026 and does not reconstruct the model ecosystem students encountered in earlier cohorts.
- MOSS measures surface overlap; its results depend on task structure, preprocessing, reference sampling and the minimum retained match length.
- All student files came from one university's introductory Python course, and exact-form analysis covered five purposively selected functions.
What to watch next
- Independent replication using multiple institutions, programming languages and task structures.
- Preregistered, held-out evaluation with consented process histories and task-specific error rates.
- Assessment designs that make reasoning, tests and intermediate work visible without intrusive surveillance.
Evidence trail
Sources used for this report
Links checked 3 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Education
Does the way students use AI matter?
A peer-reviewed survey of 713 Saudi undergraduates found that supportive and creative AI use was associated with stronger self-reported creativity and metacognitive awareness, while substitutive use showed negative associations. The one-time self-report design cannot establish what AI caused—or whether skills improved.
8 min · 1 source
Education
Can AI avoid overreacting to one wrong answer?
A diffusion-based knowledge-tracing model led four educational benchmarks, including difficult response reversals. The evaluation reuses folds for early stopping and scoring and does not test classroom decisions.
8 min · 2 sources
Education
Will a curriculum-grounded AI assistant actually save teachers time?
Sanoma has launched Sanna for a free trial in seven European countries after a two-month pilot involving more than 1,500 teachers. The company describes co-design and guardrails, but publishes no measured time saving or pupil-outcome result.
6 min · 2 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.