What does the evidence say about ‘vibe coding’ in university courses?
A scoping review of 37 studies found that educational value depended on planning, verification and assessment design more than tool access alone. The literature is young, heterogeneous and often self-reported, so it does not establish one best teaching model.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The review searched four bibliographic databases plus arXiv, screened 118 unique records and included 37 studies published from 2022 to July 2026.
- 2Stronger outcomes were associated with planning, exploration, testing and critique; weaker patterns involved direct solution requests, repeated execution and limited inspection of generated code.
- 3The evidence cannot support a universal teaching prescription: seven sources were preprints, many studies were short or self-reported, and one reviewer coded the final set.
Research topic
Scoping review of AI-assisted programming interaction, learning outcomes, instructional design, assessment and curriculum in university computing education
The answer: teach verification and reasoning, not just prompt-to-code speed
The clearest finding from a new scoping review is that access to an AI coding tool was not enough to produce learning. Across 37 studies, stronger performance tended to accompany detailed specification, exploratory interaction, testing and evaluation. Weaker patterns involved asking for a complete solution, pasting errors back to the model, rerunning code and inspecting little beyond whether it appeared to work.
For university teachers, that points toward assignments that make planning, independent competence and verification visible. It does not prove that one assessment design or chatbot causes better learning. The studies used different tools, learners, tasks and outcomes, and many relied on perceptions or short interventions. The review maps a rapidly changing evidence base rather than delivering a pooled effect estimate.[1]
The Impact Brief · Free
Follow the evidence in education.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
How 183 records became 37 included studies
The author searched Web of Science, Scopus, IEEE Xplore and the ACM Digital Library on 7 July 2026, with arXiv as a complementary source. The searches returned 183 records: 34 from Web of Science, 92 from Scopus, 16 from IEEE Xplore, 18 from ACM and 23 from arXiv. Deduplication left 118 unique records. Title and abstract screening excluded 50, full-text assessment excluded 31 tool-development papers without substantive teaching or learning context, and 37 studies remained.
The included evidence was very recent. Thirty-three of the 37 studies appeared in 2025 or 2026; 19 were conference papers, 11 journal articles and seven preprints. Twenty-four involved university students, four practitioners and two faculty members. Twenty-four studies, or 64.9%, focused on undergraduates, while 15 did not clearly report the course level and 14 did not report prior programming experience.
The review followed PRISMA guidance for scoping reviews and published its search strategies. Its protocol was registered on the Open Science Framework on 20 August, after the July search. One reviewer coded all 37 papers after piloting the framework on ten. That provides a consistent documented process, but no independent second coding or inter-rater agreement check.[1]
What students actually did with coding assistants
The interaction evidence repeatedly showed shallow verification. In one study of 732 remediation episodes, 78% were pure retesting and 86.3% of the clarification attempts happened in the first round. Another mixed-level study found that 63.6% of coded actions involved running the prototype; 91.7% of those tests covered common cases, 2.2% covered edge cases and none used unit tests. Debugging accounted for 61% of prompts, while direct engagement with course code or execution logs accounted for 7.4% of actions.
Expertise changed the pattern but did not remove risk. The review describes beginners who trusted output they could not verify and more experienced students who overlooked generated boilerplate. Among advanced learners, higher-performing students were more exploratory, while lower-performing students moved more directly from the assignment prompt to a generated answer. Detailed specifications correlated positively with final grade in one course, while pasting unprocessed console errors correlated negatively.
These are associations within varied studies, not proof that a particular prompting behaviour causes a grade increase. Better-prepared students may already possess the knowledge needed to write clearer specifications and evaluate results. A useful curriculum therefore has to teach those underlying capabilities rather than grade superficial prompt style.[1]
Learning gains appeared when support was deliberately scaffolded
The review grouped reported outcomes into conceptual learning in 21 studies, productivity in 13 and affective outcomes in nine; the categories overlapped. In a randomised study of 122 undergraduates, 62 received AI-supported prompts and 60 served as controls. The intervention group moved from trial-and-error toward planning, and planning and debugging precision were associated with post-test performance. Another comparison reported that 38 students using AI increased mean proficiency scores from 24.34 to 34.34, while 42 students in human pair programming did not show a comparable rise.
Other results were more cautious. In one study, 76% of participants considered the tools beneficial, but 51.2% reported no improvement in understanding loops and arrays, compared with 41.5% who perceived improvement. A 31-person hackathon showed that beginners and mixed-experience participants could produce working prototypes in a day, which demonstrates access and completion rather than durable programming knowledge.
Several teaching patterns recurred: planning hints before code generation, adaptive support that withholds full answers, assignments completed first without AI and then with it, and compulsory critique of generated output. One study involving 102 students found that planning hints drew the most engagement and were the only hint type consistently associated with higher performance. These approaches are plausible and testable, but the review does not establish a single optimal sequence.[1]
Assessment needs evidence beyond the final program
The review organises assessment responses into three broad strategies. Teachers can resist trivial generation by setting tasks that a straightforward model output cannot complete; segregate unaided and AI-assisted phases; or diversify evidence through oral explanations, version history, interaction records, reflective work and context-specific tasks. The practical goal is to distinguish a functioning artefact from the student's ability to understand, test and defend it.
Process evidence is not automatically valid. A course that assigned 30% of the grade to prompt logs and additional marks to reflection received favourable student ratings, but the study acknowledged that students could also use a model to generate the reflection. Keystroke recording was poorly accepted in another study. The review recommends keeping process data under student control and using it as a learning resource rather than default surveillance.
For students, the fairest policy is explicit before an assignment starts: what assistance is permitted, what must be completed independently, which records are collected and how judgment will be demonstrated. Oral checks and small AI-free tasks may be less intrusive than continuous monitoring. Accessibility needs also matter; an assessment designed to detect misuse should not inadvertently disadvantage students who rely on assistive technology.[1]
What remains unproven
The term ‘vibe coding’ covered different tools, levels of autonomy, prompts and learning settings. The 37 studies were heterogeneous, and seven were unreviewed preprints at the search date. Many used self-report, observational designs or short interventions. English-language selection may have excluded relevant evidence, and the literature was concentrated in undergraduate and introductory contexts. These features prevent a reliable pooled estimate and limit causal conclusions.
The review found that 31 studies reported at least one identifiable empirical evidence source. Of those 31, 21 used surveys, ten prompt or chat logs, eight grades, six code artefacts and five interaction logs; categories could overlap. Survey enthusiasm, task completion and course grades answer different questions. Longer follow-up is needed to know whether students retain debugging, architecture and security skills when AI help is removed.
The study received no external funding, and the author declared no conflicts of interest. The paper disclosed using GPT-5 to improve writing clarity, with the author reviewing the output and taking responsibility. The underlying extraction dataset is available from the author on request rather than as an immediately downloadable public dataset.[1]
What would change the assessment
Confidence would rise with preregistered, multisite comparisons that assign students to well-defined teaching approaches and test both immediate performance and later unaided competence. Studies should report participant experience, course level, model version, permitted tool use, attrition and the exact outcome denominator. Independent coding and open extraction tables would make future reviews easier to verify.
For teachers acting now, the evidence supports a cautious design principle rather than a mandate: preserve independent practice, require specifications and tests, and assess explanation and verification as well as working code. Review the burden on staff and students, and do not equate faster completion with learning. The assessment would weaken if better controlled studies show that these scaffolds do not improve durable knowledge, or if process-heavy assessment adds surveillance and workload without better validity.
Professional practice is also moving, but industry demand is not by itself an educational outcome. Courses can teach students to supervise generated software while still protecting the foundations needed to notice insecure, incorrect or brittle output. The research question is no longer simply whether students use AI; it is which learning design helps them remain responsible for what the software does.[1]
What this means for people
- Students need clear rules and opportunities to practise both independent programming and responsible AI-assisted work.
- Teachers can assess specifications, tests and explanations instead of relying only on the final code artefact.
- Universities should avoid intrusive monitoring when oral checks or student-controlled process evidence can answer the same question.
- Employers may gain graduates better prepared to verify generated software, but durable competence still needs direct testing.
Global context
The review was conducted by a Singapore-based author but synthesised an international English-language literature. The evidence was not mapped as a representative sample of countries, and local curricula, assessment rules, tool access and privacy law differ. Institutions should test approaches in their own courses rather than treating one rapidly changing evidence map as a universal policy.
What the evidence does not yet show
- The 37 studies varied in tools, populations, designs and outcomes, so the review did not estimate a pooled effect.
- Seven included sources were preprints, and many studies relied on surveys, short interventions or observational evidence.
- One reviewer coded the final set; no inter-rater reliability statistic was available.
- The evidence was concentrated in recent undergraduate and introductory settings, with course level and prior experience often unreported.
- English-language selection and a July 2026 search cut-off may omit relevant or later evidence.
What to watch next
- Randomised or strong quasi-experimental comparisons of specific scaffolding and assessment designs.
- Delayed, unaided tests of programming knowledge rather than task completion alone.
- Transparent reporting of model versions, student experience, permitted use and attrition.
- Evidence on workload, accessibility, privacy and student acceptance of process-based assessment.
- Replication beyond undergraduate introductory courses and English-language literature.
Living evidence record
Impact record IAI-1TA3TPM
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
10 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what AI in Education published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 10 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Education
Are Kurdistan universities ready to govern generative AI?
Not visibly, according to a peer-reviewed audit of 30 university websites. Twenty-one institutions were in the study's lowest readiness stage and only two showed direct public GenAI guidance—but the audit measured published evidence in June 2026, not confidential policy or actual classroom practice.
9 min · 2 sources
Education
What should design students know about AI?
Researchers in Guangzhou developed a 23-item scale covering technical skills, tool use, ethics and originality, perceived value, and independent evaluation of AI output. Two Chinese student samples supported the five-factor structure, but the instrument still needs cross-cultural and outcome validation.
8 min · 1 source
Education
Can expert checks make AI study materials useful?
A two-cohort economics-course study associated tutor-verified AI materials with fewer marks below a UK degree boundary, but the design cannot isolate verification, rule out cohort differences or prove that use caused the result. It is an unreviewed preprint.
8 min · 1 source
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.