Back to the news portal
EducationNew analysis today · source 7 October 2026Research paperResearchSource analysisInternationalUnited KingdomGlobal education

Does AI help students think about learning?

A 33-study meta-analysis found a positive average for technology-supported metacognition, but almost all effects came from self-report and the review did not show that AI outperformed other educational technology.

By The Impact of AI Editorial DeskReleased 8 October 2026 at 07:11 BST8 min read1 source

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The review synthesised 33 studies and 94 effect sizes, estimating a positive average standardised association of Hedges’ g = 0.653 with a 95% confidence interval from 0.520 to 0.786.
  • 2Technology type did not significantly explain differences: 20 studies used non-AI technology, four non-generative AI and nine generative AI, so the review does not establish an AI or generative-AI advantage.
  • 3Ninety of 94 effects relied on self-report, no included study was rated at low overall risk of bias, and the 95% prediction interval ran from about -0.02 to 1.32.
Key themesMetacognitionGenerative AIEducational technologySelf-regulated learningEvidence qualityStudent agency

Research topic

Whether technology-supported educational interventions improve learners’ metacognitive outcomes, and whether effects differ by AI category, implementation design or educational context

The Impact of AI cover showing a conceptual learner reviewing a learning plan, reflection prompts and feedback loops, with the evidence limit that 95.7% of effects relied on self-report.
AI-generated editorial illustration. The learner and interface are conceptual and do not depict a study participant, an evaluated product or measured classroom results.

The direct answer: a positive average, not an AI advantage

Technology-supported educational interventions were associated with better metacognitive outcomes on average in a new peer-reviewed review, but the result cannot be read as proof that AI uniquely improves how students plan, monitor or evaluate their learning. Yijia Yuan and Yiran Du of the University of Cambridge combined 94 effect sizes from 33 studies in a three-level random-effects model. Their pooled estimate was Hedges’ g = 0.653, with a 95% confidence interval from 0.520 to 0.786. That is a moderate standardised difference across the included literature, not a common benefit that every learner or product should be expected to deliver.

The technology comparison is especially important for current AI claims. Twenty studies examined non-AI tools, four examined non-generative AI and nine examined generative AI. Technology type did not significantly moderate the effect, with an omnibus p value of 0.087. Generative-AI studies had the largest descriptive subgroup estimate, g = 0.831, compared with 0.643 for non-AI tools and 0.346 for non-generative AI, but the subgroups were small and heterogeneous. The authors explicitly say those values do not demonstrate comparative effectiveness. A school or university therefore cannot use this paper to claim that adding a chatbot will strengthen metacognition more than a well-designed portfolio, prompt, dashboard or feedback system.[1]

What the review counted and how it handled dependence

The searches covered Web of Science, Scopus, ERIC, IEEE Xplore and PsycINFO through 9 July 2026, with backward and forward citation checks. From 5,326 records, the researchers removed 2,532 duplicates, screened 2,794 unique records and assessed 168 full texts. Thirty-three studies met the criteria: a randomised or quasi-experimental comparison, students in formal K–12 or post-secondary education, a quantitative metacognition outcome and enough data to calculate a standardised between-condition effect. Two reviewers independently screened and coded studies, with reported agreement ranging from substantial to very high.

Nine studies were randomised and 24 quasi-experimental. Nine involved K–12 learners and 24 post-secondary students; 15 were in STEM, 17 in non-STEM subjects and one was not readily classified. Because individual studies often supplied more than one outcome, the authors nested 94 effects within 33 studies instead of treating every estimate as independent. They used restricted maximum likelihood, cluster-robust checks and sensitivity models that assumed different correlations among outcomes from the same participants. The pooled estimate remained close to 0.65 across those alternatives and after removing the two studies rated at serious risk of bias.[1]

The strongest limit is what the outcomes actually measured

Almost all of the evidence concerns what learners said about their own regulation. Self-report questionnaires produced 90 of 94 effects, or 95.7%. Only three effects used performance, judgement or calibration measures, and one used trace or process data. No included study contributed a researcher-rated task product. Metacognitive regulation—such as planning, monitoring, reflection or strategy use—accounted for 77 effects, while metacognitive knowledge accounted for 11 and inseparable composite outcomes for six. A learner reporting greater awareness of study strategies can be meaningful, but it is not the same as accurately detecting misunderstanding, changing strategy at the right moment or transferring a skill to a new unaided task.

Risk of bias compounds that measurement problem. None of the 33 studies received a low overall risk rating. The nine randomised studies all had some concerns; 22 quasi-experimental studies had moderate risk and two had serious risk. The pooled result survived exclusion of the two serious-risk studies, but that leaves a literature still vulnerable to confounding, expectancy, imperfect allocation and outcome-measure limitations. The paper also excluded non-English reports, and its moderator analyses were not preregistered. This is a transparent, carefully qualified synthesis, yet the evidence underneath it is less decisive than the headline average might suggest.[1]

Variation between studies is large and the design signals are exploratory

The 95% prediction interval extended from about -0.02 to 1.32, meaning a comparable future implementation could plausibly range from essentially no effect to a large positive effect. Total I-squared was 72.7%, attributed mainly to differences between studies. That variation is not a statistical footnote: it suggests that the intervention, teaching context, learner group, outcome definition and implementation quality matter more than the general label ‘technology supported’. The average is robust as a description of the reviewed literature, but it is a poor forecast for an unspecified AI product in an unspecified classroom.

Guided interventions, immediate feedback and system-directed support each produced larger estimates in separate exploratory models. None survived the authors’ correction for eight moderator tests; the adjusted p value was 0.095 for each. The features also overlapped, making their independent contributions hard to isolate. System direction can help a novice notice when to reflect, but it can also move regulatory work from the learner to the software. An interface that supplies the plan, the monitoring cue and the correction may improve a questionnaire score while it is present without building independent regulation after support is withdrawn.[1]

What educators can do now—and what would change the assessment

For educators, the practical question is not whether a tool carries an AI label but which decisions it leaves with the learner. A useful activity can require students to state a goal before using assistance, record what feedback changed their judgement, explain why they accepted or rejected a suggestion and then complete a related task without the tool. Teachers can inspect whether support is gradually reduced rather than permanently substituting for planning or evaluation. Procurement claims should distinguish engagement and self-reported awareness from observed monitoring accuracy, retention, attainment and transfer.

The evidence would become decision-ready through preregistered randomised trials that compare defined AI and non-AI supports while measuring baseline skill, observed behaviour, task quality, delayed retention and performance after scaffolding is removed. Studies should report subgroup effects and access conditions, include learners outside English-language publications, and separate the contribution of guidance, feedback timing and learner control. The review reports no funding and no competing interests; its open materials support scrutiny and replication. Its most defensible conclusion is positive but narrow: designed technology support may help learners report stronger metacognitive regulation on average. It does not establish that generative AI is the active ingredient or that students have become better independent learners.[1]

What this means for people

  • Students may benefit from prompts and feedback that preserve responsibility for planning, judging and revising their own work.
  • Teachers should not interpret a moderate pooled self-report effect as proof that a particular chatbot improves learning or independent thinking.
  • Schools and universities can demand outcome evidence on retention and transfer before buying products marketed as metacognitive AI support.

Global context

The review imposed no geographical eligibility restriction but included only English-language reports, and most studies were post-secondary. Its intervention categories span non-AI tools, conventional adaptive systems and generative AI across varied subjects and education levels. That breadth improves relevance but also produces substantial heterogeneity. Local language, curriculum, teacher support, device access and expectations about learner autonomy may change effects, so the pooled estimate should not be transplanted into national policy or procurement models without contextual trials.

What the evidence does not yet show

  • Twenty-four of 33 studies were quasi-experimental, and no included study was rated at low overall risk of bias.
  • Self-report measures produced 90 of 94 effect sizes, leaving little evidence on observed monitoring, calibration, task performance or transfer.
  • The 95% prediction interval ranged from about -0.02 to 1.32 and total I² was 72.7%, indicating substantial between-study variation.
  • The non-generative AI and generative-AI subgroups contained only four and nine studies, respectively, so descriptive subgroup rankings are not comparative proof.
  • Moderator analyses were exploratory and not preregistered; nominal signals did not survive false-discovery-rate adjustment.
  • English-language eligibility may have excluded relevant studies and reduced global representativeness.

What to watch next

  • Preregistered trials that compare specific AI, non-AI and no-technology supports under the same teaching conditions.
  • Observed monitoring accuracy, calibration, task quality, delayed retention and transfer after support is withdrawn.
  • Larger studies able to separate guidance, feedback timing and learner control instead of testing overlapping design bundles.
  • Evidence from multilingual settings and education systems under-represented in the English-language literature.
  • Whether adaptive assistance is faded in ways that strengthen independent regulation rather than making learners dependent on prompts.

Living evidence record

Impact record IAI-0CW1B5H

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

8 October 2026

Source trail

1 direct source across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Single-source reporting disclosure

This record analyses one direct source. It can establish what Educational Psychology Review published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.

Evidence trail

Sources used for this report

Links checked 8 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Education

Did ChatGPT improve writing scores in a 16-week Taiwan course?

Twelve university students used ChatGPT throughout a five-stage inquiry-writing course. Their closed-book writing scores rose by 2.79 points, but the change was not statistically significant and the study had no comparison group.

8 min · 1 source

Education

Does AI help students design climate solutions?

In a peer-reviewed UAE study, 69 students reported greater creativity and agency when using AI, but the research measured perceptions in one course—not learning gains, stronger climate solutions or real-world environmental impact.

7 min · 1 source

Education

Does the way students use AI matter?

A peer-reviewed survey of 713 Saudi undergraduates found that supportive and creative AI use was associated with stronger self-reported creativity and metacognitive awareness, while substitutive use showed negative associations. The one-time self-report design cannot establish what AI caused—or whether skills improved.

8 min · 1 source

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.