Should AI wait until people think first?
In an unreviewed study of 398 adults across five countries, a writing assistant withheld direct generation until users stated a position and supporting argument. The design shifted effort earlier and improved one conditional error-classification measure, but did not significantly improve overall judgment accuracy.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Whether delaying direct generative assistance until a writer contributes a position and argument changes effort, AI use and later source-evaluation performance

At a glance
- 1The final sample contained 398 adults recruited through Prolific in Canada, India, South Africa, the UK and the US, divided among human-only, standard-chatbot, engagement-based unlock and time-matched unlock conditions.
- 2Engage-to-Unlock users classified already-detected passage errors more accurately than standard-chatbot users, 90.43% versus 78.12%, but overall judgment and complete-diagnosis accuracy did not significantly differ.
- 3The time-matched condition was recruited in a second phase rather than concurrently randomised, and 14 of its 94 sessions had an unlock-delivery deviation; the paper is an unreviewed preprint from a Google DeepMind internship.
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Living evidence record
Impact record IAI-0WD9I45
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
3 October 2026
Source trail
2 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
The intervention changes when direct generation appears
Many workplace AI products optimise for immediate assistance: the user can ask for a draft before forming a view. Engage-to-Unlock tests a different interface rule. It begins in a guidance mode called Teach-me and reveals the direct-generation mode, Tell-me, only after the user's draft contains both a position and a developed supporting claim. The question is not whether to ban AI, but whether a small capability boundary can preserve early human sense-making before automation becomes available.
The preprint was submitted on 1 October 2026 and has not been peer reviewed. Its authors are affiliated with ETH Zurich, Google DeepMind, King's College London, Google London, Google Research, the University of Pennsylvania and Carnegie Mellon University. The paper says the work was completed during a student researcher internship at Google DeepMind in London. That institutional context matters because the intervention uses a Gemini model and concerns a product-design choice that AI providers could deploy.[1][2]
Four conditions separate engagement from a simple delay
Of 445 people who completed the session, 398 passed two attention checks and entered the final analysis. They were recruited through Prolific in Canada, India, South Africa, the United Kingdom and the United States and received US$30 for roughly one hour. The sample's mean age was 35.7; 204 participants were women, 184 men and ten were non-binary, other or unspecified. Nearly all—97.5%—reported an undergraduate degree, so the results should not be assumed to represent the wider workforce.
Phase 1 randomly assigned participants to Human-Only (99), Standard Chatbot (98) or Engage-to-Unlock (107). A separate Phase 2 recruited 94 people for Time-Matched Unlock. These users received Tell-me after the same elapsed time as a matched Engage-to-Unlock participant, regardless of their own progress. That yoked condition is important because it asks whether the trigger needs to reflect genuine contribution or whether simply waiting produces similar effects. But it was not randomised concurrently with Phase 1, leaving open recruitment-phase differences.[2]
The writing and evaluation tasks used the same evidence
Participants first wrote at least 200 words arguing about a four-day workweek using summaries of four source documents—two supportive and two critical. They then evaluated eight 150-word AI-generated passages based on the same summaries, without AI assistance. Four passages were sound and four contained either evidence-reporting or reasoning errors. Across 398 people, that produced 3,184 judgments, including 1,592 flawed-passage trials. The design therefore measures both what people wrote and whether they later spotted errors in related material.
In Engage-to-Unlock, an LLM-based discourse classifier watched for a position and at least one distinct argumentative development. On a separate 30-essay annotated corpus, the classifier achieved 91.5% micro-F1 and 83.6% macro-F1. Those figures are useful but not a guarantee that every unlock was fair: a false negative could delay access for a writer who had already thought through the problem, while a false positive could open generation prematurely. Fewer than half of Engage-to-Unlock participants who obtained Tell-me access subsequently prompted it.[2]
One conditional accuracy result improved; overall accuracy did not
Overall passage-judgment accuracy ranged from 61.61% to 66.36% across conditions, with no statistically significant difference. On the four flawed passages, the share both detected and correctly classified was 44.16% for Engage-to-Unlock, 38.27% for Standard Chatbot and 44.68% for Time-Matched Unlock; neither planned comparison was significant. These outcomes are the closest to the practical question of whether someone actually catches and diagnoses unreliable AI content, so the study does not establish a general accuracy gain.
A narrower result did differ. Among flaws participants had already detected, Engage-to-Unlock users classified the error correctly 90.43% of the time, compared with 78.12% for Standard Chatbot; the Holm-adjusted p-value was .003. The corresponding 83.58% in Time-Matched Unlock was not significantly different from Engage-to-Unlock after correction, at p=.051. Because the denominator excludes missed flaws, the result describes diagnosis conditional on detection. It should not be reported as a 90% overall ability to police AI errors.[2]
The design moved time and effort rather than saving it
Engage-to-Unlock participants spent more time on the writing task and less on the later evaluation than Standard Chatbot users. Their mean evaluation time was 17.41 minutes, versus 22.21 minutes under Standard Chatbot and 20.29 under Time-Matched Unlock. Combined time across the two tasks did not differ significantly. The strongest interpretation is therefore effort reallocation: requiring an initial contribution moved some work earlier rather than making the whole session shorter.
Participants in Engage-to-Unlock also reported more effort than Standard Chatbot users and sent more prompts, mostly in Teach-me mode. The study found no significant adjusted differences in perceived process control, satisfaction, document ownership or workflow disruption. These measures counter a simple claim that friction necessarily makes an interface feel worse. Yet they also show that the mechanism did not deliver broad subjective benefits. Whether users would accept the same constraint under deadlines, in expert work or after months of repeated use remains unknown.[2]
What product teams and employers can responsibly test
The useful design idea is capability staging. A system might begin with retrieval, questions and critique; unlock drafting after a worker records their goal and evidence; and reserve rewriting or automation for a later stage. That could be tested in education, policy analysis, legal review or clinical documentation where independent framing matters. It should not become a surveillance device that scores whether an employee has shown enough thinking. Workers need to know what the trigger measures, how to override it and whether their draft is being retained or used for evaluation.
Organisations should measure the outcome they actually care about: factual quality, later independent performance, error detection, completion time, user autonomy and uneven effects across language backgrounds or disability. This experiment measured one short English writing task and a related evaluation immediately afterwards. It did not measure skill retention, transfer to a new domain, workplace productivity, employer outcomes or whether users learn to game the unlock signal.[2]
Company context and study limits constrain the claim
The accessible manuscript does not present a formal competing-interests statement. It does say the work was completed during a Google DeepMind internship, lists several Google or DeepMind affiliations, and used Gemini 3.1 Pro Preview for both assistant modes. That does not invalidate the experiment, but independent replication with other models and teams would reduce the risk that product assumptions, prompting choices or unpublished system behaviour shaped the result.
Fourteen of the 94 Time-Matched Unlock sessions were affected by a timer constraint that partially disrupted the planned unlock schedule. Because that group was also recruited in a separate phase, the near-significant comparison with Engage-to-Unlock is especially difficult to interpret. Confidence would rise with a preregistered, fully concurrent randomised trial; a larger and more educationally diverse sample; repeated tasks over weeks; blinded scoring; and outcomes that test later unaided judgement. For now, the preprint supports a promising interface experiment, not a rule that every AI assistant should force people to wait.[2]
What this means for people
- Writers may benefit when an assistant supports early reflection before offering a full draft, but the current evidence does not show a broad accuracy gain.
- Employers could test staged assistance for high-stakes knowledge work while measuring quality and autonomy rather than assuming more friction is better.
- Workers need disclosure, override options and privacy protections if a system analyses drafts to decide when generative functions unlock.
Global context
Participants came from Canada, India, South Africa, the UK and the US, giving the experiment cross-national reach. It was still an English-language Prolific sample dominated by degree holders, and country-level effects were not the central analysis; local workplace and accessibility evidence remains necessary.
What the evidence does not yet show
- The manuscript is an unreviewed preprint and reports one short English argumentative-writing task followed immediately by a related evaluation.
- Time-Matched Unlock participants were recruited in a separate phase, and 14 of 94 sessions had an unlock-delivery deviation.
- Overall judgment accuracy and complete diagnosis of flawed passages did not significantly improve; the strongest accuracy result was conditional on a flaw already being detected.
- The sample was recruited through Prolific across five countries and 97.5% reported an undergraduate degree, limiting workforce representativeness.
- The study did not measure long-term learning, skill retention, transfer, productivity, unequal access or behaviour under real workplace stakes.
- The work was completed during a Google DeepMind internship and used a Gemini model; no formal competing-interests statement appears in the accessible manuscript.
What to watch next
- Peer review and independent replication using non-Google models and fully concurrent random assignment.
- Longitudinal tests of later unaided performance, skill retention and transfer to new tasks.
- Transparent, user-controllable capability staging that does not turn contribution signals into worker surveillance.
Evidence trail
Sources used for this report
Links checked 3 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Work & Skills
Hiring experiment finds AI skills can improve interview prospects
A study with 1,700 recruiters in the UK and US reports that AI skills increased interview invitations across office, software and design roles, sometimes offsetting age or education disadvantages.
4 min · 1 source
Work & Skills
Does an AI label bias code review—or does job rank matter more?
In a randomized vignette experiment, 447 Microsoft engineers did not penalize an AI-use disclosure on identical code, but rated the same code and author more favourably when the fictional author was labelled Principal rather than Level 1. One AI-normalized US workplace cannot settle how other teams respond.
9 min · 2 sources
Work & Skills
Who can claim the UK's free AI and technology training—and what does it prove?
The Government Digital Service has opened more than 200 free learning and certification pathways to eligible UK civil and public servants until 31 December. Last year's programme recorded over 10,000 registrations and 1,200 certifications, but the government has not published completion, skill-gain or workplace-outcome data.
6 min · 2 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.