Can people reliably spot AI-written text?
Not in this Japanese experiment. Across 3,120 judgements, four reported reading strategies produced only 41% to 45% accuracy. The result cautions against accusing a person from prose style alone, but one human text, one prompt and retrospectively grouped explanations sharply limit its reach.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The study analysed 3,120 judgements from 390 Japanese adults: every participant assessed the same seven LLM-generated texts and one human-written public comment.
- 2Four clusters of retrospectively reported strategies achieved mean accuracy between 41% and 45%; the largest, labelled higher-order pattern recognition, reached 45%, still roughly chance performance in this balanced task.
- 3The experiment cannot validate a universal detection rate: it used one prompt, one human comparison, a 2025 data collection and weakly separated strategy clusters. It does support a practical warning against attributing authorship from style alone.
Research topic
How self-reported judgement strategies relate to adults' accuracy when distinguishing AI-generated from human-written text
The direct answer: none of the four strategies worked reliably
People in this study did not reliably distinguish AI-written from human-written prose. The researchers grouped participants' explanations into four judgement strategies, from intuition to what they called higher-order pattern recognition. Average accuracy ranged from 41% to 45%. The largest and ostensibly most sophisticated group achieved 45%, which is not a dependable basis for deciding whether a person wrote a particular passage.
That finding matters wherever an 'AI-like' tone can trigger suspicion: a teacher reviewing coursework, an editor considering a submission, a recruiter reading an application or a moderator assessing a post. The paper does not prove that detection is impossible in every context. It does show that confidence in stylistic tells can exceed their demonstrated value. A consequential authorship decision needs provenance, drafting history, disclosed tool use and a fair review process—not a hunch about polished syntax, repetition or emotional flatness.[1]
The Impact Brief · Free
Follow the evidence in society & media.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
What 390 adults were asked to judge
The analysis used data from a Japanese online survey conducted through Questant and the Macromill panel on 24 and 25 January 2025. The final sample included 390 adults aged 22 to 69, with a mean age of 47.8; 227 were men and 163 women. The researchers excluded 13 responses—10 without a gender response and three whose explanations were considered meaningless or uninterpretable. This was a secondary analysis and was not preregistered.
Each participant saw the same eight target texts. Seven were generated by different large language models—GPT-4o, o1, Claude 3.5, Gemini, Copilot, Llama 3.1 and Perplexity—and one was a human-written public comment. Participants rated whether each text was AI- or human-authored on a seven-point scale. The researchers later reduced those ratings to a binary classification, producing 3,120 judgements. After all eight decisions, participants gave an open-ended explanation of how they had decided.
The denominator is important. The study did not sample many human authors, genres, prompts or editing practices. It tested repeated judgements over a fixed, artificial set with seven AI texts and one human text. Although the binary scoring treated AI and human answers as two classes, the underlying set was not a broad or balanced representation of everyday writing. The results therefore describe this task, not an intrinsic human detection ability across school essays, news, workplace messages or literature.[1]
How the researchers derived four reading strategies
The team converted the open-ended Japanese explanations into sentence embeddings using a Japanese BERT model, then applied hierarchical clustering. They interpreted four clusters as unsubstantiated intuition, stylistic cues, syntactic cues and higher-order pattern recognition. The first captured answers resting mainly on a feeling; the next two referred to vocabulary, tone or sentence structure; the last described broader coherence, contextual naturalness or patterns across a whole passage.
These labels are researcher interpretations of retrospective comments, not experimentally assigned behaviours. Participants explained their approach only after completing all eight judgements, so the account may compress, rationalise or forget what happened text by text. The study also interprets the strategies along a System 1–System 2 continuum, but it did not directly measure cognitive processing, attention or response time. It cannot establish that one group genuinely used a stable mental system while another did not.
The clusters themselves were not sharply separated. The reported silhouette score was 0.1467 with two clusters, 0.0635 with three and 0.0496 with four; higher values ordinarily indicate clearer separation. Four clusters were retained partly because they were interpretable. An alternative embedding model changed cluster composition. That does not make the exercise useless, but it means the exact four-part taxonomy should be treated as an exploratory description rather than a validated scale of detection strategies.[1]
The detailed results show inconsistency, not a winning cue
Mean accuracy was 0.41 for the unsubstantiated-intuition group, 0.43 for the stylistic-cue group, 0.44 for the syntactic-cue group and 0.45 for the higher-order group. Group sizes were 61, 65, 48 and 216 respectively. Hierarchical Bayesian logistic regression adjusted for gender, age and prior LLM use, but the central practical result remained modest: the strategies did not yield a reliable separator.
Performance also varied substantially by model. Participants correctly identified 56.4% of Claude texts and 55.9% of Llama 3.1 texts, compared with 24.9% for o1 and 32.6% for Copilot. The corresponding figures were 42.3% for GPT-4o, 41.3% for Gemini and 52.3% for Perplexity. Across the seven AI texts, accuracy was 43.7%. The single human passage was correctly identified by 45.4% of participants.
Those model-specific percentages should not be turned into a league table. Each model contributed one response to one task, and generation settings, prompt wording and subsequent editing can all change detectable surface features. A different prompt or genre could reverse the pattern. Because every participant judged the same material, the study can compare reactions to those eight texts, but it cannot estimate how often a named model is detectable in the wider world.[1]
What teachers, editors and workers should do differently
For educators, the humane implication is procedural. A student's vocabulary, tidy structure or lack of idiom should not by itself become evidence of misconduct. Such cues can also reflect translation, disability, formal training, anxiety or a writer working outside their first language. Institutions should publish what assistance is allowed, ask for process evidence proportionate to the assessment and provide an appeal path. Drafts, cited sources and a short discussion of reasoning are more informative than a style-based accusation.
Editors and employers face the same risk. An automated or intuitive judgement can unfairly penalise people whose writing happens to resemble a stereotyped model output. If authorship matters, request disclosure and preserve document history; if quality matters, assess accuracy, originality, evidence and fitness for purpose directly. Moderators should separate the question 'was a model involved?' from 'does this violate a rule or cause harm?' Model assistance does not make a claim false, and human authorship does not make it true.
The study does not evaluate commercial AI detectors, watermarking or cryptographic provenance systems, so it cannot rank those tools. Its result is narrower but still consequential: unaided readers did poorly here, and naming a more elaborate reading strategy did not solve the problem. Human review remains essential, but reviewers need evidence beyond stylistic confidence.[1]
What remains unproven
The data were collected in early 2025, while the paper was published in October 2026. Models and public familiarity may have changed during that interval. The sample came from one Japanese web panel and the task was presented in Japanese, limiting transfer to other languages, countries and media environments. Prior LLM use was included as a covariate, but the design cannot determine whether training people on particular cues would improve performance or simply make them confidently wrong in new situations.
Reducing a seven-point judgement to two categories also discards confidence and ambiguity. A participant who chose the midpoint and one who expressed extreme certainty may count the same after dichotomisation. With only one human text, the study cannot describe how strategy performance varies among different human writers. Nor can it separate knowledge of a model's characteristic output from general skill at recognising authentic human context.[1]
Funding, disclosure and what would change the assessment
The work was funded by JSPS KAKENHI grant JP23K11107. The authors reported that the funder had no role in the study beyond supporting an English-editing fee and declared no commercial or financial conflict of interest. They disclosed using generative AI for English editing, followed by review by a native English speaker. Those disclosures are relevant to transparency but do not alter the evidential limits of the design.
A stronger assessment would require preregistered studies with many independently sampled human and AI texts across prompts, languages, genres and editing conditions. Researchers should preserve confidence ratings, test strategies prospectively, compare trained and untrained readers, and measure false accusations as well as successful detections. Field studies should examine fair procedures in schools and workplaces. Until then, the most defensible conclusion is practical rather than universal: in this controlled task, people and their stated strategies were not reliable enough to determine authorship on style alone.[1]
What this means for people
- Students and workers should not face consequential authorship accusations based only on prose that sounds 'AI-like'.
- Teachers, editors and moderators need proportionate process evidence, disclosed rules and human appeal routes.
- People writing in a second language or with atypical styles may be especially exposed to false suspicion when stereotype replaces evidence.
Global context
This study provides peer-reviewed evidence from Japan, where the language and web-panel context are important parts of the task. Detection cues do not transfer automatically across languages, school systems or workplace rules. International policy should therefore focus on transparent permitted-use standards and verifiable process evidence, not a universal list of supposedly human or machine stylistic traits.
What the evidence does not yet show
- All participants judged the same seven AI outputs and one human comment, so the text sample cannot represent the range of models, people, prompts, genres or editing practices.
- Strategies were clustered from retrospective explanations, not assigned or observed prospectively, and the four-cluster solution had weak separation.
- The study was exploratory, secondary and not preregistered; its System 1–System 2 interpretation was not directly measured.
- The seven-point judgement scale was dichotomised, losing information about confidence and ambiguity.
- Japanese web-panel data collected in January 2025 may not generalise to other languages, populations or rapidly changing 2026 models.
What to watch next
- Preregistered replications with many human authors, prompts, genres and current model versions.
- Prospective tests that assign or teach detection strategies instead of inferring them from later explanations.
- False-accusation rates and appeal outcomes in schools, publishing and workplaces.
- Comparisons with provenance records, watermarking and independently validated detection tools.
- Results that retain confidence ratings and test performance across languages and disability-related writing differences.
Living evidence record
Impact record IAI-1OYQYCN
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
9 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Frontiers in Artificial Intelligence published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 9 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Society & Media
Can people spot AI-generated images?
A preregistered experiment with 2,091 adults in Denmark found that people were slightly more accurate than random probability guesses, yet on average would have scored better by assigning every image a 50% chance of being AI-generated. The test used one 2024 image tool and does not measure belief, sharing or harm.
10 min · 1 source
Society & Media
How did fake experts reach real newsrooms?
OpenAI says two covert influence operations used false journalist identities and a front research centre to place material in real outlets. AI mostly helped with drafting, translation and internal reporting; the harder failure was identity and source verification, and claimed reach remains only partly corroborated.
10 min · 2 sources
Society & Media
What is changing in ChatGPT for teens?
OpenAI says a US College Planner for grades 10–12 is coming, alongside flashcards, easier quizzes and multi-photo note capture. The company also released large usage counts, but no study protocol, denominators for several comparisons or evidence that the tools improve learning or admissions outcomes.
7 min · 1 source
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.