Can people spot AI-generated images?
A preregistered experiment with 2,091 adults in Denmark found that people were slightly more accurate than random probability guesses, yet on average would have scored better by assigning every image a 50% chance of being AI-generated. The test used one 2024 image tool and does not measure belief, sharing or harm.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1In a preregistered experiment, 2,091 adult Danes each rated seven images drawn from a 52-image pool containing authentic and AI-generated versions of 26 real events.
- 2Participants beat simulated random probability responses, but their mean Brier score was 0.034 worse than the simple strategy of assigning every image a 50% chance of being AI-generated.
- 3The study used one image generator in June 2024, primed participants that some images were synthetic, and omitted source and engagement cues; it cannot establish present-day global detection rates or downstream harm.
Research topic
How accurately do people distinguish authentic from AI-generated images across countries, current generators and realistic platform contexts, and which interventions improve calibrated judgment without making people distrust genuine evidence?

The direct answer: not reliably in this controlled 2024 test
Most participants could not reliably tell the AI-generated images from authentic ones. Their probability judgements were slightly better than a simulation that produced random answers between 0 and 100, but they were worse than a deliberately uninformative rule: giving every image a 50% probability of being AI-generated. That second comparison is the practical headline because it shows that confidence in visual clues often failed to add useful information.
The finding does not mean every person was wrong half the time, nor that all synthetic images are undetectable. The researchers used Brier scores, which reward both correct discrimination and well-calibrated confidence. A confident wrong answer is penalised more than a cautious one. The study therefore tests whether people's probability judgements tracked the true image source, not only whether a forced real-or-fake label happened to be correct.
It also does not measure whether anyone believed the depicted event, shared the image or changed a political view. Every authentic and synthetic image referred to a real event, and participants were told that some images had been generated. The evidence supports caution about unaided visual inspection; it does not quantify misinformation exposure or prove a particular policy intervention works.[1]
How 2,091 adults judged a 52-image pool
The polling company Epinion ran the survey from 17 to 24 June 2024 among Danish adults aged 18 and over. The final sample contained 2,091 respondents, with an 8.4% dropout rate, and was constructed to represent the adult population on age, gender and region. The experiment was preregistered and approved by the University of Copenhagen's institutional review board.
Each respondent saw seven images, one at a time, randomly drawn from a pool of 52. The pool covered 26 events; each event had one authentic image and one AI-generated version. A person could not see the same event twice. The images spanned four categories: everyday scenes, people, crises and political subjects. After each image, respondents rated from 0 to 100 how likely it was to have been generated with AI.
The authentic images came from Wikimedia Commons. The synthetic images were made with OpenArt.ai, taking roughly five to 15 minutes per image, in an attempt to represent the resources of a small online actor rather than a state-backed production team. Participants received a short description of the event but no account identity, source reputation or engagement metrics. The median viewing time was seven seconds, reflecting rapid social-media browsing more closely than forensic examination.
That design isolates visual discernment and avoids one major confound: recognising a famous false image from memory. But it also creates an artificial task. People were warned that some images were synthetic and all examples depicted true events. On a real feed, users must also judge whether the event itself is real, whether a caption is misleading, and whether the account deserves trust.[1]
What the scores show—and what 'chance' means here
The researchers calculated a Brier score for each participant from the seven probability estimates. A score of zero would be perfect; lower is better. Participants' average was 0.05 lower than the distribution produced by random probability estimates, a statistically significant difference with a 95% confidence interval from -0.057 to -0.042. In that limited sense, the group did extract some source information from the images.
Yet a constant 50% answer produces a Brier score of 0.25. Participants' mean score was 0.034 higher, and therefore worse, than that benchmark; the 95% confidence interval was 0.028 to 0.039. The distinction matters. Randomly moving a probability slider is a weak comparator. A constant 50% response represents honest uncertainty and, on average, outperformed attempts to infer authenticity from the available visual cues.
No image category was statistically distinguishable from the others in discernment performance. Images of people were not reliably easier than ordinary scenes. Participants who scored highly on a three-question cognitive-reflection test did better than those who scored poorly, with an average Brier-score improvement of 0.04. Their average of 0.26 still did not beat the 0.25 constant-uncertainty benchmark. The only subgroup reported below 0.25 was people aged 18 to 24.
Women and respondents who reported experience with generative AI also had better Brier scores in the main models, while older respondents did worse. These are subgroup associations inside one experiment, not universal demographic traits. The paper notes that the gender difference was no longer statistically significant when using raw probability estimates rather than Brier scores, suggesting that calibration, not necessarily discrimination, contributed to the result.[1]
Why the result should change verification practice
For readers, the useful lesson is not to become suspicious of every polished or imperfect image. It is to treat visual intuition as one weak signal. Reverse-image search, provenance information, corroboration from independent sources and examination of the original account can provide evidence that a seven-second glance cannot. A visible oddity may prompt a check, but the absence of an oddity is not proof of authenticity.
For publishers and platforms, the result argues against making users the final detection layer. Content credentials, accessible provenance, clear labelling, rapid correction routes and professional verification can reduce reliance on individual eyesight. None is a complete solution: metadata can be removed, labels can be misread and automated detectors can fail after generators change. Effective systems need multiple signals and a way to express uncertainty rather than a single real-or-fake badge.
There is a symmetric risk. Teaching people that synthetic media are common without supplying reliable verification tools can make them dismiss genuine evidence—the so-called liar's dividend. A useful intervention should therefore improve discrimination and calibration while measuring false suspicion of authentic images. Raising doubt alone is not success.
The study's policy discussion points to watermarks, transparency standards, platform monitoring and stronger independent media and fact-checking. Those proposals were not evaluated in the experiment. They remain hypotheses whose effects should be tested against evasion, accessibility, user comprehension and the risk that bad actors simply move to unlabelled tools.[1]
Important limits: one country, one older generator and a primed task
The data are Danish and were collected more than two years before publication. Denmark has high internet use and digital literacy, but its media system, institutional trust and population do not represent the world. The authors argue that these characteristics may make Denmark a demanding test for the pessimistic claim, yet that inference still needs direct replication across languages, cultures, ages and levels of connectivity.
All synthetic images came from one freely available 2024 tool and were manually selected after multiple prompts. Current models, editing workflows and provenance systems differ. Better generation could make detection harder; new visible artefacts or labels could make it easier. A result about one stimulus set should not be converted into a timeless percentage for all AI imagery.
The experiment also told participants they would see authentic and AI-generated material. That may encourage unusually close attention to pixels while removing the source cues people normally use. Conversely, participants spent little time and saw relatively low-quality images, which may resemble casual feeds. These design choices pull external validity in different directions, so the study is most persuasive as evidence of a problem, not a precise forecast of real-world performance.
Data collection was conducted as part of a grant from Denmark's Agency for Culture and Palaces and Agency for Digital Government. The authors separately reported no financial support for the research, authorship or publication of the article and declared no potential conflicts of interest. Data and code are available through a linked Dataverse repository, supporting independent reanalysis.[1]
What evidence would change the assessment
The next decisive studies would repeat the preregistered design with current generators, multiple tools and images produced at several effort levels. They should include countries with different languages, media environments and levels of digital access; report accuracy and calibration separately; and predefine how subgroup comparisons will be handled. Reusing some common stimuli across countries would help distinguish cultural effects from image effects.
Realistic feed experiments should then add account identity, captions, engagement signals and competing demands on attention. They should measure not only source classification but also belief, sharing, reporting and memory. A promising intervention would improve correct identification of synthetic images without increasing rejection of authentic ones, and its benefit would persist when labels are incomplete or adversaries adapt.
Until those tests exist, the careful conclusion is narrow but consequential: in this large Danish experiment, unaided visual judgement was not a dependable authentication system. People, newsrooms and platforms need evidence that travels with or can be independently connected to an image, plus verification processes that remain useful when certainty is impossible.[1]
What this means for people
- Readers should not treat visual confidence as authentication; source checks and corroboration remain important even when an image looks plausible.
- Publishers and platforms should not shift the entire burden to users, who performed poorly even when warned that synthetic images were present.
- Poorly designed warnings can also make genuine evidence easier to dismiss, so interventions must measure false distrust as well as detection.
Global context
Synthetic-image governance is a global information-integrity problem, but this evidence comes from one high-connectivity European country. Detection ability may vary with media literacy, trust, platform use, language and access to verification tools. International standards for provenance could help evidence travel across borders, while implementation must work for low-bandwidth settings and should not assume that every publisher or user can perform forensic analysis.
What the evidence does not yet show
- The experiment was conducted only in Denmark, so the population result cannot be assumed to hold across cultures, languages and media systems.
- Data were collected in June 2024 with one image-generation service; the study does not benchmark current or multiple generators.
- Participants were primed that some images were synthetic and saw no source, account or engagement cues, making the task different from a real feed.
- Each respondent rated seven images, and the stimuli all depicted true events; the study does not measure false claims, belief, sharing, persuasion or harm.
- Subgroup findings are modest and model-dependent; they should not be used to label demographic groups as inherently reliable or unreliable.
What to watch next
- Preregistered multi-country replications using several current image generators and effort levels.
- Tests that separate discrimination from confidence calibration and report false suspicion of authentic images.
- Realistic feed studies incorporating source identity, captions, engagement cues and competing attention.
- Randomised evaluations of provenance labels, content credentials, media-literacy prompts and platform interventions.
- Longitudinal evidence showing whether any detection skill survives generator changes and adversarial adaptation.
Living evidence record
Impact record IAI-0VH4JY1
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
9 October 2026
Source trail
1 direct source across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Single-source reporting disclosure
This record analyses one direct source. It can establish what Social Media + Society published or reported, but it is not independent corroboration of every performance claim or predicted outcome. The confidence label will change only when broader evidence is added.
Evidence trail
Sources used for this report
Links checked 9 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Society & Media
How did fake experts reach real newsrooms?
OpenAI says two covert influence operations used false journalist identities and a front research centre to place material in real outlets. AI mostly helped with drafting, translation and internal reporting; the harder failure was identity and source verification, and claimed reach remains only partly corroborated.
10 min · 2 sources
Society & Media
UK Electoral Commission reviews AI's role in the 2026 elections
The Electoral Commission reports on digital campaigning and AI around the 2026 elections, examining voter confidence, campaign transparency and the evidence available to regulators.
2 min · 1 source
Society & Media
Why do people see ChatGPT so differently?
Convenience samples in Germany and Serbia differed sharply in AI literacy and in whether they saw ChatGPT as useful or risky. The study maps associations among knowledge, beliefs and attitudes, but its self-report design cannot show that literacy caused trust or that the samples represent either country.
10 min · 1 source
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.