Back to the news portal
Science & ResearchResearch paperResearchMulti-source analysisCanadaUnited StatesItalyInternational

Why did AI vision miss a human illusion?

New analysis today of a peer-reviewed 2 October Current Biology experiment. Motion adaptation shifted human judgements and position codes decoded from macaque inferior-temporal cortex, while nine tested artificial-vision networks did not reproduce the effect on their own; this is a targeted benchmark, not proof that the models cannot localise objects.

By The Impact of AI Editorial DeskReleased 4 October 2026 at 06:04 BST8 min read3 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesComputer visionNeuroAIVisual perceptionMotion adaptationBenchmarkingAnimal research

Research topic

Whether object-position representations in primate inferior-temporal cortex and artificial vision networks follow perceived position rather than unchanged pixel coordinates after motion adaptation

The Impact of AI research cover asking why AI vision missed a human illusion, with conceptual eyes, moving gratings and neural-network layers.
AI-generated editorial illustration. The eyes, gratings and network layers are conceptual; they do not depict a study participant, a macaque, a brain recording, a measured result or a particular commercial system.

At a glance

  • 1The work involved 79 online human participants across three separate conditions, two adult male rhesus macaques and nine pre-trained artificial neural networks.
  • 2After horizontal motion adaptation, people reported an opposite-direction position shift and macaque IT decoders shifted in the same general direction; the tested networks encoded position but did not spontaneously reproduce the aftereffect.
  • 3The comparison is narrow and mechanistic. It does not show that all AI vision systems fail at dynamic perception, that a perceptual illusion improves engineering performance or that the biological and behavioural effects were identical in magnitude.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-0I4LNIN

Explore the full tracker

Evidence stage

Studied

Confidence

Corroborated

Reporting basis

Multi-source analysis

Independent support

Present

Record status

Monitoring

Last checked

4 October 2026

Source trail

3 direct sources across 2 source types.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

A test where the pixels stay still but perception moves

The study asks whether a visual system’s estimate of where an object is follows the physical image alone or incorporates recent experience. The researchers used the motion aftereffect: after prolonged exposure to movement in one direction, a stationary test object can appear displaced in the opposite direction. Because the test image itself is unchanged, the design separates perceived position from pixel position more cleanly than an ordinary object-localisation benchmark.

That distinction matters for claims that an artificial network is brain-like. A model can decode an object’s coordinates accurately from a static picture without using the history-dependent computations that shape biological perception. The paper therefore does not reward a network merely for getting the final coordinate right. It asks whether its internal representation changes after the same kind of adaptation that changes human judgements and primate neural activity.[1][2]

The denominator spans people, macaques and nine networks

Seventy-nine people aged 18 to 45 completed online tasks through Amazon Mechanical Turk, with different participants assigned to each condition: 35 performed the no-adapter object-position task, 22 followed rightward motion adaptation and 22 followed leftward adaptation. Participants viewed objects for 100 milliseconds and clicked the perceived centre. The adaptation groups first viewed a moving grating for 30 seconds, with three-second top-ups before later trials. This is a between-group online design, not repeated clinical testing of one representative population.

The neural experiment used two adult male rhesus macaques with chronically implanted electrode arrays in inferior-temporal, or IT, cortex. Neural activity was recorded while the animals passively viewed the stimuli. The modelling arm tested nine networks trained elsewhere for engineering objectives, including feedforward image-recognition models and architectures with temporal or recurrent processing. The authors extracted features from layers selected for alignment with macaque IT data rather than retraining every system on the experimental task.[1][2]

What changed after motion adaptation

Human position judgements moved opposite to the adapting motion along the horizontal axis, as the illusion predicts. The paper reports a mean horizontal change of minus 0.20 degrees after rightward adaptation and plus 0.13 degrees after leftward adaptation; vertical estimates did not shift systematically. The response pattern remained reliable across conditions, suggesting that adaptation changed the estimate rather than simply adding random clicking noise.

Decoders trained on pre-adaptation macaque IT activity and then applied without retraining to post-adaptation activity showed a direction-opponent pattern of the same general kind. The effect was asymmetric: the leftward decoded shift following rightward adaptation was statistically significant, while the rightward trend following leftward adaptation was not. That asymmetry is important. The neural result qualitatively aligned with human perception, but it was not a perfectly balanced replication of the behavioural effect.[1][2]

The networks knew position but missed the history-dependent shift

Several tested networks encoded object position strongly. For example, the paper reports high correlations between decoded and ground-truth position for VGG-16 and ResNet-18 in the baseline task. Yet standard feedforward systems, a recurrent model and a video model did not generate the adaptation-induced displacement on their own. Adding a generic suppression process reduced information in model features but still did not create the direction-specific change seen in the biological data.

The authors could induce a similar bias by transforming model features using empirically derived IT adaptation changes. That is a constructive clue rather than proof of a finished solution: the missing ingredient may involve how motion history reshapes the geometry of later object representations, not simply lower activation after repeated stimulation. A model built or trained to capture that mechanism would still need independent tests on new stimuli and tasks.[1][2]

What this means for people using computer vision

Most deployed computer-vision systems are evaluated on whether outputs are useful and safe, not on whether they reproduce a human illusion. Human-like error is not automatically a design goal. But history dependence can matter when a system works in video, robotics, assistive technology or a changing physical environment. A model that treats each frame as effectively independent may respond differently from a person when recent motion alters perceived location.

The practical contribution is therefore a stress test. Developers can ask whether a system’s position code remains stable for good reasons, or because the architecture lacks a dynamic computation needed in a particular setting. For human–machine collaboration, comparing the conditions under which people and models diverge may be more informative than averaging accuracy across ordinary images. The study does not establish a safety failure in any deployed product.[1][2][3]

Limits keep the conclusion narrower than the headline

The human conditions used separate online groups and a mouse-click task, while the macaque comparison relied on passive viewing and decoded neural activity. The two macaques do not provide a species-wide denominator, and the neural effect’s asymmetry cautions against claiming a one-to-one match. The nine networks are diverse but cannot represent every current or future vision system, training objective or architecture.

The paper is strongest as evidence that a specific form of history-dependent spatial coding is present in the tested biological systems and absent from the tested networks under this protocol. It does not show that the networks cannot localise objects, that IT alone causes the illusion or that adding an IT-derived transformation will improve real-world performance. The authors report no competing financial interests; disclosed support includes Canadian public research programmes, the Simons Foundation Autism Research Initiative, Brain Canada and European Horizon funding.[1][2]

What evidence would change the assessment

Confidence would rise with preregistered replication in larger and more diverse human samples, additional animals, within-participant comparisons and recordings across connected visual areas. Testing unfamiliar objects, motion speeds and natural video would show whether the effect generalises beyond the chosen gratings and 40-image aftereffect set. A causal intervention in IT or connected motion pathways would help distinguish correlation from mechanism.

On the AI side, the decisive test is prospective: train models with explicit history-dependent objectives or biological constraints, freeze them, and evaluate them on held-out illusion and real-world motion tasks. If independent models reproduced the direction-opponent shift without an experiment-specific transformation—and if that behaviour improved human compatibility or task safety—the case for a missing general computational ingredient would become stronger.[1][2]

What this means for people

  • Designers of robotics, assistive vision and video systems gain a concrete benchmark for checking whether recent motion changes a model’s spatial representation.
  • People working alongside vision systems may benefit when developers study predictable human–machine disagreements instead of reporting only average accuracy.
  • The study used invasive recordings from two macaques, making the animal denominator, welfare oversight and limits central to interpreting the evidence.

Global context

The research links York University in Canada, prior neural data collection and animal-care oversight involving US and Canadian standards, and modelling collaboration in Italy. Its engineering implications are international because widely reused vision architectures are deployed across borders. The result is not a regional performance ranking: it is a laboratory comparison of biological and artificial representations under a controlled motion-adaptation task.

What the evidence does not yet show

  • The 79 human participants were recruited online and split across three separate conditions rather than forming a representative or within-participant cohort.
  • Only two adult male rhesus macaques contributed the intracortical data, limiting biological generalisation.
  • The decoded neural response after leftward adaptation was a non-significant trend, so the macaque effect was directionally aligned but asymmetric.
  • Nine tested networks do not exhaust current computer vision; architecture, training data and objectives may materially change performance.
  • Reproducing a human illusion is not itself evidence of better accuracy, safety or usefulness in a deployed system.

What to watch next

  • Independent replication with larger human and animal samples and recordings across additional visual areas.
  • Models trained prospectively for history-dependent spatial coding, then tested on held-out motion and real-world video tasks.
  • Evidence that biological alignment improves a practical human–machine outcome rather than merely reproducing an illusion.

Evidence trail

Sources used for this report

Links checked 4 October 2026

This report is labelled multi-source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Science & Research

People accepted 78% of wrong AI actions when uncertainty stayed hidden

New analysis today of a 30 September preprint: in a controlled puzzle study, people often approved incorrect AI moves when the system did not reveal ambiguity. Targeted warnings helped, but the best-performing warning relied on oracle knowledge that a real product would not have.

7 min · 1 source

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.