Back to the news portal
AI Risks & SafetyResearch paperResearchSource analysisFranceEuropeInternational

Can a robot infer what a person intends?

A peer-reviewed video-language model matched or exceeded reported human scores on five-choice intention questions. When answer options disappeared, text-overlap scores fell below 20—leaving open-vocabulary claims, cultural bias and surveillance risk unresolved.

By The Impact of AI Risks & Safety DeskReleased 4 October 2026 at 12:55 BST7 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesHuman intentionVideo-language modelsRoboticsBenchmarkingSurveillanceCultural bias

Research topic

Whether a two-stage video-language model can infer human goals beyond fixed labels, and how multiple-choice benchmark performance differs from open-ended intention description

The Impact of AI research cover asking whether a robot can infer what a person intends, with a conceptual camera view separating observed action from possible goals.
AI-generated editorial illustration. The person, robot and inferred goal cards are conceptual; they do not depict a participant, care setting, surveillance system or measured deployment.

At a glance

  • 1IntentVLM fine-tunes two Qwen3-VL modules with LoRA: one proposes possible goals from a video and question, and the second selects the most plausible candidate.
  • 2On four five-choice IntentQA categories, the 4B model reported test accuracies from 78.49% to 88.60% and averaged roughly 85%, above earlier reported baselines and human benchmark values.
  • 3The open-ended setting was much weaker: ROUGE scores stayed below 20 on the test set, and the study did not deploy a robot or test whether inferred intentions were fair, safe or useful in real interactions.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-1F8U05V

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

4 October 2026

Source trail

2 direct sources across 2 source types.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

Guessing a goal is harder than recognising an action

A camera may show a person reaching for a kettle. The visible action is not the same as the intention: they might make tea, clean the appliance or move it out of the way. Robots designed to assist proactively need to reason about that hidden goal, yet acting on a confident mistake can be intrusive or dangerous. The challenge grows when a system must describe an unforeseen intention rather than choose from a fixed menu.

IntentVLM, developed at Sorbonne University’s Institute of Intelligent Systems and Robotics, divides the task into two stages inspired by forward–inverse models of cognition. One video-language model generates candidate goals from the video and question. A second model evaluates those candidates and selects the likeliest one. Both use frozen Qwen3-VL backbones with trained low-rank adaptation modules, reducing the number of parameters changed during fine-tuning.[1][2]

The decisive accuracy result comes from five choices

The main intention benchmark is IntentQA. Each example supplies a video, a question, five candidate answers and a labelled correct answer. Questions span four reasoning types: what intention explains an action, how an intention leads to an outcome, what happens next and what earlier intention explains an observed outcome. The authors train on the benchmark’s training set and report validation and test accuracy, comparing their system with earlier graph-based, language-based and video-question-answering models reported by the benchmark creators.

On the test set, IntentVLM scored 84.10% for Causal What, 88.60% for Causal How, 83.95% for Temporal Next and 85.15% for Temporal Past. The corresponding human figures reproduced from the original benchmark were 77.76%, 80.22%, 79.05% and 78.49%, although the new paper does not rerun or describe that human study. The strongest earlier model listed averaged 57.64% on the reported test columns. These are substantial gains in a constrained five-answer task, but the answer choices themselves provide context and eliminate many possible goals.[2]

Removing the choices changes the result

The paper also asks the model to generate a free-form intention without candidate answers. Here, the picture is far less decisive. IntentVLM’s test-set ROUGE-1 score was 19.18 and ROUGE-L 18.91, while cosine similarity was 34.67 and BERTScore F1 87.20. A zero-shot 4B Qwen3-VL baseline actually scored slightly higher on ROUGE and BERTScore. Fine-tuning a direct-answer model reduced several metrics rather than solving the problem.

The authors interpret this as evidence that candidate options constrain and disambiguate the problem, allowing models to exploit structure that vanishes in a truly open response. That is the essential qualification to the paper’s “open-vocabulary” framing. The two-stage system can generate candidates before selecting one, but its strongest evidence still comes from a benchmark with supplied options and a known correct label. High BERTScore alongside low lexical overlap also shows how strongly conclusions depend on the chosen automated metric.[2]

A second benchmark checks retained scene understanding

To test whether specialised fine-tuning damaged broader visual understanding, the authors used Inst-IT Bench. Its image split contains 1,036 question–answer pairs across 338 images; its video split contains 1,001 pairs across 206 videos. Each is available in multiple-choice and open-ended form. The comparison includes similarly sized video-language models and larger alternatives, with the paper reporting that the 4B model remained close to its zero-shot counterpart after intention training.

That result addresses catastrophic forgetting on one external benchmark, not general understanding in the world. Benchmark clips are curated and have fixed questions, while a robot sees continuous, noisy scenes and must decide when evidence is sufficient. The study reports no physical robot experiment, live interaction, latency, calibration, abstention rate or consequence of a wrong inference. It therefore does not establish that the model knows when it is uncertain or that acting on its answer improves human–robot collaboration.[2]

Intention inference can become surveillance with a psychological label

The authors’ responsible-innovation statement explicitly warns that continuous video monitoring can be experienced as surveillance. It also notes that limited public egocentric datasets may encode cultural bias and that errors in assistive or caregiving settings could harm vulnerable people. Those are not peripheral concerns. Intention is an interpretation of behaviour, not a directly observed fact, and different cultures, disabilities and contexts can make the same movement mean different things.

A system deployed in a workplace, shop, school or care home could turn uncertain inferences into interventions or records about motivation. People may change their behaviour if they feel watched, and those least represented in training data may be misread most often. Meaningful safeguards would include visible consent, strict purpose limits, local processing where possible, short retention, access to logs, human confirmation before consequential action and a genuine option not to be profiled.[2]

What would justify a claim about real human intention

A stronger evaluation would first publish the exact train, validation and test denominators used, test on people and environments not represented in the source data, and report calibration and abstention as well as accuracy. Independent annotators from different cultural backgrounds should assess whether multiple intentions are plausible rather than assuming one label is ground truth. Open-ended answers should be judged for usefulness and harm, not only text similarity.

The next step after that is a consented live study in which a robot observes without acting, followed only later by tightly bounded assistance with human confirmation. Researchers should measure false interventions, missed opportunities, participant comfort, subgroup performance and whether the system improves an outcome people value. The accessible paper did not state a specific funding acknowledgement or competing-interest declaration. For now, IntentVLM is a strong multiple-choice benchmark result and a candid demonstration of how much harder free-form intention inference remains—not evidence that cameras can read minds.[1][2]

What this means for people

  • Robots could offer help sooner when a person’s goal is clear, but a mistaken inference can interrupt, patronise or endanger them.
  • Workers, residents and shoppers may be profiled through video without knowing that a system is assigning goals to their behaviour.
  • People whose gestures or routines differ from benchmark norms may experience more errors unless evaluation includes their contexts and gives them control.

Global context

The research is from France and relies on public video benchmarks assembled from broader visual datasets. Reported averages do not establish performance across languages, cultures, disabilities or legal regimes. The EU AI Act and data-protection rules may constrain biometric and workplace uses, while governance and surveillance norms differ elsewhere. Cross-regional evidence is needed before a benchmark score can support deployment.

What the evidence does not yet show

  • The highest accuracies come from five-choice benchmark questions; free-form generation performed much less convincingly.
  • Human scores were reproduced from the original IntentQA study rather than collected under the new model’s protocol.
  • No robot, live participant, continuous scene, latency test or real-world intervention was evaluated.
  • One labelled answer may not capture multiple plausible intentions, and public egocentric video datasets may encode cultural and demographic bias.
  • The accessible paper did not state a specific funding source or competing-interest declaration.

What to watch next

  • Pre-registered evaluations with exact split denominators, calibration, abstention and multiple plausible labels.
  • Independent cross-cultural annotation and subgroup analysis across disability, age and interaction context.
  • Consented observational studies before any system is allowed to act on inferred intention.
  • Governance that prevents intention inference from becoming persistent workplace, retail, school or care surveillance.

Evidence trail

Sources used for this report

Links checked 4 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

AI Risks & Safety

Can hashes make AI conversations auditable without publishing them?

A peer-reviewed experiment converted nearly four million public chatbot interactions into cryptographic commitments and detected eight induced ledger manipulations. It is a promising integrity mechanism, not proof that a conversation is true, authorised or safe.

9 min · 1 source

AI Risks & Safety

Can hidden image text mislead dental AI?

A peer-reviewed German stress test found that adversarial text placed inside 270 dental radiographs could flip four vision-language models from an abnormal to a normal finding. OCR sanitisation sharply reduced the measured attacks, but the experiment used a permissive prompt, a pathology-heavy benchmark and no live clinical system.

7 min · 3 sources

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.