Can a robot infer what a person intends?
A peer-reviewed video-language model matched or exceeded reported human scores on five-choice intention questions. When answer options disappeared, text-overlap scores fell below 20—leaving open-vocabulary claims, cultural bias and surveillance risk unresolved.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
Research topic
Whether a two-stage video-language model can infer human goals beyond fixed labels, and how multiple-choice benchmark performance differs from open-ended intention description

At a glance
- 1IntentVLM fine-tunes two Qwen3-VL modules with LoRA: one proposes possible goals from a video and question, and the second selects the most plausible candidate.
- 2On four five-choice IntentQA categories, the 4B model reported test accuracies from 78.49% to 88.60% and averaged roughly 85%, above earlier reported baselines and human benchmark values.
- 3The open-ended setting was much weaker: ROUGE scores stayed below 20 on the test set, and the study did not deploy a robot or test whether inferred intentions were fair, safe or useful in real interactions.
The Impact Brief
Keep the evidence trail, not the noise.
Get the most consequential AI developments with direct sources and clear limits.
Living evidence record
Impact record IAI-1F8U05V
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent support
Present
Record status
Monitoring
Last checked
4 October 2026
Source trail
2 direct sources across 2 source types.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Guessing a goal is harder than recognising an action
A camera may show a person reaching for a kettle. The visible action is not the same as the intention: they might make tea, clean the appliance or move it out of the way. Robots designed to assist proactively need to reason about that hidden goal, yet acting on a confident mistake can be intrusive or dangerous. The challenge grows when a system must describe an unforeseen intention rather than choose from a fixed menu.
IntentVLM, developed at Sorbonne University’s Institute of Intelligent Systems and Robotics, divides the task into two stages inspired by forward–inverse models of cognition. One video-language model generates candidate goals from the video and question. A second model evaluates those candidates and selects the likeliest one. Both use frozen Qwen3-VL backbones with trained low-rank adaptation modules, reducing the number of parameters changed during fine-tuning.[1][2]
The decisive accuracy result comes from five choices
The main intention benchmark is IntentQA. Each example supplies a video, a question, five candidate answers and a labelled correct answer. Questions span four reasoning types: what intention explains an action, how an intention leads to an outcome, what happens next and what earlier intention explains an observed outcome. The authors train on the benchmark’s training set and report validation and test accuracy, comparing their system with earlier graph-based, language-based and video-question-answering models reported by the benchmark creators.
On the test set, IntentVLM scored 84.10% for Causal What, 88.60% for Causal How, 83.95% for Temporal Next and 85.15% for Temporal Past. The corresponding human figures reproduced from the original benchmark were 77.76%, 80.22%, 79.05% and 78.49%, although the new paper does not rerun or describe that human study. The strongest earlier model listed averaged 57.64% on the reported test columns. These are substantial gains in a constrained five-answer task, but the answer choices themselves provide context and eliminate many possible goals.[2]
Removing the choices changes the result
The paper also asks the model to generate a free-form intention without candidate answers. Here, the picture is far less decisive. IntentVLM’s test-set ROUGE-1 score was 19.18 and ROUGE-L 18.91, while cosine similarity was 34.67 and BERTScore F1 87.20. A zero-shot 4B Qwen3-VL baseline actually scored slightly higher on ROUGE and BERTScore. Fine-tuning a direct-answer model reduced several metrics rather than solving the problem.
The authors interpret this as evidence that candidate options constrain and disambiguate the problem, allowing models to exploit structure that vanishes in a truly open response. That is the essential qualification to the paper’s “open-vocabulary” framing. The two-stage system can generate candidates before selecting one, but its strongest evidence still comes from a benchmark with supplied options and a known correct label. High BERTScore alongside low lexical overlap also shows how strongly conclusions depend on the chosen automated metric.[2]
A second benchmark checks retained scene understanding
To test whether specialised fine-tuning damaged broader visual understanding, the authors used Inst-IT Bench. Its image split contains 1,036 question–answer pairs across 338 images; its video split contains 1,001 pairs across 206 videos. Each is available in multiple-choice and open-ended form. The comparison includes similarly sized video-language models and larger alternatives, with the paper reporting that the 4B model remained close to its zero-shot counterpart after intention training.
That result addresses catastrophic forgetting on one external benchmark, not general understanding in the world. Benchmark clips are curated and have fixed questions, while a robot sees continuous, noisy scenes and must decide when evidence is sufficient. The study reports no physical robot experiment, live interaction, latency, calibration, abstention rate or consequence of a wrong inference. It therefore does not establish that the model knows when it is uncertain or that acting on its answer improves human–robot collaboration.[2]
Intention inference can become surveillance with a psychological label
The authors’ responsible-innovation statement explicitly warns that continuous video monitoring can be experienced as surveillance. It also notes that limited public egocentric datasets may encode cultural bias and that errors in assistive or caregiving settings could harm vulnerable people. Those are not peripheral concerns. Intention is an interpretation of behaviour, not a directly observed fact, and different cultures, disabilities and contexts can make the same movement mean different things.
A system deployed in a workplace, shop, school or care home could turn uncertain inferences into interventions or records about motivation. People may change their behaviour if they feel watched, and those least represented in training data may be misread most often. Meaningful safeguards would include visible consent, strict purpose limits, local processing where possible, short retention, access to logs, human confirmation before consequential action and a genuine option not to be profiled.[2]
What would justify a claim about real human intention
A stronger evaluation would first publish the exact train, validation and test denominators used, test on people and environments not represented in the source data, and report calibration and abstention as well as accuracy. Independent annotators from different cultural backgrounds should assess whether multiple intentions are plausible rather than assuming one label is ground truth. Open-ended answers should be judged for usefulness and harm, not only text similarity.
The next step after that is a consented live study in which a robot observes without acting, followed only later by tightly bounded assistance with human confirmation. Researchers should measure false interventions, missed opportunities, participant comfort, subgroup performance and whether the system improves an outcome people value. The accessible paper did not state a specific funding acknowledgement or competing-interest declaration. For now, IntentVLM is a strong multiple-choice benchmark result and a candid demonstration of how much harder free-form intention inference remains—not evidence that cameras can read minds.[1][2]
What this means for people
- Robots could offer help sooner when a person’s goal is clear, but a mistaken inference can interrupt, patronise or endanger them.
- Workers, residents and shoppers may be profiled through video without knowing that a system is assigning goals to their behaviour.
- People whose gestures or routines differ from benchmark norms may experience more errors unless evaluation includes their contexts and gives them control.
Global context
The research is from France and relies on public video benchmarks assembled from broader visual datasets. Reported averages do not establish performance across languages, cultures, disabilities or legal regimes. The EU AI Act and data-protection rules may constrain biometric and workplace uses, while governance and surveillance norms differ elsewhere. Cross-regional evidence is needed before a benchmark score can support deployment.
What the evidence does not yet show
- The highest accuracies come from five-choice benchmark questions; free-form generation performed much less convincingly.
- Human scores were reproduced from the original IntentQA study rather than collected under the new model’s protocol.
- No robot, live participant, continuous scene, latency test or real-world intervention was evaluated.
- One labelled answer may not capture multiple plausible intentions, and public egocentric video datasets may encode cultural and demographic bias.
- The accessible paper did not state a specific funding source or competing-interest declaration.
What to watch next
- Pre-registered evaluations with exact split denominators, calibration, abstention and multiple plausible labels.
- Independent cross-cultural annotation and subgroup analysis across disability, age and interaction context.
- Consented observational studies before any system is allowed to act on inferred intention.
- Governance that prevents intention inference from becoming persistent workplace, retail, school or care surveillance.
Evidence trail
Sources used for this report
Links checked 4 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
AI Risks & Safety
Can hashes make AI conversations auditable without publishing them?
A peer-reviewed experiment converted nearly four million public chatbot interactions into cryptographic commitments and detected eight induced ledger manipulations. It is a promising integrity mechanism, not proof that a conversation is true, authorised or safe.
9 min · 1 source
AI Risks & Safety
Can hidden image text mislead dental AI?
A peer-reviewed German stress test found that adversarial text placed inside 270 dental radiographs could flip four vision-language models from an abnormal to a normal finding. OCR sanitisation sharply reduced the measured attacks, but the experiment used a permissive prompt, a pathology-heavy benchmark and no live clinical system.
7 min · 3 sources
AI Risks & Safety
Can an AI-designed protein carry a detectable watermark without losing its function?
A peer-reviewed Nature paper reports lab-tested watermarks in designed protein binders and a separate mark in predicted structures. The result is a proof of concept for provenance—not a safety certificate or a universal detector for synthetic biology.
7 min · 2 sources
Reader discussion
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.