Back to the news portal
Society & MediaResearch paperResearchSource analysisGlobalUnited StatesNorth America

What do people actually ask image-upload AI to do?

An unreviewed Microsoft-led study analysed 42,617 de-identified Copilot image-upload sessions and checked its taxonomy against 23,413 ChatGPT sessions. Users often chained recognition into writing, code and data tasks—but private, model-generated summaries limit what outsiders can verify.

By The Impact of AI Research DeskReleased 3 October 2026 at 09:50 BST8 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesMultimodal AIUser behaviourAI benchmarksPrivacyHuman-computer interaction

Research topic

Which capabilities people combine when uploading images to general-purpose AI assistants, whether image-based tasks differ from text-only use, and how closely existing benchmarks match observed workflows

The Impact of AI research cover asking what people ask image-upload AI to do, with a conceptual upload flowing through recognition, OCR and reasoning into generated text, code and a data table.
AI-generated editorial illustration. The upload, processing stages and benchmark paths are conceptual; the image does not reproduce a user's photo, raw conversation, product interface or study chart.

At a glance

  • 1The primary sample contained 42,617 Copilot image-upload sessions from August and September 2024; a separate 23,413-session ChatGPT dataset from November 2024 and January 2025 tested whether the pattern transferred.
  • 2The model-generated taxonomy assigned ten capabilities, and 74.17% of Copilot sessions received more than one capability label. A 200-session human audit judged 60 tasks, or 30%, not reasonably completable from text alone.
  • 3Across 253 benchmark papers and 866 eligible static-image tasks, recognition-plus-reasoning was strongly represented, while common writing and code-generation combinations were only incidentally covered.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-061Y6GV

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

3 October 2026

Source trail

2 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

A usage study, not a model leaderboard

When people upload a photograph, screenshot, chart or document to an AI assistant, they may want more than identification. They may want text extracted, a decision explained, code written from a mock-up or a report drafted from a chart. The Microsoft-led research team asked whether existing multimodal taxonomies and benchmarks reflect those end-to-end jobs. It did not test which assistant answered correctly, compare vendor accuracy or measure downstream outcomes. Its unit of analysis was the user's apparent task within a conversation.

The paper is an unreviewed preprint submitted on 30 September 2026. Its authors include Microsoft and Microsoft Research staff, and the first author's work was completed during an internship at Microsoft's IDEAS group and Microsoft Research. That commercial connection is directly relevant: the primary data came from Copilot, the raw records remain private, and the taxonomy may inform product and benchmark design. An independent ChatGPT dataset provides a useful cross-platform check but does not remove the access asymmetry.[1][2]

Two assistants, four collection months

The primary dataset was an anonymous random sample of Microsoft Copilot conversations from August and September 2024. Enterprise, education and commercial accounts were excluded. A session counted as multimodal if the user uploaded at least one image; multimodal outputs such as text-to-image generation were outside scope. The sample contained 42,617 image-upload sessions. The validation dataset contained 23,413 ChatGPT image-upload sessions collected in November 2024 and January 2025. Equal-sized text-only samples from each assistant supported the task-space comparison.

Privacy constraints shaped every later step. Automated systems removed names, contact information and other sensitive attributes, then converted each session into a task summary and image summary of at most 50 words. Researchers performed the taxonomy work on those summaries without direct access to raw images—the paper calls this an eyes-off design. Processing occurred in access-controlled environments, and results were aggregated. That reduces disclosure risk but also prevents outsiders from auditing the original conversations or checking whether the summaries erased subtle context.[2]

How the ten-capability taxonomy was built

The team used the TnT-LLM procedure with GPT-4o-mini. It drew 10,000 Copilot sessions for taxonomy induction, processed 50 batches of 200 sessions and refined seed categories from earlier multimodal benchmarks until convergence. GPT-4o-mini then assigned each session a primary capability and, where applicable, a secondary one. The ten final categories covered recognition, knowledge, OCR and text extraction, language and text generation, reasoning, data analysis, code generation, document processing, image generation and editing, and spatial awareness.

A random 200-session validation set was independently labelled by a researcher and a graduate student. Human-to-human multi-label F1 was 0.783 with a 0.960 hit rate. The model's F1 against the two annotators was 0.763 and 0.773, with hit rates of 0.965 and 0.955. A second 200-session audit compared labels derived from conversation text with labels derived from summaries; human F1 was 0.73 and the hit rate 0.91. These checks support broad category consistency, not perfect reconstruction of user intent.[2]

Most image-upload tasks were compositional

Recognition was the largest capability label, appearing in 12,095 Copilot sessions, followed by knowledge in 7,226 and OCR or text extraction in 7,136. Crucially, 74.17% of image-upload sessions received more than one capability label. Knowledge plus language generation was the most frequent pair, followed by recognition plus reasoning and knowledge plus recognition. The pattern suggests that reading an image is often an entry point: the requested outcome is an explanation, document, code artefact, analysis or edit rather than a label.

The researchers also asked whether users truly needed the image. Two people audited 200 sessions and judged 60, or 30%, not reasonably completable from text alone. This denominator matters: the majority of sampled tasks could potentially have been expressed through text, perhaps because screenshots save retyping or preserve layout. Image upload can therefore be both an accessibility and convenience feature and a genuinely new source of visual evidence. The study does not measure whether the assistant interpreted that evidence correctly or whether uploading it improved the outcome.[2]

Where benchmarks match—and miss—observed use

The team retrieved 253 benchmark papers from a survey catalogue updated in May 2026 and extracted 1,173 evaluation tasks. GPT-5.4 and Gemini 3.1 Pro independently labelled static-image requirements and capability combinations, with GPT-5.6 Sol adjudicating disagreements. After non-image tasks and 17 very broad umbrella tasks were removed, 866 tasks remained. A blinded human check of 71 disputed cases found mean F1 of 0.761 and macro kappa of 0.579 for the adjudicator—useful but far from error-free.

Recognition plus reasoning was strongly supplied relative to observed use, with a supply-to-demand ratio of 3.12. Yet knowledge plus language generation—the most common real-world pair at 8.4% of sessions—had a ratio of 0.22 and was classed as incidental coverage. Code generation plus recognition scored 0.09; code generation plus reasoning scored 0.11. The authors' point is not that visual perception benchmarks are obsolete. It is that an assistant can pass tests that end with a label or short answer while still being weak at the longer artefacts users expect from the same visual input.[2]

What this means for users and product teams

For users, a polished document or working-looking code file can hide an early perception error. Product teams should expose extracted text and detected structure, keep the original upload available for comparison, and let people correct what the system read before downstream generation. Evaluations should trace error propagation from OCR or recognition into calculations, decisions and artefacts. A complete workflow benchmark might score whether a generated spreadsheet reconciles, whether code runs, whether a report attributes the image correctly and whether the result supports a real decision without inventing detail.

The study does not justify uploading sensitive material casually. De-identification occurred inside controlled research systems, while ordinary users face different retention, account and workplace policies. The research also excludes enterprise and education Copilot accounts, audio and video inputs, and multimodal outputs. Its conversations are from 2024, so usage may already have changed as models and interfaces improved. The results describe selected platforms and periods, not global users, cultures, languages or accessibility needs.[2]

Commercial interests and the next evidence

The paper lists Microsoft affiliations for four of five authors and notes that the first author's work was done during a Microsoft internship. It does not provide a separate funding or competing-interests declaration in the version reviewed. The authors report internal privacy review and informed consent for the independent validation data, but the private datasets cannot be inspected or redistributed. That limits reproducibility and gives platform operators more visibility into public behaviour than independent researchers can obtain.

Confidence would rise with peer review, transparent sampling weights, language and geography breakdowns, independently governed secure access, and replications on other assistants and newer periods. Outcome studies should connect task types to correctness, completion, user effort, accessibility and harm. Benchmarks also need full visual-to-artifact workflows with executable or decision-relevant scoring. Until then, the study is best read as evidence about what people attempt—not proof that image-upload AI succeeds at those attempts.[1][2]

What this means for people

  • Image uploads can reduce retyping and preserve visual structure, but an unnoticed perception error can contaminate every downstream artefact.
  • Users need visible extraction steps and easy correction before the system turns an image into code, analysis or advice.
  • Platform telemetry can improve evaluations, yet private access creates a research-power imbalance and requires strong privacy governance.

Global context

The samples come from two globally used assistants but the paper does not publish country, language or demographic denominators. Global generalisation is therefore unproven, especially where scripts, connectivity, device use and visual conventions differ.

What the evidence does not yet show

  • This is an unreviewed preprint led largely by Microsoft-affiliated authors and based on private platform data.
  • The analysis uses model-generated summaries rather than raw images or conversations, and independent researchers cannot audit the original records.
  • The Copilot and ChatGPT data come from four months in 2024–25; user behaviour, interfaces and model capabilities may have changed.
  • Capability labels describe attempted tasks, not answer accuracy, completion, usefulness or downstream harm.

What to watch next

  • Peer review and replication with newer, multilingual and independently governed datasets.
  • Benchmarks that test complete visual-to-document, visual-to-code and visual-to-decision workflows.
  • Studies linking multimodal task categories to correctness, accessibility, user effort, privacy and consequential outcomes.

Evidence trail

Sources used for this report

Links checked 3 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Society & Media

Can image-aware AI catch fake news before it spreads?

A new peer-reviewed model improved two fixed benchmark tests by combining article text, images and an AI-generated image description. It did not verify claims or face a live, changing news stream.

7 min · 2 sources

Society & Media

Will patients avoid doctors who say they use AI?

A preregistered experiment with 1,030 US adults found lower ratings and appointment intentions for fictional family doctors who said they used AI. The effect is relevant to patient trust, but an advert-based intention is not a real healthcare choice.

10 min · 2 sources

Society & Media

What should AI companies prove before giving conversational agents to children?

A new World Economic Forum white paper sets out child-focused practices spanning design, privacy, learning, harmful interactions and accountability. It is useful guidance, but not a binding standard or an outcome study, and its public summary reports no sample or systematic-review method.

6 min · 1 source

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.