Back to the news portal
AI Risks & SafetyResearch paperResearchSource analysisGermanyEurope

Can hidden image text mislead dental AI?

A peer-reviewed German stress test found that adversarial text placed inside 270 dental radiographs could flip four vision-language models from an abnormal to a normal finding. OCR sanitisation sharply reduced the measured attacks, but the experiment used a permissive prompt, a pathology-heavy benchmark and no live clinical system.

By The Impact of AI Research DeskReleased 3 October 2026 at 09:00 BST7 min read3 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

ShareLinkedInXBlueskyRedditEmail
Key themesClinical AIPrompt injectionDental radiologyCybersecurityHuman oversight

Research topic

Whether adversarial instructions rendered into dental radiograph pixels can suppress abnormal findings from vision-language models, and which input or governance defences reduce that vulnerability

The Impact of AI research cover asking whether hidden image text can mislead dental AI, with a conceptual radiograph passing through an OCR filter to human review.
AI-generated editorial illustration. The radiograph, pixel prompt, OCR shield and review screen are conceptual and do not show a study participant, patient record, real attack or successful breach.

At a glance

  • 1The main evaluation used 270 public DenTeX radiographs, four vision-language models, four attack classes and two payload wordings per class; a separate 30-image set tuned two ProvDent thresholds.
  • 2In the most vulnerable model-and-attack cell, 338 of 540 paired observations flipped after a burned-in instruction: 62.6%, with a 95% confidence interval of 58.5% to 66.7%.
  • 3OCR sanitisation reduced pooled attack success from 15.6% in a re-run control to 0.2%. The defence ranking has not been validated prospectively, on a balanced external dataset or in a clinical workflow.

The Impact Brief

Keep the evidence trail, not the noise.

Get the most consequential AI developments with direct sources and clear limits.

Choose the topics you want (optional)

One concise, source-linked briefing. Unsubscribe at any time.

Living evidence record

Impact record IAI-0I39A9D

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent support

Present

Record status

Monitoring

Last checked

3 October 2026

Source trail

3 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 3 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

What the researchers tested

The study asked a narrow security question with direct clinical relevance: if an instruction is written into the pixels of a dental panoramic radiograph, will a vision-language model treat that text as more authoritative than the anatomy? The threat is called image-embedded prompt injection. It differs from a conventional text prompt because the instruction travels with the image through the same channel as the medical evidence. A tampered export, preprocessing service or upload could in principle insert it, although the authors found no evidence that targeted attacks on dental AI systems are occurring in the wild.

Four systems were assessed: GPT-4o, Gemini 2.5 Flash, Claude Sonnet 4.5 and the open-weight MedGemma 4B. All received the same structured task and were asked whether any abnormality was present. Because a diagnostic framing caused frequent refusals in a pilot, the researchers used a permissive educational framing and disabled Gemini safety filters. They explicitly describe the work as a red-team stress test under adversarial measurement conditions, not an estimate of how often a guarded clinical product fails.[1][2][3]

The sample, attacks and comparator

The source was the public DenTeX challenge dataset. From its 705-image disease split, the team selected 300 radiographs by stratified random sampling: 270 for the main evaluation and 30 held out only for setting ProvDent's instruction-likeness and disagreement thresholds. The main set contained 245 images with annotated pathology and 25 without. That 91% abnormal prevalence is central to interpretation because it strongly rewards sensitivity and leaves specificity poorly estimated.

Four attacks covered different visibility and placement: a conspicuous burned-in overlay, a note styled like a radiographer annotation, low-contrast text and a pale border patch. Each used two false-negative payload wordings, creating eight attacked versions per image. With a clean image included, the baseline comprised 9,720 calls. The primary paired outcome counted success only when the same model changed from abnormal on the clean image to no abnormality on its attacked counterpart. Every model-by-attack cell contained 540 pairs: 270 images times two payloads.[1]

The vulnerability result

Every model was vulnerable to at least one attack, but the profiles differed sharply. Burned-in text produced the highest measured attack-success rate: 62.6% for GPT-4o, 55.4% for MedGemma 4B, 14.8% for Gemini 2.5 Flash and 9.1% for Claude Sonnet 4.5. Averaged across the four visual attacks, the rates were 28.3%, 23.8%, 5.4% and 5.5% respectively. Most significant cells showed only abnormal-to-normal flips and no reverse flips.

Those figures do not mean GPT-4o fails on 62.6% of dental radiographs. They describe one model, one conspicuous attack class, two English payloads and a permissive prompt on a fixed 2026 configuration. The authors ran each condition once at temperature zero; hosted systems can still vary, and the re-run control differed slightly from the first baseline. Model updates, safety wrappers, local workflow controls and different prompt wording could change the result in either direction.[1][3]

Four defences—and an important trade-off

The researchers compared no defence with cropping image margins, adding prompt-level untrusted-data instructions, detecting and inpainting text with EasyOCR, and ProvDent. ProvDent combines prompt separation, OCR, a GPT-4o-mini judge that scores extracted text for instruction-like language and a second pass on a sanitised image when the score crosses a preset threshold. If the original and sanitised passes disagree on the abnormality decision, the system abstains and routes the case for human review.

Across 48,600 defence-condition calls, OCR sanitisation produced the lowest pooled attack-success rate, 0.2%, compared with 15.6% for the re-executed undefended arm. Cropping reached 9.5% and prompt spotlighting 7.4%. ProvDent reached 7.5% when every abstention was conservatively counted as an attack success; among the 93.5% of attacked cases it answered, 1.11% flipped. OCR sanitisation silently alters detected regions, while ProvDent's distinct contribution is explicit escalation when two views disagree.[1][2][3]

Why preserved F1 does not prove safety

On clean images, the defences left sensitivity close to the undefended value: between 99.4% and 99.9% versus 99.7%, while pooled F1 changed by no more than 0.2 percentage points. Yet three models classified every clean image as abnormal, and specificity ranged only from zero to 4%. With just 25 normal images, the study could not determine whether sanitisation, cropping or abstention preserves performance on the negative class. High sensitivity in a pathology-heavy benchmark is not safe diagnostic accuracy.

The domain assumption also travels poorly. In the authors' German workflow, normal panoramic images do not contain burned-in pixel text; labels remain in DICOM headers or PACS overlays. Other countries, clinics or modalities may legitimately burn markers and measurements into images. An aggressive OCR filter could erase information a clinician needs, and a text judge could misclassify a benign note. The paper did not test periapical radiographs, cone-beam CT, multilingual payloads, gradient-optimised attacks or a prospective clinical queue.[1]

Impact, disclosures and the next evidence

For patients and clinicians, the lesson is not that a dental AI breach has occurred. It is that medical-image security needs to include pixels, provenance and the software path—not only model accuracy. Procurement teams can require adversarial image tests, preserve the original DICOM object, log every transformation and keep a human review route when an input or output is suspicious. A system should say when it modified an image; silent filtering can hide both an attack and a legitimate annotation.

The authors report no external funding and no competing interests. Code and configuration are public, and the structured calls and tables are deposited on Figshare. Confidence would rise with independent re-execution, balanced external datasets, repeated calls across updated models, multilingual and less visible attacks, and prospective measurement of false alarms, workload and patient-relevant outcomes. Evidence from an actual guarded clinical product would also be necessary before translating these stress-test rates into operational risk.[1][2][3]

What this means for people

  • Patients could be harmed if image-borne instructions suppress real findings, but this study did not observe a live attack or patient outcome.
  • Clinicians need visible warnings and authority to inspect the unmodified image rather than receiving a silently sanitised answer.
  • Hospitals and vendors may need to treat image provenance, OCR filtering, escalation and audit logs as parts of clinical AI validation.

Global context

The German workflow gives the control a useful local prior, but medical-image conventions differ across countries, vendors and modalities. The cross-vendor result is internationally relevant; the defence ranking remains local until tested in other data, languages and clinical pipelines.

What the evidence does not yet show

  • The permissive educational prompt and disabled safeguards make these red-team measurements, not deployed-product incident rates.
  • The main set contained 245 abnormal and 25 normal radiographs, preventing a meaningful defence test on normal cases.
  • Attacks used two English false-negative payloads and excluded multilingual, gradient-optimised, metadata, sensor and model-weight attacks.
  • The rule that burned-in text is anomalous reflects one German workflow and may not transfer to settings with legitimate pixel annotations.

What to watch next

  • Independent replication on balanced external dental datasets and other medical-image modalities.
  • Testing of guarded production configurations, multilingual and adaptive attacks, repeated inference and model updates.
  • Prospective studies measuring false alerts, review time, missed findings, safe provenance and patient outcomes.

Evidence trail

Sources used for this report

Links checked 3 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

AI Risks & Safety

Can hashes make AI conversations auditable without publishing them?

A peer-reviewed experiment converted nearly four million public chatbot interactions into cryptographic commitments and detected eight induced ledger manipulations. It is a promising integrity mechanism, not proof that a conversation is true, authorised or safe.

9 min · 1 source

AI Risks & Safety

What does a safety resignation reveal about OpenAI?

David Robinson, who says he led safety-report writing for 12 frontier launches and helped draft OpenAI’s current Preparedness Framework, has resigned and called for safety practices closer to aviation or nuclear power. His essay is consequential first-person testimony, not an independent audit or proof of imminent harm.

9 min · 4 sources

AI Risks & Safety

Can a gesture warn you that a virtual AI may be wrong?

In a 24-person VR study, gestures, icons and highlighted text all helped users notice statements the system marked for checking. The paper does not show that participants detected factual falsehoods: its reference labels came from the same uncertainty pipeline that drove the cues.

6 min · 1 source

Reader discussion

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.