Back to the news portal
Science & ResearchResearch paperResearchSource analysisSouth KoreaEast AsiaKorean-language AI

Can Korean medical AI be tested on leaked questions?

An audit found that 3,333 of 3,448 supplied validation records shared identifiers with training files. The benchmark is valuable, but those files cannot support an independent claim about model performance.

By The Impact of AI Editorial DeskReleased 11 October 2026 at 13:00 BST7 min read2 sources

Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern

Share
Social links
LinkedInXBlueskyRedditEmail

At a glance

  • 1The two resources contained 222,073,009 source-data tokens and 34,487 labelled question-answer pairs across 17 domains, but the distribution was concentrated in an effective 5.62 domains and 78.7% of items were multiple-choice.
  • 2Among 3,448 distributed validation records, 3,333—96.7%—shared qa_id values with model-training files; median normalized question similarity was 0.885.
  • 3On a separate supplied 2,494-row evaluation file containing 2,489 unique normalized questions, essential-care and specialised-care Qwen2.5-14B models scored 64.35% and 64.03%; the paired difference was not significant.
Key themesMedical question answeringBenchmark integrityData leakageKorean languagePatient safetyModel evaluation

Research topic

Whether Korean medical question-answering datasets provide independent, specialty-aware tests for accompanying fine-tuned language models

The answer: no—overlapping validation files cannot provide an independent test

A peer-reviewed audit of two Korean medical question-answering resources found a basic evaluation problem: 3,333 of the 3,448 records distributed as validation data shared qa_id values with files used to train the accompanying fine-tuned models. That is 96.7% of the validation records. The median normalized similarity between the questions was 0.885. A score calculated on those files would therefore risk measuring recognition of training material rather than performance on genuinely unseen questions.

This does not make the underlying resources useless. They contain a large Korean-language corpus and tens of thousands of labelled question-answer pairs organised around medical domains. The paper's contribution is to separate dataset value from benchmark validity. Developers can use training material to build models, but they need an independently constructed, decontaminated test set before making credible comparisons. In medicine, that distinction matters because an inflated benchmark result can travel into procurement documents or clinical claims long before anyone checks what the model had already seen.[1][2]

The Impact Brief · Free

Follow the evidence in science & research.

Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.

Choose your topics (optional)

One concise, source-linked briefing. Unsubscribe at any time.

What the researchers audited

The study examined the AI Hub Specialized Medical Knowledge and Essential Medical Knowledge packages. Together they contained 222,073,009 source-data tokens and 34,487 labelled question-answer pairs across 17 labelled domains. The authors characterised the source material, answer format, domain structure and relationship between the supplied training and validation files. They also implemented a package-specific baseline using two Qwen2.5-14B models fine-tuned for specialised-care and essential-care content.

The audit found that the benchmark was not evenly spread across medicine. Its Herfindahl-Hirschman concentration index was 0.178, equivalent to 5.62 equally represented domains despite 17 labels. Multiple-choice questions accounted for 78.7% of items. Question-type distributions differed statistically between the two resources, but the reported Cramér's V of 0.069 indicates a small association. Those details matter because an overall accuracy can hide which specialties and response formats dominate the result.[1]

The separate baseline test found no meaningful package advantage

Because the distributed validation records were not independent, the researchers did not treat them as a clean comparison. Instead they used an AI Hub-supplied KorMedMCQA-derived evaluation file with 2,494 rows and 2,489 unique normalized questions. On that file, the essential-care model answered 64.35% correctly and the specialised-care model 64.03%. The absolute gap was only 0.32 percentage points.

The authors used an exact McNemar test to compare paired answers from the two models. The resulting p value was 0.440, so the data did not support a reliable difference between them on this evaluation file. That finding is narrower than saying the two training packages are equivalent. The evaluation set may not represent every specialty, rare condition, free-text response or real clinical query. It shows that the paper's particular baseline did not reveal a significant advantage—not that specialty-specific training can never matter.[1]

Why clinicians, developers and public buyers should care

For clinicians and patients, the practical risk begins when benchmark language becomes a claim about safe assistance. Roughly 64% multiple-choice accuracy would not justify autonomous advice even if the test were perfectly independent. Medical questions often require missing context, differential diagnosis, escalation and an explanation of uncertainty. A benchmark answer can be marked correct while a clinically important omission remains invisible.

For developers, the study offers a straightforward governance lesson: dataset cards should map every record across train, validation and test partitions using stable identifiers and near-duplicate checks. For hospitals and government buyers, headline accuracy should not be accepted without asking who constructed the test set, whether it post-dates training, how duplicates were removed and whether errors were reviewed by specialty. A clean test can still be narrow; a contaminated test cannot establish generalisation at all.[1]

Korean-language evaluation needs local breadth as well as clean splits

Korean medical AI should be evaluated in Korean because terminology, abbreviations, patient phrasing and care pathways do not transfer perfectly from English. The paper advances that goal by documenting local resources instead of relying solely on translated international exams. Yet the concentration measure shows why language localisation alone is not enough: a benchmark can be Korean and still underrepresent many clinical specialties or question types.

A stronger test would be built from sources not used in fine-tuning, include prespecified specialty quotas and distinguish knowledge recall from safe reasoning. It would also include free-text and ambiguous cases, identify questions where several actions are defensible, and ask Korean clinicians to grade clinically consequential mistakes. Results should be reported per domain with confidence intervals, not only as a pooled percentage. Patient-facing claims would then require separate prospective evidence in an actual service.[1]

Funding, limits and the evidence that would change the assessment

The authors reported no specific funding and declared no competing interests. The journal describes the publication as an early peer-reviewed accepted version that is citable but may receive further editorial changes. The study audits supplied data and runs a baseline comparison; it does not evaluate clinician use, patient outcomes, hallucination management, calibration or deployment security. Its independent evaluation file was supplied through AI Hub and remains a multiple-choice-style benchmark rather than a live clinical trial.

Confidence would rise with a separately authored and date-separated test set whose questions and source documents were never available during model development. The evaluation plan should be registered before scoring, include per-specialty denominators, duplicate and semantic-overlap reports, calibrated uncertainty and blinded clinical error review. External teams should be able to reproduce the result. The decisive step for patient care would be a prospective clinician-in-the-loop study measuring harmful errors, overrides, escalation and outcomes—not a higher score on recycled questions.[1][2]

What this means for people

  • Clinicians and patients need assurance that a medical model is being tested on unseen questions rather than material it may recognise from training.
  • Developers gain a concrete method for detecting identifier overlap and near-duplicate questions before reporting benchmark results.
  • Public buyers should demand clean splits, specialty denominators and clinical error review before treating an accuracy figure as evidence of readiness.

Global context

The audit focuses on Korean resources, but data leakage is a global benchmark problem. English-language medical exams, translated datasets and local-language corpora can all contain duplicates or material that entered model training. Local evaluation is essential for terminology and care context, while independent splits, transparent denominators and clinically meaningful review are universal requirements.

What the evidence does not yet show

  • The paper evaluates dataset structure and benchmark integrity, not clinical deployment or patient outcomes.
  • The supplied validation files overlapped extensively with training data and were unsuitable as independent tests of the accompanying models.
  • The separate evaluation file contained 2,494 rows and remained a limited question-answer benchmark rather than a representative sample of clinical conversations.
  • Domain concentration and the predominance of multiple-choice items limit what pooled accuracy says about specialty breadth and free-text reasoning.
  • A roughly 64% score cannot support autonomous medical advice, and the study did not test calibration, abstention or clinician oversight.

What to watch next

  • A newly constructed, decontaminated Korean medical test set that is unavailable during training.
  • Prespecified results by specialty, question type and clinical consequence, with uncertainty intervals.
  • Independent replication and blinded Korean clinician review of errors.
  • Prospective evidence that supervised use improves work without creating harmful reassurance or delay.

Living evidence record

Impact record IAI-0W4NI3X

Explore the full tracker

Evidence stage

Studied

Confidence

Supported

Reporting basis

Source analysis

Independent or research support

Present

Record status

Monitoring

Last checked

11 October 2026

Source trail

2 direct sources across 1 source type.

People impact

Documented in this record.

Uncertainty

Limits and next checks are explicit.

Stages describe the evidence available—not whether a technology is good or bad. See the public method.

Related-source reporting disclosure

This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.

Evidence trail

Sources used for this report

Links checked 11 October 2026

This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.

Continue the story

Related reporting

All reports

Science & Research

Can AI keep living cells under the nanoscope longer?

Two reconstruction strategies cut light dose tenfold or increased frame rate fourfold in RESOLFT experiments. They extend what researchers can observe, but computational restoration still needs experiment-specific validation.

7 min · 2 sources

Science & Research

Can more data fix measurement errors in AI?

No, not in these simulations. Across five model families, noisy or misassigned input features reduced predictive performance and distorted feature-importance rankings; larger samples narrowed variation but did not remove the bias.

7 min · 1 source

Science & Research

Can a general-purpose LLM search crystal compositions?

In a closed computational benchmark, GPT-5.4 recovered 95.65% of 3,740 low-energy Elpasolite targets within 5,000 proposals. Iterative evaluator feedback drove the result; no new material was synthesised, and the approximate energy labels are not proof of stability or usefulness.

8 min · 2 sources

Reader commentary

Add evidence, experience or a question

No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.

Explore commentary across the portal →

Do not include personal, confidential or unlawful information.

Published reader notes

0

No published reader notes yet. You can start the evidence-led discussion above.

Prefer a private correction or response? Contact the newsroom.