Can Korean medical AI be tested on leaked questions?
An audit found that 3,333 of 3,448 supplied validation records shared identifiers with training files. The benchmark is valuable, but those files cannot support an independent claim about model performance.
Editorial responsibility: The Impact of AI Editorial Desk · Report a factual concern
At a glance
- 1The two resources contained 222,073,009 source-data tokens and 34,487 labelled question-answer pairs across 17 domains, but the distribution was concentrated in an effective 5.62 domains and 78.7% of items were multiple-choice.
- 2Among 3,448 distributed validation records, 3,333—96.7%—shared qa_id values with model-training files; median normalized question similarity was 0.885.
- 3On a separate supplied 2,494-row evaluation file containing 2,489 unique normalized questions, essential-care and specialised-care Qwen2.5-14B models scored 64.35% and 64.03%; the paired difference was not significant.
Research topic
Whether Korean medical question-answering datasets provide independent, specialty-aware tests for accompanying fine-tuned language models
The answer: no—overlapping validation files cannot provide an independent test
A peer-reviewed audit of two Korean medical question-answering resources found a basic evaluation problem: 3,333 of the 3,448 records distributed as validation data shared qa_id values with files used to train the accompanying fine-tuned models. That is 96.7% of the validation records. The median normalized similarity between the questions was 0.885. A score calculated on those files would therefore risk measuring recognition of training material rather than performance on genuinely unseen questions.
This does not make the underlying resources useless. They contain a large Korean-language corpus and tens of thousands of labelled question-answer pairs organised around medical domains. The paper's contribution is to separate dataset value from benchmark validity. Developers can use training material to build models, but they need an independently constructed, decontaminated test set before making credible comparisons. In medicine, that distinction matters because an inflated benchmark result can travel into procurement documents or clinical claims long before anyone checks what the model had already seen.[1][2]
The Impact Brief · Free
Follow the evidence in science & research.
Get a five-minute weekday briefing on what changed, why it matters and where the evidence comes from. Choose the topics you care about.
What the researchers audited
The study examined the AI Hub Specialized Medical Knowledge and Essential Medical Knowledge packages. Together they contained 222,073,009 source-data tokens and 34,487 labelled question-answer pairs across 17 labelled domains. The authors characterised the source material, answer format, domain structure and relationship between the supplied training and validation files. They also implemented a package-specific baseline using two Qwen2.5-14B models fine-tuned for specialised-care and essential-care content.
The audit found that the benchmark was not evenly spread across medicine. Its Herfindahl-Hirschman concentration index was 0.178, equivalent to 5.62 equally represented domains despite 17 labels. Multiple-choice questions accounted for 78.7% of items. Question-type distributions differed statistically between the two resources, but the reported Cramér's V of 0.069 indicates a small association. Those details matter because an overall accuracy can hide which specialties and response formats dominate the result.[1]
The separate baseline test found no meaningful package advantage
Because the distributed validation records were not independent, the researchers did not treat them as a clean comparison. Instead they used an AI Hub-supplied KorMedMCQA-derived evaluation file with 2,494 rows and 2,489 unique normalized questions. On that file, the essential-care model answered 64.35% correctly and the specialised-care model 64.03%. The absolute gap was only 0.32 percentage points.
The authors used an exact McNemar test to compare paired answers from the two models. The resulting p value was 0.440, so the data did not support a reliable difference between them on this evaluation file. That finding is narrower than saying the two training packages are equivalent. The evaluation set may not represent every specialty, rare condition, free-text response or real clinical query. It shows that the paper's particular baseline did not reveal a significant advantage—not that specialty-specific training can never matter.[1]
Why clinicians, developers and public buyers should care
For clinicians and patients, the practical risk begins when benchmark language becomes a claim about safe assistance. Roughly 64% multiple-choice accuracy would not justify autonomous advice even if the test were perfectly independent. Medical questions often require missing context, differential diagnosis, escalation and an explanation of uncertainty. A benchmark answer can be marked correct while a clinically important omission remains invisible.
For developers, the study offers a straightforward governance lesson: dataset cards should map every record across train, validation and test partitions using stable identifiers and near-duplicate checks. For hospitals and government buyers, headline accuracy should not be accepted without asking who constructed the test set, whether it post-dates training, how duplicates were removed and whether errors were reviewed by specialty. A clean test can still be narrow; a contaminated test cannot establish generalisation at all.[1]
Korean-language evaluation needs local breadth as well as clean splits
Korean medical AI should be evaluated in Korean because terminology, abbreviations, patient phrasing and care pathways do not transfer perfectly from English. The paper advances that goal by documenting local resources instead of relying solely on translated international exams. Yet the concentration measure shows why language localisation alone is not enough: a benchmark can be Korean and still underrepresent many clinical specialties or question types.
A stronger test would be built from sources not used in fine-tuning, include prespecified specialty quotas and distinguish knowledge recall from safe reasoning. It would also include free-text and ambiguous cases, identify questions where several actions are defensible, and ask Korean clinicians to grade clinically consequential mistakes. Results should be reported per domain with confidence intervals, not only as a pooled percentage. Patient-facing claims would then require separate prospective evidence in an actual service.[1]
Funding, limits and the evidence that would change the assessment
The authors reported no specific funding and declared no competing interests. The journal describes the publication as an early peer-reviewed accepted version that is citable but may receive further editorial changes. The study audits supplied data and runs a baseline comparison; it does not evaluate clinician use, patient outcomes, hallucination management, calibration or deployment security. Its independent evaluation file was supplied through AI Hub and remains a multiple-choice-style benchmark rather than a live clinical trial.
Confidence would rise with a separately authored and date-separated test set whose questions and source documents were never available during model development. The evaluation plan should be registered before scoring, include per-specialty denominators, duplicate and semantic-overlap reports, calibrated uncertainty and blinded clinical error review. External teams should be able to reproduce the result. The decisive step for patient care would be a prospective clinician-in-the-loop study measuring harmful errors, overrides, escalation and outcomes—not a higher score on recycled questions.[1][2]
What this means for people
- Clinicians and patients need assurance that a medical model is being tested on unseen questions rather than material it may recognise from training.
- Developers gain a concrete method for detecting identifier overlap and near-duplicate questions before reporting benchmark results.
- Public buyers should demand clean splits, specialty denominators and clinical error review before treating an accuracy figure as evidence of readiness.
Global context
The audit focuses on Korean resources, but data leakage is a global benchmark problem. English-language medical exams, translated datasets and local-language corpora can all contain duplicates or material that entered model training. Local evaluation is essential for terminology and care context, while independent splits, transparent denominators and clinically meaningful review are universal requirements.
What the evidence does not yet show
- The paper evaluates dataset structure and benchmark integrity, not clinical deployment or patient outcomes.
- The supplied validation files overlapped extensively with training data and were unsuitable as independent tests of the accompanying models.
- The separate evaluation file contained 2,494 rows and remained a limited question-answer benchmark rather than a representative sample of clinical conversations.
- Domain concentration and the predominance of multiple-choice items limit what pooled accuracy says about specialty breadth and free-text reasoning.
- A roughly 64% score cannot support autonomous medical advice, and the study did not test calibration, abstention or clinician oversight.
What to watch next
- A newly constructed, decontaminated Korean medical test set that is unavailable during training.
- Prespecified results by specialty, question type and clinical consequence, with uncertainty intervals.
- Independent replication and blinded Korean clinician review of errors.
- Prospective evidence that supervised use improves work without creating harmful reassurance or delay.
Living evidence record
Impact record IAI-0W4NI3X
Evidence stage
Studied
Confidence
Supported
Reporting basis
Source analysis
Independent or research support
Present
Record status
Monitoring
Last checked
11 October 2026
Source trail
2 direct sources across 1 source type.
People impact
Documented in this record.
Uncertainty
Limits and next checks are explicit.
Stages describe the evidence available—not whether a technology is good or bad. See the public method.
Related-source reporting disclosure
This record analyses 2 linked source records around the same underlying development. The extra records add method, date or context, but they do not by themselves constitute independent replication of every performance claim or predicted outcome.
Evidence trail
Sources used for this report
Links checked 11 October 2026
This report is labelled source analysis. We summarise and analyse source material in our own words; company statements remain attributed claims until independently supported. Translated summaries preserve the meaning of the original source and link back to it. Read our editorial standards.
Continue the story
Related reporting
Science & Research
Can AI keep living cells under the nanoscope longer?
Two reconstruction strategies cut light dose tenfold or increased frame rate fourfold in RESOLFT experiments. They extend what researchers can observe, but computational restoration still needs experiment-specific validation.
7 min · 2 sources
Science & Research
Can more data fix measurement errors in AI?
No, not in these simulations. Across five model families, noisy or misassigned input features reduced predictive performance and distorted feature-importance rankings; larger samples narrowed variation but did not remove the bias.
7 min · 1 source
Science & Research
Can a general-purpose LLM search crystal compositions?
In a closed computational benchmark, GPT-5.4 recovered 95.65% of 3,740 low-energy Elpasolite targets within 5,000 proposals. Iterative evaluator feedback drove the result; no new material was synthesised, and the approximate energy labels are not proof of stability or usefulness.
8 min · 2 sources
Reader commentary
Add evidence, experience or a question
No account is required. Reader notes are published after a brief civility, relevance and safety check; disagreement is welcome.
Explore commentary across the portal →Published reader notes
0No published reader notes yet. You can start the evidence-led discussion above.
Prefer a private correction or response? Contact the newsroom.