Eigen RadarAI
Analysis

Korean medical AI questions overlap across training and validation files

An audit by Mi Ae Yang with Kang Su Ha found that 96.7 per cent of validation rows in a Korean medical question-answer resource shared identifiers with training questions. The audit also found uneven specialty coverage. A separate comparison of two adapted Qwen models showed no statistically significant accuracy difference, highlighting the distinction between a benchmark score and an independent evaluation.

Artificial Intelligence··Night
A medical data reviewer behind two displays, with a stethoscope and Korean flag on the desk.

Validation questions return in training

An audit by Mi Ae Yang with Kang Su Ha found extensive question reuse in Korean medical question-answer data supplied through AI Hub, a data-resource platform. In the validation file of 3,448 rows, matching identifiers linked 3,333 rows to training questions: 96.7 per cent of the validation total. Question-text similarity had a median value of 0.885. The researchers concluded that this partition cannot independently test models fine-tuned on the same source.[1]

Specialties receive uneven coverage

The audit covered specialized and essential medical knowledge resources containing 222,073,009 tokens and 34,487 question-answer pairs across 17 domains. Tokens are the text units processed by a language model. Multiple-choice questions made up 78.7 per cent. A concentration index of 0.178 gave 5.62 as the effective number of domains, showing how heavily material clustered despite the larger nominal domain total.[1]

Separate model scores remain close

A comparison on KorMedMCQA, a Korean medical multiple-choice benchmark, used 2,489 distinct questions across 2,494 rows. Two Qwen2.5-14B adaptations reached accuracies of 64.35 per cent and 64.03 per cent for essential and specialized knowledge respectively. The paired McNemar test gave a p value of 0.440, leaving the difference statistically insignificant. The researchers called for test-set decontamination and clinical-error review against prespecified criteria; the scores establish no clinical readiness.[1]

References

  1. News sourceScientific ReportsKorean medical-question audit finds training–validation overlap↩1↩2↩3