16+
DOI: 10.18413/2518-1092-2026-11-3-0-6

VALIDITY ASSESSMENT OF PUBLIC MRI DATASETS FOR NEURO-ONCOLOGICAL ABNORMALITY DETECTION: UNCOVERING CRITICAL DATA QUALITY ISSUES

The reliability of machine learning models in high-stakes applications like brain tumor diagnosis depends critically on the quality of training data. This paper presents a forensic audit of four widely used brain MRI datasets: Chakrabarty, Br35H, Bhuvaji, and Figshare. We evaluate these datasets against key validity criteria, including data provenance, label integrity, sample independence, representativeness, and ethical compliance. Our findings reveal that only the Figshare dataset satisfies most essential quality standards. The remaining three exhibit profound flaws – such as label inaccuracies, undisclosed duplication, and ethical lapses – that compromise their scientific validity. Controlled experiments further demonstrate that these data-quality issues can artificially inflate model performance by significant margins, undermining benchmark reliability. We conclude that datasets like Chakrabarty, Br35H, and Bhuvaji require substantial structural revision before they can be responsibly used in clinical AI development.

Number of views: 25 (view statistics)
Количество скачиваний: 34
Full text (PDF)Скачать XMLTo articles list
  • User comments
  • Reference lists

While nobody left any comments to this publication.
You can be first.

Leave comment: