Sports medicine · International

Knee-imaging study tests a classifier in a small external patient group

An imaging-model evaluation concerns adults with knee pain rather than a screened athlete population.

17 Jan 2025 Asia Shanghai; research publication
Research figure: MRI processing and classification pipeline.
Research figure: MRI classification pipeline.Lyu, Liangjing; Ren, Jing; Lu, Wenjie; Zhong, Jingyu; Song, Yang; Li, Yongliang; Yao, Weiwu · CC BY 4.0 · resized, uncroppedFigure licence

Study publication:

What the January paper found

Published 17 January, the MRI study analysed 215 anterior-knee-pain patients, mean age about 54. A classifier distinguished patellofemoral from non-patellofemoral osteoarthritis. Training and internal-test groups contained 109 and 73 patients; external testing contained 33. External area under the curve was 0.885, with a wide interval. The small external sample and single MRI machine limit generalisation. This was not an athlete-screening cohort and does not establish improved treatment decisions, clinical readiness or individual diagnosis.

Imaging AI needs a documented path from data to model

The original Checklist for Artificial Intelligence in Medical Imaging, known as CLAIM, asks authors to describe the clinical question, data sources, reference standard, image preparation and model evaluation. It also addresses how datasets are divided and how performance is reported. This is useful context for radiomics because the result depends on several stages before the classifier produces an output. Image acquisition, segmentation, feature selection and the chosen labels can all affect what the model learns. Readers need enough information to distinguish a reproducible procedure from an impressive number without a visible method. Clear separation of training and test information matters because using test data to select a model can inflate the apparent performance. Reporting guidance does not certify a model's safety or usefulness in practice. It makes the evidence easier to inspect. For a knee-imaging study, the useful questions include how the condition was defined, who supplied the images and whether evaluation data were kept separate from the process that created the classifier. Those details help identify what has actually been tested.

External validation tests a defined population and decision

Riley and colleagues' 2024 methodological article describes external validation as applying an existing prediction model to a different, relevant dataset that was not used for development. The validation sample should represent the population and setting where the model is intended to be used. Their framework examines overall fit, calibration and discrimination, with attention to relevant subgroups. Calibration concerns agreement between predictions and observed outcomes; discrimination concerns separation between outcome groups. These are related but different descriptions of performance. A strong result on one measure does not automatically settle the others. Where predictions are intended to direct decisions, the article also calls for evaluation of clinical usefulness, such as net benefit. This gives readers a boundary between assessing a model's output and assessing the consequences of using it. Validation is a continuing process because performance can differ between populations, settings and time periods. A research report therefore needs to specify what was tested and where further evaluation is needed. The framework does not declare any particular model ready for care or supply a diagnosis for an individual.

This research report is for education and professional discussion. Personal diagnosis, treatment and return-to-sport decisions require a qualified clinician.

Original sources

Read the underlying records.

Explore the official results, organiser reports or original research behind this story.