Athletics · International
Running-watch study shows strong correlations alongside distance-specific forecast errors
A single-model evaluation compares displayed predictions with measured running times.

Study publication:
What the January paper found
Published 29 January, the study involved 154 amateur runners, 123 men and 31 women, who wore a HUAWEI WATCH GT Runner for at least six weeks. Displayed forecasts were compared with performances in 288 tests over three distances. Predictions correlated strongly with results, but agreement ranges differed across distances. Tests were repeated observations, not additional independent runners. Findings concern this device, sample and protocol; they cannot verify other watches, software versions or conditions. Correlation does not mean every forecast is accurate for every runner, and no experiment established that following a watch improved race performance.
Agreement matters beyond correlation
Bland and Altman's original method-comparison paper explains why a correlation coefficient is an inadequate test of agreement between two measurement methods. Correlation can be high when values differ consistently, especially when the people being measured cover a broad range. Their approach examines the differences between paired values, the average bias and the spread of those differences. This gives a practical interpretation to a prediction comparison: a method can track which runners are faster while still placing an individual forecast several minutes from the measured result. Looking at the difference in the outcome's units is therefore more informative than relying on a strong association alone. The acceptable size of an error depends on the purpose of the estimate and needs to be considered separately. The method does not decide that purpose for a reader. It supplies a transparent way to show variation. For a running forecast, a group summary, distance-specific bias and the likely range of individual differences answer related but distinct questions about what the prediction can support.
Prediction models need a defined population and evaluation
The original TRIPOD statement provides guidance for reporting studies that develop or validate multivariable prediction models. It calls for clear information about the participants, predictors, outcome, modelling methods and performance assessment. It also distinguishes model development from evaluating a model in other data. These principles are useful context for a wearable forecast, even when a commercial algorithm's internal details are unavailable. A reader needs to know who was tested, how the outcome was measured and whether the evaluation matches the intended use. A result in one group does not automatically transport to another population or to a later implementation of a product. Reporting guidance cannot make an opaque algorithm transparent, and it is not an endorsement of a device. It helps identify the information needed to judge an evaluation. The useful next research question is therefore whether performance is maintained under clearly described new conditions. For readers of forecast studies, this keeps a measured comparison separate from marketing language and from the untested claim that using a prediction changes training or competitive results.
This research report is for education and professional discussion. Personal diagnosis, treatment and return-to-sport decisions require a qualified clinician.
Original sources
Read the underlying records.
Explore the official results, organiser reports or original research behind this story.
