mannheim.au

Waiting for runners

29 September 2026

Most trail-running events publish live results during a race, with predictions of when a runner will reach the next checkpoint. These predictions are usually linear extrapolations: they take the runner's average speed and extend it to the next stop.

This isn't predictive analytics and the results can be annoying. The website might tell supporters to wait 20 minutes but the runner turns up an hour later. They might miss their runner entirely because the prediction misjudged the runner's pace.

These race trackers should be much better because it's so easy to make more accurate predictions.

As a case study, consider Ultra-Trail Australia's 102 km event in the Blue Mountains. I examined the 2025 race data then predicted every 2026 finisher's arrival at each checkpoint as they left the one before. The chart below shows a runner's actual position, alongside four methods of predicting their speed.

The dark dot is a real runner from the 2026 race, moving at their real pace. The orange dot is the linear extrapolation often used by events, the blue dot uses multipliers for each leg (discussed below), the green dot is a linear regression model and the yellow dot is a regression fitted separately for each leg, each showing where it thinks the runner is. The predictions diverge sharply on the climb from the Six Foot Track checkpoint to the Katoomba Aquatic Centre, and again on the Furber Steps at the finish. The linear extrapolation reaches the finish about 40 minutes before the runner does.

Leg multipliers

I used the 2025 results to calculate leg multipliers: ratios of a leg's speed to the speed on the previous leg. For each leg after the first timing mat, I regressed every runner's speed on that leg against their speed on the leg before it, with no intercept. That gives one multiplier per leg: the runner's speed on the last leg, times it, is the prediction.

Leg Multiplier
Foggy Knob to the Six Foot Track 0.97
Six Foot Track to the Aquatic Centre 0.87
Aquatic Centre to Echo Point 1.13
Echo Point to Gordon Falls 0.65
Gordon Falls to the Fairmont 1.29
Fairmont to Queen Victoria Hospital 0.87
Queen Victoria Hospital to the emergency aid station 1.21
Emergency aid station to the finish 0.62

The smallest multipliers belong to the leg into and out of the Leura valley to Gordon Falls, and the last leg up the Furber Steps. Runners cover each at about two-thirds of the pace of the leg before.

Multiple linear regression

The multipliers use one thing about the runner: their speed on the previous leg. A multiple linear regression can use more information. It predicts a leg's speed as a constant plus each predictor times a weight. The weights are the ones that make its predictions miss the 2025 results by the least.

I gave this model five predictors: the leg's climb and descent per kilometre, the runner's average speed so far, their speed $ $ on the last leg and the hours they have already run. Fitted, it reads

\[ \hat{y} = 3.19 + 0.56x_1 + 0.14x_2 - 0.030x_3 - 0.036x_4 + 0.0029x_5 \]

where:

Average speed so far is the most influential predictor, the hours already run take a little speed off and every 100 metres of climb per kilometre costs about 3.6 km/h, everything else held equal.

None of these methods predicts the first leg, because nothing is known about the runner until the first timing mat.

Per-leg regression

The model above uses the same equation for every leg. A second regression model fits a separate equation for each leg, with four predictors: average speed so far, speed in the previous leg, hours since the start and minutes spent at aid stations so far. Climb and descent drop out, because within one leg they are the same for every runner. This model has eight equations and 40 numbers where the first has one equation and six.

Comparing predictions

Each box covers the middle half of the 2026 finishers, showing how far the prediction missed their real time. The whiskers reach the fastest and slowest 5% of the field, and the dots are the runners beyond them. Linear extrapolation is in orange, the leg multipliers in blue, the regression model in green and the per-leg regression in yellow, fitted on 2025 and applied to 2026.

Linear extrapolation is 37 minutes early at the Aquatic Centre, 25 minutes early at Queen Victoria Hospital and 43 minutes early at the finish, for the median runner. The multipliers are 5, 10 and 5 minutes late at the same checkpoints.

Method Average miss Within 30 minutes
Linear extrapolation 21.5 minutes 72%
Leg multipliers 11.2 minutes 94%
Regression model 15.0 minutes 86%
Per-leg regression 10.3 minutes 94%

The multipliers, fitted a year earlier, are eight numbers and a multiplication. They halve the extrapolation's average miss and put 94% of arrivals within half an hour, over 100 kilometres. The per-leg regression is better again by under a minute. The single-equation regression is worse than the multipliers: one set of terrain weights doesn't suit a 3-kilometre leg and a 24-kilometre one.

This lesson recurs in predictive analytics: a simple model built on the right variable (such as the leg multipliers) is usually about as good as an elaborate one, and far easier to explain. Finding the right variable is the trick.

Let me fix this problem

Linear extrapolation needs nothing, which is why it's used on race websites everywhere.

The multipliers need previous years' splits for the course. From these I derive eight numbers, which are enough for simple but reliable predictions. A regression model needs the same splits.

There are more accurate models I didn't describe in this article. All of this is very little work. The result is supporters standing at the checkpoint when their runner comes through. If a timing business wants to use a more accurate approach, I'll build it for free.