Probabilities

Two models, one race, the same matrix from each. Twenty drivers against twenty finishing positions, committed to git before the cars line up. The darker the cell, the more probability sits there.

Reading the next call.

How to read a row

The matrix

Every row is one driver and sums to one, because the driver finishes somewhere. A row with one dark cell and nineteen faint ones is a confident call. A row smeared evenly across ten cells is the model saying it does not know, which is a real answer and not a failure to produce one.

The three columns on the right are the same row added up in the ways people actually argue about: the first cell is win, the first three are podium, the first ten are points. They are not separate predictions and they cannot disagree with the matrix.

Nothing here is a tip. A driver at 8 percent to win is a driver who wins roughly one of these weekends in twelve, and the honest test of that number is a whole season of them, not this race.

Why the direct model looks flat

Plackett-Luce

The direct model knows two things: where each driver starts and how they have gone this season. With that little information it will rarely put more than about a tenth of the probability on any one driver, so its rows spread wide and its favourites are separated by two or three points rather than by twenty.

Read it as a ranking with error bars rather than a prediction of a result. Its ordering is usually sensible and its confidence is deliberately low, which is what a two-feature model should look like. A model this simple that claimed 40 percent on a winner would be lying about how much it knows.

Where it fails is anything that is not in its features. It cannot see that a fast car starts out of position, because the only thing it knows about the grid is the number.

Why the simulator looks certain

Monte Carlo

The simulator runs the race ten thousand times and counts. When the same driver wins four thousand of those runs, the number says forty percent, and it means it: pace, tyre wear, the pit stops each car chooses and the safety cars that reshuffle them all pointed the same way in most runs.

That sharpness is earned when the car really is quicker, and it is a liability when it is not. The simulator has no overtaking model, so it is at its worst where position is hard to change and its confidence should be discounted most: Monaco, Singapore, Hungary. It is at its best where a tyre advantage converts into a place, which is most of the calendar.

Its wide rows are informative too. A driver spread across P6 to P14 is usually a driver whose race hangs on one strategy branch or on whether a safety car falls in their window.

When the two disagree

The head-to-head

Disagreement is the useful part of publishing both. The bars above put the two side by side per driver, and the gaps have a grammar worth learning.

The simulator far above the direct model on a driver usually means race pace the grid does not reflect: the sim watched them come through on tyre life in run after run. The direct model above the simulator usually means the sim has buried a driver whose recovery depends on passing cars it refuses to model. Both agreeing is the strongest signal on the page, because the two models share no machinery and almost no assumptions.

Neither is trusted in advance. Every race is scored against the other and against the grid-order baseline the evening after it runs, and the running total is on the record. How both models are built, and where each one is known to break, is on the about page.