How do skipper ratings work?
A rating is a log-strength fitted to every race at once, smoothed week to week. Frozen at the start of a season, it orders 70.4% of boat pairs correctly.
A skipper rating is a number on a log scale: the chance one boat finishes ahead of another is 1 / (1 + e−gap). It is fitted to all 103,570 fleet races in the database at once, by reading each race as a sequence of choices, and then smoothed from week to week so it follows a sailor's form.
For what a point of rating means in finishing chances, and how ratings map to percentiles, see What a rating point is worth. This note is about where the numbers come from.
A race is a sequence of choices
Give every boat a strength erating. To decide who wins, pick one boat from the fleet with probability proportional to its strength. Remove it, and pick the second-place boat from those left, again in proportion to strength. Continue to the end. The probability of a finishing order is the product of the probabilities of the picks. Here is a five-boat race where the middle-rated boat wins, then the top two follow.
This is why only gaps matter: multiply every strength by the same number and nothing changes. It is also why strength of schedule takes care of itself. Finishing ahead of strong boats is a less likely event than finishing ahead of weak ones, so it moves a rating more.
Fit everyone at once
The model's likelihood is the product of that probability over every race in the data. The ratings are the numbers that make the observed results as likely as possible. Each sailor's rating is pushed up when they finished better than their rating and their rivals' ratings predicted, and down when worse, until the pushes balance. Everyone's rating depends on everyone else's, which is how a sailor who only raced in one region ends up on the same scale as one who raced everywhere: the sailors who raced both link them. There is no K-factor to set and no order in which races are processed.
Ratings move through time
A sailor is not the same skipper in their first fall as in their fourth. Each sailor gets a rating for every regatta week, written as a starting level plus the week-to-week changes. Left alone, those changes would chase every bad race, so the fit pays a price for them: a log-cosh penalty on each change, which is gentle for small moves and only grows linearly for large ones, plus a small penalty on how quickly the changes themselves change. There is an extra cost on drops, so ratings fall more reluctantly than they rise. The weights were retuned against week-ahead forecasts when I fixed the trainer in September.
Starting a career
A sailor's first rating is anchored toward zero with a weak prior, so a newcomer with two races does not get an extreme number. On top of that there is team pooling: each sailor's career-average rating is pulled toward the average of their teammates that season, with a fixed strength, so the pull fades as their own results accumulate. That one change lowered season-ahead log-loss in every held-out season. For forecasts, a sailor with no race history at all starts from the typical debut rating of their school rather than from zero.
What is not in a rating
- Boats that did not finish are not in the finishing order. They are handled by a separate model, so a DNF does not count as last place.
- Singlehanded and match-race (keelboat) events are rated separately, since the skills are different.
- Crews are not in the skipper rating. Their contribution is estimated separately and is small; see How does a crew rating work?
How good is it?
I score the model only on races it never saw. For each of four seasons I refit using races before the season began, freeze the ratings, and predict every race. Pooled, it orders 70.43% of boat pairs correctly (pairwise log-loss 0.5644; a coin flip is 50% and 0.6931). At the 166 championship regattas, refit the week before each one, it gets 72.4% (95% interval 71.6 to 73.3).
One honest weakness: it is overconfident at the top. When it says a favourite has a 97% chance of finishing ahead in a race a season later, it happens 94% of the time. That is the price of frozen ratings and a model that cannot know who will improve over the summer.
Caveats
- A rating is relative to the sailors in the data. It says nothing about conditions, boats or tactics beyond what shows up in results.
- Ratings are estimates with error, largest for sailors with few races.
- Why this model rather than Elo: see Why not Elo?
Every number on this page is produced by analysis/blog/skipper_ratings.py and analysis/plackett_luce/ in the project repository, from the same database that powers the rest of the site. Corrections welcome: quinnbrighton2005@gmail.com.