Data-Driven Sailing Home Under the hood ICSA ISSA Notes
October 7, 2026 · data through October 4, 2026 · All notes

How do skipper ratings work?

A rating is a log-strength fitted to every race at once, smoothed week to week. Frozen at the start of a season, it orders 70.4% of boat pairs correctly.

A skipper rating is a number on a log scale: the chance one boat finishes ahead of another is 1 / (1 + e−gap). It is fitted to all 103,570 fleet races in the database at once, by reading each race as a sequence of choices, and then smoothed from week to week so it follows a sailor's form.

For what a point of rating means in finishing chances, and how ratings map to percentiles, see What a rating point is worth. This note is about where the numbers come from.

A race is a sequence of choices

Give every boat a strength erating. To decide who wins, pick one boat from the fleet with probability proportional to its strength. Remove it, and pick the second-place boat from those left, again in proportion to strength. Continue to the end. The probability of a finishing order is the product of the probabilities of the picks. Here is a five-boat race where the middle-rated boat wins, then the top two follow.

pick 1CABD13% for Cpick 2ABD61% for Apick 3BDE70% for Bpick 4DE69% for Dboats (rating): A (+2.0), B (+1.2), C (+0.6), D (+0.0), E (-0.8)chance of this exact finishing order: 0.13 × 0.61 × 0.70 × 0.69 = 3.8%
One race as four picks. Each bar splits the chance among the boats still in the race; the bright segment is the boat that actually finished there. The order is a real possibility, not a likely one: about 4%.

This is why only gaps matter: multiply every strength by the same number and nothing changes. It is also why strength of schedule takes care of itself. Finishing ahead of strong boats is a less likely event than finishing ahead of weak ones, so it moves a rating more.

Fit everyone at once

The model's likelihood is the product of that probability over every race in the data. The ratings are the numbers that make the observed results as likely as possible. Each sailor's rating is pushed up when they finished better than their rating and their rivals' ratings predicted, and down when worse, until the pushes balance. Everyone's rating depends on everyone else's, which is how a sailor who only raced in one region ends up on the same scale as one who raced everywhere: the sailors who raced both link them. There is no K-factor to set and no order in which races are processed.

Ratings move through time

A sailor is not the same skipper in their first fall as in their fourth. Each sailor gets a rating for every regatta week, written as a starting level plus the week-to-week changes. Left alone, those changes would chase every bad race, so the fit pays a price for them: a log-cosh penalty on each change, which is gentle for small moves and only grows linearly for large ones, plus a small penalty on how quickly the changes themselves change. There is an extra cost on drops, so ratings fall more reluctantly than they rise. The weights were retuned against week-ahead forecasts when I fixed the trainer in September.

04080120-0.6-0.300.30.6a 0.3 dropa 0.3 gainpenalty on a one-week change in ratingchange in rating from one regatta week to the nextpenalty (log-likelihood units)
The price of a one-week change in rating at the trainer's weights. Small moves are cheap, big moves are expensive, and drops cost more than equal gains.
0.51.01.52.02.52020202120222023202420252026one skipper's weekly ratingdaterating
One ICSA skipper's rating over a career, a typical one: 29 rated weeks, the median among the 1,991 sailors with at least 20. Each dot is a regatta week; the line joins them.

Starting a career

A sailor's first rating is anchored toward zero with a weak prior, so a newcomer with two races does not get an extreme number. On top of that there is team pooling: each sailor's career-average rating is pulled toward the average of their teammates that season, with a fixed strength, so the pull fades as their own results accumulate. That one change lowered season-ahead log-loss in every held-out season. For forecasts, a sailor with no race history at all starts from the typical debut rating of their school rather than from zero.

What is not in a rating

How good is it?

I score the model only on races it never saw. For each of four seasons I refit using races before the season began, freeze the ratings, and predict every race. Pooled, it orders 70.43% of boat pairs correctly (pairwise log-loss 0.5644; a coin flip is 50% and 0.6931). At the 166 championship regattas, refit the week before each one, it gets 72.4% (95% interval 71.6 to 73.3).

One honest weakness: it is overconfident at the top. When it says a favourite has a 97% chance of finishing ahead in a race a season later, it happens 94% of the time. That is the price of frozen ratings and a model that cannot know who will improve over the summer.

Caveats

Every number on this page is produced by analysis/blog/skipper_ratings.py and analysis/plackett_luce/ in the project repository, from the same database that powers the rest of the site. Corrections welcome: quinnbrighton2005@gmail.com.