Data-Driven Sailing Home Under the hood ICSA ISSA Notes
October 7, 2026 · data through October 4, 2026 · All notes

Why not Elo?

Both track form. The difference is propagation: in a joint fit every result informs every connected rating, both ways; in Elo a race touches only its boats, once.

Elo tracks form; so does the site model. On the same forward test, the weekly Plackett–Luce model beats a tuned Elo by 0.0179 of pairwise log-loss a season ahead and 0.0214 at championships. The difference is not the logistic curve, which they share. It is that a joint fit lets every result inform every rating it is connected to, in both directions, where Elo lets a race touch only the boats in it, once, and never again.

They are close relatives

Elo's update after a race, for a boat that beat S of its rivals when its rating said to expect E, is a step of size K/(m−1) × (S − E). That is one stochastic-gradient step on the pairwise logistic log-likelihood, taken once, in race order. Plackett–Luce fits a listwise likelihood, the whole finishing order, to convergence over all the data at once. Same logistic form for a pair of boats, a different way of getting to the numbers.

1st: boat D (+0.0)2nd: boat A (+1.5)3rd: boat B (+0.8)4th: boat C (+0.3)5th: boat F (-1.0)6th: boat E (-0.4)Elo: boats beaten minus expected+2.75-0.05-0.27-0.64-0.06-1.74Plackett–Luce: score of the order+0.90+0.06+0.05-0.14+0.33-1.21
One six-boat race where the fourth-rated boat wins, seen by each method (ratings in logits, shown in brackets). Both push the winner up and the last boat down. They weigh the middle differently: Elo charges a boat for every pairwise loss; Plackett–Luce charges it only its share of each pick it was passed over in, which is small for a weak boat, and credits it fully for edging the one boat left behind.

Elo moves information one way

In Elo a race changes the ratings of the boats in it and nothing else. Once applied, the update is final. If a boat you beat last month turns out to be much stronger than its rating said, your rating does not get the credit; the information arrived too late and has nowhere to go. In a joint fit the same race is a constraint that holds at the solution together with every other race. When new results move your rival's rating, every race you sailed against them re-weighs, and your rating moves with it, and so on outward through the sailors you have shared a start line with.

Here is that in a constructed network: two regions of twenty sailors who race only among themselves, linked by six travellers from each who met once. Then one more race is added among four of the travellers.

Elo: 4 of 40 ratings moveregion Aregion Bjoint Plackett–Luce fit: 33 of 40 ratings moveregion Aregion Bone new race among the four outlined boats; dot size is the change (green up, red down, grey none)
After one new race, Elo changes 4 ratings. The joint fit changes 33: the four boats most, then everyone who has raced them, in proportion to how much racing they share.

It shows up in real ratings

The clearest test is sailors whose own results can no longer change. Take everyone in the fit made before spring 2026 who has not raced since, and compare their rating at their last race in that fit with the same rating in today's fit, which has seen a further season of other people's results. Their own races are identical in both. Any revision is information that flowed back to them through rivals.

0%12%25%38%50%39%<1 yr18%1-2 yrs10%2-4 yrs7%4+ yrsshare revised by more than 0.1share of retired sailors
22,186 sailors with no race since the 2026-W03 cutoff, by time since their last race. Share whose rating at that last race was revised by more than 0.1 once a further season of other sailors' results was fitted.

Among sailors whose last race was within a year of the cutoff, 39% were revised by more than a tenth of a rating point, with a typical revision (standard deviation) of 0.19. Four or more years out it falls to 7% and 0.12, because their rivals have also stopped racing. Under Elo every one of those revisions is zero.

Why it matters for sailing

College sailing is regional. Most regattas are within a conference, and the regions meet a few times a year at interconference events and the championships. To compare a New England sailor with one from the Pacific coast you need a path between them: the travellers who raced both. Simulated below, with the bridging regatta coming at the end of the season, as a national championship does.

0.000.300.600.901.201.000.160.62gap between the regions0.430.31error, sailors who never travelledtruthElojoint fitrating points
Two regions whose true average ratings differ by one point, linked by one late regatta of travellers. Mean over 200 simulations. Elo recovers 0.16 of the gap; the joint fit 0.62. For sailors who never left their region, the joint fit's error is 0.31 against Elo's 0.43.

Elo sees the bridge regatta as a few updates to the travellers. The regions' other sailors never raced an outsider, so their ratings stay where their home races put them, on two scales that happen to share a number. The joint fit pushes the whole region through the travellers, because every home race connects them.

The test

For four seasons (fall 2024 to spring 2026) every model is fit on races before the season, then predicts every race in it with ratings frozen. Elo's K and probability scale are tuned on the previous season, never the one being scored. A separate test refits the week before each of 166 championship regattas. Lower log-loss is better.

0.520.540.560.580.600.62season aheadchampionships, 95% intervalseason / champsAverage finish0.6185 / 0.6061Elo (K tuned)0.5823 / 0.5577Plackett–Luce, one rating per sailor0.5815 / 0.5571Plackett–Luce, weekly (earlier version)0.5704 / 0.5451Plackett–Luce, weekly (site model)0.5644 / 0.5363pairwise log-loss on races never seen in fitting (lower is better)
Out-of-sample pairwise log-loss. The site model beats Elo in every bootstrap resample of championship regattas, by 1.1 to 1.7 points of accuracy.

One row deserves a word. A Plackett–Luce fit with a single rating per sailor and no notion of time ties Elo (0.5815 against 0.5823). That is not the comparison to make, since Elo's K lets it follow form and the static model cannot; it is in the table because it separates the two ingredients. Propagation alone roughly pays for the loss of form-tracking. Add a weekly path with smoothing, and team pooling, and the joint fit pulls ahead.

By league, and at the championships

The championship test is the one that matters for the site, since its win chances are made the week before. It covers 166 regattas: 75 college (conference championships, showcases, and the 12 national semifinal and final rounds) and 91 high school (district qualifiers, district championships and the nationals). Elo and the site model, by league:

66%68%70%72%74%Elosite modelcollege, season ahead+1.3 to +1.7 ptscollege, championships+0.7 to +1.4 ptshigh school, season ahead+1.7 to +2.1 ptshigh school, championships+1.3 to +2.1 ptsshare of boat pairs ordered correctly (95% bootstrap interval of the margin, by regatta)
Pairwise accuracy of Elo and the site model, by league, season ahead and at championships. The site model wins in every bootstrap resample of every row.

The margin is larger in high school than in college on both tests: 1.3 to 2.1 points at high-school championships against 0.7 to 1.4 at college ones.

Caveats

Every number on this page is produced by analysis/blog/elo_vs_pl.py, network_effects.py and analysis/plackett_luce/evaluate_models.py in the project repository, from the same database that powers the rest of the site. Corrections welcome: quinnbrighton2005@gmail.com.