Why not Elo?
Both track form. The difference is propagation: in a joint fit every result informs every connected rating, both ways; in Elo a race touches only its boats, once.
Elo tracks form; so does the site model. On the same forward test, the weekly Plackett–Luce model beats a tuned Elo by 0.0179 of pairwise log-loss a season ahead and 0.0214 at championships. The difference is not the logistic curve, which they share. It is that a joint fit lets every result inform every rating it is connected to, in both directions, where Elo lets a race touch only the boats in it, once, and never again.
They are close relatives
Elo's update after a race, for a boat that beat S of its rivals when its rating said to expect E, is a step of size K/(m−1) × (S − E). That is one stochastic-gradient step on the pairwise logistic log-likelihood, taken once, in race order. Plackett–Luce fits a listwise likelihood, the whole finishing order, to convergence over all the data at once. Same logistic form for a pair of boats, a different way of getting to the numbers.
Elo moves information one way
In Elo a race changes the ratings of the boats in it and nothing else. Once applied, the update is final. If a boat you beat last month turns out to be much stronger than its rating said, your rating does not get the credit; the information arrived too late and has nowhere to go. In a joint fit the same race is a constraint that holds at the solution together with every other race. When new results move your rival's rating, every race you sailed against them re-weighs, and your rating moves with it, and so on outward through the sailors you have shared a start line with.
Here is that in a constructed network: two regions of twenty sailors who race only among themselves, linked by six travellers from each who met once. Then one more race is added among four of the travellers.
It shows up in real ratings
The clearest test is sailors whose own results can no longer change. Take everyone in the fit made before spring 2026 who has not raced since, and compare their rating at their last race in that fit with the same rating in today's fit, which has seen a further season of other people's results. Their own races are identical in both. Any revision is information that flowed back to them through rivals.
Among sailors whose last race was within a year of the cutoff, 39% were revised by more than a tenth of a rating point, with a typical revision (standard deviation) of 0.19. Four or more years out it falls to 7% and 0.12, because their rivals have also stopped racing. Under Elo every one of those revisions is zero.
Why it matters for sailing
College sailing is regional. Most regattas are within a conference, and the regions meet a few times a year at interconference events and the championships. To compare a New England sailor with one from the Pacific coast you need a path between them: the travellers who raced both. Simulated below, with the bridging regatta coming at the end of the season, as a national championship does.
Elo sees the bridge regatta as a few updates to the travellers. The regions' other sailors never raced an outsider, so their ratings stay where their home races put them, on two scales that happen to share a number. The joint fit pushes the whole region through the travellers, because every home race connects them.
The test
For four seasons (fall 2024 to spring 2026) every model is fit on races before the season, then predicts every race in it with ratings frozen. Elo's K and probability scale are tuned on the previous season, never the one being scored. A separate test refits the week before each of 166 championship regattas. Lower log-loss is better.
One row deserves a word. A Plackett–Luce fit with a single rating per sailor and no notion of time ties Elo (0.5815 against 0.5823). That is not the comparison to make, since Elo's K lets it follow form and the static model cannot; it is in the table because it separates the two ingredients. Propagation alone roughly pays for the loss of form-tracking. Add a weekly path with smoothing, and team pooling, and the joint fit pulls ahead.
By league, and at the championships
The championship test is the one that matters for the site, since its win chances are made the week before. It covers 166 regattas: 75 college (conference championships, showcases, and the 12 national semifinal and final rounds) and 91 high school (district qualifiers, district championships and the nationals). Elo and the site model, by league:
The margin is larger in high school than in college on both tests: 1.3 to 2.1 points at high-school championships against 0.7 to 1.4 at college ones.
Caveats
- This is the standard multi-boat Elo with a tuned K; Glicko and TrueSkill were not tested. The two-region and one-race examples are simulations; the retired-sailor check and the championship test are real data.
Every number on this page is produced by analysis/blog/elo_vs_pl.py, network_effects.py and analysis/plackett_luce/evaluate_models.py in the project repository, from the same database that powers the rest of the site. Corrections welcome: quinnbrighton2005@gmail.com.