Club & League Strength
How every club is placed on one global strength scale across divisions and borders — a margin-aware base rating with explicit draws and home advantage, calibrated to an independent reference at the top, and a cross-league exchange rate learned from the players who move between divisions.
Oraca puts every club on one global strength scale — from the top of the pyramid down to the fourth tier, and across borders — so any side can be compared with any other. Most strength ratings work well inside a single league, where everyone plays everyone. Ours is built to solve the two cases they can't: how a lower-league club compares to a top-flight one, and how leagues in different countries compare.
In one line: match results tell you who is stronger within a connected group of teams; they cannot tell you how a fourth-tier club compares to a top-flight one, because the two almost never meet. We fix that by calibrating the top of the pyramid against an independent reference, and pegging everything below it using the players who move between divisions.
The two problems a results-only rating can't solve
A rating built purely on who-beat-whom is excellent inside a densely connected league. It has two structural blind spots:
- The lower-league island problem. A club's rating is only meaningful against the opponents it has actually played, and the football pyramid is a chain of weakly connected islands — the only games linking, say, the third tier to the top are a handful of cup ties a year. With so few connecting matches, a results-only rating compresses: it can't confidently say how far apart the divisions really are, so lower-league sides drift toward the middle and the true gap is understated.
- The cross-border problem. Top divisions in different countries barely meet outside continental competition, so a results-only rating can't pin whether one country's league is genuinely stronger than another's.
How it is built
The base rating. Every club carries a rating that moves toward the result after each match, with a bigger winning margin moving it more than a narrow one. Unlike a textbook rating, we model draws and home advantage explicitly rather than treating a draw as half a win and home advantage as a fixed fudge — football is draw-heavy, and handling that properly measurably improves how well the rating predicts the next result.
Anchoring the top. The top divisions are calibrated against an external, independently established, cross-border reference, which fixes the absolute scale and the country-vs-country comparisons our own match graph can't resolve. Crucially we re-apply that calibration every season, not once — because a league's average strength drifts over the years as its clubs do well or badly in continental competition, and a single all-time calibration leaves the most recent season floating. We caught exactly this in our own data: one major league's recent average had drifted hundreds of points above its true level, sinking famous clubs far down the global order. Re-calibrating each season restored them — and, tested on results the rating had never seen, the per-season calibration improved it.
The cross-league exchange rate — the inventive piece. To bridge divisions that rarely play each other, we use a completely different signal: player movement. For every player who features in one division one season and a different one the next, we measure the change in his production and difference him against himself, which cancels his individual ability and leaves the league-difficulty effect. To defeat selection bias — clubs buy up the best of a lower league, which makes the gap look smaller than it is — we anchor on promotion and relegation, where whole squads cross the boundary together with no cherry-picking (see Cross-League Projection for the same idea applied to a player's profile). Thin league pairings are pooled toward the average so they aren't overfit, each crossing is weighted by playing time, and only the production stats that genuinely reflect difficulty are used — a diagnostic showed some stats point the wrong way, because they track team style rather than league strength, so they're excluded.
How we keep it honest
Estimated from box-score movement alone — with no knowledge of results or reputation — the exchange rate recovers the consensus order of Europe's "big five" (the English, Spanish, Italian, German and French top flights, in that order) and a clean, monotonic English pyramid from the top tier down to the fourth. It agrees strongly with our results-based rating while clearly adding information of its own — it is not a restatement of the league table, it's a second, independent measurement that happens to confirm it.
What it can't do
- Style inflation in some leagues. A possession-heavy league inflates raw production, so the drop in a player's numbers when he leaves understates that league's true strength. This is the classic difficulty of any movement-based method, and it is not yet fully corrected.
- Deliberately conservative spread. Season-to-season production is noisy, so we shrink hard: the rankings are robust, but the absolute gaps are intentionally understated rather than overstated.
- Residual team selection. Promotion and relegation remove individual cherry-picking, but a promoted squad did over-perform its old tier as a unit; mixing in transfers softens this without erasing it.
- Coverage, not model. The rating only sees the leagues we cover.
The research behind it
The base rating draws on the football-rating literature; the exchange rate on cross-domain league translation (hockey, baseball, econometrics), since the football-specific prior art on selection-corrected league strength is thin.
- Hvattum, L. M. & Arntzen, H. (2010). Using ELO ratings for match result prediction in association football. International Journal of Forecasting. — The margin-of-victory-weighted rating our base uses, with out-of-sample forecasting validation.
- Davidson, R. R. (1970). On extending the Bradley–Terry model to accommodate ties. — The explicit draw and home-advantage model.
- Csató, L. (2021). Tournament Design: How Operations Research Can Improve Sports Rules. Palgrave Macmillan. — Ranking from incomplete, unbalanced schedules — directly the lower-league "island" problem.
- Desjardins, G. — NHL Equivalency. — The canonical movement-based cross-league translation; it acknowledges selection bias without correcting it — the gap our promotion-and-relegation anchor closes.
- Schuckers, M., Lopez, M. & Macdonald, B. (2022). Estimating player aging curves. — The within-player "delta method" behind differencing a mover against himself.
- Heckman, J. (1979). Sample selection bias as a specification error. — The textbook selection correction; considered as a robustness check rather than the primary estimator.
Keep reading
- Player Rating
- Market Values
- Cross-League Projection
- Trajectory & Ceiling Forecasting
- Career Outlook
- Source Models
- Recommendations
- Tactical & Realistic Fit
- Connection Degree
- Work-Permit Eligibility
- Identity Resolution
- How the Data Stays Correct
Oraca is in private beta with a small number of clubs.
Request early access