Methodology

How Team Ratings Work

A team rating is one number that answers a narrow question: given everything this team has done and everyone it has done it against, where does it belong? This page is the long version of how that number is produced, including the parts that are known to be imperfect.

The loop

Rating a team means judging its performances against the quality of its opponents, and the quality of those opponents is itself unknown until they have been rated. That circularity is solved by repetition rather than by cleverness.

Every team starts at the same neutral baseline. The model re-scores every game in the season using the ratings it currently holds, which produces a new set of ratings; it then re-scores every game again using those, and keeps going. Each pass is a better estimate than the one before. The loop stops when the largest change to any rating falls below a small threshold, with a hard iteration cap as a backstop so a pathological schedule cannot spin forever.

Once it settles, every rating contains the schedule that produced it. This is why there is no separate strength-of-schedule column anywhere on the site: it is not an adjustment bolted onto the end, it is the mechanism. A team's number already accounts for who it played, because that is what it was computed against.

Offense and defense are solved separately

The loop does not produce a single value per team. It maintains an offensive and a defensive rating for each channel it tracks, and they constrain each other: a team's offensive rating is built from what it scored relative to what each opponent's defense typically allows, and its defensive rating from what it allowed relative to what each opponent's offense typically produces.

Keeping the two sides apart is what lets the model distinguish a 7-5 team that wins with defense from a 7-5 team that wins with offense, and it is what the offense/defense scatter on each sport's analytics section is plotting. The performance side of the rating averages the two once both have converged, and the published offensive and defensive numbers on a team page are those two halves.

What the model reads

For college football, the rating is built from several independently-solved channels. Each one runs the full loop above on its own before any of them are combined.

Scoring margin is the heaviest input, because points are what the game is actually decided by. Yardage margin is a lighter second channel, and it is there to catch what points miss - a team that moved the ball at will and lost on two turnovers is not as bad as its score, and yardage is the cheapest available correction for that.

Win quality is a third channel, and it is what makes this a ranking rather than a power rating. It is not a raw win percentage - an unadjusted win rate rewards a weak schedule, which is the failure mode where a team from a small conference tops the board on the strength of playing nobody. Instead, each win is credited according to how good the beaten opponent is, and that credit is run through the same opponent-adjustment loop as everything else.

Supporting box-score channels round it out: yards per play, rushing and passing production, first downs, tackles for loss and special teams. Each is solved and opponent-adjusted the same way, and each carries far less weight than scoring margin. They exist to separate teams whose points and yards are similar, not to outvote the scoreboard.

Home field is removed before any of this happens. Each game's home team has a fixed home-field allowance subtracted from its points and yards, so the loop sees neutral-site-equivalent performances and a road win is not quietly penalised. The size of that allowance is a setting; how large home advantage actually turns out to be in the ingested data is measured separately and published on each sport's analytics section.

A win is not worth 1.0

Treating every win as a full win and every loss as a zero throws away real information. A three-point win in overtime and a five-touchdown win are not the same event, and a model that says they are will rank a team that keeps surviving above one that keeps dominating.

So win credit is scaled by margin. A win by about two possessions or more - comfortably out of doubt - earns full credit. Below that, credit slides toward a coin flip as the margin narrows to nothing.

With one exception, which matters: the narrow-win discount only applies when the opponent was below average. Beat an average-or-better team by a field goal and it counts as a full win, because a close game against a good team is not evidence of luck, it is evidence of a good game. Grinding out a three-point win over a bad team is a different fact, and the model treats it as one.

Blowouts are capped, twice

Running up the score is the easiest way to fool a margin-based model, and there are two separate defences against it.

The first is a cap on the scoring channel. Beyond a certain margin, additional points stop counting - winning by 45 and winning by 70 arrive at the loop as the same result. There is no information in the last four touchdowns of a rout; the opponent stopped competing and the second string was on the field.

The second defence is newer, and it exists because of a specific case where the model was wrong.

The case that produced the second defence

When the supporting box-score channels were first switched on for college football, a sanity check on the 2025 season turned up a result that was hard to defend: BYU rated below Utah, despite having the better record and having beaten Utah head to head. The gap was not just wrong in direction - it was wider than it had been before those channels were added.

Decomposing the rating channel by channel explained it. Every one of those new channels was reading full-game raw totals with no blowout protection at all - unlike the scoring channel, which has the cap described above precisely because margin inflates. Utah's season included five games decided by 27 points or more; BYU's included one. So six channels that were supposed to be six independent signals were largely re-measuring the same thing: yardage accumulated in games that were already over.

The fix was a garbage-time discount, applied to all of those channels at once. It keys off the game's actual point margin rather than each statistic's own margin, because garbage time is a property of the scoreboard and not of whichever number you happen to be looking at. A competitive game counts fully; beyond that, a game's contribution decays smoothly as the final margin grows, so a result decided by sixty points counts for a fraction of one decided by ten. It is a decay rather than a cutoff on purpose - a hard threshold would make one point of margin decide whether a whole game counted.

The reason this is on a public page rather than buried in a commit message is that it is the honest version of what building a rating model is like. The channels were added, they measured something other than what they were supposed to measure, a real case caught it, and the diagnosis changed the model. The current supporting weights are inherited from the tuning done before that fix, and re-tuning them against garbage-time-protected data is outstanding work rather than finished work.

Thin records are pulled toward average

After two games, a team's numbers are mostly noise. Left alone, an iterative model will happily rate a team that won its opener by 50 as one of the best in the country, and then spend six weeks walking it back.

College football and NFL ratings handle this by adding a small number of neutral phantom games to every team's record inside the loop - results exactly at league average, against nobody. A team with two real games is then averaging two real results against several neutral ones, so its rating sits close to average until it has earned otherwise. By midseason the real games dominate and the effect fades on its own. It is the same device as a regularisation term in a regression, and it is the reason early-season rankings here look conservative rather than chaotic.

College basketball does not currently apply this, which is a gap rather than a decision - it matters less there, since a basketball team has played a dozen games by the time anyone is looking closely, but the machinery is not present.

Lower-division opponents

Most college football teams play a non-FBS opponent, usually early, usually winning comfortably. These games are a genuine problem: the opponent is not in the rating pool, so there is no rating to adjust against.

The two obvious answers are both wrong. Dropping the game entirely means a team that won it and a team that never played it look identical, when in fact one has a result and the other does not - and it deletes a real performance from a season that may only have twelve. Treating the opponent as an ordinary team means inventing a rating for a team the model has no data on.

Instead, every non-FBS opponent is collapsed into a single synthetic opponent with a deliberately poor but permanently fixed rating. It is an anchor rather than a participant: the model never solves a rating for it and it can never appear in a rankings table. The game then flows through the same opponent-adjustment as any other, which is the point - beating it earns little credit for the same reason beating any team rated that badly does, and losing to it is costly through the same asymmetry that already makes losing to a bad real team costly. No special rule for either case. The stand-in is wired into the scoring, yardage and win-quality channels; the supporting box-score channels skip those games, which slightly undercounts them rather than inventing a baseline per statistic.

How bad that stand-in should be turned out to matter far more than it looks. The first version was set far too low, and it shipped a genuine regression: teams that blew out a lower-division opponent were rated down for it. The schedule adjustment for having played someone that bad was larger in magnitude than everything a dominant, margin-capped win earns on its own, so the game's net contribution came out negative. The win channel was the worst of it, since win credit only spans zero to one to begin with and the penalty swamped it outright.

The correction was to shrink the stand-in by roughly a factor of four, working backwards from the typical single-game range of each channel so the schedule adjustment stays smaller than what a convincing win already earns. A blowout of a lower-division opponent should net modestly positive; an unconvincing one should net close to neutral or slightly negative, because struggling there is real information; a loss should read as clearly worse than losing to an average team. Those values are reasoned from the arithmetic rather than backtested, and if a real blowout ever nets negative again the numbers need to move further, not the mechanism.

College football: the live ranking is resume-first

Everything above describes machinery that several versions of the college football model share. The version currently driving the rankings, team pages and movement makes one further choice, and it is the most consequential one on the site.

Earlier versions split the rating evenly between win quality and on-field performance, and that split was not a guess - it was arrived at by searching for the weighting that best predicted real game margins across several seasons and checking it out-of-sample on a later one. It is the right answer to the question “what predicts the next game?”.

It is not the right answer to the question a ranking asks. The BYU/Utah case above is the clean illustration: the win-quality channel did favour BYU, correctly, but an even split let Utah's stronger raw production outweigh a head-to-head loss and a worse record. No resume-based system - not the AP Poll, not a playoff committee - would come out that way, because a ranking is a claim about what a team has earned rather than a forecast of what it will do next.

So the live model weighs win quality substantially more heavily than performance. It keeps the win channel opponent-adjusted, because a resume ranking still has to account for who was beaten - swapping in a raw win percentage would reintroduce exactly the weak-schedule inflation the adjustment exists to prevent.

Getting that to actually work took a second change that is worth describing, because the first attempt did not do what it looked like it did. Raising the win weight alone barely moved the ratings. The reason was structural: in the earlier formula the win weight only traded against the scoring and yardage channels, while the supporting box-score channels kept their fixed weight no matter what - so even taking the win weight to its theoretical maximum left a substantial block of performance signal untouched. The live model collapses everything non-win into a single composite first and then blends that against win quality, which makes the win weight mean what it appears to mean: how much of the entire non-record signal to trust.

Two caveats, stated plainly. First, this trades away measured predictive accuracy on purpose - which is why score predictions deliberately do not use this model, and why a prediction will sometimes favour the lower-ranked team. Second, the specific weighting has not been validated the way the predictive weights were, and it cannot be validated by the same test: a margin-prediction backtest is precisely the measurement that argues for the even split. A resume model needs a different target - agreement with committee and poll behaviour, or a hand-built set of cases where one team clearly belongs above another - and that work has not been done yet.

Elsewhere: basketball and the NFL

College basketball runs the same loop on four channels instead of two: points, rebounds, assists and turnovers. There is no basketball equivalent of yardage, and pace varies enough between teams that no single statistic stands in for volume produced, so the blend covers it instead. Turnovers are sign-flipped before iterating, since fewer of your own is good and more forced is good - the opposite of the scored/allowed convention everything else uses. Home-court advantage is removed from points only: there are well-established figures for scoring and nothing comparably solid for the other three, and fabricating one would be worse than leaving it at zero and saying so.

Two known simplifications on the basketball side. Rebounds are treated like points - your own count for your offence, the opponent's count against your defence - which is less defensible than it is for scoring, since one team's board is the other team's lost opportunity. And the model does not yet distinguish offensive from defensive rebounds, which do genuinely different things. Both are on the list.

The NFL is ranked on its record first, by a deliberately simpler method than the college football loop. The record part is the Colley method: each win counts for more the better the team it came against, and each loss costs more the worse that team is, with margin and venue left out entirely. The rest is points: each team's points scored and allowed per game against the league average, adjusted for the offences and defences it faced. The record carries 70% of the overall rating and points 30%, the same record-first balance college football's ranking uses; the offence and defence columns are the points part on its own.

Basketball does not yet include a win-quality channel. That is a real difference from college football and the NFL, and it means the basketball rankings are power ratings in the strict sense - descriptions of how a team has played - rather than resume rankings. A team can be rated above another it lost to.

The honest status report on settings: the college football numbers have been moved repeatedly in response to real output across several seasons. The NFL's 70/30 split is a judgement about what a ranking should reward, not a number fitted to results, because there is no poll or committee to measure an NFL ranking against. The basketball weights are a first pass. Those rankings are correct arithmetic on real data, and they are less tested than the college football ones. Both statements are true and the second one is the reason this paragraph exists.

What it does not do

No injury or availability information enters the rating - a team that lost its quarterback in September carries those results at full weight. Nothing below the box score is read: no play-by-play, no drive charts, no field position, no expected-points accounting. Games are weighted equally regardless of when in the season they happened, so a team that improved sharply rates as the average of both versions of itself. And a rating carries no error bar - the shrinkage described above is how uncertainty is handled, rather than by publishing a range.

The methodology overview has the full list of limitations that apply across every model here, and explains why the weights themselves are not published.

See it applied

College football rankings · College basketball rankings · NFL rankings

Each team page shows the schedule its rating was computed from. The polls comparison puts the model next to the human polls, and the analytics section measures things like home advantage and conference bias directly from the data rather than assuming them.