← All Articles

August 22, 2026

Conference Bias in the AP Poll

We measured how many poll points each conference gets beyond what performance stats predict.

The idea of conference bias has been hotly contested for years. With talks of the SEC splitting off to form it's own league, those debates have been sparked again. The statistics normally used in these debates are extremely simplistic such as number of national champions, CFP appearances, preseason rankings etc. This is a perfect question for a sports and statistic enthusiast to tackle with real data and real statistical methods.

This project being as new as it is, we only used results from the 2025 season, but as we continue to work backwards (or forwards) the results will only get more definitive. Here's a deep dive in the tests we've run.

Test 1: does performance alone predict poll points?

Method: OLS regression + permutation test

The starting point is a regression model. Every FBS team, every week of the 2025 season, has a set of on-field numbers: win percentage, points scored and allowed per game, yards gained and allowed per game, turnover margin, and strength of schedule. We fit a regression model that predicts AP poll points from those numbers alone — pure performance.

That model isn't perfect, and it doesn't need to be. What matters is the leftover: the gap between the poll points it actually got and the poll points the model predicted from performance alone. Average that gap within a conference, and you get a number in real poll points — how much of a bump or penalty that conference's teams get that their stats don't explain.

Averaged across the full season, here's what that gap looked like, sorted highest to lowest:

ConferenceTeamsMean bias (poll points)
SEC16+204.8
FBS Independents2+177.0
Mid-American13+65.2
Big Ten18+21.3
Conference USA12+15.6
Sun Belt14−31.2
ACC17−43.1
Mountain West12−76.6
Big 1216−79.0
Pac-122−99.8
American Athletic14−106.9

For scale, the #1 team in a given week can get up to 1,650 points in this dataset. A gap over 200 points, held for a full season, is not a rounding error.

How Real Is This Result? Our first instinct was to ask "could a spread this size happen by chance?" using a permutation test — shuffle which conference gets credit for which team's gap, thousands of times, and see how often you get a spread this large randomly. That test came back at p = 0.0001: a spread this big showed up once in ten thousand shuffles.

That number was real, but it was answering the wrong question. It treated every team-week as an independent data point — meaning the SEC's 16 teams, each showing up in around 14 rows, counted as roughly 224 independent observations. They aren't. A team's bias one week is heavily correlated with its bias the next week. Treating repeated appearances by the same team as independent evidence can be misleading — it can make an effect look far more certain than the data supports. To fix this, we clustered by team: every team's weekly gaps get averaged into one number first, and every conference-level statistic after that is built from those team-level averages, not the raw team-week rows.

Test 2: which conferences hold up once you fix that

Method: team-level clustering, bootstrap confidence intervals & ANOVA F-test

With clustering fixed, the honest picture is smaller. Re-running the same significance test on team-level data, the overall spread across all eleven conferences is no longer distinguishable from chance (p = 0.46). A companion test — a standard one-way ANOVA F-test, the textbook version of "do these group averages actually differ" — lands right on the edge of conventional significance (p = 0.052).

So does that mean there's no bias after all? Not quite — it means the broad, all-conferences-at-once claim doesn't hold up on its own. To find out which specific conferences are real, we built a 95% confidence interval for each conference individually, resampling its teams with replacement thousands of times. A conference's interval "excluding zero" means we can be reasonably confident its true bias isn't actually zero — the gap survives the added skepticism.

Only two conferences clear that bar. The SEC's interval runs from +40.9 to +374.9 points — solidly positive. The American Athletic's runs from -185.0 to -23.7 — solidly negative. Every other conference, including the Big Ten at a seemingly-real +21.3, has a confidence interval that includes zero once you account for having only a dozen or so teams to work with. Their point estimates might be directionally right, but we can't currently tell them apart from noise.

Test 3: is this a weekly thing, or something else?

Method: week-over-week control regression

Next question: are voters actively favoring SEC teams every single week, or is something else going on? We reran the model with one more input added — the team's own AP poll points from the previous week's release. Controlling for where a team already stood measures the marginal bias in a single week's vote, on top of wherever it already was.

Once you control for that, the SEC's advantage collapses from +204.8 points to +14.4, and it's no longer statistically distinguishable from zero (its interval now includes zero). The overall ANOVA test on this version isn't significant either (p = 0.53). Two smaller, unexpected effects do survive: Mid-American teams get a small but real +13.5-point bump week-to-week (interval: +2.1 to +25.6), and ACC teams take a small real hit of -12.2 (interval: -25.9 to +0.5, right at the edge).

The takeaway: whatever's driving the SEC's season-long gap, it may not be as strong towards the end of the season as it was at the beginning of the season. At least conference wide anyways.

Test 4: does the gap build up over the season, or start there?

Method: early-vs-late-season split

This is the test that changed our mind about our own story. Our first draft of this piece assumed the SEC's advantage builds up gradually — small, defensible nudges each week compounding into a big number by December. That's a reasonable guess. It's also not what the data shows.

We split the season at its midpoint and reran the season-long model separately on the first half (weeks 3-9) and second half (weeks 10-16). In the first half, the effect is at its strongest: the ANOVA test is significant (p = 0.041), and the SEC's interval clearly excludes zero (+42.3 to +405.5). Mountain West, American Athletic, and the Big 12 all show significant negative bias in this half too. In the second half, the effect weakens: the ANOVA test is no longer significant (p = 0.193), and the SEC's interval now just barely includes zero (-6.3 to +373.9). The SEC's raw point estimate actually drops from +225.6 early to +179.6 late.

That's the opposite of a slow build-up. The more likely explanation: early-season polls lean more heavily on preseason expectations and reputation, before enough actual results exist to correct them. As the season goes on, results start to matter more — the performance-only model's R² (how much of the poll it explains) rises from 0.31 in the first half to 0.45 in the second — and the gap narrows, even if it never fully closes. The bias is front-loaded, not accumulated. That doesn't mean there's no bias at the end, just that it gets harder and harder to defend a team who only wins half it's games, no matter how much the voteres may want to.

Test 5: is this just an artifact of how many teams get zero votes?

Method: zero-inflation robustness check

One more honest check. Roughly 80% of team-weeks in this dataset get zero AP points — most FBS teams aren't ranked in a given week. A regression fit across a dataset that lopsided could be distorted by that boundary in ways that don't reflect real "bias" among teams actually in the conversation.

So we reran the season model restricted to only the 45 teams that received any votes at all in 2025 — a much smaller, harder test. And this is where the story shifts again. Among just those teams, the SEC's bias is no longer statistically distinguishable from zero (interval: -149.0 to +228.8) — its edge, such as it is, seems to live mostly in getting into the poll conversation, not in extra credit once it's there. The American Athletic (-330.8, interval: -425.6 to -232.6) and Big 12 (-244.7, interval: -325.6 to -144.7) both come out significantly underrated even among teams that are already earning support — and this test's ANOVA is significant too (p = 0.031). Two conferences in this subset — Independents and Sun Belt — reduced to a single team each once restricted to vote-getters, so we're treating those two as too small to mean anything here, not as evidence either way.

Test 6: how much of this depends on the SEC and AAC specifically?

Method: conference-level cluster bootstrap

One more question worth asking directly: if a different set of conferences existed, would this same story point somewhere else? Tests 1 through 5 all treat the eleven real conferences as a fixed population. This test resamples which conferences appear — drawing 11 conferences with replacement from the real 11, keeping each drawn conference's actual teams intact — 5,000 times, and tracks which conference comes out most-overrated and most-underrated in each draw.

The SEC is the most-overrated conference in 64.7% of those resamples. FBS Independents — which, remember, is really just the Notre Dame effect with UConn along for the ride — takes that spot instead in 24.2% of draws. Mirroring that, the American Athletic is the most-underrated conference in 64.8% of resamples, with the two-team Pac-12 the main rival at 23.9%. The real observed spread (311.7 points) sits right at the top edge of this bootstrap's own 95% interval ([128.2, 311.7]) — about as extreme a result as this resampling process can produce, which is reassuring. But "roughly two times out of three" is a meaningfully honest number to sit with: the SEC/AAC framing is the most likely story, not a guaranteed one. About a third of the time, one of the small-sample wildcard conferences would just as easily take the headline.

Test 7: do conferences disagree with each other more than noise would predict?

Method: meta-analytic heterogeneity test (Cochran's Q)

The last test borrows a technique from meta-analysis. Instead of asking "is any one conference's interval away from zero" (Test 2) or working directly with team-level data (the ANOVA tests throughout), this treats each conference's own team-clustered mean and standard error as a single data point — the way a meta-analysis combines separate studies — weights each conference by how precisely it's been measured, and asks whether the eleven conference estimates disagree with each other by more than their own individual uncertainty would predict. That's Cochran's Q, a standard heterogeneity test.

The result: Q = 17.19 (10 degrees of freedom), with an estimated 41.8% of the variation across conference estimates looking like real heterogeneity rather than sampling noise. But run through the same permutation test used throughout this piece, that comes back at p = 0.42 — not significant. This lines up with, rather than contradicts, everything else: the broad claim that conferences differ from each other remains genuinely uncertain, exactly consistent with Test 2's borderline ANOVA (p = 0.052) and non-significant overall spread test (p = 0.46). It's a second, independently-built test landing in the same place as the first.

What actually survives

Pulling all seven tests together: the SEC does get a real, measurable edge — it shows up clearly in the full-season and early-season tests, and although it diminishes under different perspectives, the gap definitely exists. The broad "conferences differ" claim itself seems to be weaker than we originally thought, but the next step will now to see how much certain teams gain. Top to bottom a single conference may not get as big of a bump as expected, but there may be a core set of teams that consistently get the benefit of the doubt. There's also the very real criticism that Strength of Schedule itself is a biased metric which may have contributed to a biased regression algorithm.

A few more honest limits, as stated at the beginning this is one season of correlational data, not a controlled experiment — the performance stats we used are a reasonable but imperfect stand-in for "how good is this team," and something the model can't see (roster quality, injuries, style matchups) could be doing some of the work we're attributing to bias. Independents and the Pac-12 are each just two teams (Notre Dame + UConn; Oregon State + Washington State) — not a real sample for a "conference" claim. We'd want to see this same battery of tests run across several more seasons before treating any specific number here as fixed rather than a 2025 snapshot.

For a better visual representation of these tests check out factor-importance breakdown, which already showed conference membership carrying a real share of AP poll point variance, on par with a team's own win percentage. That analysis asked how much conference matters overall; this one asked how much, for which conferences specifically, and whether the effect survives real scrutiny. The full breakdown — every test, every number, methodology included — is also up on the site below the existing factor analysis.