Model Track Record

Pre-registered backtest

We publish our misses, too. Every number below comes from a held-out backtest — the model never saw the games it is graded against — and we show the results that don't clear the bar right next to the ones that do.

Pre-registered, not retrofitted: the pass/fail bars and confidence intervals on this page were written down before we looked at the outcomes they grade. That is what lets a result like “95% confidence interval crosses zero — not yet demonstrated” survive onto this page instead of getting explained away.

Weekly Projection Calibration

As of 2026-07-18 · source: _bmad-output/planning-artifacts/projection-band-calibration-backtest-1877.md

Every week we publish a projected floor-to-ceiling range plus a single point projection for every skill-position player, built walk-forward — the projection for week 8 has only ever seen weeks 1–7. We checked how often the published range actually covered what happened, on two independent held-out replays.

2024 → 2025 (clean, headline)
64.5%
of outcomes covered vs. 65% nominal — well-calibrated
2023 → 2024 (supporting)
59.9%
of outcomes covered vs. 65% nominal — well-calibrated (within ±0.07)

Coverage by position (2024 → 2025)

QB
59.1%
RB
63.0%
WR
64.9%
TE
70.0%
PositionOur projection (RMSE)Naive trailing-average
Overall6.817.18
QB8.498.95
RB6.787.15
WR6.556.89
TE6.036.45

Pooled RMSE, our leak-free point projection vs a naive season-to-date trailing average (lower is better). Lower RMSE is better — our projection beats the naive baseline at every position.

QB bands under-cover the most (59.1% vs the 65% target) — the residual miss is mild and concentrated in the low tail (bench/DNP weeks), not a width bug. Effective sample size is roughly 68 independent week×position clusters per leg, not the raw row count — read the second decimal as suggestive, not a tight season-level claim.

Guillotine Survival Mechanism

As of 2026-07-19 · source: data/survival_skill_scoreboard.json

In guillotine leagues, the last-place team is eliminated every week — the format punishes a shallow roster harder than any other. Our guillotine draft board orders players by a projected floor (a realistic worst week), not average projected points. We backtested that ordering against the market's own ADP as the baseline to beat, across four seasons (2022, 2023, 2024, 2025), with both sides matched to the same information vintage — no lookahead on either side.

Survival edge (point estimate)
-0.78 wks
weeks alive beyond the market baseline
95% confidence interval
-3.87 to +2.18 wks
crosses zero
Not yet statistically demonstrated
SeasonSurvival edge (weeks)
2022+2.71
2023+1.53
2024-4.94
2025-2.41

Positive in 2 of 4 backtested seasons (one-sided sign-test p=0.6875 — not significant at this sample size). Weakest week in the backtest: week 11, where the market baseline slightly out-lasted our board (-0.3088 weeks of alive-rate edge).

The entire claim rests on n_eff = 4 independent seasons. The 18 draft slots inside each season collapse to a few distinct survival outcomes (shared score realisation), so they add resolution, not independence. Survival edge was positive in 2/4 seasons (exact one-sided sign-test p = 0.6875); even a perfect sweep at this n is not significant at 0.05. The season-clustered 95% CI on the survival edge is [-3.8684, 2.1768] weeks and CROSSES zero; with so few clusters this CI is wide and should be read as directional, not a certified magnitude. These numbers are the board_order='published' (served, #2114) headline; the 'floor' board scores -1.0147wk vs this one's -0.7794wk on the same seasons/market — a reminder that floor-only evidence does not certify a served change.

The miss: within-position finish-rank skill

A second, harder metric checks how well our board's internal player order (within each position) matches how the season actually finished, compared with ADP. On this metric our board currently scores worse than the market — and unlike the survival edge above, this gap does clear statistical significance (95% CI -0.200 to -0.110, excludes zero).

QB
-0.212
RB
-0.172
WR
-0.163
TE
-0.066

Negative = worse ranking skill than ADP for that position; positive = better. Weakest: QB (-0.212), weakest season 2025 (-0.205).

We are not claiming our guillotine board currently outperforms the market: the primary survival-edge interval crosses zero, and the finish-rank-skill metric is currently negative and statistically significant. This is a mechanism we are testing in the open, including the metrics that don't clear our own bar yet.

Tested and closed

As of 2026-07-18 · source: data/qb_floor_v2_evidence_1798.json

Not every idea we test ships. Here's one we closed rather than arm, with the numbers that closed it.

Quarterback scoring-floor refinement

We tried a refinement to how we estimate a quarterback's floor value on the guillotine draft board. It did not clear our own guardrails: the effect on within-position draft order moved the wrong direction on net (pooled change -0.0101, confidence interval crosses zero — not demonstrated), and the effect on season-long survival regressed in the season we could measure it in. We closed the idea rather than ship a change we can't show helps.

How to read this page

  • Everything above is leak-free by construction: every backtest here is checked to confirm the model never used information from the game, week, or season it is being graded on.
  • Small samples are flagged, not hidden. The guillotine backtest spans four NFL seasons — four independent data points for anything season-level, not thousands. We report effective sample size, not just row count, and mark a result “not yet demonstrated” whenever its confidence interval crosses zero.
  • No raw third-party proprietary values appear on this page. Every accuracy metric is our own model's output, graded against publicly available actual results.

For our always-on, week-by-week graded hit-rate ledger (start/sit rankings and streaming picks), see our Track Record page.

This page covers pre-registered backtest validation, not live in-season grading. Numbers are regenerated as new backtests complete; each section states its own sample size and confidence interval so you can judge how much weight to put on it.