TRANSACTION · MLBDiamondbacks: Acquired OF Lars Nootbaar from St (Aug 3)
TRANSACTION · MLBReds: Acquired RHP Alejandro Rivera and 2B Juan Brito from Cleveland Guardians for DH Nat…
TRANSACTION · MLBNationals: Reinstated RHP Yovany Cruz from the IL (Aug 3)
Free during beta —to track favorites + alerts

Model reliability · audited

The Calibration
Grade

When the model says 70%, do those picks win about 70%? One number for how true its stated confidence is — graded on every completed game, wins and losses alike.

89out of 100

Calibrated & verified · 7 of 9 buckets calibrated within their 95% band

Record 14361–7158 (66.7%, 95% CI 66.1%–67.4%)ECE 0.011 · 7,004 graded

The reliability line

Each dot is a confidence bucket: where the model's stated probability (x) met the rate those picks actually won (y). A perfectly calibrated model rides the diagonal; the 95% band shows where small samples are still noise.

7,004 picks · 9 buckets

Model calibration: predicted win probability vs actual win rate by bucket, with 95% Wilson confidence intervals, over 7,004 graded picks.
PredictedActual win rateSample95% CIVerdict
50 to 55%52.8%1380 of 26165155%calibrated
55 to 60%55.3%1241 of 22435357%overconfident
60 to 65%61.6%771 of 12525964%calibrated
65 to 70%67.8%353 of 5216472%calibrated
70 to 75%74.2%167 of 2256879%calibrated
75 to 80%78.0%64 of 826886%calibrated
80 to 85%91.5%43 of 478097%calibrated
85 to 90%94.1%16 of 177399%calibrated
90% or higher0.0%0 of 1079%overconfident

The harder it commits, the more it's right

Win rate by confidence tier — the climb is the proof this is real signal, not favorite-picking. Tier thresholds are a cross-sport presentation convention.

Toss-up
52.6% (3,927)
Lean
57.5% (4,804)
Edge
67.7% (6,101)
Highest confidence
80.8% (6,687)

Last 30 days

Daily accuracy vs. the 50% coin-flip baseline. The cold days stay in the line.

30-day accuracy versus the 50% coin-flip baselineCOIN FLIP 50%

How it's computed

  • For each confidence bucket we take the sample-weighted gap between the mean stated probability of the picks in it and the rate they actually won — the Expected Calibration Error (ECE). It uses the mean stated prob, not the bucket midpoint, so the number is reproducible from the published CSV.
  • Grade = round(100 − ECE × 1000): an ECE of 0.00 → 100, 0.05 → 50, 0.10 or worse → 0. The raw ECE sits beside the grade so the scaling is auditable.
  • It grades the model's raw stored win probability against the realized outcome on every completed game in the five fitted leagues — no window, no win-filter, cold streaks included.
  • Below 1,000 graded games or 4 populated buckets the grade reads “—” rather than publish a noisy number.

This grades calibration — whether our stated confidence is true — not profit.

Download the full bucket-level CSV →

Get the graded slate in your inbox

Last night's record — wins and losses alike — plus tonight's slate. Free, double opt-in, one click to leave.

Support this work →