Model reliability · audited
When the model says 70%, do those picks win about 70%? One number for how true its stated confidence is — graded on every completed game, wins and losses alike.
Calibrated & verified · 7 of 9 buckets calibrated within their 95% band
Each dot is a confidence bucket: where the model's stated probability (x) met the rate those picks actually won (y). A perfectly calibrated model rides the diagonal; the 95% band shows where small samples are still noise.
7,004 picks · 9 buckets
| Predicted | Actual win rate | Sample | 95% CI | Verdict |
|---|---|---|---|---|
| 50 to 55% | 52.8% | 1380 of 2616 | 51–55% | calibrated |
| 55 to 60% | 55.3% | 1241 of 2243 | 53–57% | overconfident |
| 60 to 65% | 61.6% | 771 of 1252 | 59–64% | calibrated |
| 65 to 70% | 67.8% | 353 of 521 | 64–72% | calibrated |
| 70 to 75% | 74.2% | 167 of 225 | 68–79% | calibrated |
| 75 to 80% | 78.0% | 64 of 82 | 68–86% | calibrated |
| 80 to 85% | 91.5% | 43 of 47 | 80–97% | calibrated |
| 85 to 90% | 94.1% | 16 of 17 | 73–99% | calibrated |
| 90% or higher | 0.0% | 0 of 1 | 0–79% | overconfident |
Win rate by confidence tier — the climb is the proof this is real signal, not favorite-picking. Tier thresholds are a cross-sport presentation convention.
Daily accuracy vs. the 50% coin-flip baseline. The cold days stay in the line.
This grades calibration — whether our stated confidence is true — not profit.
Last night's record — wins and losses alike — plus tonight's slate. Free, double opt-in, one click to leave.