When the model says 60%, does it hit 60%? Every decisive ticket with a model probability (547 of 567), bucketed by what was predicted. A band is honest when its predicted rate sits inside the interval its actual record supports.
| Predicted band | n | Predicted | Hit | 95% CI | Verdict |
|---|---|---|---|---|---|
| 0-10% | 6 | 7.0% | 0/6 | 0–39% | INSUFFICIENT EVIDENCE |
| 10-20% | 39 | 17.3% | 10.3% | 4–24% | WITHIN INTERVAL |
| 20-30% | 44 | 24.0% | 25.0% | 15–39% | WITHIN INTERVAL |
| 30-40% | 21 | 35.1% | 38.1% | 21–59% | WITHIN INTERVAL |
| 40-50% | 43 | 46.4% | 46.5% | 33–61% | WITHIN INTERVAL |
| 50-60% | 234 | 55.9% | 46.6% | 40–53% | OUTSIDE INTERVAL |
| 60-70% | 124 | 63.9% | 62.9% | 54–71% | WITHIN INTERVAL |
| 70-80% | 35 | 72.4% | 57.1% | 41–72% | OUTSIDE INTERVAL |
| 80-90% | 1 | 81.1% | 1/1 | 21–100% | INSUFFICIENT EVIDENCE |
| 90-100% | 0 | — | — | — | INSUFFICIENT EVIDENCE |
The dashed line bet the identical tickets at a level $5 — no Kelly, no conviction sizing. The gap is what sizing decisions were worth: +$100.74.
A standing daily column, keyed to the date so a re-publish prints the same edition. Each question gets two answers. Ten thousand replays drawn at the model's own probabilities set the ceiling it promised — every ticket here cleared a +5pp edge floor to get bet at all, so those probabilities are optimistic by construction and a good season still places low against them. Ten thousand more drawn at the prices actually paid set the neutral line: the same slate, the same stakes, every ticket assumed exactly as likely as its price implied. That second population breaks even on average, and it flatters us — it takes the book's vigged price as the truth. An honest season sits between the two. Today's question: where does our projected finish land?
Two panels refuse to backfill. Historical verdicts survive only in prose on a self-selected subset of tickets — not evidence — and vetoed plays were never graded. Rather than dress a biased sample as an audit, both start counting from today.
| Verdict | Tickets | Record | Net | ROI | 95% CI |
|---|---|---|---|---|---|
| Survives | 118 | 62-56 | +$1.32 | +0.2% | 43.6-61.3% |
| Softened | 130 | 69-61 | +$4.64 | +0.7% | 44.5-61.4% |
Every ticket carries its gauntlet verdict into the tracker at bet time. Survives vs Softened, head-to-head, graded flat-$5 to strip out the Softened 1-unit sizing cap — the cap's own recalibration data.
From today, every play the gauntlet rejects is graded anyway, by gate — edge floor, inviolables, the Owl's vetoes — a league table of which refusal earns its keep. First verdict at n=30 vetoes.