# Benchmarks: what "good" means here
*Published 2026-07-18. Goals on this page change only by dated public edit (see §9).
All performance numbers for our own models are quoted from the served public artifacts
at /data/ (backtests generated 2026-07-17/18 UTC on the production droplet); local
re-runs of the same code and data reproduce them to ±0.002 log-loss and ±1.5
percentage points of per-line ECE (cross-machine gradient-boosting nondeterminism —
seed and row order are pinned; floating-point reduction order is not). Anything
labeled [DERIVED] is arithmetic on cited numbers, not a published measurement.
Anything labeled [INDUSTRY] is a practitioner source, not peer-reviewed.*
---
## 1. Brier scores are base-rate dependent — never compare them across markets
The Brier score (mean squared error of a probability) depends on the event's base
rate: always predicting the base rate p̄ scores p̄(1−p̄) — the "climatology"
baseline, the standard of reference in forecast verification. A Brier of 0.11 on a
13% event (HR) is *worse relative to available headroom* than 0.24 on a 55% event.
Every market on this site is therefore anchored to its own climatology:
| Market | Base rate (source) | Climatology Brier |
|---|---|---|
| Game winner (home) | 53.5% home wins, 2025–26 store, n=3,875 reg-season games | **0.2488** (FiveThirtyEight's independently computed "unskilled" baseline on 2016–20: 0.2489) |
| K over 4.5 | 55.2% (holdout 6/1–7/16/26) | 0.2473 |
| K over 5.5 | 37.8% | 0.2351 |
| K over 6.5 | 25.5% | 0.1900 |
| Hits over 0.5 | 61.9% | 0.2359 |
| Hits over 1.5 | 22.2% | 0.1724 |
| HR 1+ | 13.1% | 0.1136 |
| Game total runs | mean 8.95, sd 4.58 (2025–26 store) | climatological **CRPS 2.57** (see §6) |
The marginal-distribution baseline reported in every backtest of ours is exactly
this climatology computed per line; our "skill" is always stated against it.
Cross-check on the elite public reference: FiveThirtyEight's pooled 0.2401 against
baseline 0.2489 is a Brier skill score of 3.5% — matching their published 0.0351.
## 2. The elite public evidence: what the best publicly graded MLB models score
- **FiveThirtyEight, "Checking Our Work"** (Jay Boice & Gus Wezerek; archived:
http://web.archive.org/web/20221222160221/https://projects.fivethirtyeight.com/checking-our-work/mlb-games/ ;
raw data CC-BY-4.0: https://github.com/fivethirtyeight/checking-our-work-data).
MLB game winners, 2016–2020, ~10,800 games: **Brier 0.2401 pooled**
(95% CI 0.2389–0.2413), seasons ranging **0.2365–0.2428**, vs unskilled baseline
0.2489. Calibration essentially on the diagonal at every well-populated bin
(forecasts binned at 60% won 61.3%, n=2,729; at 45% won 45.5%, n=3,890).
A companion article (Neil Paine, Feb 2021, archived:
http://web.archive.org/web/20210608104749/https://fivethirtyeight.com/features/how-well-did-our-sports-predictions-hold-up-during-a-year-of-chaos/)
quotes a slightly different 2020 figure (0.243) from a different snapshot; we
treat the Checking-Our-Work table, whose raw data is public, as canonical. Same
article: MLB favorites win only ~57–59% of games (their NFL model: favorites
68.6%, Brier 0.208 — baseball's floor is structurally higher).
- **Choe & Ramdas, "Comparing Sequential Forecasters"** (CMU; arXiv:
https://arxiv.org/abs/2110.00115). All 25,165 MLB games 2010–2019: Vegas closing
odds beat FiveThirtyEight by an average Brier margin inside (0.00062, 0.00265)
per game (95% confidence sequence), and "none of the other forecasters,
including 538, have outperformed vegas from 2010 to 2019."
[DERIVED] Combining with the 0.2401 above puts closing-line Brier at roughly
**0.237–0.240** — note the derivation crosses two different evaluation windows
(2010–19 vs 2016–20); treat as an anchor, not a measurement.
- **Wolfson & Koopmeiners** (arXiv: https://arxiv.org/abs/1501.07179): models
using only game outcomes "never exceeded 58%" accuracy on MLB even with 7/8 of a
season of training data. **Cui** (Wharton thesis, 2020:
https://fisher.wharton.upenn.edu/wp-content/uploads/2020/09/Thesis_Andrew-Cui.pdf)
surveys published public models at 55–60% accuracy; his richer feature set
reached 61.77% (AUC 0.671) on 9,700+ games, 2016–19. Public MLB winner accuracy
lives in the 55–62% band, with the low end for outcome-only features.
**No equivalent public evidence exists for player props.** We found no
peer-reviewed or independently auditable public record of prop-line probabilities
graded with proper scoring rules and calibration curves (see §8).
## 3. The variance ceiling: why nobody scores much better than this
- **Lopez, Matthews & Baumer**, *Annals of Applied Statistics* 12(4), 2018
(arXiv: https://arxiv.org/abs/1701.05976; ~26,728 MLB games of market lines,
2006–16): "The median probability of the best team winning a neutral site game
is highest in the NBA (67%), followed ... and MLB (56%)." "It is rare that the
best MLB team is *ever* given a 70% probability of winning, with the middle 50%
of games ranging from 57% to 63%." "MLB game outcomes remain lightly-weighted
coin flips."
- **Wolfson & Koopmeiners** (above): MLB games are the least informative of the
four major sports about team strength — "not surprising ... given the
significant role that the starting pitcher plays."
- **Richards**, SABR *Baseball Research Journal*, Fall 2014
(https://sabr.org/journal/article/probabilities-of-victory-in-head-to-head-team-matchups/):
the log5 formula (James, 1981). A .600 team beats a .500 team exactly 60% of the
time; beats a .400 team 69.2%. [DERIVED from that formula] An average team beats
a .600 team 40% of the time, and beats a .650 juggernaut 35% of the time.
- Season ceiling (Wikipedia, "List of best MLB season win–loss records":
https://en.wikipedia.org/wiki/List_of_best_Major_League_Baseball_season_win%E2%80%93loss_records):
the best team of the last century, the 2001 Mariners (.716), still lost 28% of
its games. FiveThirtyEight (Paine, archived above): "the top-ranked Dodgers
would have only a 72 percent chance of beating the bottom-ranked Pirates at a
neutral field, and that's about as extreme as mismatches can get in MLB."
- [DERIVED — corrected arithmetic] If true single-game probabilities live in
[0.35, 0.65] (the cited range), a perfect-information forecaster's expected
Brier E[p(1−p)] cannot go below 0.2275, and that floor requires every game at
the extremes; a realistic mix consistent with the Lopez distribution gives
≈ 0.244. **Total headroom between climatology (0.2489) and omniscience is on
the order of 0.005–0.02 Brier.** That is the entire playing field, which is why
a 3.5% skill score (§2) is elite and why our goals are stated in thousandths.
## 4. The closing line is the standard — and props carry extra friction
- **Levitt**, *The Economic Journal* 114(495), 2004 (NBER WP:
https://www.nber.org/papers/w9422): "little evidence that there exist bettors
who are systematically able to beat the bookmaker."
- **Sauer**, *Journal of Economic Literature* 36(4), 1998
(https://ideas.repec.org/a/aea/jeclit/v36y1998i4p2021-2064.html): betting
prices are "to a first approximation ... efficient forecasts of outcomes."
- **Woodland & Woodland**, *Journal of Finance* 49(1), 1994 (MLB-specific:
https://onlinelibrary.wiley.com/doi/10.1111/j.1540-6261.1994.tb04429.x):
baseball's deviations from efficiency "are shown to be insufficient to allow
for profitable betting strategies when commissions are considered."
- **Buchdahl**, Football-Data.co.uk
(https://www.football-data.co.uk/blog/pinnacle_efficiency.php): across 87,960
closing prices, closing-line-implied value predicts realized returns with slope
≈ 1.00 — the empirical basis of "closing line value" as the benchmark.
- [INDUSTRY] Hold by market tier (Wizard of Odds:
https://wizardofodds.com/article/player-props-understanding-the-math-behind-the-lines/ ;
mainline context: https://www.nfeloapp.com/tools/sportsbook-hold-calculator/):
MLB moneylines ~3–4%; major two-way props 4–6%; secondary props 6–10%. Prop
markets carry roughly 2–3x the friction of moneylines — thinner, less sharp.
**Our position (unchanged):** the goal is calibration, not beating Vegas. When the
v1.1 closing-odds column ships, the standard we adopt is *proximity to* the
de-vigged closing line (§7, goal CL-1), with the elite reference being the
0.0006–0.0027 Brier gap that separated FiveThirtyEight from Vegas closing. A model
that matches the closing line with zero edge, reported straight, is the product.
## 5. Calibration norms — and why small-n curves get uncertainty bands
- The tradition: **Murphy & Winkler**, *JASA* 79(387), 1984
(https://www.tandfonline.com/doi/abs/10.1080/01621459.1984.10478075) — the US
National Weather Service's precipitation-probability program is the canonical
proof that large-sample, near-diagonal reliability is achievable in production.
**Gneiting, Balabdaoui & Raftery**, *JRSS-B* 69(2), 2007
(https://sites.stat.washington.edu/raftery/Research/PDF/Gneiting2007jrssb.pdf):
the goal is to maximize sharpness *subject to* calibration.
- ML practice: **Guo et al.**, ICML 2017 (https://arxiv.org/abs/1706.04599):
uncalibrated deep nets commonly show ECE 5–16%; ~1–2% after recalibration is
considered strong. (Their ECE uses 15 equal-width bins; ours uses 10 —
comparisons are approximate.)
- Small-n honesty: **Bröcker & Smith**, *Weather and Forecasting* 22(3), 2007
(https://www.lse.ac.uk/CATS/Assets/PDFs/Publications/Papers/2007/73-IncreasingReliabilityDiagrams-2006-Brocker-Smith.pdf):
observed bin frequencies wander off the diagonal even for perfectly reliable
forecasts; reliability claims at small n require consistency bands.
**This applies to OUR live curves** and to our goals. Our estimator (10
equal-width bins, sample-weighted mean |observed − predicted|) has a nonzero
noise floor under *perfect* calibration. Simulated on our own probability mix
with our own code:
| Sample | Perfect-calibration measured ECE: mean / 95th pct |
|---|---|
| K, one line, n=500 | 3.3–4.1% / 5.3–6.3% |
| K, one line, n=2,000 | 1.6–2.0% / 2.5–3.1% |
| K pooled, n=3,000 | 1.7% / 2.6% |
| Hits pooled, n=1,000 | 2.6% / 4.1% |
| Hits pooled, n=10,000 | 0.8% / 1.3% |
| HR, n=500 | 1.8% / 3.4% |
| HR, n=5,000 | 0.6% / 1.1% |
Every ECE goal in §7 is set above the 95th-percentile noise floor at its stated
n — a goal below that floor would fail a perfect model, which is theater, not
measurement. Binning-choice caveats: Dimitriadis, Gneiting & Jordan, *PNAS*
118(8), 2021 (https://arxiv.org/abs/2008.03033).
- Context for our current numbers: at backtest scale (19,892 hit pairs) our
pooled hits ECE of 0.50% sits *below* the 0.57% that a perfectly calibrated
model would measure on average at that n — i.e., statistically
indistinguishable from calibrated. Per-line K ECE at backtest n (≈1,035/line)
ranges 2.6–4.4% across re-runs, which is *inside* the noise band at that n:
we cannot yet distinguish it from calibrated, and we say so rather than claim it.
## 6. Scoring rules: which market uses what, and why
- **Log-loss** (multiclass, natural log) for the full K/hits/HR count
distributions: strictly proper, punishes overconfident tails, and is the metric
of our persistence challenge.
- **Brier per line**: strictly proper for the binary over/under events, bounded,
and comparable to book-implied probabilities once the odds column exists.
- **CRPS** for total-runs distributions: the proper score for full predictive
distributions. Roots: Matheson & Winkler, *Management Science* 22(10), 1976
(https://ideas.repec.org/a/inm/ormnsc/v22y1976i10p1087-1096.html); modern
treatment: Gneiting & Raftery, *JASA* 102(477), 2007
(https://sites.stat.washington.edu/raftery/Research/PDF/Gneiting2007jasa.pdf);
count-data adaptation (our exact setting): Czado, Gneiting & Held, *Biometrics*
65(4), 2009 (https://onlinelibrary.wiley.com/doi/abs/10.1111/j.1541-0420.2009.01191.x);
operational precedent at scale: Bracher et al., *PLOS Comp. Biol.* 17(2), 2021
(https://arxiv.org/abs/2005.12881). We found no published CRPS evaluation of
pitcher-strikeout distributions; to our knowledge this scoreboard is the first
public one (see §8).
- **Climatological CRPS reference for totals, computed from our store**
(regular-season games, empirical total-runs distribution 2025 through
2026-05-31, n=3,317, scored on the 2026-06-01..07-16 holdout, n=558):
**2.5678 runs**. Pooled in-sample variant: 2.53. A point forecast at the
climatological median scores CRPS 3.63 on the same holdout — distributions
beat point estimates by construction, which is why we publish strips, not
numbers. Note: the holdout scored 9.35 runs/game vs 8.88 in training — run
environments drift within a season, and even climatology is nonstationary;
the reference will be recomputed (by dated edit) each time the reference
window rolls.
## 7. The goals — falsifiable, dated, current standing stated
Accumulation rates as of 7/18/26: ~90 K line-grades/day (~30 starts), ~540 hit
pairs/day, ~270 HR grades/day, live since 7/17; season ends ~9/28 (~72 slates).
Dates below assume those rates and carry ±4 days for slate size, postponements,
and voids. "Backtest" = the served out-of-sample artifacts (train < 6/1/26,
holdout 6/1–7/16/26). Live verdicts are computed on the live record only —
backtest results never count toward a live goal.
| # | Goal (metric, population, threshold) | Gate n / date | Standing 7/18/26 |
|---|---|---|---|
| K-1 | **The persistence challenge:** live multiclass log-loss (classes 0–12+) below the shrunk-Poisson persistence baseline (2/3 × pitcher trailing-5 prior-start mean K + 1/3 × league mean; exact formula in `pipeline/backtest.py`) on identical graded starts. Verdict with paired bootstrap CI. | checkpoint n≥1,000 starts ~8/20; verdict season end 9/28 (n≈2,100) | **NOT MET in backtest, but the gap is not statistically distinguishable from zero**: **figures corrected 2026-07-20** — these quoted a local re-run; the SERVED artifact at `/data/backtest_calibration.json` reports 2.2525 vs 2.2370, deficit **+0.0155**. The 0.0062 spread between the two runs exceeds the ±0.002 tolerance stated at the top of this page, so the served figure is authoritative and the site now derives this sentence from it rather than restating it with a paired-bootstrap 95% CI of **[−0.0103, +0.0371]** — the interval crosses zero. The honest statement is that we cannot distinguish the model from persistence on log-loss, not that persistence beats us (corrected 7/18 after the quant gate; the earlier wording erred in the self-critical direction, which is still an error). The model wins all three line Briers (0.2289/0.2130/0.1715 vs 0.2350/0.2183/0.1726). Stays posted until decisively beaten at the n≥1,000 checkpoint. |
| K-2 | Pooled K ECE ≤3.5% at n≥3,000 live line-grades; ≤3.0% at season end (n≈6,500). Noise floor (95th pct, perfect model): 2.6% / 1.8%. | ~8/20; 9/28 | **Standing corrected 2026-07-20** (the figures below were stale against the served artifact): backtest pooled ECE **2.22%** at **n=3,270**, per `/data/backtest_calibration.json`. The previous text quoted 1.87% at n=3,105, which the served file no longer supports. On track against the ≤3.5% gate, with less headroom than the old number implied. |
| K-3 | Per-line K ECE ≤4.0% on every line at n≥2,000 per line. Noise floor at that n: ≤3.1%. | season end (~9/22–9/28) | **Standing corrected 2026-07-20 — the previous text understated this, and in the direction that flattered us.** Served per-line ECE is **2.16% / 3.87% / 2.62%** at the 4.5 / 5.5 / 6.5 lines (`/data/backtest_calibration.json`), a range of **2.16–3.87%**, not the 2.6–2.9% previously published. The 5.5 line sits **0.13 points** below K-3's own 4.0% failure threshold — not comfortably inside it. Still inside the ≤3.1% noise band only for two of the three lines, so the honest statement is that K-3 is at genuine risk on the 5.5 line rather than merely 'uncertain'. |
| H-1 | Live hits log-loss below BOTH the marginal (climatology) and persistence baselines on identical graded rows. | n≥10,000 pairs ~8/5; season end | Met in backtest: 1.1992 vs 1.2052 / 1.2144. |
| H-2 | Pooled hits ECE ≤1.5% at n≥10,000 live pairs; stretch ≤1.0% at n≥30,000. Noise floor: 1.3% / ~0.7%. (Tightened from the draft "≤2% at n≥1,000", which was below the n=1,000 noise floor — see §5.) | ~8/5; ~9/10 | Met in backtest at scale: 0.50% at 19,892 pairs (indistinguishable from perfect at that n). Live degradation allowance is built into the 1.5%. |
| HR-1 | HR ECE ≤2.0% at n≥5,000 live grades; ≤1.5% at n≥15,000. Noise floor: 1.1% / ~0.6%. | ~8/5; ~9/11 | Met in backtest (1.14%) but AT RISK: confirmed −0.93pt summer-drift under-prediction is most of that ECE; if drift worsens before the weather features ship, this can fail. Stated now so a failure later is a data point, not a surprise. |
| HR-2 | Weather-feature lift disclosed: publish holdout Δlog-loss, ΔBrier, and how much of the summer bias temperature absorbs — whatever the sign. | shipped 2026-07-19 | **MET — published in §10 below.** Discrimination barely moved (Δlog-loss −0.0006, ΔBrier −0.0001, both inside the noise band); the honest result is that weather bought calibration, not skill: it absorbed **64% of the summer under-prediction** (−0.85pt → −0.31pt) and cut ECE 1.08% → 0.61%. |
| HR-3 | Live HR log-loss below both baselines. | n≥5,000 ~8/5 | Met in backtest: 0.4154 vs 0.4200 / 0.4209 — thin margins, honestly thin. |
| CL-1 | **Dormant — activates by dated edit when the v1.1 closing-odds column ships.** Per prop market: cumulative live Brier within **+0.005** of the de-vigged (proportional two-way normalization, method published with the column) closing line on matched grades, n≥1,000 matched per market. Reference: the elite-public-to-Vegas gap on winners was 0.0006–0.0027 (§2); prop friction 4–10% [INDUSTRY]. Beating the close is NOT the goal and will not be claimed. | n≥1,000 matched per market after activation | Dormant. |
| W-1 | Runs-engine launch gate (winners): backtest Brier < 0.2488 (store climatology, home rate 53.5%, 2025–26, n=3,875) on the standard holdout, before anything publishes. Engine builds only after the weather workstream clears its gate. | pre-launch | **FAILED — winners is NOT published (dated edit 2026-07-20).** A single fit scores 0.2486, which clears. That is not a pass: refitting across 100 seeds clears the gate on only **73**, margin 0.65 seed-sigma, so the reported result is seed-dependent. Discrimination is the real problem — **AUC 0.539** against ~0.60 for elite public models — and the gap vs holdout climatology is −0.0024 with a 95% CI of **[−0.0060, +0.0011]** that crosses zero. Root cause is structural, not tuning: deriving P(win) from two runs-*scored* marginals fights the ninth-inning truncation. The home team forfeits its last at-bat in 44.8% of games and does so *precisely when winning* (an endogenous stopping rule), so home runs-scored understates home strength — model mean P(home) 0.505 vs a 0.540 training rate. A separate direct win model would score better and is **rejected**: it breaks the one-engine coherence the design exists to provide. v1.1 candidates: 9-inning-equivalent runs, or carrying the truncation explicitly. **Basis note:** 0.2488 is the *pooled in-sample* variance over all 3,875 games (0.5350 × 0.4650), not a holdout-scored figure — unlike T-1, which is holdout-scored. Threshold left unchanged; see the clarification below this table. |
| W-2 | Live winners: paired Brier gap vs slate-matched climatology < 0 at n≥300 games (~launch + 20 slates at ~15 games/day), with 95% CI shown. Honesty note: SE of the gap at n=300 is ≈0.004, so only roughly 538-pooled-size skill (−0.009) clears zero decisively; a smaller true edge will still be inside the CI and we will say so rather than claim it. | launch + ~20 slates | **Stays dormant (dated edit 2026-07-20).** W-1 failed, so winners does not publish and there is no live record to score. Does not activate until W-1 is cleared robustly. |
| W-3 | Elite approach: live Brier ≤0.245 at n≥300; ≤0.243 (inside FiveThirtyEight's 0.2365–0.2428 seasonal band, at its weak end) at n≥600. If launch lands after ~8/18, n≥600 crosses the season boundary and the goal carries into 2027 unchanged. | launch + ~20 / ~40 slates | **Stays dormant (dated edit 2026-07-20).** Same reason as W-2. Worth stating plainly: at AUC 0.539 this engine is not close to the 538 band, and reaching it needs the truncation fix, not more tuning. |
| T-1 | Totals launch gate: mean discrete CRPS < **2.5678** (climatological reference defined in §6) on the same holdout. | pre-launch | **MET — totals published 2026-07-20.** Model CRPS **2.5466**, clearing on **100 of 100** refits at 4.0 seed-sigma. Three things ship with that number and are not optional: (1) **the gate is a floor, not a skill test** — an `is_home`-only null model, one binary feature and zero information, scores **2.5604** against the same 2.5678 gate, so clearing it means "not worse than knowing nothing"; (2) **any edge over climatology is marginal** — gap −0.0212, 95% CI **[−0.0413, −0.0007]**, excluding zero by 0.0007, which is *less than this engine's own seed-to-seed spread*, and the same interval crossed zero on a different store vintage; read as suggestive, not established; (3) **an undisclosed level miss would be disqualifying, so here it is** — model mean total 9.06 vs 9.35 actual (−0.29 runs/game, −1.5σ), the model tracking the training run environment while the holdout ran hotter. Same failure family as the HR summer drift behind HR-1/HR-2. |
| T-2 | Live totals: paired mean CRPS below slate-matched climatology at n≥300 graded games; line Briers published at posted totals vs the climatology anchors (7.5: 0.2430, 8.5: 0.2504, 9.5: 0.2458 on the holdout). | launch + ~20 slates (~15 games/day ⇒ **checkpoint ≈ 2026-08-09**); season end 9/28 | **ACTIVATED 2026-07-20** — totals is live and grading nightly. Verdict reported with a paired-bootstrap 95% CI, and the CI governs the claim: at n=300 the SE on a CRPS gap of this size is large enough that a true edge of the backtest's magnitude (−0.02) will **not** clear zero decisively, so the expected honest outcome at the checkpoint is "cannot distinguish", and that is what will be written. Backtest numbers do not count toward this goal (§9). |
### Clarification: W-1 and T-1 were stated on different bases (dated 2026-07-20)
Caught by making reference-reproduction an *assertion* rather than a
printout — `pipeline/backtest_runs.py` refuses to report a gate until it
first reproduces the frozen climatology numbers from the store. Four
reproduced to four decimals (CRPS 2.5678 and all three line Briers). The
fifth did not, and the reason is a documentation defect, not a data one:
- **T-1's 2.5678 is holdout-scored** — a climatology fit on the training
window and scored on the 558-game holdout, exactly as §6 states.
- **W-1's 0.2488 is the pooled *in-sample* variance** over all 3,875
games: literally 0.5350 × 0.4650, the "home rate 53.5%, n=3,875" in its
own parenthetical. It is not a holdout number.
**The thresholds are unchanged.** Renegotiating a gate while measuring
against it is how gates stop meaning anything, and W-1 failed on seed
stability regardless of basis. What changes is only the description.
One consequence is worth keeping, because it makes W-1 a *better* gate
than its wording suggests: **home-field advantage collapsed in this
holdout**, 53.96% in training to 50.72% over 6/1–7/16. Against that, no
constant forecast can reach 0.2488 — even an oracle constant that knows
the holdout rate scores 0.2499. W-1 therefore demands genuine per-game
discrimination rather than a well-chosen base rate, which is precisely
what the runs engine failed to supply.
## 8. Prior art
We searched for public, independently verifiable records of player-prop
probability models graded with proper scoring rules. Findings, stated exactly:
- **Academic literature:** game-level and event-level evaluation exists (§2;
Bayesian batter-pitcher matchup models; strikeout classification work), but we
found no peer-reviewed evaluation of prop-line probability forecasts with
proper scores and calibration curves.
- **KSplit Analytics** (https://ksplitanalytics.com/) is the nearest prior art:
it publishes MLB pitcher strikeout probability distributions and states that
"accuracy is measured using CRPS and bucket-level calibration across the
entire distribution," with distributions "archived with the full probability
table and actual outcome." As of 2026-07-18, no graded performance numbers,
calibration curves, or timestamped archive were publicly visible without an
account. Those are methodology statements; whether a public record follows is
worth watching.
- **PropsBot** (https://propsbot.ai/best-ai-for-mlb-props/) advertises "Brier
score of 0.1903 vs Vegas 0.1947" and "31.7% ROI on 101,881 graded MLB picks."
We found no published methodology, no stated evaluation population or period,
and no timestamped, independently auditable log behind either number. We note
without further comment that the market-efficiency literature in §4 documents
no persistently profitable strategy of remotely that magnitude.
- The defensible claim, and the one we make: **we found no public,
independently verifiable, pre-registered record of player-prop probabilities
graded with proper scoring rules and calibration analysis.** If one exists or
emerges, we will link it here.
Our differentiation is the record itself: probabilities timestamped before first
pitch, immutable once published, graded nightly against Statcast (voids shown,
never deleted), with proper scores and uncertainty-banded calibration curves.
## 10. Weather in the HR model — what it actually bought (2026-07-19)
Goal HR-2 promised this number whichever way it landed. It landed
split: weather did almost nothing for **discrimination** and did
something real for **calibration**.
Same split as every other backtest (train < 2026-06-01, holdout
6/1–7/18, **10,462 batter-games in 580 games**), one model trained with
the weather features and one without, everything else identical.
| Metric | Without weather | With weather | Delta |
|---|---|---|---|
| Log-loss | 0.4147 | 0.4141 | **−0.0006** |
| Brier | 0.1121 | 0.1120 | **−0.0001** |
| AUC | 0.5883 | 0.5898 | **+0.0015** |
| ECE | 1.08% | 0.61% | **−0.47pt** |
| Rate bias | −0.85pt | −0.31pt | **+0.54pt** |
**Read it straight.** The log-loss and Brier movements are third- and
fourth-decimal — smaller than the run-to-run nondeterminism this page
already discloses (±0.002 log-loss), so they are *not* evidence that
weather made the model better at telling which batters homer. AUC moved
+0.0015 on a 0.588 baseline, which is nothing. If the question is "does
knowing the weather help you rank hitters?", the answer from this
holdout is **no, not measurably**.
What weather did do is fix a **level** problem. The model was
systematically under-predicting home runs in the summer holdout by 0.85
percentage points; with weather that shrinks to 0.31, so roughly
**64% of the seasonal bias is absorbed** — mostly by air density, the
composite of temperature, humidity and elevation that actually governs
how far a ball carries. Calibration error more than halved. For a market
whose whole claim is honest probabilities rather than edge, a level fix
is the more valuable of the two — but it is a narrower claim than
"weather improved the model," and we are not making the broader one.
**Caveats, stated once and plainly.**
- An earlier draft of the internal lift note claimed air density
multiplied the effect "roughly sixfold" and credited weather with the
full calibration gain. The quant gate refuted both: the specification
advantage was seed noise, and ~70% of the ECE gain was reproducible
with a one-parameter intercept shift and no weather at all. The
numbers above are the corrected ones. Weather is retained because the
conditional version generalises where a constant does not, and because
the runs engine needs the layer regardless — not because it was the
only way to move ECE.
- **Vintage skew:** the backtest trains on ERA5 reanalysis (what the
weather *was*); live serving uses the forecast API (what it was
*predicted to be*). Those are different products, and a day-ahead
forecast gets wind direction wrong often enough that the
blowing-out/blowing-in sign flips in roughly one game in five. The
bias fix rides on air density, which forecasts well; the
discrimination sliver rides on wind, which does not. Expect live
discrimination at or below the already-negligible numbers above.
- **Retractable roofs** are recorded as a flag, not a state: whether the
roof was actually open for a given game is not available in free data,
so those games carry real outdoor weather plus a flag saying "this may
be wrong."
- **Park bearings** are ±7.5° for 26 parks and ±22.5° for four; nothing
here assumes finer precision.
- **Reproducibility note:** the production host cannot fetch historical
schedule data from MLB's API — every historical endpoint returns 406
from that datacenter IP while same-day calls succeed continuously. The
historical first-pitch times underpinning this backfill were therefore
fetched from a different machine and seeded in. The data is public
either way and the pipeline is reproducible; that specific host cannot
perform the historical fetch itself.
## 11. The runs engine — methods note (2026-07-20)
One model over a team's full run distribution, paired into a joint, from
which the totals strip derives. Winners derives from the same joint and
**is not published** — it failed W-1 (see §7).
### The model
`HistGradientBoostingClassifier` over run counts 0–14 (class 14 = "14 or
more"), one row per team per game, both sides sharing one model with
`is_home` as a feature. 18 features, all strictly prior: trailing-30 team
offense (runs/game, xwOBA, K%, BB%, HR/PA), opponent run prevention,
opposing-starter season priors (K, HR, xwOBA), opponent bullpen rates and
a 3-day workload proxy, park run prior, and four weather terms (air
density, temperature, wind-out-to-centre, dome). Leakage discipline is
the house rule: every join is `merge_asof(..., allow_exact_matches=False)`
or a `shift(1)` within entity, so a game's features never see itself.
Train/serve skew is prevented structurally rather than by convention —
the training frame and the live serving path call the *same*
`attach_features` function, and `scripts/verify_totals_serving.py`
rebuilds real played games through the serving path and asserts every
feature matches at max|diff| 0.
### The pairing, and why independence
Two marginals are combined into a joint over (home runs, away runs). The
naive choice is the outer product, which assumes independence; that is
the convenient assumption and it was **measured before any pairing code
was written**, not after:
| scale | rho | 95% CI |
|---|---|---|
| raw home vs away runs | −0.0029 | [−0.0355, +0.0297] |
| residual, conditioned on the model | −0.0054 | [−0.0365, +0.0267] |
| **copula (normal scale)** | **−0.0160** | **[−0.0465, +0.0149]** |
Estimated on out-of-fold predictions across all 3,875 games, folds
grouped by `game_pk` so a game's two team-rows never straddle a fold. The
design doc predicted "mildly positive, ~0.05–0.15"; **that prediction was
wrong**, and it is recorded as wrong rather than quietly amended. The
interval does more than fail to reject zero — it *bounds* |rho| under
0.047, so a Gaussian copula could not move the joint materially. A copula
at the CI's own upper bound costs +0.00115 CRPS, ~7.5% of the measured
edge, and the gate clears even at rho = +0.15.
Two traps sit in that measurement and both are documented in the code:
the 558-game holdout alone carries a CI of ±0.08, too wide to separate
"independent" from "mildly coupled" (hence out-of-fold across all games);
and **"did the home team bat last" is a collider** — caused by both
teams' runs — so conditioning on it manufactures correlation, +0.39/+0.45
within groups against ~0 pooled.
### Ties, and the defect that had to be fixed
A baseball game cannot end tied, and none of the 3,875 in the store did.
The independent outer product nonetheless placed **9.84% of its mass on
H == A**. Because a tie forces an *even* total, that mass landed entirely
on even totals and left a visible sawtooth: model P(even) 0.5006 against
0.4178 store-wide, a 2.9σ miss. The deeper problem was that the two
markets were not in fact reading the same joint — winners reallocated the
diagonal via an extras constant while totals silently kept it.
Both now derive from a joint **conditioned on H ≠ A**. Measured on one
store, the fix was worth **−0.0058 CRPS** (2.5525 → 2.5467) — about 38%
of the entire measured edge over climatology — and moved P(even) to
0.4461 against 0.4391 observed and the mean total from 8.876 to 9.061.
(Production figures in §7 come from the droplet's own store and differ in
the fourth decimal; the delta above is quoted from a single machine so
the before/after is like-for-like.) The extras constant is no longer
applied anywhere — conditioning splits that mass by the model's own
asymmetry rather than a league average — and is retained only as a
published diagnostic.
The check that should have caught this did the opposite: the old
coherence suite asserted `E[T] = E[H] + E[A]`, an identity that holds
**only while the impossible mass stays in**, so it would have failed the
correct fix. Coherence checks are now split into construction guards
(algebraic identities, labelled as regression guards and explicitly *not*
evidence) and falsifiable checks — the load-bearing new one being
**parity**, model P(even) against observed, which is what detects tie
mass leaking into a strip.
### Scoring and serving
Totals grade by discrete CRPS, `Σ_k (F(k) − 1{y ≤ k})²` (Czado, Gneiting
& Held 2009), verified against hand-computed cases in
`scripts/verify_crps.py` before any headline number was believed —
including the edge case where an actual above the distribution's support
must still be penalised. P(over L) at 7.5/8.5/9.5 is always a
**summation over the published strip**, never a separate model, which is
what makes the boards structurally unable to contradict each other.
Publishing is lineup-triggered on the existing half-hourly sweep, first
write wins, and a prediction stamped at or after first pitch voids. Worth
being explicit, because the trigger's name overstates what it does:
**this model does not read the lineup card.** Its offense inputs are
team-level, so a posted nine does not change the strip. What a posted
lineup buys is *confirmation of the starting pitcher*, which matters
because opposing-starter priors are 3 of the 18 features. Void taxonomy
is postponement and late-stamp only — there is no scratch concept at game
level — and voids store NULL, never 0.
### Known limits
- **Right-tail censoring.** Class 14 is a "14 or more" lump, so the
totals strip is slightly short in the right tail. Measured cost:
+0.0009 CRPS, i.e. it costs the model rather than flattering it.
- **Weather vintage.** The backtest uses reanalysis (what the weather
was); live serving uses the forecast at publish time. Live lift will be
smaller and **that gap is not yet measured** — closing it needs a
forecast/reanalysis paired sample at first-pitch hour.
- **Level drift.** −0.28 runs/game on the holdout (§7, T-1).
- **Early stopping** splits rows at random, so a game's two team-rows can
straddle train and validation. Left as-is because the bias can only
cost us: an optimistic validation score stops *later*, stopping
currently halts at n_iter=10, and every longer fit is monotonically
worse on the holdout (100 iters 2.5888, 400 iters 2.7646, both failing
T-1).
## 9. Revision policy
- Goals change only by dated, public edit to this page, with the reason stated.
- A failed goal is marked FAILED and stays visible; it is never deleted or
retro-edited. Metric definitions and estimator code (10-bin ECE, log-loss
bases, CRPS form) are frozen per goal; changing a definition closes the old
goal (with its verdict at close) and opens a new one.
- Backtest numbers never satisfy a live goal. Live goals are computed from the
graded public log only.
*Further reading (not load-bearing, no numbers cited from these):* Tango,
Lichtman & Dolphin, "The Book: Playing the Percentages in Baseball" (2007);
Kain & Logan, *Journal of Sports Economics* 15(1), 2014 (paywalled).
Source: BENCHMARKS.md — the original, unrendered, for anyone who wants to diff it.