How ValuCast Works
← Back to boardAs of June 2026 · ValuCast H+P v1
Our display rule: Missing data stays missing; estimates are labeled. Resolve a metric in the glossary.
Scorecard generated JULY 30, 2026
| Component | Provenance |
|---|---|
| Hitting model | Built & validated by ValuCast |
| Hitting inputs | Baseball Savant xBA/xSLG |
| Pitching model | Built & validated by ValuCast — no third-party projections |
| Pitching inputs | Public MLB statistics |
| Steamer board | External comparison board; matching historical benchmark pending |
What ValuCast is
ValuCast is an independent baseball model, not a public-ranking average. For prospects and ValuCast H+P, outside rankings, market values, and consensus lists are comparison context only; they do not generate the score.
ValuCast turns player projections into values tuned to your league's categories, weights, and roster rules. You can value players from two sources: Steamer (an external comparison board, and the default) or ValuCast H+P, our own in-house projection.
The two boards
The default board values current-season actuals + Steamer rest-of-season. The ValuCast board values our own full-season projection. Flipping the toggle is useful for eyeballing differences, but it is not an apples-to-apples backtest; our real check is the held-out validation below.
Promotion posture: ValuCast H+P is available as an opt-in projection source so users can compare it directly. It will not become the default board until its track record and live behavior clear the same honesty gates shown here.
What the Dynasty Value number means
Dynasty Value is a 0-100 score, not a real-world unit. 100 anchors to the top of the board — the single most valuable dynasty asset — and every other player is placed relative to that top. A 95 is a near-elite asset; a 72 is a solid regular; the gap between two numbers is the only thing that carries meaning, not the digits themselves.
- MLB vs. prospects. MLB players and prospects are each run through their own 0-100 normalization and then aligned at the top of the board. We do not claim a single unit-reconciled calibration across the two universes — treat a prospect's number and a big-leaguer's number as comparable in ballpark, not to the decimal.
- The $ column. The dollar figure is a replacement-adjusted auction price: total payout equals teams × budget, spread across the pool by value, and driven by your roster / teams / budget settings. Change those knobs and the dollars re-scale; the 0-100 score does not.
Prospect Rank v1
The Prospects board is ValuCast's own prospect ordering. Top prospects are generated from a universal baseball profile first, then translated into a public rank and value. The score is built from factual current performance, age/level context, draft/signing investment, historical outcome patterns, and availability/sample risk.
How top prospects are generated
- Start with the current eligible prospect universe and MLBAM identity.
- Build each player's factual baseball profile: role, age, level, current MiLB stat line, sample size, draft/signing facts, ETA, and availability context.
- Score the profile through ValuCast's prospect model and universal outcome index.
- Apply sample, availability, and bucket calibration rules by group, not by name.
- Rank the final scores into one universal prospect board.
Category Fit is a separate league-settings view. It can help a user understand roster fit, but it does not generate the public prospect rank.
What can and cannot affect a prospect score
- Can affect score: Prospect Model v0.6, universal outcome context, MiLB performance rates and sample reliability, draft/signing facts, age/level context, and factual availability status.
- Cannot affect score: outside dynasty rankings, outside values, value history, public prospect rankings, and market signals. Those are comparison context only.
- Bucket calibration: applied by rule, not by name. Current rules cover lower-minors pedigree compression, thin upper-level pitcher samples, and upper-level hitters with full samples but limited game impact.
Shape comps (measured, not vibes)
Hitter prospect cards show Closest MLB Shapes: the real MLB seasons (2000–2025, age 26 or younger, 400+ PA) nearest the prospect's MLB-translated K% / BB% / ISO. Rates are compared era-relative — z-scored within each season's population of regulars — so a strikeout rate means what it meant in that run environment. Alongside the named matches we show how the nearest matches with a complete five-season follow-up window actually aged (match seasons 2000–2020 — the ones with a complete five-season follow-up by 2025), deduped by player so no career votes twice. Outcome tiers are measured, with published cut points: 450+ PA/yr with a 110+ era-relative OPS ("regular with an above-average bat"), 400+ PA/yr ("everyday regular"), 150+ ("limited playing time"), under 150 ("faded out") — playing time is the yardstick, so an injury year or a light-workload position (catchers) counts against a player, and 2020 is pro-rated to 162 games.
The same translated axes also produce three transparent component matches: Power (ISO), Contact (K%), and Approach (BB%). Each card names the closest historical season and its measured distance on that single era-normalized axis. We do not turn distance into a match percentage.
Pitcher shape comps use translated K-BB%, K/9, and BB/9. Historical starter and reliever seasons are normalized and matched in separate role pools; mixed historical seasons and prospects with ambiguous current usage are suppressed. Pitcher cards show names, rates, and measured distance only — no outcome odds or future-role claim.
The bias we can't remove, named: comps can only be made to players who actually earned a qualifying young-MLB season. Players who reached that bar and then declined are counted (the "faded out" tier), but a prospect with this exact shape who never got a real MLB look is absent from the pool by construction. That structural survivorship is exactly why a shape comp is a descriptive lens, never an outcome forecast. Comps never feed a ValuCast score, know nothing about defense or speed, and only render for prospects whose translation is high-confidence with a real sample. Rebuilt nightly by the public data pipeline.
How the model works
For each player we weight recent seasons and regress toward the league average (more for small samples). Hitters are then age-adjusted and de-noised toward Statcast expected stats (Savant xBA/xSLG) before projecting. Pitchers are projected per batter faced with a continuous starting-vs-relieving usage blend (no hard starter/reliever cliff). Age adjustment applies to hitters only — the v1 pitching model has no aging curve, a planned future improvement.
Under the hood
A Marcel-style method: recent seasons weighted 5.0,4.0,3.0 (most recent first), regressed to the league mean, with per-component rate (events ÷ opportunities) reconstructed into counts. Hitters are age-adjusted and add Statcast input de-noising (blend actual contact/power toward Savant xBA/xSLG, redistributing into 1B/2B/3B/HR by the player's own extra-base mix). Pitchers use a continuous starter-probability blend so swingmen and converted arms aren't miscategorized; the v1 pitching model has no age curve (a planned improvement) and consumes no third-party projection.
Model equations
Season weighting is done per component as weighted
events ÷ weighted opportunities (not a weighted rate). For a component with
event count E and opportunities PA, weights 5.0,4.0,3.0:
rate = (5·E₋₁ + 4·E₋₂ + 3·E₋₃) / (5·PA₋₁ + 4·PA₋₂ + 3·PA₋₃).
Regression to the mean over N opportunities with a regression
constant n_reg (hitters 1200, pitchers 300):
projected_rate = (rate·N + league_rate·n_reg) / (N + n_reg), then the
projected count is rebuilt as projected_rate × projected_PA.
Age adjustment (hitters only) nudges hitter projections along an empirical aging curve. The v1 pitching model applies no aging curve.
Hitter de-noising (knobs α): blend the actual rate toward the
Savant expected rate, rate* = (1−α)·rate + α·x (x = xBA/xSLG), then
redistribute the extra hits/bases preserving the player's hit-type mix.
Pitcher role blend: with psp the starter
probability, each component is shifted continuously by
f[c]^(h_sp − p_sp) — no hard SP/RP split.
Worked example (HR rate). An age-29 hitter (near peak, age factor
≈ 1.0) with no Statcast movement: 30 HR / 600 PA,
26 / 580, 20 / 520
over the last three seasons (weights 5/4/3). Weighted
events ÷ weighted opportunities =
(5·30 + 4·26 + 3·20) / (5·600 + 4·580 + 3·520) = 314 / 6880 ≈ 0.046
HR per PA. Regressing toward a 0.033 league HR rate with
n_reg = 1200 gives
(314 + 0.033·1200) / (6880 + 1200) ≈ 0.0438.
Projected PA = 0.5·600 + 0.1·580 + 200 = 558,
so projected HR ≈ 0.0438 × 558 ≈
24.4 HR.
Validation details
We hold out future seasons (2020–2025), project forward from prior seasons only (no peeking), and score against simple baselines. Reported as a mean-absolute-error (MAE) ratio — below 1.0 means lower error than the baseline.
- Hitting vs classic Marcel, n = 1770 qualified hitter-seasons, correlation-win rate 0.519 — essentially even on ranking (a near coin-flip); the edge is in error magnitude on rate stats, not in re-ordering players. qualified hitters (>= 200 actual PA, projectable, has prior season).
- Pitching vs persistence (last-year carry-forward), n = 2234 qualified pitcher-seasons, skill correlation-win rate 0.694 — a modest ranking edge; most of the improvement is smaller error on the rate stats below, not re-ordering pitchers. qualified pitchers (role-specific IP floor, projectable, has prior season).
Per-stat MAE ratios:
Hitting (vs classic): AVG 0.942 · OBP 0.959 · SLG 0.96 · OPS 0.951 — counting stats ≈ 1.0 (de-noising doesn't touch them).
Pitching (vs persistence): ERA 0.632 · WHIP 0.751 · K_9 0.834 · BB_9 0.715 · IP 0.981 · K 0.932.
W, SV and QS are reported separately and treated as lower-confidence — they depend heavily on team decisions and opportunity, not pitcher skill.
MLB Projection Track Record — held-out scorecard
On seasons the model never saw, ValuCast made smaller errors than the standard baselines in the areas shown below. The gains are mostly in rate stats; playing-time and opportunity stats remain harder and are labeled that way.
| Step | Held-out result |
|---|---|
| Pitching vs persistence (skill) | 0.807 MAE ratio (~19.3% lower error, n = 2234 pitcher-seasons) — IP and K roughly neutral; the signal is concentrated in ERA / WHIP / K-9 / BB-9. The pitching gain concentrates in ERA and BB-9 — the noisiest, most luck-driven stats — and the hold-out includes the 60-game 2020 season, so read the point estimate as directional, not precise. |
| Hitting vs classic Marcel | 0.979 aggregate MAE ratio (~2.1% lower error, n = 1770 hitter-seasons) — signal concentrated in AVG/OBP/SLG/OPS. Aggregate over 1770 seasons; the win comes from four correlated rate stats, so it is a narrow edge, not an across-the-board one. |
| reliability-weighted regression | tie — not shipped |
| in-house expected-stat model (own xBA) | shortfall — not shipped |
Ahead of the Curve — track-record rules (pre-registered 2026-07-02)
Every prospect we flag as ranked meaningfully ahead of the public field is tracked to one of exactly seven outcomes — including the ones where we backed off our own call, which count against us. No call ever leaves the ledger silently.
| Rule | Definition |
|---|---|
| Success rate | wins (field moved to us, or fully caught up) over every decided call — field-moved-away and our own retreats sit in the denominator; undecided calls are excluded, never hidden. |
| Noise floor | a board move only counts past 10 spots or 15% of the call's initial gap — small jitter hits flagged and unflagged players at identical rates, so it counts for nobody. |
| Matched controls | every call is compared against never-flagged players at the same rank range, role, and date — the published lift is how much more often the field comes toward our calls than toward theirs. |
| Maturity | headline cohort = calls at least 14 days old (one full deep-board refresh cycle); the aggregate publishes only past a 30-day horizon with 25+ decided calls. |
| Retreat attribution | when a divergence closes, whichever side moved more gets the credit — ties count against ValuCast. |
| Targets | pre-registered: 50% decided-rate and 1.5x control lift. The model never optimizes toward this scorecard — it stays blind to consensus by construction. |
These definitions are frozen ahead of the publish gate. The full ledger is human-readable at /ledger — misses included — and the raw artifact is public at /aotc-scorecard.json for anyone who wants to snapshot it themselves.
Adversarial audit log — dated
Scoring notes — 2026-07-12, the night before first publish
Before the gate matured we ran an independent adversarial audit of the scorecard math (a different AI model family from the ones that built and reviewed it, then every finding re-verified against the raw archive). Everything it found is disclosed here, dated, before the first number published — including what cuts against us.
- Compliance fix, disclosed: the headline decided-rate and
control lift had been computed on the all-calls cohort, contradicting
the registered maturity rule above ("headline cohort = calls at least 14 days
old"). Pinned to the matured cohort before first publish. On that day's data
the decided rate moved up (24.6% → 28.7%) and the lift
moved down (1.36x → 1.31x); both still miss our
pre-registered targets. Artifact version 0.2.0 → 0.2.1; decided calls
younger than 14 days now publish alongside as
immature_decided, never hidden. - Known contaminant, published as registered: "we backed off" is attributed mechanically — our rank moved more than the field's. Model retraining nights move our whole board at once, and 50 of the current 65 retreats happened without the field moving away (in 20 of them the field actually moved toward us). The registered rule counts every one against us, so that is how we publish it. A re-baseline-aware retreat rule will be pre-registered as a dated v0.3 and applied forward only — rewriting the rule after seeing the score is exactly what pre-registration exists to prevent.
- Statuses recompute; nothing is hand-frozen: every build re-derives every call from the full archive, so a "caught up" call whose gap re-opens later re-opens with it and the headline moves both ways. Ledger copy that previously described those outcomes as "final" has been corrected.
- Conservative choices we are keeping: controls are drawn from players never flagged as of the latest build (an as-of-call-date pool would read ~1.40x on all calls instead of 1.36x), and controls are matched on rank band, role, and date but not development level (a same-level-only sensitivity reads ~1.65x matured instead of 1.31x — above target). Both alternatives would raise our self-grade, so neither ships as a post-hoc change; they live here as published sensitivities.
Stabilized read — added 2026-07-13 (the day after first publish)
The first day the headline went live it read one thing; the next build it read another, off a single day of new board data. That is the metric working as designed — it re-scores every open call against the newest field snapshot each night — but it means a single build is a noisy reading, not a track record. Trumpeting one good day would be cherry-picking; trumpeting one bad day would be its own kind of dishonesty. So from 2026-07-13 forward we also publish a trailing 14-build average of the exact same daily numbers, replayed from the committed archive so anyone can reproduce it, alongside the min and max so the day-to-day noise is visible rather than hidden. The pre-registered daily number is unchanged — this is a summary laid on top of it, not a rewrite of it. We report the average, we show the range, and we let the number be as modest and as noisy as it honestly is.
Sensitivity Analysis — how much do our assumptions matter?
The full study — every dial, measured
Our rankings depend on a bunch of judgment-call dials — how much to trust a hot three-week stretch with no track record behind it, how much a bad current season should outweigh a great last year, how much a prospect's odds of reaching the majors count versus his fantasy ceiling once he's there. None of these has a single right setting. So instead of handing you one ranking and asking you to trust it, we grabbed each dial, nudged it to a reasonable alternative, re-ran all 2,796 ranked prospects, and counted how many actually changed spots. It's us showing our work — which assumptions move the board a lot, and which barely matter. This is a real re-scoring of the live model, not a simulation; every figure below regenerates nightly from the public data pipeline.
| Assumption changed | Measured effect |
|---|---|
| Outcome-vs-impact weighting how much a prospect's odds of reaching the majors matter vs. his fantasy value once there (0.58 / 0.42 → 0.50 / 0.50) |
164 move 25+ spots, 1,295 move 10+, average 12.8 spots; the top 100 never moves 25+. |
| Thin-sample confidence penalty how far an unproven hot streak should sort behind a proven line (28 → 20 (looser) or 35 (tighter)) |
20 (looser): 1,450 move 25+ spots, 2,250 move 10+, average 53.6 spots; the top 100 never moves 25+. 35 (tighter): 871 move 25+ spots, 1,680 move 10+, average 31.4 spots; the top 100 never moves 25+. |
| Sample regression strength how many PA/IP before a hot line is trusted (200 → 150 (trust thin samples more) or 300 (regress harder)) |
150 (trust thin samples more): 575 move 25+ spots, 1,000 move 10+, average 16.0 spots; the top 100 never moves 25+. 300 (regress harder): 744 move 25+ spots, 1,147 move 10+, average 22.8 spots; 3 of them from the top 100. |
| Stale-line pull weight how hard a bad current-season line overrides an inflated prior year (0.40 / 0.60 → 0.50 / 0.75) |
12 move 25+ spots, 17 move 10+, average 1.5 spots; the top 100 never moves 25+, all moving down. |
The pattern that holds across every dial: the top 100 barely moves. Your blue-chip prospects are locked in no matter how we set things — almost all the churn is deep in the board, where prospects are nearly tied and any nudge just reshuffles a crowd.
Two things this analysis does not tell you: whether any of these alternate settings would have predicted real outcomes better — that's a separate, harder question we validate with held-out backtests before changing anything live — and it isn't recomputed on every page load, since a full board rebuild takes real compute. This is a dated study (run 2026-07-07 against that day's committed board) that we re-run and republish as the model changes.
What we have and haven't proven
Our validation is against internal baselines (persistence and classic Marcel), which it beats on held-out data. ValuCast has not yet proven it beats Steamer or ZiPS: we lack matching archived preseason projections for a fair, apples-to-apples historical backtest. So Steamer is an external comparison board, with a matching historical benchmark pending — not a benchmark we've beaten.
Forward rate check — advisory only. The evaluator now rebuilds actual rates from post-freeze counting-stat deltas rather than comparing frozen projections with cumulative season rates. As of 2026-07-30, Hitter rate error: 1.0416x Steamer. Pitcher rate error: 1.0662x Steamer. The hitter and pitcher checks are separate: both roles must clear independently. Publication remains held.
Plate discipline: measured vs. estimated
Prospect cards show plate-discipline rates counted directly from MLB's own public play-by-play — the same feed the league uses to score games. There are two kinds of number here, and we keep them visibly separate.
- Measured (exact): Swing%, Whiff%, SwStr%. These come straight from each pitch's outcome description (swing, foul, swinging strike, ball in play), so they are deterministic and reproducible. We validated them against a public reference to the decimal — for one Double-A hitter over 41 games our Swing% and SwStr% matched the outside source exactly. No estimation is involved.
- Estimated: Chase%, Z-Swing%, Z-Contact%, Zone%. These depend on where each pitch crossed the zone. Modern tracked coordinates are common at Triple-A and rare below it — but coverage varies game by game, not level by level, so we label each player's level sample individually. Pitches without tracked coordinates get a pixel-to-feet calibration (fit on Triple-A games that carry both) to place them in or out of the zone. A sample's zone metrics are labeled est. whenever any of its zone calls used that calibration — which is why some Triple-A samples carry the tag and some lower-level samples measured on real coordinates do not. The flag follows the data, not the level.
Why there is no exit velocity or contact-quality data. Public minor-league play-by-play carries no batted-ball tracking, so ValuCast does not show — and will not estimate — exit velocity, launch angle, or hard-hit rate at these levels. We would rather leave a box empty than fabricate a number.
Cohorts and sample. Percentiles compare a player only to others at the same level, and a bucket must clear a minimum pitches-seen sample before any bar renders — below that floor we show no percentile at all, the same small-sample discipline the skill bars use. A player who changed levels shows a separate row set per level; we never blend a percentile across two levels. These are observe-only card context: they never feed a player's value, rank, or projection.
Discipline Leaders ranks only hitters who clear that minimum pitches-seen floor, within their current level. Only Chase%, Whiff%, SwStr%, and Z-Contact% receive leader framing; Swing%, Zone%, and Z-Swing% remain descriptive context. Estimated labels follow the underlying sample exactly as they do on player cards, and the cohort can change as players are promoted.
AAA Statcast: measured pitch shape and contact quality
At Triple-A, MLB tracks pitches and batted balls with the same Statcast system it uses in the majors, and Baseball Savant publishes the pitch-level data. For Triple-A players we show a second, separate source built straight from those tracked measurements: measured AAA Statcast (Savant), AAA only, display context — not a model input.
- Pitch shape (pitchers): velocity, induced vertical break, horizontal break, spin, extension, and usage per pitch type. These come directly from the tracked release and movement measurements — not estimates.
- Contact quality (hitters): average and max exit velocity, hard-hit rate, and launch angle. These are tracked batted-ball measurements, shown for Triple-A only.
Because these are measured values, they carry no “est.” tag. This is a parallel source to the plate-discipline section above, not a replacement: plate discipline estimates zone metrics below Triple-A from a pixel calibration, whereas this AAA source measures the zone from real tracked pitch locations. We do not restate either source in terms of the other.
Why exit velocity appears here but not on the plate-discipline section. Public minor-league play-by-play carries no batted-ball tracking, so the play-by-play plate-discipline layer will never show or estimate exit velocity at any level — that limit stands. This AAA source is different data: Statcast's own Triple-A batted-ball tracking, which does measure exit velocity. It is shown for Triple-A only, and never for Double-A or High-A. Like everything on this page, it is observe-only: it never feeds a player's value, rank, or projection.
Model verdicts: two layers of accountability
We keep two accountability layers, not one. The call-level layer is The Ledger and Call-Up Receipts — every individual call we make, tracked to an outcome. The model-level layer is the Model Verdicts page: every ValuCast model and subsystem carries an honest four-way verdict — VALIDATED, PROVISIONAL, DEPRECATED, or REJECTED — and every verdict links the committed evidence behind it. The published failures (a forward rate evaluator held for repair, a Wins target the labels can't train) are first-class rows there, not hidden. We publish what we haven't proven rather than wait for it to look better.
Methodology changelog
Append-only changes to how ValuCast measures, validates, or explains the product.
- 2026-07-30 — Prospect confidence re-baseline. Corrected how thin minor-league samples are discounted, reversing an unintended broad re-score while retaining protection against abrupt transition cliffs. Ranking changes on this date reflect the re-baseline, not new player performance.
- 2026-07-10 — Plate-discipline layer. Added measured Swing%, Whiff%, and SwStr%, with estimated zone metrics labeled separately.
- 2026-07-09 — Methodology honesty polish. Added the live forward-gate disclosure, including when the ValuCast projection is losing.
- 2026-07-02 — Ahead-of-the-Curve scorecard v2. Pre-registered controlled, exit-accounted targets before scoring public divergence calls.
- 2026-07-01 — Saves and holds split. Separated saves and holds targets before combining them for SV+HLD formats.
The Archives: committed boards, replayed
The Archives replays ValuCast's prospect board exactly as it stood on any archived date. These are the exact daily board snapshots the build commits every morning — the page is a replay of the committed artifact as published that day (“as-shown”), never a current-model re-scoring of the past (“as-reconstructed”). The as-of date on the page is the snapshot date, and the archive runs from ValuCast's first committed day to today: we do not reconstruct dates before the archive begins, and a date with no readable snapshot says so plainly rather than showing a neighboring date in its place.
Comparability. Dates before the JULY 30, 2026 scoring re-baseline are flagged pre-baseline: they were scored on an older scale, comparable among themselves but not directly to current numbers. The boundary follows the model's re-baseline epoch, so it moves automatically whenever the scale is re-baselined.
Consensus. The consensus column on a historical board is the aggregate median rank + board count across outside boards (minimum 2 boards) — never individual board ranks, the same rule the live board follows.
