against a 26% background rate — 2.1× better than chance
1,061 calls across 4 scored roster updates, measured against 4,889 scored cards. Radar flags candidates before an update is announced; it does not guarantee upgrades.
Signal 70 and above: 76% upgraded — 3.0× better than chance, on n=165.
What actually happened, by signal strength
Every scored card sorted into 10 signal bands, and the share of each band that San Diego Studio actually upgraded. The dashed line is the background rate — the share of all scored cards that upgraded. Bars above it did better than chance.
Share of the band that upgradedBackground rate 26%
Faded bars have fewer than 50 cards behind them — signal 80–89 (n=39), signal 90–99 (n=28). That is not enough to separate a real edge from luck, so we do not lead with them. signal 90–99 posts the highest rate on the chart, on one of the smallest samples.
Show the numbers
| Signal band | Upgraded | Cards | vs chance |
|---|---|---|---|
| 0–9 | 2% | 1,040 | 0.1× |
| 10–19 | 9% | 917 | 0.3× |
| 20–29 | 23% | 908 | 0.9× |
| 30–39 | 32% | 765 | 1.3× |
| 40–49 | 45% | 573 | 1.8× |
| 50–59 | 58% | 350 | 2.2× |
| 60–69 | 68% | 171 | 2.7× |
| 70–79 | 73% | 98 | 2.9× |
| 80–89 | 79% | 39(thin) | 3.1× |
| 90–99 | 82% | 28(thin) | 3.2× |
Accuracy detail
Mixed methods46%
Upgrade recall
Of real upgrades, how many DR caught
4
Windows scored
May 8 – Aug 14
1,061
Total predictions
584 correct
60%
Avg per-window
Across 4 scored windows
Methodology improvement
Precision
65%
Hand-crafted attribute-gap heuristic + frequency-learned weights + hand-tuned boosts.
Precision
41%▼ −23.4
Served valueGap baseline re-anchored (diamondops-7f2g2): the downside is measured from the first half of the CURRENT approved-major window instead of a never-refreshed window pinned to the card's earliest-ever signal crossing. Restores a real downside to 65% of degraded rows (incomplete estimates 39% -> ~16%) and abstains where no trustworthy baseline exists.
Each version ran on different roster windows — an observed trend, not a controlled comparison.
Diamond Radar flags players who may be upgraded at the next roster update — it doesn't guarantee upgrades. Recall measures how many actual upgrades DR identified; precision measures how often flagged players were actually upgraded.
Accuracy by tier
recall not reported — only 0 real upgrades at this tier
18 of 59 calls correct · recall 56% of the 32 real upgrades at this tier
27 of 75 calls correct · recall 34% of the 80 real upgrades at this tier
89 of 167 calls correct · recall 41% of the 215 real upgrades at this tier
282 of 457 calls correct · recall 45% of the 626 real upgrades at this tier
168 of 301 calls correct · recall 55% of the 307 real upgrades at this tier
Precision — how often a flagged card at this tier was actually upgraded. Recall — how many real upgrades at this tier DR caught.
Tier-jump recall
54% — caught 193 of 357 real tier jumps
Of cards that crossed a rarity boundary, how many DR flagged for an upgrade.
Precision by confidence
How often a flagged card was actually upgraded, split by the confidence attached to each call. Bands are independent — a higher band does not always score better.
Miss profile
The two ways a prediction can be wrong: calling an upgrade that did not happen, and missing one that did.
477
False alarms — we called it, it did not happen
93 of those were strong-signal calls — the model called them loudly and was wrong. Signal strength is a separate measure from the confidence bands above — the two are not the same thing.
676
Missed movers — it happened, we did not call it
- Near miss — we came close to calling it
- 264
- Weak signal — we saw something, but did not call it
- 311
- Blind — we did not see it coming
- 101
Rating accuracy
When a card was upgraded, how close the predicted overall was. Separate from whether the upgrade was called at all.
21%
Exact OVR
50%
Within 1
1.93
Avg error (OVR)
Across 584 scored upgrades. We predicted +1.2 on average against an actual +3.0 — the model under-calls how far a card moves.
By roster update window
15 windows with no major SDS changes hidden.
What a “call” is
A call is a prediction Diamond Radar makes before San Diego Studio publishes a roster update — not an explanation written afterwards. Radar flags a card as an upgrade candidate while the update is still unannounced, and that flag is timestamped and frozen. When the update lands, the card either went up or it didn’t.
That ordering is the whole point, and it is what makes this page checkable. Anyone can explain a rating change after the fact. The calls above were on record first, and every scored update below shows them next to what San Diego Studio actually did.
How Radar decides a card is a candidate
The core idea is a gap. Every rated attribute on a card maps to something the real player actually does on a baseball field — contact and power map to how he hits, a pitcher’s ratings map to what he gets out of hitters. Radar reads the player’s real MLB production and compares it against what his card currently claims. A player producing far above his card’s ratings has an upward gap, and a persistent gap is what a roster update tends to correct.
A single hot week is noise, so the gap is measured over several time horizons at once — the last few days, the current update window, the past month, the season — and blended, weighted toward recent form. Horizons with no data are dropped and the rest re-weighted between them, so a player who missed a month is judged on what exists rather than penalised for the gap in the record.
Each horizon carries a confidence from sample size, and the bar adapts to the horizon rather than being fixed: a typical hitter accumulates a few dozen plate appearances in a week and several hundred across a season, and starters and relievers are held to their own workloads. A part-time player’s hot stretch counts for less than an everyday player’s, because it should.
The threshold is also tier-aware. High-overall cards have less headroom — a 95 has fewer attributes that can plausibly move — so their natural gaps are smaller, and a signal that means nothing on a Bronze card can be significant on a Diamond. A single fixed cutoff would flag low-rated players constantly and elite ones almost never. Radar predicts a direction and a size, and it can flag cards trending the other way too.
How to read this page
The headline is a pair, and it is meant to be read as one: the share of the cards Radar flagged that were upgraded, next to the share of all scored cards that were upgraded. The second number is what makes the first one mean anything. A hit rate on its own tells you nothing about whether the signal helped, because it does not say what would have happened without it.
The curve sorts every scored card by how strongly Radar flagged it, and shows what actually happened to each group. It is a count of outcomes, not a claim about them.
Nothing here is held back for a paid tier. A number can still be missing — either there is not yet enough behind it to report honestly, or it was not recorded for that stretch — and the page says which, in its place, rather than filling the gap with a figure we would not stand behind.
Roster updates are a human decision at San Diego Studio, and real-life production is one input into that decision, not the whole of it. A player can rake for a month and get nothing; another can be adjusted for reasons no model can see from a box score.
There is also a ceiling, and it is worth stating plainly: the strongest signal Radar produces does not reach certainty. No honest reading of this page should suggest a card is guaranteed, and anyone promising that is guessing.
Why most roster updates aren’t on this page
A roster update is San Diego Studio re-rating players in MLB The Show based on real-world performance. They arrive regularly, but most carry no meaningful attribute changes — nothing to predict and nothing to score, so no scorecard exists for them and none is invented here. Of the 19 update windows on record, 4 produced changes worth scoring and 15 carried none.
That is also why this record grows in months rather than weeks, and why a scored update is worth reading in full: each one is a batch of predictions resolved at a single moment, against a set of ratings that had not moved in weeks.
Every scored roster update
One page per update — the cards Diamond Radar called before San Diego Studio published it, and what they actually got.