DiamondOps
Back to News
Feature· 6 min read

Est. Profit Was Counting a Card's Failure as a Win. And 40 Cards Had the Wrong Player's Stats.

Two things were wrong on Diamond Radar, and both of them were our fault rather than the market's. We found them, we measured how bad they were, and this release fixes them.

Fair warning up front: every Est. Profit on the board changes today. Most get smaller. Some Strong Buy badges are gone. That is the fix working, not the board getting worse.

Est. Profit was treating a card not upgrading as a good outcome

Est. Profit answers one question: is holding this card worth more than selling it now? To do that it weighs what you gain if the card upgrades against what you lose if it doesn't.

The "what you lose" half was broken.

To work out what a card falls back to if nothing happens, we compared it to what the card cost before people started speculating on it. Reasonable idea. The problem was when we looked that price up: once, the very first time the card ever appeared on the radar, and never again.

A Live Series card keeps the same identity all season while its rating and its price move constantly. So on a card that had changed since, we were measuring it against a price for what was effectively a different card — sometimes one from months earlier. And when that old price happened to be higher than today's price, the maths inverted: the downside stopped being a cost and started being added to the estimate.

The result was a number that quietly rewarded a card for failing.

We measured it across 200 cards on the live board:

  • 164 of 200 estimates were more than half made up of that error
  • 68 of 200 were telling you to buy a card that our own model expected to lose value

That second one is the part that bothers us most. Sixty-eight recommendations pointing the opposite direction to our own prediction, on a Pro feature.

What changed: the downside can no longer exceed what the card is actually worth, so it can never turn into a bonus. And a card we predict will lose value can no longer carry a Strong Buy or Undervalued badge at all — that combination is now impossible by construction rather than merely unlikely.

The fallback price is now measured from the current update window

Capping the bad number stopped it being absurd. It didn't make it right — the price we were comparing against was still stale.

So we changed where it comes from. Speculation on a card resets when a roster update resolves: a card people got excited about two updates ago that didn't upgrade, then got excited about again this week, has had two separate runs at it. The price that matters is where it settled after the last update, not what it cost in April.

Est. Profit now measures the downside from the current update window. On the live board that restored a real downside figure to about two-thirds of the cards that had lost it — incomplete estimates dropped from roughly 39% of the board to around 12%.

When we can't work out the downside, Est. Profit now says so

Some cards still don't have a trustworthy fallback price. Rather than quietly showing you half an answer dressed as a whole one, Est. Profit now tells you which one you're looking at: a complete estimate, or an upside-only figure with the downside missing.

If you see the upside-only marker, read that number as optimistic. It isn't a full picture and we'd rather say so than let you assume it is.

There are also two moments where Est. Profit is now absent instead of guessing:

  • For a couple of days after a roster update. The market is still violently re-pricing and there isn't enough settled history in the new window to say what anything falls back to.
  • Before the season's first major update. There's no resolved window to measure from yet.

A blank there is deliberate. We would rather show you nothing than a number we can't stand behind. Once an update is confirmed on our side, the board now refreshes within minutes rather than waiting for the overnight run.

Forty cards were showing a different real person's statistics

This one is worse in kind, even though it touched fewer cards.

Diamond Radar reads real-world MLB performance, so every Live Series card has to be matched to the actual player. For most cards that's straightforward. For minor-leaguers, prospects and free agents — a big chunk of the low-rated board — there's often no MLB record to match against at all.

When that happened, our matching fell back to something it should never have used: the jersey number and the club. It never checked whether the name agreed.

For a player who has never appeared in an MLB game, any match found that way is, by definition, somebody else. And because established big-leaguers hold the low jersey numbers, the somebody-else was frequently a star.

We audited all 1,653 matched Live cards. Forty were wrong — including a 55-rated reliever displaying Justin Verlander's career line as if it were his.

These are real, named people. A card showing another player's production is a false statement about both of them, and it was sitting on a paid surface.

What changed:

  • Matching now requires the name to agree. Full stop.
  • It tolerates the genuine spelling differences between the game's data and MLB's own — cards like Christian Encarnacion-Strand and Leo Rivas still match correctly.
  • Every run re-checks the cards it matched earlier instead of trusting old work forever. That was the second half of the bug: a wrong match, once made, was never looked at again.
  • Around 31 cards now show no real-world stats at all, because we can't confidently say who they are. That's the right answer. Showing nothing beats showing you somebody else's numbers.

Every affected card in production has already been corrected.

"Sure Things" is now "Best Odds", and it actually has cards in it

The top-confidence board asked for a 90% chance of an upgrade before a card qualified.

We checked whether Diamond Radar can produce a 90% call. It can't — and not because of a bug. Graded against three completed scoring windows and thousands of cards, the model's real hit rate flattens out around the low 80s no matter how strong the signal gets. There is no band where nine out of ten cards upgrade.

So the board was asking for a certainty that doesn't exist in the data. It returned nothing, every time.

The bar is now set where the model actually performs, and the board is renamed Best Odds — because at that hit rate roughly one pick in five won't upgrade, and calling it "Sure Things" was a promise we couldn't keep at any threshold. Getting to 90% needs a genuinely better model, which is a different piece of work, not a slider we can move.

Why these posts keep looking like this

Both of this release's headline fixes came from checking our own numbers and disliking the answer. Neither was reported by a crash or an error log — the board looked perfectly healthy while quietly rewarding failure and attributing an All-Star's season to a Double-A reliever.

We would rather find these ourselves and tell you plainly than have you find them and wonder what else is off. That's the trade: occasionally a release note reads like a confession, and in exchange the numbers on the board mean what they say.