Sentiment Engine · Validation · First Read · Cost = $0

Does the Judge Tell the Truth?

Our grades lean on an AI that reads reviews and scores how a set builds, plays, and feels worth it. This is the first time we've checked its homework against real humans — for free. Good news up top: it's not making things up.

2026-07-14 · read-only check · no LLM calls · ground truth = Brickset human ratings · scripts/validate-judge.ts
The finding, in one line

The AI judge's scores move in the same direction as real human ratings — strongest on the two things that matter most (is it worth it, is the build fun). It also grades a touch harsher than Brickset's super-fans. It's promising, not yet proven, and the way to prove it is something we were already going to do.

0.59
how tightly the AI's "worth it?" score tracks human ratings (0 = random, 1 = perfect)
$0
this check cost nothing — no AI calls, it scores answers we already had
32
of 634 mined sets had enough human reviews to compare — the catch, explained below

01What we're actually checking

Behind every Drop Score is a small AI "judge." You feed it a LEGO review — a YouTube transcript, a blog post — and it reads the whole thing and answers questions like: how fun is the build? how good does it look? is it worth the money? It turns opinions into numbers the grade can use.

That's powerful, but it raises a fair question: can we trust what the AI decides a reviewer meant? Until today the answer was "we hope so." Every grade built on mined sentiment quietly wears an "unvalidated" badge for exactly that reason — we'd never checked the judge against a known-correct answer.

The plain-English version

Think of the AI as a new employee who reads reviews and grades sets. They seem sharp — but you've never audited their grades against anyone else's. Today we finally pulled a stack of their grades and compared them to grades we know came from real people.

02The trick: a free human answer key

Here's what makes this cost nothing. Brickset — the big LEGO catalog site — lets its members rate sets by category: Building Experience, Playability, Value for Money. Those are human scores, and three of them line up almost exactly with dimensions our AI judges.

Brickset humans rate……which matches our AI'sIn plain words
Value for MoneyfeelsWorthItIs it worth the price?
Building ExperiencebuildFunIs the build enjoyable?
PlayabilityplayabilityIs it fun to play with / do things move?

So the check is simple and free: for every set we've mined that also has Brickset member ratings, put the AI's number next to the humans' number and see how close they are. No new AI calls — the AI already did this scoring weeks ago and we cached it. No BrickLink (that's the blocked key — different service). Just a read-only comparison.

Why this matters for cost

You asked whether validating would spend money. It doesn't. We're grading answers the AI already gave against a human answer key that's free to fetch. The only "cost" of going deeper is human labeling effort — never dollars.

03The scorecard

Across the 32 sets where both the AI and enough Brickset members had weighed in, here's how the judge did. Two numbers matter most: correlation (do they move together?) and bias (does the AI run high or low?).

DimensionSetsCorrelationAvg gapAI runs…Verdict
feelsWorthIt
worth the money
310.591.7 pts−1.4 lowerBest
buildFun
enjoyable build
320.491.5 pts−1.1 lowerSolid
playability
play value
280.311.9 pts−1.6 lowerWeakest
How to read "correlation"

Imagine plotting each set with the human score on one axis and the AI score on the other. 0 means total scatter — a coin flip. 1 means a perfect straight line — they agree every time. We landed around 0.5 on the things that matter. That's a genuine, moderate relationship: when humans liked a set more, the AI usually did too.

04What it means, in plain terms

Is the AI just making up numbers?
NoIt clearly tracks reality. A 0.49–0.59 correlation on build-fun and worth-it means the judge is reading the room, not rolling dice. The dimension that drives your headline "Will you love it?" read — feelsWorthIt — is the best-validated of the three. That's the reassuring part: the thousands of grades we just refreshed are directionally honest.

Why does the AI grade lower than the humans?
Mostly expected The AI sits about 1.3 points below Brickset across the board. But consider who writes Brickset reviews: the people who liked a set enough to buy it and log in to rate it. That crowd grades generously. Our AI reads the whole conversation — including the critics on YouTube. So a chunk of that gap isn't the AI being wrong; it's the AI being less rose-tinted than a room full of super-fans.

Can we stamp these dimensions "validated" now?
Not yet Thirty-ish sets is enough to say "this isn't noise," not enough to make a promise on. Stamping "validated" is a trust claim, and I won't overclaim it on a small, noisy sample. There's a middle setting built for exactly this moment — see the recommendation below.

05The catch — our reviews are too new

Only 32 of 634 mined sets had enough human ratings to compare. That's not a bug — it's a timing mismatch, and it points straight at what to do next.

What we've mined
Almost all 2025 & 2026 sets — that's the batch we just distilled. Brand-new sets. The reviews exist on YouTube the week they launch.
Where humans rate
Brickset ratings pile up over years as owners slowly log in. A 2018 set has hundreds of ratings; a 2026 set has almost none yet.

So the two datasets barely overlap today. The fix isn't a clever tweak — it's mining older sets, which is the exact back-catalog work we already scoped (the cheap, Supadata-free blog run for 2017–2024). Those older sets are the ones flush with Brickset ratings, so mining them hands us the validation sample for free.

The convergence

Validating the engine and filling in the 2017–2024 back catalog turn out to be the same job. Do the older-catalog mining, and this scorecard grows from 32 sets to hundreds — enough to move from "promising" to "proven" — at no extra cost beyond that mining itself.

06What I'd do next

  1. Hold off on "validated."A 0.5 correlation on 30 sets is real but not a guarantee. Stamping it now would be overselling.
  2. Optionally soften two dimensions to "still earning trust."There's a gray middle badge built for exactly this — honest about the directional agreement without overclaiming. It'd fit feelsWorthIt and buildFun. It's a free one-line change (no AI), and it's your call since it changes what shows on live grades.
  3. Let the back-catalog mining do the heavy lifting.The 2017–2024 blog run (whenever you green-light the spend) multiplies the validation sample as a side effect. Then we re-run this exact check and see if we can honestly stamp "validated."
  4. Keep the harness — it's free and re-runnable.scripts/validate-judge.ts is committed. Any time coverage grows, one command regenerates this scorecard.
Bottom line

The sentiment engine passes its first real audit: it tracks human judgment, best on the read that matters most, and it cost nothing to confirm. It's a bit stern — largely because it isn't a fan club. It's not "proven" yet, but the road to proven is the back-catalog mining you already have on deck.

The Drop Score · Sentiment Validation · 2026-07-14