Our grades lean on an AI that reads reviews and scores how a set builds, plays, and feels worth it. This is the first time we've checked its homework against real humans — for free. Good news up top: it's not making things up.
The AI judge's scores move in the same direction as real human ratings — strongest on the two things that matter most (is it worth it, is the build fun). It also grades a touch harsher than Brickset's super-fans. It's promising, not yet proven, and the way to prove it is something we were already going to do.
Behind every Drop Score is a small AI "judge." You feed it a LEGO review — a YouTube transcript, a blog post — and it reads the whole thing and answers questions like: how fun is the build? how good does it look? is it worth the money? It turns opinions into numbers the grade can use.
That's powerful, but it raises a fair question: can we trust what the AI decides a reviewer meant? Until today the answer was "we hope so." Every grade built on mined sentiment quietly wears an "unvalidated" badge for exactly that reason — we'd never checked the judge against a known-correct answer.
Think of the AI as a new employee who reads reviews and grades sets. They seem sharp — but you've never audited their grades against anyone else's. Today we finally pulled a stack of their grades and compared them to grades we know came from real people.
Here's what makes this cost nothing. Brickset — the big LEGO catalog site — lets its members rate sets by category: Building Experience, Playability, Value for Money. Those are human scores, and three of them line up almost exactly with dimensions our AI judges.
| Brickset humans rate… | …which matches our AI's | In plain words |
|---|---|---|
| Value for Money | feelsWorthIt | Is it worth the price? |
| Building Experience | buildFun | Is the build enjoyable? |
| Playability | playability | Is it fun to play with / do things move? |
So the check is simple and free: for every set we've mined that also has Brickset member ratings, put the AI's number next to the humans' number and see how close they are. No new AI calls — the AI already did this scoring weeks ago and we cached it. No BrickLink (that's the blocked key — different service). Just a read-only comparison.
You asked whether validating would spend money. It doesn't. We're grading answers the AI already gave against a human answer key that's free to fetch. The only "cost" of going deeper is human labeling effort — never dollars.
Across the 32 sets where both the AI and enough Brickset members had weighed in, here's how the judge did. Two numbers matter most: correlation (do they move together?) and bias (does the AI run high or low?).
| Dimension | Sets | Correlation | Avg gap | AI runs… | Verdict |
|---|---|---|---|---|---|
| feelsWorthIt worth the money | 31 | 0.59 | 1.7 pts | −1.4 lower | Best |
| buildFun enjoyable build | 32 | 0.49 | 1.5 pts | −1.1 lower | Solid |
| playability play value | 28 | 0.31 | 1.9 pts | −1.6 lower | Weakest |
Imagine plotting each set with the human score on one axis and the AI score on the other. 0 means total scatter — a coin flip. 1 means a perfect straight line — they agree every time. We landed around 0.5 on the things that matter. That's a genuine, moderate relationship: when humans liked a set more, the AI usually did too.
feelsWorthIt — is the best-validated of the three. That's the reassuring part: the thousands of grades we just refreshed are directionally honest.Only 32 of 634 mined sets had enough human ratings to compare. That's not a bug — it's a timing mismatch, and it points straight at what to do next.
So the two datasets barely overlap today. The fix isn't a clever tweak — it's mining older sets, which is the exact back-catalog work we already scoped (the cheap, Supadata-free blog run for 2017–2024). Those older sets are the ones flush with Brickset ratings, so mining them hands us the validation sample for free.
Validating the engine and filling in the 2017–2024 back catalog turn out to be the same job. Do the older-catalog mining, and this scorecard grows from 32 sets to hundreds — enough to move from "promising" to "proven" — at no extra cost beyond that mining itself.
feelsWorthIt and buildFun. It's a free one-line change (no AI), and it's your call since it changes what shows on live grades.The sentiment engine passes its first real audit: it tracks human judgment, best on the read that matters most, and it cost nothing to confirm. It's a bit stern — largely because it isn't a fan club. It's not "proven" yet, but the road to proven is the back-catalog mining you already have on deck.