Before we let the review-reading AI decide what a set is worth, we make it prove itself to you. Here's the whole plan — in plain English, no code.
Right now an AI reads every LEGO review and scores it — how fun the build is, whether it's worth the money, six other things. Those AI scores make up most of a set's grade. But nobody has ever checked the AI's homework. We don't actually know if it's any good at this.
You wanted to prove the AI before giving it real power over the value score — and you're right. So this tool turns you into the answer key: you watch a stack of real reviews, score them yourself, and we measure how often the AI agreed with you. Trust gets earned, not assumed.
Your whole job: watch ~120 reviews and score them in a clean web page, across 4 easy batches, at your own pace. We hand back a report showing how much trust each category has earned — and the moment "Worth it?" makes the grade, the new value math is unblocked.
The single most important thing feeding a Drop Score is an AI's read on what real reviewers said. And we've been trusting it on faith.
When the engine grades a set, it sends every review it can find to an AI and asks: on a 0–10, how fun is the build? how good does it look? is it worth the money? Those answers pour into roughly two-thirds of the final grade. It's the heart of the whole thing — the part that makes us more than a piece-counter.
Here's the uncomfortable bit: we've never measured whether the AI is actually good at it. Does it catch a sarcastic "oh great, another grey spaceship" and not mark it as praise? When it says a set is a 7 for build fun, would you have said 7 — or 4? We genuinely don't know. The engine even stamps a quiet "unproven" warning on every grade because of exactly this.
The new value calculation (Smart Price v2) wants to give the AI's "worth it?" verdict real muscle — enough to overrule the math and call a $400 set a scam when reviewers do. That's the right design. But handing that much power to an AI we haven't checked would be backwards. Prove it first. You said it yourself: validate before we move further.
There's no shortcut. The only way to know if the AI reads reviews like a real LEGO person would is to compare it to a real LEGO person. That's you.
So we run a head-to-head. We gather a pile of real reviews, you score them yourself — your honest gut read, 0–10 — and the AI scores the same ones. Then we lay the two side by side and ask, category by category: did the AI agree with the human? Where it consistently did, we start trusting it. Where it didn't, it stays on probation.
Hiring someone to score reviews for you. You don't hand them the keys on day one — you sit them down with 120 reviews you've also scored, compare their answers to yours, and only let them work unsupervised on the kinds of calls they kept getting right. The ones they fumbled, you keep checking. That's this, exactly.
The beautiful part: once it's built, it's reusable forever. Every time the AI changes, or you want to re-check it, you re-run the comparison against the same pile. Prove-the-AI becomes a button, not a project.
This is the part that's easy to get wrong and ruins everything if you do: you must score each review before you ever see the AI's score.
Picture the scoring page showing you "AI said: 7" and then asking for your number. You'd drift toward 7 without meaning to — nobody's immune. And then the test wouldn't be measuring how good the AI is. It'd be measuring how easily you get nudged. The agreement number would look great and mean nothing.
The AI's answers are not in the scoring page at all. They sit in a separate, sealed file you never open. You score from a blank sheet — just you, the review, and your gut. The two only meet after you've locked in your numbers, when the scorer lines them up. You couldn't peek even if you tried, because there's nothing to peek at.
Three small parts, in order. You only ever touch the middle one.
A mock, not the real thing — but that's the shape: watch, score what's discussed, skip the rest, progress saved automatically. Notice there's nowhere the AI's guess could hide.
Trust shouldn't snap from "no" to "yes." A category that's almost there is not the same as one that's way off — so instead of a hard pass/fail, the AI earns trust along a dial, with an honest gray middle.
The top of the dial — full trust — still means clearing three bars at once. These are the engine's existing bars, not new ones I invented:
| The bar | What it means in plain terms | The number |
|---|---|---|
| Enough evidence | You and the AI both scored the same review enough times that it isn't down to luck. | 30+ reviews |
| Close enough | Across those reviews, the AI's score sits near yours on average. | miss ≤ 1.5 of 10 |
| Consistently close | It doesn't swing wildly — most of the time it's right next to your score. | 70%+ within 2 pts |
But here's the gray — and it matters, because it's how you've always wanted this tool to think. Between "unproven" and "fully trusted" there's a real, labeled band: Earning it. A category with plenty of reviews that lands close but not quite over the line isn't branded a failure and thrown out. Its read still counts — at a lighter touch — wearing an honest "still earning trust" label instead of a stark warning, and we keep watching it. Four honest states, not two:
So the report isn't a stamp — it's a dial per category, showing exactly how far the AI got:
Read the dial left-to-right as "how much trust earned." Past the TRUSTED line = full trust, warning off. Just short of it, with the data to back it = Earning it (the gray). Far short = the AI genuinely struggles there. Greyed-out = not enough reviews to judge yet — different problem, fixed by gathering more.
An "Earning it" worth-it score is exactly the kind of thing Smart Price v2 can let nudge the value a little — without handing it the full vote. That's the same throttle idea from before, now with a real home: the gray isn't just nicer to look at, it's a genuine third setting between off and on. It's "live in the gray," made concrete.
Everything else is automated. Your contribution is the one thing no script can fake: your honest eye on a real review.
"Worth the money?" is the priority. It's the exact verdict the new value math wants to lean on, so the moment it reaches the trusted end of the dial, the big build is unblocked — even if a couple of rarely-discussed categories need a second round. Score every category as you go (you're already watching), but that's the one we're racing to get there.
Two things I want you to know going in, so nothing's a surprise.
A few things barely come up in reviews — "playability" on a display set, "good for parts" on a licensed playset. Those might not hit 30 scored reviews in one pass, so they'll land in Needs more for now. That's fine and expected. The scorer names exactly which ones fell short, and we top them up with a focused batch later — we never guess or fudge a result.
The reviews, the sealed answer key, and your scores all live as plain files in the project. No database changes, nothing to approve, nothing that can break the live engine. And because it's files, re-checking the AI later is just re-running the comparison — cheap and repeatable.
You've always said: don't let thin or unproven data fake a confident number. This is that, pointed at the AI. We keep its opinions live but labeled until it earns the label off — same move you made with the "unproven" badge and living grades. Nothing gets switched off; trust just gets earned in daylight, with room for the gray in between.
When the dial comes back, the path is short.
You glance at the categories that reached Trusted, and paste one line the scorer prints — that's the official "we trust the AI here now." The "unproven" warning lifts off grades that only use trusted categories. And then, with the AI finally proven, we unpause Smart Price v2 and build the new value math on a foundation we can actually stand behind. The whole reason we stopped.
1 · The "loved & loathed" seed list. To make sure you're not just scoring a pile of "it's fine" reviews, I'll deliberately include some sets known to sit at the value extremes — the Helicarrier and other "scam"-reviewed / GWP-criticized sets on the low end, and clear part-out bargains like the N-1 Starfighter and Pikachu on the high end. Any sets you'd insist are in or out?
2 · How deep should the gray go? The dial above always shows in the report. The real question: should "Earning it" also become a true setting inside the engine — a category in that band counts at a lighter touch with a softer "still earning trust" label (and feeds the Smart Price throttle) — or should the engine stay strict on/off for now and the gray live only in this report? My rec: make it real in the engine — it's your whole philosophy, and it's the natural home for the throttle.
Settled already: 4 batches of 30. ✓
Bottom line: we're not changing the grade or the engine's spine — we're making the AI earn your trust before it gets a vote, with honest room for "almost." You watch ~120 reviews across 4 batches, score them blind, and we prove it in daylight. The day "Worth it?" reaches Trusted, the new value math comes off the bench.