The Brick Drop · Build Spec · Plain-Language Edition

Grading the grader

Before we let the review-reading AI decide what a set is worth, we make it prove itself to you. Here's the whole plan — in plain English, no code.

Judge-validation tooling · 2026-06-23 · what it is · why now · exactly what I need from you
The 60-second read

Right now an AI reads every LEGO review and scores it — how fun the build is, whether it's worth the money, six other things. Those AI scores make up most of a set's grade. But nobody has ever checked the AI's homework. We don't actually know if it's any good at this.

You wanted to prove the AI before giving it real power over the value score — and you're right. So this tool turns you into the answer key: you watch a stack of real reviews, score them yourself, and we measure how often the AI agreed with you. Trust gets earned, not assumed.

Your whole job: watch ~120 reviews and score them in a clean web page, across 4 easy batches, at your own pace. We hand back a report showing how much trust each category has earned — and the moment "Worth it?" makes the grade, the new value math is unblocked.

120
reviews you'll score, in 4 batches of 30 — the test the AI has to pass
8
things the AI scores per review (build fun, worth-it, looks…) — each judged on its own
30
reviews you both score before a category can be judged at all (no lucky guesses)
1.5
biggest average miss (out of 10) the AI is allowed before it falls short of full trust
70%
how often the AI must land within 2 points of your score for full trust
0
database changes — it's all files, nothing to migrate, re-runnable anytime

01Why bother — the AI's on the honor system

The single most important thing feeding a Drop Score is an AI's read on what real reviewers said. And we've been trusting it on faith.

When the engine grades a set, it sends every review it can find to an AI and asks: on a 0–10, how fun is the build? how good does it look? is it worth the money? Those answers pour into roughly two-thirds of the final grade. It's the heart of the whole thing — the part that makes us more than a piece-counter.

Here's the uncomfortable bit: we've never measured whether the AI is actually good at it. Does it catch a sarcastic "oh great, another grey spaceship" and not mark it as praise? When it says a set is a 7 for build fun, would you have said 7 — or 4? We genuinely don't know. The engine even stamps a quiet "unproven" warning on every grade because of exactly this.

Why this is the gate, right now

The new value calculation (Smart Price v2) wants to give the AI's "worth it?" verdict real muscle — enough to overrule the math and call a $400 set a scam when reviewers do. That's the right design. But handing that much power to an AI we haven't checked would be backwards. Prove it first. You said it yourself: validate before we move further.

02The idea: you become the answer key

There's no shortcut. The only way to know if the AI reads reviews like a real LEGO person would is to compare it to a real LEGO person. That's you.

So we run a head-to-head. We gather a pile of real reviews, you score them yourself — your honest gut read, 0–10 — and the AI scores the same ones. Then we lay the two side by side and ask, category by category: did the AI agree with the human? Where it consistently did, we start trusting it. Where it didn't, it stays on probation.

Think of it like

Hiring someone to score reviews for you. You don't hand them the keys on day one — you sit them down with 120 reviews you've also scored, compare their answers to yours, and only let them work unsupervised on the kinds of calls they kept getting right. The ones they fumbled, you keep checking. That's this, exactly.

The beautiful part: once it's built, it's reusable forever. Every time the AI changes, or you want to re-check it, you re-run the comparison against the same pile. Prove-the-AI becomes a button, not a project.

03The one rule that keeps it honest

This is the part that's easy to get wrong and ruins everything if you do: you must score each review before you ever see the AI's score.

Picture the scoring page showing you "AI said: 7" and then asking for your number. You'd drift toward 7 without meaning to — nobody's immune. And then the test wouldn't be measuring how good the AI is. It'd be measuring how easily you get nudged. The agreement number would look great and mean nothing.

So the sheet is blind by design

The AI's answers are not in the scoring page at all. They sit in a separate, sealed file you never open. You score from a blank sheet — just you, the review, and your gut. The two only meet after you've locked in your numbers, when the scorer lines them up. You couldn't peek even if you tried, because there's nothing to peek at.

04How it works — three pieces

Three small parts, in order. You only ever touch the middle one.

🎟️
The Builder
picks the reviews,
seals the answer key
→
✍️
The Sheet
you watch & score,
blind, at your pace
→
📊
The Scorer
grades the AI,
hands you the dial
  1. The Builder picks the reviews.It pulls ~120 reviews spread across the whole spectrum — cheap and expensive sets, licensed and not, YouTube and blogs. On purpose, it mixes in sets people loved and sets people called a scam, so the AI has to prove it can read a pan, not just nod along to praise. It runs the AI on all of them, then splits the result into two files: a sealed answer key (the AI's scores — you never see it) and a blank scoring sheet for you.
  2. You score the sheet.You open a web page — one card per review, the video embedded right there to watch (or the article linked to read). You type your honest 0–10 for each thing the review actually talks about, and leave the rest blank. It saves as you go, so you can stop after ten and come back tomorrow. When a batch is done, one "Export my scores" button.
  3. The Scorer grades the AI.It lines your scores up against the sealed answer key and, for each of the 8 categories, shows how close the AI was and how much trust it earned on the dial (next section). For categories that reach full trust, it prints one line for you to paste that flips them on — and the "unproven" warning lifts off every grade that leans on them.
label-sheet · batch 2 · review 7 of 30
75192 · Millennium Falcon (UCS)
2017 · Star Wars · YouTube review
▶ watch the review right here
Build fun8
Worth the money?6
How it looks9
Playability— skip
Good for parts— skip
37 / 120
saved ✓

A mock, not the real thing — but that's the shape: watch, score what's discussed, skip the rest, progress saved automatically. Notice there's nowhere the AI's guess could hide.

05How trust is earned — a dial, not a switch

Trust shouldn't snap from "no" to "yes." A category that's almost there is not the same as one that's way off — so instead of a hard pass/fail, the AI earns trust along a dial, with an honest gray middle.

The top of the dial — full trust — still means clearing three bars at once. These are the engine's existing bars, not new ones I invented:

The barWhat it means in plain termsThe number
Enough evidenceYou and the AI both scored the same review enough times that it isn't down to luck.30+ reviews
Close enoughAcross those reviews, the AI's score sits near yours on average.miss ≤ 1.5 of 10
Consistently closeIt doesn't swing wildly — most of the time it's right next to your score.70%+ within 2 pts

But here's the gray — and it matters, because it's how you've always wanted this tool to think. Between "unproven" and "fully trusted" there's a real, labeled band: Earning it. A category with plenty of reviews that lands close but not quite over the line isn't branded a failure and thrown out. Its read still counts — at a lighter touch — wearing an honest "still earning trust" label instead of a stark warning, and we keep watching it. Four honest states, not two:

Trusted — full weight, warning off Earning it — close; counts lighter, kept honest Struggling — plenty of data, AI disagreed Needs more — too few reviews yet

So the report isn't a stamp — it's a dial per category, showing exactly how far the AI got:

AI trust dial — example output
Worth the money?
41 revs · miss 1.1 · 80% close
✓ Trusted
Build fun
52 revs · miss 0.9 · 85% close
✓ Trusted
How it looks
47 revs · miss 1.4 · 73% close
✓ Trusted
Instructions
33 revs · miss 1.7 · 66% close
◐ Earning it
Playability
31 revs · miss 2.3 · 52% close
△ Struggling
Good for parts
only 12 revs — too few yet
○ Needs more

Read the dial left-to-right as "how much trust earned." Past the TRUSTED line = full trust, warning off. Just short of it, with the data to back it = Earning it (the gray). Far short = the AI genuinely struggles there. Greyed-out = not enough reviews to judge yet — different problem, fixed by gathering more.

Why the gray actually pays off

An "Earning it" worth-it score is exactly the kind of thing Smart Price v2 can let nudge the value a little — without handing it the full vote. That's the same throttle idea from before, now with a real home: the gray isn't just nicer to look at, it's a genuine third setting between off and on. It's "live in the gray," made concrete.

06Your part, step by step

Everything else is automated. Your contribution is the one thing no script can fake: your honest eye on a real review.

  1. Open the sheet.A single web page — double-click it, or I'll give you a local link. No install, no login, works offline.
  2. Watch, then score — blind.For each card: watch the embedded video or read the linked article, then type your gut 0–10 for each thing it discusses. Didn't mention playability? Leave it blank — a blank is useful, it tells us the review just didn't cover that.
  3. Work in 4 batches of 30.Each batch is a clean finish line — knock out 30, export, done for the day if you like. It saves after every card, so you can stop and resume mid-batch too. ~120 total across the four.
  4. Hit "Export my scores," send it back.That's the finish line for you. I run the scorer and bring you the trust dial.
The one category that matters most

"Worth the money?" is the priority. It's the exact verdict the new value math wants to lean on, so the moment it reaches the trusted end of the dial, the big build is unblocked — even if a couple of rarely-discussed categories need a second round. Score every category as you go (you're already watching), but that's the one we're racing to get there.

07The honest caveats

Two things I want you to know going in, so nothing's a surprise.

Some categories may not reach the dial this round

A few things barely come up in reviews — "playability" on a display set, "good for parts" on a licensed playset. Those might not hit 30 scored reviews in one pass, so they'll land in Needs more for now. That's fine and expected. The scorer names exactly which ones fell short, and we top them up with a focused batch later — we never guess or fudge a result.

It's all just files — nothing to migrate

The reviews, the sealed answer key, and your scores all live as plain files in the project. No database changes, nothing to approve, nothing that can break the live engine. And because it's files, re-checking the AI later is just re-running the comparison — cheap and repeatable.

This is your "fail closed" instinct, applied to the AI

You've always said: don't let thin or unproven data fake a confident number. This is that, pointed at the AI. We keep its opinions live but labeled until it earns the label off — same move you made with the "unproven" badge and living grades. Nothing gets switched off; trust just gets earned in daylight, with room for the gray in between.

08What happens after — & the calls for you

When the dial comes back, the path is short.

You glance at the categories that reached Trusted, and paste one line the scorer prints — that's the official "we trust the AI here now." The "unproven" warning lifts off grades that only use trusted categories. And then, with the AI finally proven, we unpause Smart Price v2 and build the new value math on a foundation we can actually stand behind. The whole reason we stopped.

Two calls before I write the build plan

1 · The "loved & loathed" seed list. To make sure you're not just scoring a pile of "it's fine" reviews, I'll deliberately include some sets known to sit at the value extremes — the Helicarrier and other "scam"-reviewed / GWP-criticized sets on the low end, and clear part-out bargains like the N-1 Starfighter and Pikachu on the high end. Any sets you'd insist are in or out?

2 · How deep should the gray go? The dial above always shows in the report. The real question: should "Earning it" also become a true setting inside the engine — a category in that band counts at a lighter touch with a softer "still earning trust" label (and feeds the Smart Price throttle) — or should the engine stay strict on/off for now and the gray live only in this report? My rec: make it real in the engine — it's your whole philosophy, and it's the natural home for the throttle.

Settled already: 4 batches of 30. ✓


Bottom line: we're not changing the grade or the engine's spine — we're making the AI earn your trust before it gets a vote, with honest room for "almost." You watch ~120 reviews across 4 batches, score them blind, and we prove it in daylight. The day "Worth it?" reaches Trusted, the new value math comes off the bench.

— Build Spec · Judge Validation · The Brick Drop · 2026-06-23