The Brick Drop · From the labeling floor

Sharpening the signals

Four metric upgrades from your first six reviews — plus the head-to-head you asked for: where you and the AI agree, and where it needs your eye.

2026-06-26 · 6 reviews labeled · your read vs the AI · what I'd change & why
The 60-second read

Your four calls are right, and they're the best kind of feedback — straight off the labeling floor, not a whiteboard. The plan: ✂ drop Instructions (you scored it 0 times in 6 reviews; the one time the AI scored it, it was really grading the stickers), ✏ Playability → Functions & Functionality (your AFOLs talk about working features, not "play"), ✏ Good-for-parts → Unique Pieces & Parts, and ➕ add Minifigs — three of your six notes were about minifigs.

Your kid-set instinct is even better than a rename — it's a whole lens, and the engine already has the machinery for it (more in §3).

And the head-to-head: on the objective sets the AI is sharp — it pegged the Helicarrier "$400 scam" and basically tied you on the DC-3. Where it drifts is AFOL enthusiasm: on the N-1 you rated the build and accuracy 9–10; it said 5–6, because it weighed the loudest complaint (the grey-vs-chrome gripe) heavier than the craft. That gap is the most useful thing your labels have surfaced.

6
reviews you've labeled — enough to already see the patterns
0/6
times you scored Instructions — it's a phantom metric
3/6
of your notes were about minifigs — they need their own metric
~3
point gap on the N-1 (you 9–10, AI 5–6) — the enthusiasm blind spot
~0.5
avg gap on the DC-3 — the AI nailed an objective set
~30–40
reviews to validate the dims that matter — not all 120

01Your six, head-to-head with the AI

Exactly what you asked to see: the AI's score (and the quote it leaned on) next to yours, per dimension. Now that you've labeled these blind, seeing the AI's read can't bias you — so react away.

🟢 Δ≤1 — agree🟠 Δ1.5–2.5 — drifting🔴 Δ≥3 — real gap⚪ one side left it blank (a coverage call)AI · You

02What the head-to-heads reveal

Four takeaways — and each one points at a fix you already proposed.

A · On objective sets, the AI is genuinely sharp

It pegged the Helicarrier (worth-it AI 2 / you 3 — both calling out the "$400 scam," 500+ likes) and basically tied you across the board on the DC-3 (avg gap ~0.5). When the community verdict is clear, the judge reads it well. That's the encouraging half.

B · It under-rates AFOL enthusiasm on beloved sets

On the N-1 you rated build & accuracy 9–10; the AI said 5–6. Reading its evidence, you can see why: it latched onto the loudest gripe — "the body should be metallic… inexcusable for a UCS set" — and let that drag the craft scores down. You, an AFOL, see past the color miss to the shaping and design. This is the real question your labels raise: should a dimension reflect the loudest complaint in the comments, or an expert's read of the craft? You're the ground truth, so the AI is "wrong" here — and that's a calibration we can teach it.

C · It over-eagerly scores things you (correctly) left blank

On the Helicarrier the AI scored instructions, build fun, accuracy, and parts — you left all four blank, because the review didn't really cover them. Same on the DC-3 (it scored playability + parts from passing "you could MOC this" comments). This is the exact "the judge hallucinated coverage" failure our diagnostic was built to catch — and your blanks are the gold standard for fixing it. Keep leaving boxes blank when a review's silent; it's signal, not laziness.

D · Instructions is a phantom — and the AI proves your point

You scored it 0 of 6. The one time the AI gave it a number (Helicarrier, 3), read the quote: it's about a "disgusting sticker sheet" and "criminal runway stickers." That's not instruction quality — it's stickers. The metric is so empty the AI fills it with the nearest complaint. Your call to cut it is the right one.

03The metric upgrades

Your four changes, my take on each, and how each one actually plugs into the engine.

RemoveInstructions

Agreed — cut it. LEGO instructions are the gold standard of the toy world; reviewers almost never grade them, and when stuff does come up it's sticker pain or "I'd reorder these steps" (your Pikachu note nailed this) — not instruction quality. It's the dimension least likely to ever reach a trustworthy sample, and it tempts the AI into category errors. Where the real signal goes: sticker pain and build annoyances belong in Build fun — a brutal sticker sheet genuinely makes a build less fun, so that's its honest home.

RefocusPlayability → Functions & Functionality

Agreed — and your kid-set idea is the best part. "Playability" is a kid-toy frame; your reviewers are adults asking does the landing gear retract, do the doors open, is the mechanism clever — that's functions & functionality. Your own Creel House note ("a functionality that gives it a cool factor") is exactly the new framing.

Your kid-set instinct = a lens, and we already built lenses

You said: if a set's clearly for kids, "playability" is the better word, and that could power a parent filter. That's not a tweak — it's the lens system the engine already has (Universal / Display / Builder / Investor… and now Parent). The underlying mined signal stays "functions & features." The Parent lens re-labels it "Playability" and weights it for kid-suitability — and because a set's age rating is a known fact (LEGO marks sets 18+ vs 6+), we can auto-pick the framing: an 18+ UCS shows "Functions," a City set shows "Playability." One signal, two honest faces. Genuinely smart — let's build it.

RefocusGood for parts → Unique Pieces & Parts

Agreed. "Good for parts" is vague; AFOLs care about rare, new, and unique pieces — exclusive colors, fresh molds, elements worth buying the set to harvest. The AI's instinct is already close (its N-1 and DC-3 "parts" scores came from "I'd mod/color-swap this" comments) — the sharper name just points it at uniqueness, not generic MOC-ability. Bonus: this pairs cleanly with the engine's existing fact-based rare-parts signal (new-mold/recolor counts), so the mined sentiment and the hard data reinforce each other.

AddMinifigs

Agreed, and overdue. Three of your six notes were about minifigs, and right now they get awkwardly stuffed into "how it looks" / "accuracy." They deserve their own metric. Your three sub-factors are the right spec, and they map to a clean hybrid score (part hard-fact, part review sentiment — same recipe as Smart Price):

Your sub-factorHow we measure it
Enough figs for the size & priceFact minifig count vs pieces & price — already in the data (a $300 set with 5 figs reads thin; the engine knows the count).
Print & piece qualitySentiment mined from reviews — "gorgeous prints" (DC-3) vs "underwhelming, thrown together from spare parts" (Creel House's Vecna). This is the part only humans/reviews can judge.
Exclusive to this set?Fact + Sentiment — the catalog flags set-exclusive figs; reviewers call out "you can only get X here." Exclusivity lifts the score; army-builder reuse lowers it.

So Minifigs becomes a mined dimension (the reviewer's overall figure verdict) anchored by the hard count/exclusivity facts — and it lands in the Minifigures cluster the rubric already has a slot for.

04You don't have to do all 120

You said this would take forever — and you're right that 120 is a slog. Here's the honest math, because it's a lot less than it looks.

Validation needs about 30 head-to-head scores per dimension — not 120. I built a 120-review pile only so the rarely-discussed dimensions could also reach 30. For the ones you actually care about (Worth it?, build, looks, accuracy, and now Minifigs), they show up in almost every review, so ~30–40 labeled reviews gets them over the line. You're 6 in — so think ~25–30 more, not 114.

The lighter path

Label a batch → I run the dial → we only top up what hasn't passed. The moment "Worth it?" clears the bar, Smart Price v2 is unblocked — you don't need the rest validated to move. And the cleanup helps: cutting Instructions removes a "do I even score this?" decision on every card, and blanks are free (leave them when a review's quiet). Quality over quantity — your 6 are already worth more than 60 rushed ones.

05What I'll do — & the one ask

If you're happy with the four calls above, here's the build. It's real engine work, so I'll TDD it and keep your existing labels intact.

  1. Reshape the mined dimensions.Drop Instructions; rename Playability → Functions & Functionality and Good-for-parts → Unique Pieces & Parts; add Minifigs (hybrid fact + sentiment). Update the AI's distiller prompt so it scores the new set — and stops grading stickers as "instructions."
  2. Wire the Parent lens.Re-label Functions → "Playability" and re-weight for kid-suitability when a set's age rating says it's for kids — your parent-filter idea, made real.
  3. Re-read the corpus with the new metrics.Re-distill the 120 reviews against the sharper dimensions (~15 min), and migrate your 6 labels so nothing's lost — your Playability scores become Functions, your parts scores become Unique Parts; you'll just add Minifigs going forward.
  4. Refresh the sheets at thebrickdrop.com/judge.Same cross-device sync, new metric set + a Minifigs box. You pick up right where you left off.
The one ask

Give me the go-ahead (or tweak any of the four). One open question for you: for Minifigs, should the quantity sub-factor lean more on the hard count-vs-price, or on what reviewers say about whether there are enough? My instinct is to anchor on the fact and let the sentiment adjust it — but you know the AFOL ear better than I do.


Bottom line: six reviews in, you've already made the metrics sharper and caught the AI's one real blind spot. Cut the phantom, refocus two, add the one that matters, teach the lens to speak parent — and you're labeling a tool that actually sounds like an AFOL.

— From the labeling floor · The Brick Drop · 2026-06-26