Four metric upgrades from your first six reviews — plus the head-to-head you asked for: where you and the AI agree, and where it needs your eye.
Your four calls are right, and they're the best kind of feedback — straight off the labeling floor, not a whiteboard. The plan: ✂ drop Instructions (you scored it 0 times in 6 reviews; the one time the AI scored it, it was really grading the stickers), ✏ Playability → Functions & Functionality (your AFOLs talk about working features, not "play"), ✏ Good-for-parts → Unique Pieces & Parts, and ➕ add Minifigs — three of your six notes were about minifigs.
Your kid-set instinct is even better than a rename — it's a whole lens, and the engine already has the machinery for it (more in §3).
And the head-to-head: on the objective sets the AI is sharp — it pegged the Helicarrier "$400 scam" and basically tied you on the DC-3. Where it drifts is AFOL enthusiasm: on the N-1 you rated the build and accuracy 9–10; it said 5–6, because it weighed the loudest complaint (the grey-vs-chrome gripe) heavier than the craft. That gap is the most useful thing your labels have surfaced.
Exactly what you asked to see: the AI's score (and the quote it leaned on) next to yours, per dimension. Now that you've labeled these blind, seeing the AI's read can't bias you — so react away.
Four takeaways — and each one points at a fix you already proposed.
It pegged the Helicarrier (worth-it AI 2 / you 3 — both calling out the "$400 scam," 500+ likes) and basically tied you across the board on the DC-3 (avg gap ~0.5). When the community verdict is clear, the judge reads it well. That's the encouraging half.
On the N-1 you rated build & accuracy 9–10; the AI said 5–6. Reading its evidence, you can see why: it latched onto the loudest gripe — "the body should be metallic… inexcusable for a UCS set" — and let that drag the craft scores down. You, an AFOL, see past the color miss to the shaping and design. This is the real question your labels raise: should a dimension reflect the loudest complaint in the comments, or an expert's read of the craft? You're the ground truth, so the AI is "wrong" here — and that's a calibration we can teach it.
On the Helicarrier the AI scored instructions, build fun, accuracy, and parts — you left all four blank, because the review didn't really cover them. Same on the DC-3 (it scored playability + parts from passing "you could MOC this" comments). This is the exact "the judge hallucinated coverage" failure our diagnostic was built to catch — and your blanks are the gold standard for fixing it. Keep leaving boxes blank when a review's silent; it's signal, not laziness.
You scored it 0 of 6. The one time the AI gave it a number (Helicarrier, 3), read the quote: it's about a "disgusting sticker sheet" and "criminal runway stickers." That's not instruction quality — it's stickers. The metric is so empty the AI fills it with the nearest complaint. Your call to cut it is the right one.
Your four changes, my take on each, and how each one actually plugs into the engine.
Agreed — cut it. LEGO instructions are the gold standard of the toy world; reviewers almost never grade them, and when stuff does come up it's sticker pain or "I'd reorder these steps" (your Pikachu note nailed this) — not instruction quality. It's the dimension least likely to ever reach a trustworthy sample, and it tempts the AI into category errors. Where the real signal goes: sticker pain and build annoyances belong in Build fun — a brutal sticker sheet genuinely makes a build less fun, so that's its honest home.
Agreed — and your kid-set idea is the best part. "Playability" is a kid-toy frame; your reviewers are adults asking does the landing gear retract, do the doors open, is the mechanism clever — that's functions & functionality. Your own Creel House note ("a functionality that gives it a cool factor") is exactly the new framing.
You said: if a set's clearly for kids, "playability" is the better word, and that could power a parent filter. That's not a tweak — it's the lens system the engine already has (Universal / Display / Builder / Investor… and now Parent). The underlying mined signal stays "functions & features." The Parent lens re-labels it "Playability" and weights it for kid-suitability — and because a set's age rating is a known fact (LEGO marks sets 18+ vs 6+), we can auto-pick the framing: an 18+ UCS shows "Functions," a City set shows "Playability." One signal, two honest faces. Genuinely smart — let's build it.
Agreed. "Good for parts" is vague; AFOLs care about rare, new, and unique pieces — exclusive colors, fresh molds, elements worth buying the set to harvest. The AI's instinct is already close (its N-1 and DC-3 "parts" scores came from "I'd mod/color-swap this" comments) — the sharper name just points it at uniqueness, not generic MOC-ability. Bonus: this pairs cleanly with the engine's existing fact-based rare-parts signal (new-mold/recolor counts), so the mined sentiment and the hard data reinforce each other.
Agreed, and overdue. Three of your six notes were about minifigs, and right now they get awkwardly stuffed into "how it looks" / "accuracy." They deserve their own metric. Your three sub-factors are the right spec, and they map to a clean hybrid score (part hard-fact, part review sentiment — same recipe as Smart Price):
| Your sub-factor | How we measure it |
|---|---|
| Enough figs for the size & price | Fact minifig count vs pieces & price — already in the data (a $300 set with 5 figs reads thin; the engine knows the count). |
| Print & piece quality | Sentiment mined from reviews — "gorgeous prints" (DC-3) vs "underwhelming, thrown together from spare parts" (Creel House's Vecna). This is the part only humans/reviews can judge. |
| Exclusive to this set? | Fact + Sentiment — the catalog flags set-exclusive figs; reviewers call out "you can only get X here." Exclusivity lifts the score; army-builder reuse lowers it. |
So Minifigs becomes a mined dimension (the reviewer's overall figure verdict) anchored by the hard count/exclusivity facts — and it lands in the Minifigures cluster the rubric already has a slot for.
You said this would take forever — and you're right that 120 is a slog. Here's the honest math, because it's a lot less than it looks.
Validation needs about 30 head-to-head scores per dimension — not 120. I built a 120-review pile only so the rarely-discussed dimensions could also reach 30. For the ones you actually care about (Worth it?, build, looks, accuracy, and now Minifigs), they show up in almost every review, so ~30–40 labeled reviews gets them over the line. You're 6 in — so think ~25–30 more, not 114.
Label a batch → I run the dial → we only top up what hasn't passed. The moment "Worth it?" clears the bar, Smart Price v2 is unblocked — you don't need the rest validated to move. And the cleanup helps: cutting Instructions removes a "do I even score this?" decision on every card, and blanks are free (leave them when a review's quiet). Quality over quantity — your 6 are already worth more than 60 rushed ones.
If you're happy with the four calls above, here's the build. It's real engine work, so I'll TDD it and keep your existing labels intact.
Give me the go-ahead (or tweak any of the four). One open question for you: for Minifigs, should the quantity sub-factor lean more on the hard count-vs-price, or on what reviewers say about whether there are enough? My instinct is to anchor on the fact and let the sentiment adjust it — but you know the AFOL ear better than I do.
Bottom line: six reviews in, you've already made the metrics sharper and caught the AI's one real blind spot. Cut the phantom, refocus two, add the one that matters, teach the lens to speak parent — and you're labeling a tool that actually sounds like an AFOL.