The whole process, in order, in plain language — every step the engine takes to turn a LEGO set into a letter grade — walked end-to-end on a brand-new 2026 set: The Razor Crest (75447).
Everything below is what happens when a person runs a single line:
npm run grade 75447-1 --mine --priceThree switches, three jobs: grade the set, --mine the real reviews, and --price the parts. The engine then runs the fifteen steps in this document, in this order, and prints a grade. Let's follow it.
First it has to know what it's grading. It asks two LEGO databases for the basic facts and merges them: Brickset (price, piece count, minifigure count, theme, year) and Rebrickable (the exact contents). If the two disagree on, say, piece count, Rebrickable's full inventory wins; for the price, Brickset's MSRP is used.
fetchBricksetSet + fetchRebrickableSet → mergeFacts in src/lib/pipeline/gradeSet.ts. Output is a tidy SetFacts object: name, year, pieces, minifig count, theme, MSRP.Not every category applies to every set. A flower bouquet has no minifigures; a bucket of bricks isn't trying to be a display model. So before scoring, the engine classifies the set and switches off the categories that don't apply (they're marked "not applicable" and quietly dropped — no penalty).
The types it detects: figureless (no minifigs → Minifigures off), pure display (Botanicals/Architecture/Art → Playability off), parts-pack (a bulk tub → Finished Model off, and the "padding" penalty is suppressed because identical parts are the point), and figure-anchored polybag (a tiny bag sold for one rare minifig → Smart Price off).
src/lib/scoring/archetypes.ts. The Razor Crest is a normal play/display set with minifigures, so all six categories stay live.The engine pulls the set's complete parts list from Rebrickable — every element, its colour, and how many. From that one list it derives two different things:
fetchRebrickableParts → rebrickableInventoryToSummary + rebrickablePartsToPriceable in src/lib/sources/rebrickable.ts.This is the heart of replacing "price per piece." Instead of counting pieces, the engine adds up what the parts are actually worth on the open market.
For each distinct part+colour it looks up the BrickLink price guide — but it only fetches a price it doesn't already have, because a part's value is the same in every set. Prices are saved in a shared cache, so the second set that uses a 2×4 black brick reuses the first set's lookup. It multiplies each price by the quantity, sums it, and gets a part-out value — what the box's contents are worth in parts.
Then the clever part: it computes a value ratio = part-out value ÷ what you pay (the MSRP), and ranks that ratio against a cohort — other sets of the same theme and size class. Sitting at the cohort's median scores ~6/10; a clear bargain heads toward 9+, a clear rip-off toward 3. A set is only scored if enough of its parts could be priced (≥80% coverage) and a real cohort of comparable sets exists.
priceInventory + loadCohort (src/lib/pipeline/price.ts), sumPartValue (partValue.ts), smartPriceScore/valueRatio/sizeBucket (smartPrice.ts). The shared cache is the part_prices table.The other half of "Money & Worth": content density. Are the non-figure parts printed/decorated (which builders prize) or plain? Is there genuine variety, or is the piece count puffed up with identical 1×1s and oversized baseplates? Printed parts push the score up; filler padding pushes it down (unless it's a parts-pack, where identical parts are the whole point).
whatYouGetScore in src/lib/facts/factScores.ts, fed by the inventory summary from step 3.Now the opinion half. The engine doesn't search the web on demand — that's slow and expensive. Instead a background sweep has already checked a hand-picked list of 48 trusted LEGO reviewers' channels, worked out which set each video is about, and filed them in a "review index." Grading just reads that index.
Matching a video to the right set is the genuinely hard part, because different sets share names (there are six "Razor Crest" sets, three sets named exactly "X-Wing Starfighter", and 32 "Millennium Falcon" sets). So the matcher is careful: it reads the title (not the noisy description) for the set number (rock-solid) or the full name; it ignores a title that carries a different set's number; it won't match a set that didn't exist yet when the video was posted; and each match is tagged high / medium / low confidence. It also pulls in a review aggregator's (Brick Insights) pre-matched links as a second source.
matchVideoToSets (youtube.ts) over the full 24,644-set catalog; discover.ts writes the review_index table.For each review the engine pulls the transcript (for a video) or the article text (for a blog/forum) plus the top comments, and hands it to an AI judge (Claude). The judge does two jobs:
distillReview (the coversSet gate + per-dimension JSON schema, model claude-opus-4-8) in src/lib/pipeline/distill.ts; mineSetReviews works the highest-confidence reviews first. Each resulting score is tagged with its source so it's trusted accordingly (Brickset > community > blog > YouTube).Reading a review with the AI costs money, so the engine saves the resulting scores (never the review text — just the numbers and a link back). Re-grading the same set later reuses them and skips the AI entirely, unless you force a fresh pass.
dimension_signals table; loadCachedSignals / saveCachedSignals hooks on gradeSet. Legal spine: store derived data + a link, never the raw words.Each rubric point (e.g. "build fun") may now have several opinions plus, sometimes, a hard fact. The engine blends them into one 0–10. Every opinion is weighted by three common-sense factors:
Where a fact defines a point (Smart Price, Build Length, minifig count), the fact is authoritative and opinions only nudge within bounds. One special case: the "feels worth it" crowd sentiment is applied as a small capped nudge to Smart Price — it can fine-tune the hard math but never overrule it.
blend + dimension in src/lib/scoring/; weight = log(1+volume) × trust × recency. Gates: a feeling-point needs ≥3 mined opinions before it counts at all.The points roll up into six categories (each its own 0–10), using fixed sub-weights — e.g. Money & Worth is 65% Smart Price + 35% What You Get. Any point with no data is dropped and the surviving points re-share the weight, so a missing point never silently scores zero.
| Category | Weight | What it asks |
|---|---|---|
| 💰 Money & Worth | 22% | Fair deal — judged by real part value? |
| 🔧 The Build | 19% | Fun to put together? |
| 🏆 The Finished Model | 19% | Looks great done? |
| 🧑🚀 Minifigures | 16% | Are the little people good? |
| 🛠️ For the Hobbyist | 12% | Playable + good parts donor? |
| ✨ Rare & New Parts | 12% | Exciting new moulds/colours? |
| 📈 Hold Its Value | side badge | Will it appreciate? (not in the grade) |
SUB_WEIGHTS in constants.ts; rollupCluster in src/lib/scoring/rollup.ts. Build Length's weight is gently price-scaled (a long build matters more when you paid more).The six categories combine into one 0–10 by their weights. Same fairness rule: a category with no data (or too little) drops out and the rest re-share its weight, within tier first (a missing big category's weight goes to the other big categories before the medium ones). One guardrail: if, after all that, a single category would carry more than 40% of the whole grade, the grade is flagged as a "data-light estimate" rather than letting a confident letter rest on too thin a base.
computeDropScore + rollup; the over-concentration cap is RENORM_CAP = 0.40.The 0–10 is rounded to one decimal (half-up) and mapped to a letter: 9.5+ = A+, 9.0 = A, … 6.0 = C, … below 4.0 = F. The bottom is deliberately coarse — below-average is below-average; we don't over-resolve failing sets.
LETTER_BANDS in constants.ts; toLetter in src/lib/scoring/bands.ts.Every grade ships with two separate honesty signals — they measure different things:
The two are independent: a set can have full coverage but low consensus (everything scored, but from noisy, disagreeing reviews), or high consensus on thin coverage. We surface both, and translate them into a plain badge — Locked in · Solid read · Split verdict · Early read · Too quiet — instead of a confusing raw percentage.
src/lib/scoring/score.ts: confidence = volumeFactor × agreement, coverage = scored ÷ applicable.Finally it stores the result in the database: the set's facts, the computed grade and per-category breakdown, the part prices it looked up (for the next set to reuse), and the distilled review scores (for the next re-grade to reuse). Only derived data and links — never anyone's review text.
upsertSet + upsertDropScore (+ part_prices, dimension_signals) in src/lib/sources/supabaseStore.ts.From npm run grade 75447-1 --mine --price: pulled the facts (930 pcs, 5 figs, $149.99); priced all 930 parts ($172.15 part-out); read its 10 indexed reviews, of which the AI judge distilled 6 (the rest didn't usefully cover the set); blended + rolled up to C+ (6.5), full coverage, 39% consensus.
| Category | Score | Points scored | What carried it |
|---|---|---|---|
| 💰 Money & Worth | 4.9 | 1 / 2 | What-You-Get only — Smart Price sat out (no cohort for this size yet) |
| 🔧 The Build | 7.0 | 4 / 4 | fully scored from the mined reviews |
| 🏆 The Finished Model | 7.8 | 2 / 2 | reviewers liked the look + the accuracy |
| 🧑🚀 Minifigures | 8.0 | 1 / 2 | strong on the 5-figure lineup |
| 🛠️ For the Hobbyist | 4.2 | 1 / 2 | thinner — playability/parts signal was light |
| ✨ Rare & New Parts | n/a | 0 / 1 | no novelty data → dropped, no penalty |
Read it honestly: a solid C+ set — good build and looks, strong figures — held back by a thinner value/hobbyist showing and, crucially, a "Split verdict" 39% consensus, because the six real reviewers genuinely disagreed. The engine isn't pretending to be sure. (And note Smart Price will only sharpen this once its size bucket has a priced cohort — a known, labelled gap, not a silent fudge.)
Companion docs: the friendly overview is docs/how-it-works-plain.html; the precise technical version with formulas is docs/how-it-works.html.