The Brick Drop · The Procedure

How a Drop Score Gets Made

The whole process, in order, in plain language — every step the engine takes to turn a LEGO set into a letter grade — walked end-to-end on a brand-new 2026 set: The Razor Crest (75447).

Worked example generated live by `npm run grade 75447-1 --mine --price` on 2026-06-18. The 15 steps still describe the engine; notes marked "Since this run" flag what changed by 2026-08-15.
The set we're grading 75447 — The Razor Crest (Star Wars, 2026, 930 pieces). The brand-new Mandalorian gunship. We picked it on purpose: it's new, so it tests the live review-mining; it's big, so it tests part pricing; and "Razor Crest" is a name shared by six different sets, so it tests whether the engine grabs the right one.
C+
6.5 / 10
Full coverage · 39% consensus — a "Split verdict." Every applicable point got scored (mining filled them in), but the reviewers genuinely disagreed, so the engine is honest that it's not a slam-dunk read. Everything in this document is the actual procedure that produced this — no hand-waving, no fake data.

0 · The one command

Everything below is what happens when a person runs a single line:

npm run grade 75447-1 --mine --price

Three switches, three jobs: grade the set, --mine the real reviews, and --price the parts. The engine then runs the fifteen steps in this document, in this order, and prints a grade. Let's follow it.

1 · Identify the set

1

First it has to know what it's grading. It asks two LEGO databases for the basic facts and merges them: Brickset (price, piece count, minifigure count, theme, year) and Rebrickable (the exact contents). If the two disagree on, say, piece count, Rebrickable's full inventory wins; for the price, Brickset's MSRP is used.

Under the hood fetchBricksetSet + fetchRebrickableSet → mergeFacts in src/lib/pipeline/gradeSet.ts. Output is a tidy SetFacts object: name, year, pieces, minifig count, theme, MSRP.
Razor Crest Star Wars · 2026 · 930 pieces · 5 minifigures · MSRP $149.99 — all read live from Brickset + Rebrickable.

2 · Figure out what kind of set it is

2

Not every category applies to every set. A flower bouquet has no minifigures; a bucket of bricks isn't trying to be a display model. So before scoring, the engine classifies the set and switches off the categories that don't apply (they're marked "not applicable" and quietly dropped — no penalty).

The types it detects: figureless (no minifigs → Minifigures off), pure display (Botanicals/Architecture/Art → Playability off), parts-pack (a bulk tub → Finished Model off, and the "padding" penalty is suppressed because identical parts are the point), and figure-anchored polybag (a tiny bag sold for one rare minifig → Smart Price off).

Think of it like A teacher throwing out a test question that doesn't apply to you — you're not marked wrong on it, it's just removed, and the rest counts for a little more.
Under the hood src/lib/scoring/archetypes.ts. The Razor Crest is a normal play/display set with minifigures, so all six categories stay live.

3 · Open the box (the inventory)

3

The engine pulls the set's complete parts list from Rebrickable — every element, its colour, and how many. From that one list it derives two different things:

  • A content summary — how many parts are printed/decorated (vs plain), how much variety there is, and whether the count is padded with hundreds of identical filler pieces. (Used in step 5.)
  • A priceable list — each part paired with its BrickLink id and colour id, ready to look up a market price. (Used in step 4.) Conveniently, Rebrickable already carries the BrickLink ids inside each part, so no translation table is needed.
Under the hood fetchRebrickableParts → rebrickableInventoryToSummary + rebrickablePartsToPriceable in src/lib/sources/rebrickable.ts.

4 · Price every part (Smart Price)

4

This is the heart of replacing "price per piece." Instead of counting pieces, the engine adds up what the parts are actually worth on the open market.

For each distinct part+colour it looks up the BrickLink price guide — but it only fetches a price it doesn't already have, because a part's value is the same in every set. Prices are saved in a shared cache, so the second set that uses a 2×4 black brick reuses the first set's lookup. It multiplies each price by the quantity, sums it, and gets a part-out value — what the box's contents are worth in parts.

Then the clever part: it computes a value ratio = part-out value ÷ what you pay (the MSRP), and ranks that ratio against a cohort — other sets of the same theme and size class. Sitting at the cohort's median scores ~6/10; a clear bargain heads toward 9+, a clear rip-off toward 3. A set is only scored if enough of its parts could be priced (≥80% coverage) and a real cohort of comparable sets exists.

Think of it like Appraising a used car by what its actual parts and condition are worth, not its asking price — and then judging that against other cars in its class, because 30 mpg is great for a truck and mediocre for a hatchback.
Under the hood priceInventory + loadCohort (src/lib/pipeline/price.ts), sumPartValue (partValue.ts), smartPriceScore/valueRatio/sizeBucket (smartPrice.ts). The shared cache is the part_prices table.
Razor Crest All 930 parts were priced (100% coverage, every part found on BrickLink) → a part-out value of $172.15 against a $149.99 MSRP, i.e. the parts are worth about 15% more than the box price. But the engine did not turn that into a Smart Price score — because its size class (750–1500 pieces) has no priced cohort yet to rank it against, and the engine refuses to fake a ranking without one. So Money & Worth fell back to "What You Get" alone (scored 4.9, only 1 of its 2 points). That's the system being honest about a real data gap, not hiding it.
Since this run · minifigs + True Value Two things changed here. Minifig value is now priced in — the engine prices each figure's own parts and adds them to the part-out total (this run's $172.15 counted only the loose bricks). And the set page now shows a True Value breakdown: the anchor price, the part-out value, the ratio as a plain line ("parts are worth ~N% of the price"), a Bargain/Fair/Overpriced tag, and where the value sits (plain bricks / printed / functional / figures / brand premium). It never divides by piece count.

5 · What you actually get

5

The other half of "Money & Worth": content density. Are the non-figure parts printed/decorated (which builders prize) or plain? Is there genuine variety, or is the piece count puffed up with identical 1×1s and oversized baseplates? Printed parts push the score up; filler padding pushes it down (unless it's a parts-pack, where identical parts are the whole point).

Under the hood whatYouGetScore in src/lib/facts/factScores.ts, fed by the inventory summary from step 3.

6 · Find the reviews

6

Now the opinion half. The engine doesn't search the web on demand — that's slow and expensive. Instead a background sweep has already checked a hand-picked list of 48 trusted LEGO reviewers' channels, worked out which set each video is about, and filed them in a "review index." Grading just reads that index.

Matching a video to the right set is the genuinely hard part, because different sets share names (there are six "Razor Crest" sets, three sets named exactly "X-Wing Starfighter", and 32 "Millennium Falcon" sets). So the matcher is careful: it reads the title (not the noisy description) for the set number (rock-solid) or the full name; it ignores a title that carries a different set's number; it won't match a set that didn't exist yet when the video was posted; and each match is tagged high / medium / low confidence. It also pulls in a review aggregator's (Brick Insights) pre-matched links as a second source.

Think of it like A librarian who shelves new arrivals as they arrive, so when you ask for a book it's already on the shelf — instead of ransacking the whole building every time.
Under the hood matchVideoToSets (youtube.ts) over the full 24,644-set catalog; discover.ts writes the review_index table.
Razor Crest Already had 10 indexed reviews waiting (4 high-confidence number matches + 6 lower-confidence), found by the sweep — exactly the "reviews are already on the shelf" payoff.

7 · The AI reads each review

7

For each review the engine pulls the transcript (for a video) or the article text (for a blog/forum) plus the top comments, and hands it to an AI judge (Claude). The judge does two jobs:

  1. Confirms it's the right set. A "ranking every UCS set" round-up, or a review of a different Razor Crest, still mentions the words — so the judge first decides "does this video actually review this set?" If not, it's thrown out. If it's a multi-set video that genuinely covers our set, the judge scores only our set's portion.
  2. Scores the experience. For each thing reviewers actually talked about — build fun, instructions, sturdiness, looks, accuracy, playability, parts value, worth-the-money — it gives a 0–10, with a quote for evidence, and it is sarcasm-aware ("oh great, another grey spaceship" is not praise). Anything not discussed is left blank, not guessed.
Think of it like A tireless intern who watches every review, fills out the exact same scorecard each time, never plays favourites, and instantly hears an eye-roll that a keyword-counter would read as a compliment.
Under the hood distillReview (the coversSet gate + per-dimension JSON schema, model claude-opus-4-8) in src/lib/pipeline/distill.ts; mineSetReviews works the highest-confidence reviews first. Each resulting score is tagged with its source so it's trusted accordingly (Brickset > community > blog > YouTube).

8 · Remember the scores

8

Reading a review with the AI costs money, so the engine saves the resulting scores (never the review text — just the numbers and a link back). Re-grading the same set later reuses them and skips the AI entirely, unless you force a fresh pass.

Under the hood dimension_signals table; loadCachedSignals / saveCachedSignals hooks on gradeSet. Legal spine: store derived data + a link, never the raw words.

9 · Blend every opinion per point

9

Each rubric point (e.g. "build fun") may now have several opinions plus, sometimes, a hard fact. The engine blends them into one 0–10. Every opinion is weighted by three common-sense factors:

  • How much it's based on — more reviews count more, but with diminishing returns (going 1→10 matters; 1,000→1,010 barely registers).
  • How trustworthy the source is — a structured Brickset rating outweighs a random comment.
  • How recent it is — but only for money opinions, which go stale when the price changes; "the build is fun" never expires.

Where a fact defines a point (Smart Price, Build Length, minifig count), the fact is authoritative and opinions only nudge within bounds. One special case: the "feels worth it" crowd sentiment is applied as a small capped nudge to Smart Price — it can fine-tune the hard math but never overrule it.

Under the hood blend + dimension in src/lib/scoring/; weight = log(1+volume) × trust × recency. Gates: a feeling-point needs ≥3 mined opinions before it counts at all.

10 · Add the points into categories

10

The points roll up into six categories (each its own 0–10), using fixed sub-weights — e.g. Money & Worth is 65% Smart Price + 35% What You Get. Any point with no data is dropped and the surviving points re-share the weight, so a missing point never silently scores zero.

CategoryWeightWhat it asks
💰 Money & Worth22%Fair deal — judged by real part value?
🔧 The Build19%Fun to put together?
🏆 The Finished Model19%Looks great done?
🧑‍🚀 Minifigures16%Are the little people good?
🛠️ For the Hobbyist12%Playable + good parts donor?
✨ Rare & New Parts12%Exciting new moulds/colours?
📈 Hold Its Valueside badgeWill it appreciate? (not in the grade)
Under the hood SUB_WEIGHTS in constants.ts; rollupCluster in src/lib/scoring/rollup.ts. Build Length's weight is gently price-scaled (a long build matters more when you paid more).

11 · Add the categories into one number

11

The six categories combine into one 0–10 by their weights. Same fairness rule: a category with no data (or too little) drops out and the rest re-share its weight, within tier first (a missing big category's weight goes to the other big categories before the medium ones). One guardrail: if, after all that, a single category would carry more than 40% of the whole grade, the grade is flagged as a "data-light estimate" rather than letting a confident letter rest on too thin a base.

Under the hood computeDropScore + rollup; the over-concentration cap is RENORM_CAP = 0.40.

12 · Turn the number into a letter

12

The 0–10 is rounded to one decimal (half-up) and mapped to a letter: 9.5+ = A+, 9.0 = A, … 6.0 = C, … below 4.0 = F. The bottom is deliberately coarse — below-average is below-average; we don't over-resolve failing sets.

Under the hood LETTER_BANDS in constants.ts; toLetter in src/lib/scoring/bands.ts.
Since this run · never blank + investor floor A grade is now only withheld if zero categories could be scored; any set with at least one scored category always gets a real letter (a thin one just reads "Early read"), and each headline read — Worth it? · Will you love it? · Hold value? — has its own fallback so it never shows a dash. Separately, on the Investor lens a scored grade never drops below C−, so a weak investment doesn't read like a failed set.

13 · Say how sure we are

13

Every grade ships with two separate honesty signals — they measure different things:

  • Coverage (breadth) — how many of the applicable points we could actually score, vs left blank for lack of data.
  • Consensus (depth) — how solid the evidence behind the scored points is: lots of agreeing reviews = high; a split room or one lonely data point = low.

The two are independent: a set can have full coverage but low consensus (everything scored, but from noisy, disagreeing reviews), or high consensus on thin coverage. We surface both, and translate them into a plain badge — Locked in · Solid read · Split verdict · Early read · Too quiet — instead of a confusing raw percentage.

Think of it like A weather forecast: "we have data for the whole week" is coverage; "but the models disagree about Thursday" is consensus. An honest forecaster tells you which.
Under the hood src/lib/scoring/score.ts: confidence = volumeFactor × agreement, coverage = scored ÷ applicable.

14 · Save it

14

Finally it stores the result in the database: the set's facts, the computed grade and per-category breakdown, the part prices it looked up (for the next set to reuse), and the distilled review scores (for the next re-grade to reuse). Only derived data and links — never anyone's review text.

Under the hood upsertSet + upsertDropScore (+ part_prices, dimension_signals) in src/lib/sources/supabaseStore.ts.

15 · The whole thing on one screen

npm run grade 75447-1 --mine --price │ 1 · facts ......... Brickset + Rebrickable → name, pieces, MSRP, theme 2 · archetype ..... which categories apply 3 · inventory ..... parts → content summary + priceable list 4 · Smart Price ... BrickLink part-out value ÷ MSRP, ranked vs cohort 5 · What You Get .. printed / variety / filler 6 · discover ...... read review_index (allowlist sweep + Brick Insights) 7 · AI judge ...... per review: covers this set? → 0–10 per dimension (sarcasm-aware) 8 · cache ......... save the scores (not the text) 9 · blend ......... per point: facts + opinions, weighted by volume × trust × recency 10 · categories ... points → 6 clusters (sub-weights, drop-and-renormalize) 11 · overall ...... clusters → one 0–10 (weights, renormalize, 40% cap) 12 · letter ....... round half-up → A+ … F 13 · how sure ..... coverage (breadth) + consensus (depth) → named badge 14 · save ......... sets, drop_scores, part_prices, dimension_signals
The Razor Crest, end to end — the actual run

From npm run grade 75447-1 --mine --price: pulled the facts (930 pcs, 5 figs, $149.99); priced all 930 parts ($172.15 part-out); read its 10 indexed reviews, of which the AI judge distilled 6 (the rest didn't usefully cover the set); blended + rolled up to C+ (6.5), full coverage, 39% consensus.

CategoryScorePoints scoredWhat carried it
💰 Money & Worth4.91 / 2What-You-Get only — Smart Price sat out (no cohort for this size yet)
🔧 The Build7.04 / 4fully scored from the mined reviews
🏆 The Finished Model7.82 / 2reviewers liked the look + the accuracy
🧑‍🚀 Minifigures8.01 / 2strong on the 5-figure lineup
🛠️ For the Hobbyist4.21 / 2thinner — playability/parts signal was light
✨ Rare & New Partsn/a0 / 1no novelty data → dropped, no penalty

Read it honestly: a solid C+ set — good build and looks, strong figures — held back by a thinner value/hobbyist showing and, crucially, a "Split verdict" 39% consensus, because the six real reviewers genuinely disagreed. The engine isn't pretending to be sure. (And note Smart Price will only sharpen this once its size bucket has a priced cohort — a known, labelled gap, not a silent fudge.)

Companion docs: the friendly overview is docs/how-it-works-plain.html; the precise technical version with formulas is docs/how-it-works.html.