A complete view of the logic behind the holistic LEGO set grade: the rubric, the math, the data sources, and how the AI judges real reviews — with citations for everything.
The LEGO community judges a set's worth with one lazy number: price per piece (the "$0.10/brick = good deal" rule). It's broken because it ignores part size and value (a tiny 1×1 tile counts the same as a big trans-black canopy), minifigures, build enjoyment, display quality, and everything else that actually makes a set good.
The Drop Score replaces that with a grade built from both value and joy, drawn from real human experience at scale, and always honest about how sure it is. The non-negotiable principles:
Every point on the rubric is scored 0–10 by blending up to three streams, weighted by how much evidence exists. A missing stream simply drops out — we never guess to fill it.
| Stream | What it is | Where it comes from |
|---|---|---|
| ① Facts | Objective data, computed | Brickset (price, pieces, figs), Rebrickable (full inventory, novelty), BrickLink (per-part market value) |
| ② Mined opinion | Real published reviews, distilled by AI | YouTube transcripts + comments (via Supadata + YouTube API), Brickset structured sub-ratings, Brick Insights aggregate |
| ③ Community | Visitor grades + required written "why" | The Drop Score site's no-account submission form (Layer 3, not yet built) |
The blend formula (methodology §2.1, implemented in src/lib/scoring/blend.ts):
wᵢ = log(1 + volumeᵢ) × trustᵢ × recencyᵢ recencyᵢ = decays ? exp(−ageMonths / 12) : 1 // value sentiment decays; look/build sentiment doesn't pointScore = Σ(wᵢ · scoreᵢ) / Σ(wᵢ) // 0–10
Trust weights: Brickset 1.0 · Community 0.9 · Reddit 0.8 · Blog 0.8 · YouTube 0.7 · Forum 0.7 · Amazon 0.5. Facts, where they define a point, are authoritative for that point; opinion streams enrich within the bounds each dimension sets.
Six clusters feed the grade; a seventh ("Hold Its Value") is a side badge that never moves the letter (except under the Investor lens). Tags: Fact computed · Feeling mined/community · Mixed both.
| Cluster | Weight | Points (sub-weight) |
|---|---|---|
| 💰 Money & Worth | 22% | Smart Price F (65) · What You Get F (35) · Feels Worth It M (capped nudge, not a slice) |
| 🔧 The Build | 19% | Build Fun M (55) · Build Length F (25, price-scaled) · Holds Together & QC Fl (20) |
| 🏆 The Finished Model | 19% | Display Appeal Fl (65) · Fidelity & Scale M (35) |
| 🧑🚀 Minifigures | 16% | Figure Lineup M (65) · Figure Desirability & Exclusivity M (35) |
| 🛠️ For the Hobbyist | 12% | Playability Fl (55) · Parts for Custom Building M (45) |
| ✨ Rare & New Parts | 12% | Rare & New Parts F (100) |
| 📈 Hold Its Value | side badge — 30% only under Investor | Appreciation Outlook F (50) · Value Floor F (50) — see §5b |
Anti-double-counting is enforced: minifig signals live only in Minifigures; part novelty only in Rare & New; generic element variety only in Parts for Custom Building; the inventory "filler" signal hurts only value (What You Get); Build Length adds value salience only inside The Build. Source: methodology spec §3.1.
Implemented in src/lib/scoring/score.ts (the computeDropScore() orchestrator).
smartPrice' = clamp(smartPrice + clamp(feelsWorthIt − smartPrice, ±0.75), 0, 10).| Archetype | Trigger | Effect |
|---|---|---|
| Figureless | 0 minifigures | Minifigures cluster → N/A |
| Pure display | display line (Botanicals, Architecture, Art) | Playability → N/A |
| Parts-pack / bulk | Classic tubs, assortments | Finished Model, Fidelity, Build Fun → N/A; filler penalty suppressed |
| Figure-anchored polybag | < 100 pcs AND ≥1 exclusive fig | Smart Price → N/A; Minifigures carries the grade |
| A+ | A | A− | B+ | B | B− | C+ | C | C− | D+ | D | F |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 9.5–10 | 9.0–9.4 | 8.5–8.9 | 8.0–8.4 | 7.5–7.9 | 7.0–7.4 | 6.5–6.9 | 6.0–6.4 | 5.5–5.9 | 5.0–5.4 | 4.0–4.9 | 0–3.9 |
The bottom is deliberately coarse (wide D, no D−): below-average is below-average; we don't over-resolve failing sets.
Universal is the loud default. A viewer can flip a lens that re-weights clusters to their priorities. The Investor lens is the only mode where resale enters the grade.
| Cluster | Universal | Display | Builder | Parent | Investor |
|---|---|---|---|---|---|
| Money & Worth | 22 | 18 | 18 | 20 | 22 |
| The Build | 19 | 16 | 16 | 22 | 8 |
| Finished Model | 19 | 26 | 12 | 14 | 12 |
| Minifigures | 16 | 18 | 10 | 20 | 12 |
| Hobbyist | 12 | 10 | 24 | 18 | 4 |
| Rare & New | 12 | 12 | 20 | 6 | 12 |
| Hold Its Value | — | — | — | — | 30 |
The bridge from raw facts to 0–10 dimension values (src/lib/facts/). All curves are piecewise-linear over tunable anchors.
Compute a value ratio = the set's part value ÷ what people actually pay (street price), with a fallback chain. Then rank that ratio against the set's cohort (same theme × size class) and map the percentile to 0–10.
valueRatio = partValueTotal / streetPrice // primary (Rebrickable inventory × BrickLink per-part value)
→ weightGrams / streetPrice // fallback 1 (weight predicts retail better than piece count)
→ pieces / streetPrice // fallback 2 (last resort)
score = piecewiseLinear(percentile_in_cohort, [[0,1],[0.1,3],[0.5,6],[0.9,9],[1,10]])
// cohort median → 6.0, clear bargain → 9+, clear rip-off → <3
Size buckets: <250 / 250–750 / 750–1500 / 1500+ pieces. Licensing premium is surfaced explicitly, never silently penalized.
| Dimension | How it's scored (anchors / rule) |
|---|---|
| Build Length | piece count → 0–10 [[25,1.5],[250,4.5],[750,6.5],[1500,8.0],[3000,9.0],[6000,9.7]]; its cluster weight scales with price: 25% × clamp(price/100, 0.7, 1.5) |
| What You Get | 0.6·printedShare-curve + 0.4·variety-curve − filler penalty (largest single-element share × 3) |
| Figure Lineup | minifig count → [[0,0],[1,5],[3,7],[6,8.5],[10,9.5],[20,10]] |
| Rare & New | weighted novelty (newMolds + 0.5·recolors + 0.5·exclusives) → [[0,4],[1,6],[3,7.5],[6,9],[12,10]] |
Appreciation Outlook (the scarcity dimension) | see §5b — age-through-life × how collectable the theme is |
Value Floor (the resale dimension) | see §5b — part-out ÷ price |
All anchors above are tunable defaults, centralized in src/config/constants.ts. The structure is the decision; the exact numbers get tuned against real sets.
The "Hold Its Value" cluster used to be a frozen 5.0 for every set — both of its ingredients depended on data we never populated. It was rebuilt from facts we actually have. It's a side badge that only enters the letter under the Investor lens (weight 30).
anchor = streetPrice ?? msrp
valueFloor = piecewiseLinear(partValueTotal / anchor,
[[0.5,2.0],[1.0,5.0],[1.5,7.0],[2.0,8.5],[3.0,10.0]])
// 1.0× "fully backed by parts" reads 5.0; 3.0× is a 10
Returns "insufficient" (an early read, never a fabricated 0) when there's no part-out value or no price. Note this is a floor read — distinct from Smart Price's deal-framing where ~1.75× is the neutral market point.
progress = piecewiseLinear(ageYears, [[0,0.50],[1,0.56],[2,0.68],[3,0.82],[5,0.93],[8,1.0]]) ceiling = THEME_DESIRABILITY[theme] ?? 5.5 // how collectable the theme is score = clamp(5.0 + progress × (ceiling − 5.0)) // base 5.0 = "unproven upside"
A brand-new set starts at 0.5 of the way through its life (not "failing"), and its confidence is lowered (not its score) until it's 2+ years old. Theme desirability is a curated, tunable tier — high (Star Wars, Icons, Ideas, Architecture, Harry Potter ≈ 9; Botanicals, Pokémon ≈ 8.5), mid (Marvel/Technic/Super Mario ≈ 6), low (City 3.5, Duplo/Classic 3, Gear 2.5). Star Wars has a low part-out but high appreciation — part-out alone misses collector demand, which is the whole point of keeping the two ingredients separate.
The grade answers "is this priced fairly?" in one letter; True Value shows the receipts. Rendered by deriveTrueValue() (src/lib/app/gradeToInput.ts) into src/ui/TrueValueBreakdown.tsx, it appears on any set that has a part-out value (~1,000 sets today).
anchor price = streetPrice ?? msrp (tagged "street" or "MSRP")
parts ratio = partOut / anchor → shown as "parts are worth ~N% of the price"
verdict = ratio ≥ 1.0 → Bargain
ratio ≥ 0.7 → Fair
else → Overpriced
bar width = FLOOR(6%) + (100−6)·(amount/max)^0.5 // legibility curve; the $ figure is the truth
It breaks the value into parts (plain bricks / printed & detailed / functional pieces / figures) and a premium group (brand & IP, plus a licensed royalty line). Tags surface Licensed and Retired. It never divides by piece count — price-per-piece is banned everywhere.
partValueTotal and shown as a real "Figures" line. For sets not yet re-priced, True Value shows a clearly-labeled flat estimate (≈ $2.50/fig, $3.50 licensed) tagged "Figures — est.", and the price-read grade stays on the real priced-parts basis so the number and the grade never contradict each other.This is the engine's "feelings" half. For each review, Claude (claude-opus-4-8) reads the transcript + top comments and returns structured per-dimension scores. Implemented in src/lib/pipeline/distill.ts using the Anthropic API with structured outputs (a JSON schema constrains the response).
You distill a single LEGO set review into per-dimension scores for a grading engine.
Score each dimension 0–10 ONLY if the review (or its comments) actually speaks to it;
otherwise set score to null. Dimensions:
- buildFun: how enjoyable/clever the build is (vs repetitive/tedious)
- instructions: instruction clarity, bagging, sticker pain
- integrityQC: sturdiness when handled; arrives complete/undamaged
- displayAppeal: how good the finished model looks / shelf presence
- fidelityScale: accuracy to the real thing / theme standard
- playability: fun to actually play with (features, swooshability)
- partsForBuilding: value as a parts source for custom builds (MOCs)
- feelsWorthIt: whether it's worth the money
CRITICAL: be sarcasm-aware. LEGO reviewers are ironic ("oh great, another grey spaceship",
"a steal at only $0.10 a piece"). Read the real sentiment, not the literal words. Weight the
reviewer's own assessment over comment noise. Set sarcasmDetected true if irony materially
shaped any score. Give a short evidence quote per scored dimension.
LEGO reviewers are relentlessly ironic. A keyword counter sees "oh great, another $0.10-per-piece cash grab" and scores "great" as positive. The model understands it's a dig. Verified live on a deliberately sarcastic X-Wing review:
"Absolutely riveting — if your idea of a good time is attaching 400 identical grey tiles"
→ buildFun 3 (read as ironic / tedious)
"once it's done, this thing looks genuinely stunning on the shelf"
→ displayAppeal 9
"$0.10-a-piece masterpiece ... worth it only if Star Wars is your entire personality"
→ feelsWorthIt 3
sarcasmDetected: true
Each non-null distilled score becomes one signal (volume 1) fed to the same blend as every other opinion. Value sentiment (feelsWorthIt) is marked decaying (recency matters because price changes); look/build sentiment doesn't decay. A dimension needs ≥3 mined opinions (MIN_VOLUME_FOR_SENTIMENT) before it counts toward the grade — below that it's "applicable but insufficient," which lowers coverage rather than faking a score.
Every grade carries two separate signals, measured in src/lib/scoring/dimension.ts + score.ts:
scored ÷ applicable across the in-grade dimensions.confidence) — how solid the evidence behind the scored points is: confidence = volumeFactor × agreement, where volumeFactor = min(1, log(1+totalVolume)/log(61)) and agreement = 1/(1+variance). Facts are authoritative (≥0.8). A low number means one of two things: few people weighed in (thin volume) or they disagreed (low agreement).The raw readout — C− (5.7), full coverage, 34% confidence — makes the two signals fight: "full" sounds great, "34%" sounds broken, and the reader can't tell whether to trust the grade. That 34% isn't the tool doubting its own math — it's six real reviewers who showed up and disagreed. So we name depth for what it is (how much the crowd agrees) and always surface the reason:
We graded the whole set — but six reviewers came in split. Take this as a read, not gospel.
Coverage 9/9 points · Agreement 34% · from 6 mined reviews + 18 Brickset ratings
"Data-light" retires into a small vocabulary of named states that say the why out loud — same two axes, no jargon:
Worked example: rich, agreeing data on only 3 of 9 points = high consensus, low coverage (an "Early read"). All points scored but only from noisy, disagreeing sources = full coverage, low consensus (a "Split verdict"). The live X-Wing mined grade was a Split verdict — C− (5.7), whole set graded, consensus shaky (34%) — honestly low because six real reviewers disagreed and were more critical than Brickset's 5-star ratings.
coverageTier, confidence, per-cluster scoredCount) — only the presentation changes.One command (npm run grade <set> --mine) runs the whole chain (src/lib/pipeline/gradeSet.ts):
Discovery is split from grading so we never pay search.list (100 units) per grade. A separate sweep — scripts/discover.ts — pages each allowlist channel's recent uploads via playlistItems (~1 unit/50 videos), matches each video to a known set by its number or its name (so "UCS X-Wing review" with no number still lands), and writes the origin reviews to review_index. The same index is also fed by a Brick Insights origin harvest — each set's outbound "Read reviews" links (YouTube + blogs/forums), minus retailers/affiliates and the already-ingested Brickset. Grading then reads that index and distills from it directly; search.list only fires as a cold-start fallback for an un-indexed set. Distilled signals are cached in dimension_signals, so a re-grade reuses them and skips mining (and LLM spend) entirely — pass --remine to force a fresh pass. (Live-verified: the Brick Insights harvest added 8 origin links — 7 blog + 1 forum — for the UCS Falcon; the grade distilled from the index with zero search units, and the next grade hit the cache.)
Matching, carefully (this part is easy to get wrong): name-matching is hard because different sets share words — there are many "X-wing" and "Millennium Falcon" sets across the years. So it matches on the title only (descriptions list every set in a roundup); needs every distinctive word of the set name (≥2); ignores a title that carries a different set's number; drops any set released after the video; and accepts only when exactly one known set fits. A naive first cut wrongly swept 10 unrelated videos under the X-Wing — review round-ups plus the New Republic X-Wing Starfighter 75460, a different set — which the guards now reject; a set number in the title always wins.
| Source | Used for | Endpoint / docs |
|---|---|---|
| Rebrickable | inventory, parts, minifigs, part novelty | rebrickable.com/api/v3/docs |
| Brickset | MSRP, pieces, minifig count, structured review sub-ratings | brickset.com/tools/webservices/v3 |
| BrickLink | per-part market value (Smart Price), resale | bricklink.com/v3/api.page (OAuth1) |
| Brick Insights | origin-review harvest (outbound "Read reviews" links → review_index) + bootstrap aggregate score | brickinsights.com (JSON-LD + card--cta links on /sets/{id}; index at /sitemap/sets) |
| Supadata | YouTube transcripts (managed, ToS-offloaded) | supadata.ai → GET /v1/transcript?url=… |
| YouTube Data API v3 | allowlist discovery (playlistItems, ~1u/50) + comments; search.list cold-start fallback only | developers.google.com/youtube/v3 (key from console.cloud.google.com) |
| Claude API | sarcasm-aware review distillation (the LLM judge) | platform.claude.com/docs · model claude-opus-4-8 |
The rubric was not invented — it was synthesized from how real reviewers and the community actually judge sets (a multi-agent research pass at the start of the session):
Full reasoning is in docs/superpowers/specs/2026-06-13-drop-score-methodology-design.md (and the sourcing spec alongside it).
The "Obsidian Dispatch" look + score-card layout were grounded in: Metacritic, OpenCritic, Pitchfork (the score-as-typographic-monument), Rotten Tomatoes + its Pentagram rebrand (one-accent discipline), IGN, Brick Insights, BrickEconomy, and Linear (the buildable dark/single-accent spec). Full brief: docs/design/2026-06-17-ux-inspiration-brief.md.
Two discovery paths exist; the live demo used the on-demand one.
48 hand-vetted channels, every channel ID verified via the YouTube API, in src/config/reviewer-allowlist.ts. A representative slice (full URLs let you find any video they've made):
| Channel | URL | Channel ID |
|---|---|---|
| JANG's LEGO Reviews | @JANGsLEGOreviews | UCPi-XOCT88MwgFlkLHzBNhA |
| just2good | @just2good | UCp_mZttcKNIcUBVdi0wTdIA |
| MandRproductions | @MandRproductions | UCLnr9MzQ_v_a_OPmJkye5lA |
| Solid Brix Studios | @SolidBrixStudios | UC_EahESpmOsO_5hAmsNSqlw |
| BrickVault | @BrickVault | UCrhb3SP2lZBgguLHIWWuHOQ |
| RacingBrick | @RacingBrick | UCfU8ME4_m48QwDJCUvpzqyQ |
| Tiago Catarino | @TiagoCatarino | UCqLbTAc5Mn2cJDu1-FZ2W3g |
| Brick Fanatics | @BrickFanatics | UCGLVjA9MlM0KQHySRVwAeuw |
The remaining 40 (theme specialists, MOC builders, build-alongs, value channels) are in the file with the same fields. Excluded deliberately: "Held der Steine" (German), no-talking ASMR channels (no usable transcript). Research provenance: Feedspot, Brick Insights reviewers, The Brick Land.
For a single set, the miner calls the YouTube search API and keeps results whose title mentions the set number or a name token:
GET https://www.googleapis.com/youtube/v3/search
?part=snippet&type=video&order=relevance
&q=LEGO 75355 X-wing Starfighter review
Reproduce in a browser:
https://www.youtube.com/results?search_query=LEGO+75355+X-wing+Starfighter+review
Confirmed videos used in the X-Wing demo (the live set varies per run; these two distilled successfully with full English transcripts + ~30 comments each):
Now wired: the production path is allowlist playlistItems (1 unit/50 videos) cross-fed with the Brick Insights origin-review map, both landing in review_index. Grading distills straight from that index; on-demand search.list (100 units/call) is only a cold-start fallback for an un-indexed set.
Project "DropScore" (eubnrfwibntsqetdmear, us-east-2). Six tables, all with Row-Level Security. Derived data only.
| Table | Holds | Public access (RLS) |
|---|---|---|
sets | cached objective facts per set | read |
part_prices | global per-part value cache (BrickLink), set-independent | none (internal/service-role) |
dimension_signals | derived per-dimension scores (no raw text) | read |
review_index | linkback map: set → origin review URLs (links only) | read |
drop_scores | cached engine output per set per lens | read |
community_submissions | no-account grades; a written why is required at the DB level | read + insert |
Cache writes happen via the service-role key; the public site reads via the publishable key. Mapper code: src/lib/sources/supabaseStore.ts.
All in one place — src/config/constants.ts. Change behavior here without touching logic.
| Knob | Default | Meaning |
|---|---|---|
| Cluster weights | 22/19/19/16/12/12 | relative importance of the 6 in-grade clusters |
| FEELS_WORTH_IT_CAP | ±0.75 | max nudge community value sentiment applies to Smart Price |
| BUILD_LENGTH_REF / clamp | $100 / [0.7,1.5] | how strongly price scales Build Length's weight |
| Smart Price buckets / anchors | <250/750/1500 ; median→6 | cohort size classes + percentile→score curve |
| Stream trust | Brickset 1.0 … Amazon 0.5 | how much each source counts in the blend |
| HALFLIFE_MONTHS | 12 | recency decay for value sentiment |
| MIN_VOLUME_FOR_SENTIMENT | 3 | mined opinions needed before a feeling-dimension scores |
| MIN_STRUCTURED_COUNT | 8 | Brickset ratings needed before they're a confidence booster |
| MIN_COMMUNITY_TO_MOVE | 5 | community grades needed before they move the number |
| LOW_DATA_FLOOR | 5 | overall "low data" threshold |
| RENORM_CAP | 40% | max share any one cluster can carry after renormalization |
| POLYBAG_PIECES | 100 | figure-anchored polybag trigger |
| Letter bands | see §4 | 0–10 → A+…F cutoffs |
Store only computed signals — scores, counts, dates — plus a source name and a link. Never warehouse verbatim transcripts, reviews, or comments. This single rule covers three exposures at once:
Other guardrails: never republish Brick Insights' aggregate as our headline (we recompute our own grade); community submissions require a substantive written reason (anti-spam + minable signal); a validation gate is specified before distilled scores are trusted at scale.
claude-opus-4-8, low effort, ~2K output cap). The dominant per-review cost; cache results so re-grades never re-distill.search.list = 100 units (expensive); playlistItems/commentThreads = ~1 unit (cheap). Staying cheap means allowlist discovery, not search.part_prices cache (a part's price is set-independent) — price the universe of parts once, sum locally.An honest inventory of the edges, so nothing is a surprise:
The scoring engine, the fact-scoring layer, all source adapters (live-verified), the LLM distiller, the mining orchestrator with allowlist discovery + dimension_signals caching, Supabase persistence, and — since this doc was first written — the whole app is live at thebrickdrop.com: ~5,800 sets graded across five lenses, per-set pages, search, browse-by-release, compare, saved sets, the editorial "Brick Drop Take" (generated in-voice for every set), the downloadable 4:5 share card, and the True Value breakdown. Batch grading has swept the modern catalog.
review_index and reuses cached signals; search.list is a cold-start fallback only.MIN_STRUCTURED_COUNT gate and the emphasis on mined sentiment.| Path | Responsibility |
|---|---|
src/lib/scoring/score.ts | computeDropScore() — the orchestrator (steps in §4) |
src/lib/scoring/{blend,dimension,rollup,bands,archetypes,rubric}.ts | the math primitives |
src/lib/facts/ | raw facts → 0–10 (Smart Price, Build Length, etc.) |
src/lib/sources/ | the 6 adapters + Supabase store + contracts/fakes |
src/lib/pipeline/{gradeSet,assemble,distill,mine}.ts | fetch → mine → assemble → grade |
src/config/constants.ts | every tunable knob |
docs/superpowers/specs/ | methodology + sourcing design specs |