The Brick Drop · System Documentation

The Drop Score — How It Works

A complete view of the logic behind the holistic LEGO set grade: the rubric, the math, the data sources, and how the AI judges real reviews — with citations for everything.

Originally 2026-06-17 · engine refreshed 2026-08-15 (never-blank grades, rebuilt investor value, True Value, real minifig pricing) · thebrickdrop.com · github.com/sethdesignsco/the-drop-score
One-sentence summary: every LEGO set gets an A+→F grade (with a precise 0–10) built by blending objective facts, real reviewer opinion mined from the web, and community grades — weighted by how much evidence exists, and always shown with how sure we are.

1 · The thesis — why it exists

The LEGO community judges a set's worth with one lazy number: price per piece (the "$0.10/brick = good deal" rule). It's broken because it ignores part size and value (a tiny 1×1 tile counts the same as a big trans-black canopy), minifigures, build enjoyment, display quality, and everything else that actually makes a set good.

The Drop Score replaces that with a grade built from both value and joy, drawn from real human experience at scale, and always honest about how sure it is. The non-negotiable principles:

  1. Beyond price-per-piece — value is judged by what the parts are worth and what the box delivers, never by raw piece count.
  2. Real human experience, at scale — subjective qualities (fun, looks, value) come from real people (mined reviews + community), not a formula pretending to have taste.
  3. Always honest about certainty — confidence rides beside the grade, never baked into the letter. A quiet set still gets a letter; the honesty lives in the confidence line ("Early read"), not in a blank.
  4. Never a blank score — every scorable set (and every read — Worth it? / Will you love it? / Hold value?) resolves to a real letter derived from facts, upgraded by opinion when reviews exist. A dash ("—") never ships. See §4 and §7.
  5. A community tool — anyone can contribute a grade (no account), but only with a written reason.
  6. Sounds like The Brick Drop — credible, clean data; personality on top ("The Drop Take").

2 · The three ingredient streams

Every point on the rubric is scored 0–10 by blending up to three streams, weighted by how much evidence exists. A missing stream simply drops out — we never guess to fill it.

StreamWhat it isWhere it comes from
① FactsObjective data, computedBrickset (price, pieces, figs), Rebrickable (full inventory, novelty), BrickLink (per-part market value)
② Mined opinionReal published reviews, distilled by AIYouTube transcripts + comments (via Supadata + YouTube API), Brickset structured sub-ratings, Brick Insights aggregate
③ CommunityVisitor grades + required written "why"The Drop Score site's no-account submission form (Layer 3, not yet built)

The blend formula (methodology §2.1, implemented in src/lib/scoring/blend.ts):

wᵢ = log(1 + volumeᵢ) × trustᵢ × recencyᵢ
   recencyᵢ = decays ? exp(−ageMonths / 12) : 1     // value sentiment decays; look/build sentiment doesn't
pointScore = Σ(wᵢ · scoreᵢ) / Σ(wᵢ)                  // 0–10

Trust weights: Brickset 1.0 · Community 0.9 · Reddit 0.8 · Blog 0.8 · YouTube 0.7 · Forum 0.7 · Amazon 0.5. Facts, where they define a point, are authoritative for that point; opinion streams enrich within the bounds each dimension sets.

3 · The rubric — all the points

Six clusters feed the grade; a seventh ("Hold Its Value") is a side badge that never moves the letter (except under the Investor lens). Tags: Fact computed · Feeling mined/community · Mixed both.

ClusterWeightPoints (sub-weight)
💰 Money & Worth22%Smart Price F (65) · What You Get F (35) · Feels Worth It M (capped nudge, not a slice)
🔧 The Build19%Build Fun M (55) · Build Length F (25, price-scaled) · Holds Together & QC Fl (20)
🏆 The Finished Model19%Display Appeal Fl (65) · Fidelity & Scale M (35)
🧑‍🚀 Minifigures16%Figure Lineup M (65) · Figure Desirability & Exclusivity M (35)
🛠️ For the Hobbyist12%Playability Fl (55) · Parts for Custom Building M (45)
✨ Rare & New Parts12%Rare & New Parts F (100)
📈 Hold Its Valueside badge — 30% only under InvestorAppreciation Outlook F (50) · Value Floor F (50) — see §5b

Anti-double-counting is enforced: minifig signals live only in Minifigures; part novelty only in Rare & New; generic element variety only in Parts for Custom Building; the inventory "filler" signal hurts only value (What You Get); Build Length adds value salience only inside The Build. Source: methodology spec §3.1.

4 · The math, step by step

Implemented in src/lib/scoring/score.ts (the computeDropScore() orchestrator).

  1. Detect archetype (§3.0 below) → mark which points are not-applicable.
  2. Score each point 0–10 by blending its streams (facts authoritative where present).
  3. Feels-Worth-It nudge: community/mined value sentiment pulls Smart Price within a capped band — it is not a separate additive slice (prevents double-counting): smartPrice' = clamp(smartPrice + clamp(feelsWorthIt − smartPrice, ±0.75), 0, 10).
  4. Roll up clusters = weighted average of in-scope points (sub-weights). N/A and insufficient points drop out and the surviving sub-weights renormalize.
  5. Roll up overall = weighted average of in-grade clusters, with proportional renormalization + an over-concentration cap (no cluster > 40% post-renorm; if it binds, the grade is flagged "data-light" — the number is unchanged, only the confidence tier drops).
  6. Never blank: a grade is withheld only if zero in-grade dimensions scored. Any set with at least one scored dimension always bands to a letter — a thin set reads as an honest "Early read," never a dash. (This retired the old "not enough to grade yet" abstention.)
  7. Investor floor: under the Investor lens only, a scored overall never bands below C− (5.5) — a poor investment shouldn't read like an F and drag the set's whole perception down; the hot take carries the low-end nuance.
  8. Round half-up to one decimal, map to a letter.

Set archetypes (§3.0) — they decide which points apply

ArchetypeTriggerEffect
Figureless0 minifiguresMinifigures cluster → N/A
Pure displaydisplay line (Botanicals, Architecture, Art)Playability → N/A
Parts-pack / bulkClassic tubs, assortmentsFinished Model, Fidelity, Build Fun → N/A; filler penalty suppressed
Figure-anchored polybag< 100 pcs AND ≥1 exclusive figSmart Price → N/A; Minifigures carries the grade

The scale (0–10 → letter)

B+
8.2 / 10 · the letter is the headline, the decimal is the precision, and each cluster gets its own 0–10.
A+AA−B+BB−C+CC−D+DF
9.5–109.0–9.48.5–8.98.0–8.47.5–7.97.0–7.46.5–6.96.0–6.45.5–5.95.0–5.44.0–4.90–3.9

The bottom is deliberately coarse (wide D, no D−): below-average is below-average; we don't over-resolve failing sets.

Lenses (optional re-weighting)

Universal is the loud default. A viewer can flip a lens that re-weights clusters to their priorities. The Investor lens is the only mode where resale enters the grade.

ClusterUniversalDisplayBuilderParentInvestor
Money & Worth2218182022
The Build191616228
Finished Model1926121412
Minifigures1618102012
Hobbyist121024184
Rare & New121220612
Hold Its Value————30

5 · Fact-scoring formulas

The bridge from raw facts to 0–10 dimension values (src/lib/facts/). All curves are piecewise-linear over tunable anchors.

Smart Price — the "beyond price-per-piece" core

Compute a value ratio = the set's part value ÷ what people actually pay (street price), with a fallback chain. Then rank that ratio against the set's cohort (same theme × size class) and map the percentile to 0–10.

valueRatio = partValueTotal / streetPrice          // primary (Rebrickable inventory × BrickLink per-part value)
           → weightGrams / streetPrice              // fallback 1 (weight predicts retail better than piece count)
           → pieces / streetPrice                    // fallback 2 (last resort)
score = piecewiseLinear(percentile_in_cohort, [[0,1],[0.1,3],[0.5,6],[0.9,9],[1,10]])
        // cohort median → 6.0, clear bargain → 9+, clear rip-off → <3

Size buckets: <250 / 250–750 / 750–1500 / 1500+ pieces. Licensing premium is surfaced explicitly, never silently penalized.

DimensionHow it's scored (anchors / rule)
Build Lengthpiece count → 0–10 [[25,1.5],[250,4.5],[750,6.5],[1500,8.0],[3000,9.0],[6000,9.7]]; its cluster weight scales with price: 25% × clamp(price/100, 0.7, 1.5)
What You Get0.6·printedShare-curve + 0.4·variety-curve − filler penalty (largest single-element share × 3)
Figure Lineupminifig count → [[0,0],[1,5],[3,7],[6,8.5],[10,9.5],[20,10]]
Rare & Newweighted novelty (newMolds + 0.5·recolors + 0.5·exclusives) → [[0,4],[1,6],[3,7.5],[6,9],[12,10]]
Appreciation Outlook (the scarcity dimension)see §5b — age-through-life × how collectable the theme is
Value Floor (the resale dimension)see §5b — part-out ÷ price

All anchors above are tunable defaults, centralized in src/config/constants.ts. The structure is the decision; the exact numbers get tuned against real sets.

5b · Investor value & the True Value breakdown

The "Hold Its Value" cluster used to be a frozen 5.0 for every set — both of its ingredients depended on data we never populated. It was rebuilt from facts we actually have. It's a side badge that only enters the letter under the Investor lens (weight 30).

Value Floor — what the parts are worth vs. what you pay

anchor       = streetPrice ?? msrp
valueFloor   = piecewiseLinear(partValueTotal / anchor,
               [[0.5,2.0],[1.0,5.0],[1.5,7.0],[2.0,8.5],[3.0,10.0]])
             // 1.0× "fully backed by parts" reads 5.0; 3.0× is a 10

Returns "insufficient" (an early read, never a fabricated 0) when there's no part-out value or no price. Note this is a floor read — distinct from Smart Price's deal-framing where ~1.75× is the neutral market point.

Appreciation Outlook — how much upside is left

progress = piecewiseLinear(ageYears, [[0,0.50],[1,0.56],[2,0.68],[3,0.82],[5,0.93],[8,1.0]])
ceiling  = THEME_DESIRABILITY[theme] ?? 5.5      // how collectable the theme is
score    = clamp(5.0 + progress × (ceiling − 5.0))   // base 5.0 = "unproven upside"

A brand-new set starts at 0.5 of the way through its life (not "failing"), and its confidence is lowered (not its score) until it's 2+ years old. Theme desirability is a curated, tunable tier — high (Star Wars, Icons, Ideas, Architecture, Harry Potter ≈ 9; Botanicals, Pokémon ≈ 8.5), mid (Marvel/Technic/Super Mario ≈ 6), low (City 3.5, Duplo/Classic 3, Gear 2.5). Star Wars has a low part-out but high appreciation — part-out alone misses collector demand, which is the whole point of keeping the two ingredients separate.

True Value — the price breakdown shown on the set page

The grade answers "is this priced fairly?" in one letter; True Value shows the receipts. Rendered by deriveTrueValue() (src/lib/app/gradeToInput.ts) into src/ui/TrueValueBreakdown.tsx, it appears on any set that has a part-out value (~1,000 sets today).

anchor price   = streetPrice ?? msrp        (tagged "street" or "MSRP")
parts ratio    = partOut / anchor           → shown as "parts are worth ~N% of the price"
verdict        = ratio ≥ 1.0  → Bargain
                 ratio ≥ 0.7  → Fair
                 else         → Overpriced
bar width      = FLOOR(6%) + (100−6)·(amount/max)^0.5   // legibility curve; the $ figure is the truth

It breaks the value into parts (plain bricks / printed & detailed / functional pieces / figures) and a premium group (brand & IP, plus a licensed royalty line). Tags surface Licensed and Retired. It never divides by piece count — price-per-piece is banned everywhere.

Minifig value is included. When a set's figures have been priced (each figure's own parts, via Rebrickable × BrickLink), that value is summed into partValueTotal and shown as a real "Figures" line. For sets not yet re-priced, True Value shows a clearly-labeled flat estimate (≈ $2.50/fig, $3.50 licensed) tagged "Figures — est.", and the price-read grade stays on the real priced-parts basis so the number and the grade never contradict each other.

6 · How the LLM judges reviews

This is the engine's "feelings" half. For each review, Claude (claude-opus-4-8) reads the transcript + top comments and returns structured per-dimension scores. Implemented in src/lib/pipeline/distill.ts using the Anthropic API with structured outputs (a JSON schema constrains the response).

The exact system prompt the judge runs on

You distill a single LEGO set review into per-dimension scores for a grading engine.

Score each dimension 0–10 ONLY if the review (or its comments) actually speaks to it;
otherwise set score to null. Dimensions:
- buildFun: how enjoyable/clever the build is (vs repetitive/tedious)
- instructions: instruction clarity, bagging, sticker pain
- integrityQC: sturdiness when handled; arrives complete/undamaged
- displayAppeal: how good the finished model looks / shelf presence
- fidelityScale: accuracy to the real thing / theme standard
- playability: fun to actually play with (features, swooshability)
- partsForBuilding: value as a parts source for custom builds (MOCs)
- feelsWorthIt: whether it's worth the money

CRITICAL: be sarcasm-aware. LEGO reviewers are ironic ("oh great, another grey spaceship",
"a steal at only $0.10 a piece"). Read the real sentiment, not the literal words. Weight the
reviewer's own assessment over comment noise. Set sarcasmDetected true if irony materially
shaped any score. Give a short evidence quote per scored dimension.

Why a real LLM and not keyword counting

LEGO reviewers are relentlessly ironic. A keyword counter sees "oh great, another $0.10-per-piece cash grab" and scores "great" as positive. The model understands it's a dig. Verified live on a deliberately sarcastic X-Wing review:

"Absolutely riveting — if your idea of a good time is attaching 400 identical grey tiles"
                                            → buildFun 3   (read as ironic / tedious)
"once it's done, this thing looks genuinely stunning on the shelf"
                                            → displayAppeal 9
"$0.10-a-piece masterpiece ... worth it only if Star Wars is your entire personality"
                                            → feelsWorthIt 3
sarcasmDetected: true

From judgment to engine signal

Each non-null distilled score becomes one signal (volume 1) fed to the same blend as every other opinion. Value sentiment (feelsWorthIt) is marked decaying (recency matters because price changes); look/build sentiment doesn't decay. A dimension needs ≥3 mined opinions (MIN_VOLUME_FOR_SENTIMENT) before it counts toward the grade — below that it's "applicable but insufficient," which lowers coverage rather than faking a score.

7 · Confidence & coverage — two honest signals

Confidence is where honesty lives now. Because a grade is never blank, "how sure are we" is carried entirely by these two signals and the named state below — a thin set shows a real letter with an "Early read" label, not a dash. The three headline reads (Worth it? / Will you love it? / Hold value?) each have their own fallback so they resolve too: the price read falls back to "What You Get," the love read blends a fact baseline (What You Get 0.5 / figures 0.25 / build length 0.25) with mined opinion 60/40, and the hold read falls back to the appreciation signal.

Every grade carries two separate signals, measured in src/lib/scoring/dimension.ts + score.ts:

Saying it so a human gets it

The raw readout — C− (5.7), full coverage, 34% confidence — makes the two signals fight: "full" sounds great, "34%" sounds broken, and the reader can't tell whether to trust the grade. That 34% isn't the tool doubting its own math — it's six real reviewers who showed up and disagreed. So we name depth for what it is (how much the crowd agrees) and always surface the reason:

Split verdict

We graded the whole set — but six reviewers came in split. Take this as a read, not gospel.

Coverage 9/9 points  ·  Agreement 34%  ·  from 6 mined reviews + 18 Brickset ratings

"Data-light" retires into a small vocabulary of named states that say the why out loud — same two axes, no jargon:

Locked in Solid read Split verdict Early read Too quiet

Worked example: rich, agreeing data on only 3 of 9 points = high consensus, low coverage (an "Early read"). All points scored but only from noisy, disagreeing sources = full coverage, low consensus (a "Split verdict"). The live X-Wing mined grade was a Split verdict — C− (5.7), whole set graded, consensus shaky (34%) — honestly low because six real reviewers disagreed and were more critical than Brickset's 5-star ratings.

Presentation not locked. The wording + visual above is the recommended direction, not a final decision — all five explored treatments live in docs/design/2026-06-17-confidence-coverage-presentation.html. Recommendation (revisit in Layer 3): rename the depth axis to consensus, lead the set-page card with the one-line verdict, expand into a "receipts" scatter of where each reviewer actually landed, and carry the named status as the compact badge on list/timeline views. The engine already emits everything needed (coverageTier, confidence, per-cluster scoredCount) — only the presentation changes.

8 · The data pipeline

One command (npm run grade <set> --mine) runs the whole chain (src/lib/pipeline/gradeSet.ts):

setNumber │ ├─ Brickset getSets ─────────► facts (price, pieces, figs, theme, year) ├─ Rebrickable set + parts ──► facts (pieces) + inventory summary (prints, variety, filler) ├─ Brickset getReviews ──────► structured sub-ratings → signals (buildFun, playability, value, parts) │ └─ (--mine) mineSetReviews: dimension_signals cache? ─► HIT: reuse stored signals, skip mining entirely │ MISS ↓ review_index lookup ──► origin reviews (pre-discovered by the allowlist sweep) │ (cold-start fallback: YouTube search.list, only when the index is empty) └─ per video: Supadata /transcript ──► transcript text YouTube commentThreads ─► top ~30 comments Claude distill ─────────► per-dimension 0–10 signals (sarcasm-aware) save signals → dimension_signals (so the next grade is a cache hit) merge all review signals │ facts + inventory + all signals │ buildFactInputs() ──► fact dimensions become 0–10 mergeInputs() ──► facts + opinion signals per dimension │ computeDropScore() ──► clusters → overall 0–10 → letter (+ confidence, coverage) │ Supabase upsert (sets, drop_scores) ──► read back to confirm

Discovery & caching — why mining is cheap

Discovery is split from grading so we never pay search.list (100 units) per grade. A separate sweep — scripts/discover.ts — pages each allowlist channel's recent uploads via playlistItems (~1 unit/50 videos), matches each video to a known set by its number or its name (so "UCS X-Wing review" with no number still lands), and writes the origin reviews to review_index. The same index is also fed by a Brick Insights origin harvest — each set's outbound "Read reviews" links (YouTube + blogs/forums), minus retailers/affiliates and the already-ingested Brickset. Grading then reads that index and distills from it directly; search.list only fires as a cold-start fallback for an un-indexed set. Distilled signals are cached in dimension_signals, so a re-grade reuses them and skips mining (and LLM spend) entirely — pass --remine to force a fresh pass. (Live-verified: the Brick Insights harvest added 8 origin links — 7 blog + 1 forum — for the UCS Falcon; the grade distilled from the index with zero search units, and the next grade hit the cache.)

Matching, carefully (this part is easy to get wrong): name-matching is hard because different sets share words — there are many "X-wing" and "Millennium Falcon" sets across the years. So it matches on the title only (descriptions list every set in a roundup); needs every distinctive word of the set name (≥2); ignores a title that carries a different set's number; drops any set released after the video; and accepts only when exactly one known set fits. A naive first cut wrongly swept 10 unrelated videos under the X-Wing — review round-ups plus the New Republic X-Wing Starfighter 75460, a different set — which the guards now reject; a set number in the title always wins.

9 · Sources & citations

Live data APIs (all keys verified working this session)

SourceUsed forEndpoint / docs
Rebrickableinventory, parts, minifigs, part noveltyrebrickable.com/api/v3/docs
BricksetMSRP, pieces, minifig count, structured review sub-ratingsbrickset.com/tools/webservices/v3
BrickLinkper-part market value (Smart Price), resalebricklink.com/v3/api.page (OAuth1)
Brick Insightsorigin-review harvest (outbound "Read reviews" links → review_index) + bootstrap aggregate scorebrickinsights.com (JSON-LD + card--cta links on /sets/{id}; index at /sitemap/sets)
SupadataYouTube transcripts (managed, ToS-offloaded)supadata.ai → GET /v1/transcript?url=…
YouTube Data API v3allowlist discovery (playlistItems, ~1u/50) + comments; search.list cold-start fallback onlydevelopers.google.com/youtube/v3 (key from console.cloud.google.com)
Claude APIsarcasm-aware review distillation (the LLM judge)platform.claude.com/docs · model claude-opus-4-8

Where the methodology (the rubric & weights) came from

The rubric was not invented — it was synthesized from how real reviewers and the community actually judge sets (a multi-agent research pass at the start of the session):

Full reasoning is in docs/superpowers/specs/2026-06-13-drop-score-methodology-design.md (and the sourcing spec alongside it).

Where the UI direction came from

The "Obsidian Dispatch" look + score-card layout were grounded in: Metacritic, OpenCritic, Pitchfork (the score-as-typographic-monument), Rotten Tomatoes + its Pentagram rebrand (one-accent discipline), IGN, Brick Insights, BrickEconomy, and Linear (the buildable dark/single-accent spec). Full brief: docs/design/2026-06-17-ux-inspiration-brief.md.

10 · How the review videos are found

Two discovery paths exist; the live demo used the on-demand one.

A. The curated reviewer allowlist (the canonical source)

48 hand-vetted channels, every channel ID verified via the YouTube API, in src/config/reviewer-allowlist.ts. A representative slice (full URLs let you find any video they've made):

ChannelURLChannel ID
JANG's LEGO Reviews@JANGsLEGOreviewsUCPi-XOCT88MwgFlkLHzBNhA
just2good@just2goodUCp_mZttcKNIcUBVdi0wTdIA
MandRproductions@MandRproductionsUCLnr9MzQ_v_a_OPmJkye5lA
Solid Brix Studios@SolidBrixStudiosUC_EahESpmOsO_5hAmsNSqlw
BrickVault@BrickVaultUCrhb3SP2lZBgguLHIWWuHOQ
RacingBrick@RacingBrickUCfU8ME4_m48QwDJCUvpzqyQ
Tiago Catarino@TiagoCatarinoUCqLbTAc5Mn2cJDu1-FZ2W3g
Brick Fanatics@BrickFanaticsUCGLVjA9MlM0KQHySRVwAeuw

The remaining 40 (theme specialists, MOC builders, build-alongs, value channels) are in the file with the same fields. Excluded deliberately: "Held der Steine" (German), no-talking ASMR channels (no usable transcript). Research provenance: Feedspot, Brick Insights reviewers, The Brick Land.

B. On-demand search (what the live demo used)

For a single set, the miner calls the YouTube search API and keeps results whose title mentions the set number or a name token:

GET https://www.googleapis.com/youtube/v3/search
    ?part=snippet&type=video&order=relevance
    &q=LEGO 75355 X-wing Starfighter review

Reproduce in a browser:
https://www.youtube.com/results?search_query=LEGO+75355+X-wing+Starfighter+review

Confirmed videos used in the X-Wing demo (the live set varies per run; these two distilled successfully with full English transcripts + ~30 comments each):

Now wired: the production path is allowlist playlistItems (1 unit/50 videos) cross-fed with the Brick Insights origin-review map, both landing in review_index. Grading distills straight from that index; on-demand search.list (100 units/call) is only a cold-start fallback for an un-indexed set.

11 · Storage — the Supabase schema

Project "DropScore" (eubnrfwibntsqetdmear, us-east-2). Six tables, all with Row-Level Security. Derived data only.

TableHoldsPublic access (RLS)
setscached objective facts per setread
part_pricesglobal per-part value cache (BrickLink), set-independentnone (internal/service-role)
dimension_signalsderived per-dimension scores (no raw text)read
review_indexlinkback map: set → origin review URLs (links only)read
drop_scorescached engine output per set per lensread
community_submissionsno-account grades; a written why is required at the DB levelread + insert

Cache writes happen via the service-role key; the public site reads via the publishable key. Mapper code: src/lib/sources/supabaseStore.ts.

12 · Every tunable knob

All in one place — src/config/constants.ts. Change behavior here without touching logic.

KnobDefaultMeaning
Cluster weights22/19/19/16/12/12relative importance of the 6 in-grade clusters
FEELS_WORTH_IT_CAP±0.75max nudge community value sentiment applies to Smart Price
BUILD_LENGTH_REF / clamp$100 / [0.7,1.5]how strongly price scales Build Length's weight
Smart Price buckets / anchors<250/750/1500 ; median→6cohort size classes + percentile→score curve
Stream trustBrickset 1.0 … Amazon 0.5how much each source counts in the blend
HALFLIFE_MONTHS12recency decay for value sentiment
MIN_VOLUME_FOR_SENTIMENT3mined opinions needed before a feeling-dimension scores
MIN_STRUCTURED_COUNT8Brickset ratings needed before they're a confidence booster
MIN_COMMUNITY_TO_MOVE5community grades needed before they move the number
LOW_DATA_FLOOR5overall "low data" threshold
RENORM_CAP40%max share any one cluster can carry after renormalization
POLYBAG_PIECES100figure-anchored polybag trigger
Letter bandssee §40–10 → A+…F cutoffs

14 · Cost model (what actually costs money at scale)

15 · Limits & what's NOT built (the complete picture)

An honest inventory of the edges, so nothing is a surprise:

Built, live & working (≈800 passing tests)

The scoring engine, the fact-scoring layer, all source adapters (live-verified), the LLM distiller, the mining orchestrator with allowlist discovery + dimension_signals caching, Supabase persistence, and — since this doc was first written — the whole app is live at thebrickdrop.com: ~5,800 sets graded across five lenses, per-set pages, search, browse-by-release, compare, saved sets, the editorial "Brick Drop Take" (generated in-voice for every set), the downloadable 4:5 share card, and the True Value breakdown. Batch grading has swept the modern catalog.

Shipped since the original draft

Still open / deliberately deferred

Things worth knowing

16 · File map & glossary

Where the logic lives

PathResponsibility
src/lib/scoring/score.tscomputeDropScore() — the orchestrator (steps in §4)
src/lib/scoring/{blend,dimension,rollup,bands,archetypes,rubric}.tsthe math primitives
src/lib/facts/raw facts → 0–10 (Smart Price, Build Length, etc.)
src/lib/sources/the 6 adapters + Supabase store + contracts/fakes
src/lib/pipeline/{gradeSet,assemble,distill,mine}.tsfetch → mine → assemble → grade
src/config/constants.tsevery tunable knob
docs/superpowers/specs/methodology + sourcing design specs

Glossary

The repository is the single source of truth. This document was first written 2026-06-17 and engine-refreshed 2026-08-15 (never-blank grades, rebuilt investor value, True Value, real minifig pricing); if code and doc ever disagree, trust the code.