Nine independent reviewers each tore into a different part of the grading engine and reported back honestly. This is that report — rewritten with the jargon stripped out, every term explained, and every problem walked through with a real LEGO example. Same findings, said plainly.
The grade-making machine has an excellent calculator at its heart and a set of safety rules that were never switched on. When The Drop Score has rich, clean information about a set, it does the math faithfully and hands you the right letter. But the entire point of your methodology was to stay honest when the information is thin, old, or being gamed — and that is exactly where the engine quietly overstates itself. The three named safety gates that are meant to hold back weak evidence are written into the code but never actually run. The confidence badge can be fooled by a single review into reading 100%. And Smart Price — the headline "is this set worth the money?" feature — never receives the real street-price data it was designed around, so today it silently measures against the sticker price instead. The skeleton is strong and the methodology is right; the trust layer is hollow. It is a strong prototype, not yet something to publish grades from.
Switch on the three safety gates — for real, inside the code that actually computes the score, with a test for each so they can never quietly turn off again. This was the single most-agreed finding: 5 of the 9 reviewers landed on it independently. It restores the anti-gaming and "be humble on thin data" promises at the center of your methodology, it fixes the "looks more covered than it really is" problem at the same time, and — unlike the bigger jobs (wiring up real street prices, hardening the plumbing) — it is small, local, and you can prove it works with a test.
Translating each finding meant re-reading the actual code. It held up almost everywhere — but a few small things deserved a more honest framing, and in the spirit of the whole exercise, here they are:
Jump to any part:
Audit called it: Scoring engine math
This is the calculator at the very center of The Drop Score. It takes all the raw evidence about a set — fact-based numbers (like minifig count or resale value), star ratings pulled from Brickset reviews, and sentiment squeezed out of fan reviews and community submissions — and blends it, dimension by dimension, into cluster scores and then one overall 0-10 that becomes the letter grade. It also decides how confident the grade is and whether to label it a full grade or a more cautious "data-light estimate." In short: this is where opinions and facts become the number you publish.
The raw weighted-average math here is correct and clean — when good data goes in, a sensible number comes out. The trouble is in the guardrails. Your written methodology promises three specific safety gates that are supposed to stop weak or manipulable evidence from moving the grade, and all three exist in the code as named numbers but are never actually switched on — they are decoration, not protection. On top of that, two behaviors quietly disagree with your spec: when the engine's anti-lopsidedness cap kicks in, it doesn't just cautiously re-label the grade, it actually rewrites the headline number (in the audit's example, dragging a 7.8 down to a 5.4); and a set with no usable data at all comes out as a flat 0.0, which maps to a confident F instead of an honest "we don't have enough to grade this." There are also smaller honesty gaps in how confidence is computed. None of this is broken arithmetic — it is the difference between what the code claims to do and what it actually does.
The core blending math is genuinely solid and easy to trust; the problem is that the anti-manipulation and "be humble when data is thin" promises in your methodology are written down but not wired in, so the engine is more confident and more game-able than the spec says it should be.
Each problem below is laid out the same way: what's happening, a concrete LEGO example, why it dents trust, and what the fix is.
A 'gate' here is a minimum-evidence threshold: a rule that says 'don't let this evidence move the grade until there's enough of it.' Your methodology defines three of them, and all three exist in the code as named numbers — but none is ever checked when the score is computed. They are dead constants: present, named, commented as if active, but doing nothing.
Gate 1 — MIN_COMMUNITY_TO_MOVE = 5. Your spec calls this the main anti-brigading guardrail. 'Brigading' is when a coordinated group floods your tool with submissions to swing a grade. The spec says: until at least 5 community grades exist, they should be display-only — shown on the page but contributing zero to the number. In reality the only gate the code applies to community opinion is a different, weaker one, MIN_VOLUME_FOR_SENTIMENT = 3 (dimension.ts:47). So just 3 community submissions already move the published number.
Gate 2 — MIN_STRUCTURED_COUNT = 8 (Brickset). 'Structured' means the 1-to-5 star sub-ratings Brickset reviewers leave. The spec says fewer than 8 of these should be discarded, because a handful of Brickset ratings skews old and self-selected. The code's own comment in brickset.ts:58 literally says 'The MIN_STRUCTURED_COUNT gate is applied upstream, not here' — but it is applied nowhere. So a Brickset signal built from a single reviewer enters the blend at the highest trust level of all sources (1.0), unfiltered.
Gate 3 — LOW_DATA_FLOOR = 5. The spec says a dimension needs 5 units of evidence before it counts. The code uses 3 instead (the MIN_VOLUME_FOR_SENTIMENT gate again), and the comment at dimension.ts:20 claims LOW_DATA_FLOOR gates here — it does not. So points get scored on 40% less evidence than the spec requires (3 instead of 5), and a code comment actively tells a future reader the opposite.
Picture the new Razor Crest (set 75331) just after a popular YouTuber tells fans to 'go rate it.' Four of them submit community grades averaging a glowing 9.4 for value. Under your written rules, the anti-brigading gate (5) hasn't been cleared, so those four grades should sit on the page as '4 community grades, not yet load-bearing' and contribute 0 to the number — the published grade stays put. What actually happens: the code only checks for 3, so those 4 coordinated grades sail through and pull the Razor Crest's value score upward. A grade you publish as authoritative just got nudged by exactly the kind of small coordinated push the gate was designed to absorb. Same shape with Brickset: if the Razor Crest has just 2 Brickset star-ratings, the spec says throw them out (too thin), but the engine instead lets those 2 ratings in at trust 1.0 — the heaviest weight any source gets — so two early reviewers can tilt the build score.
This is the difference between a grade that looks trustworthy and one that actually is. You can publish a Razor Crest grade believing it's protected against a handful of fans gaming it, when in truth 3 submissions or 2 stray star-ratings already moved it. The whole selling point of The Drop Score is that the number is hard to manipulate and honest about thin evidence — and right now the specific guardrails that deliver that promise are off, while the comments in the code say they're on. If you ever told a reader 'community grades don't count until 5,' that statement would be false today.
Actually enforce the gates where the score is computed: sum the community volume and drop those signals until it reaches MIN_COMMUNITY_TO_MOVE (5); discard Brickset signals built from fewer than MIN_STRUCTURED_COUNT (8) ratings; and either switch the insufficiency check to LOW_DATA_FLOOR (5) as the spec says or change the spec and the constant to match — but stop letting the comment and the spec both say 5 while the engine quietly uses 3. Add one test per gate so a dead constant can't silently come back.
There's a guardrail called the over-concentration cap (RENORM_CAP = 40%). It exists so no single cluster can carry more than 40% of the whole grade — to stop a grade from resting almost entirely on one category. Your spec is clear about what this cap is for: when it 'binds' (kicks in), it should force the set into the cautious 'data-light estimate' tier — a label saying 'we're standing on a narrow base, take this with a grain of salt.' It's meant to be a warning sticker.
What the code actually does (rollup.ts:91-105) is also re-weight the math and change the number itself. When the cap binds it redistributes the weights and recomputes the 0-10 from the flattened weights — so the headline score moves, not just the tier label.
Imagine an obscure UCS-style set where only three categories survived with data: its money-for-value scored a strong 9 and naturally carried about 80% of the surviving weight, while the build and one other category each scored a 3 with about 10% each. The honest weighted average of what we actually know is (0.8×9)+(0.1×3)+(0.1×3) = 7.8 — a solid B. But because that 9-scoring category is over the 40% cap, the engine flattens the weights to 40/30/30 and recomputes: (0.4×9)+(0.3×3)+(0.3×3) = 5.4 — a D+. That's a five-letter-grade drop (B, B-, C+, C, C-, D+), produced not because the set is worse but because the cap rewrote the arithmetic. The spec's intent was to keep the 7.8 number and simply stamp it 'data-light estimate,' warning the reader the base is narrow — not to silently turn a B into a D+.
A reader sees a confident-looking 5.4 / D+ with no hint that the 'real' weighted number was a 7.8. The grade looks like a verdict on the set when it's actually an artifact of a guardrail firing. Either direction is a trust problem: if the flattening is intentional you've never documented it, and if it isn't, you're publishing numbers that are several letter grades off from what your own evidence supports.
Compute the headline 0-10 from the honest, uncapped proportional weights, and use the cap only to decide the tier label (flip to 'data-light estimate'). If you genuinely want the cap to move the number, document that choice in the spec and add a test that pins the intended behavior so it's a decision, not a surprise.
When none of the in-grade categories have any value to contribute, the roll-up returns a score of 0 (rollup.ts:99), and the letter-band table maps 0.0 straight to an F. Nothing special-cases the 'we have literally nothing' situation. So 'no data' and 'genuinely terrible set' produce the identical output: a hard F.
This directly contradicts a headline principle of your methodology — that every set gets a grade, a full one when data is rich and a clearly-labeled data-light / factual-only estimate when it's quiet — and that no set is left looking bad merely for being undocumented.
Take a quiet new Botanicals polybag that nobody has reviewed yet and that's missing its fact fields too. The engine finds zero usable categories, returns 0.0, and the page shows a bold F. A fan glancing at it reads 'this set is a disaster.' What actually happened is 'we don't know anything about this set yet.' Your spec wanted a data-light estimate or an explicit no-data state — an honest shrug — not a confident failing grade slapped on a set whose only crime is being new and quiet.
An F is the single most damaging thing your tool can say about a product, and here it can be triggered purely by absence of data. That's reputationally risky for you (a LEGO fan community will notice a beloved-but-obscure set branded F) and for any set or theme that's simply under-covered. A grade should never look like a confident judgment when it's really a blank.
Special-case zero coverage in score.ts: when no in-grade category has a value, return an explicit withheld / null result with a visible 'no data' state, instead of 0.0 and F.
Your methodology describes several steps that are supposed to clean the evidence before it's blended. The math layer skips all of them and treats every incoming signal as already clean and already correctly paired. The missing pieces:
- Price-pairing of value opinions: the spec says a 'this is worth it' opinion only counts if it's tied to the price the person actually paid; unpaired value opinions should be discarded. (A '$50 well spent' comment means nothing for value unless we know it was $50.)
- The pre-availability lifecycle gate: sentiment posted before a set was on sale shouldn't be used to judge value or build experience (you can't judge those from a reveal trailer), though it's allowed for how the set looks.
- Outlier-dropping: throwing out a few extreme, unrepresentative ratings before averaging.
None of these exist in the scoring code; this weakness is tagged to 'design,' meaning it's an absence rather than a specific buggy line.
Suppose a Millennium Falcon UCS set has ten 'totally worth it!' reactions scraped from a reveal-day hype thread — before anyone could buy it, and with no price attached to any of them. Your spec says: drop every one of those (unpaired AND pre-availability for value). The engine instead feeds all ten straight into the value blend, so reveal-day excitement about a set nobody has actually bought inflates a value score that's supposed to reflect real, in-hand, price-aware judgment. The grade should have leaned on the cold facts until real post-purchase opinions arrived; instead it rode the hype.
These filters are what make the difference between 'what people actually concluded after buying it' and 'what people shouted on announcement day.' Without them, value and build grades can be moved by hype and by comments that were never valid evidence for those categories — so a grade can look like a considered verdict while really echoing reveal-day noise.
Build the spec's cleaning steps into the pipeline before the blend: discard value opinions with no paired price, suppress pre-availability sentiment for the value and build categories (but keep it for the look category), and drop statistical outliers before averaging. Until then, the blend is trusting evidence the methodology says should never have reached it.
Three smaller honesty gaps in how the 'confidence' figure is built:
1. Fact categories get a hard confidence floor of max(0.8, ...) — meaning the engine reports at least 80% confidence on a fact even when zero supporting signals corroborate it. But your spec (6.1) says a thin fallback (e.g. guessing from a sibling set or theme) should LOWER confidence, not pin it at a comfortable 0.8.
2. Overall confidence is a plain unweighted average (score.ts:92) of each category's confidence — a category carrying 2% of the grade counts the same as one carrying 40% when computing how sure we are.
3. The volume-to-confidence curve uses log1p(60) as a hidden divisor (dimension.ts:9) — a 'magic number' (an unexplained constant dropped straight into the formula) that doesn't appear in your spec's table of tunable constants, so nobody can see or calibrate it.
Say the Razor Crest's resale value is unknown, so the engine falls back to a rough sibling-set guess. With no corroborating evidence at all, it still reports 0.8 (80%) confidence on that number because of the floor. A reader takes the high confidence as 'they really nailed this one,' when honestly it was a soft guess that the spec says should read as low-confidence. Separately, because overall confidence is an unweighted average, a rock-solid reading on a tiny 2%-weight category can prop up the headline confidence as much as the main value category does.
Confidence is your honesty dial — it's how you tell readers 'trust this grade a lot' versus 'we're not sure.' If a soft fallback still shows 80% sure, and if a trivial category can buoy the overall confidence, then the dial reads more confident than the evidence warrants. That quietly undercuts the very thing confidence is for.
Make confidence weight-aware (average it by how much each category counts toward the grade), replace the unconditional 0.8 fact floor with a confidence that reflects how the fact was actually derived (corroborated vs. fallback guess), and lift log1p(60) into the named, documented constants list so it can be seen and tuned.
The audit is accurate and matches the code in every substantive way; I confirmed each cited line. Two tiny labeling nits, neither of which changes the audit's point: (1) The audit says the naive 7.8 maps to 'B+', but per the LETTER_BANDS table (constants.ts:151-164) 8.0 is the floor of B+, so 7.8 is actually a B. It then says the cap-distorted 5.4 is a D+, which is correct. The gap is therefore 5 letter-grade steps (B → B- → C+ → C → C- → D+), which matches the audit's own 'about five letter grades' phrasing — so the magnitude claim holds; only the 'B+' label is one notch high. (2) I reproduced the cap example's arithmetic exactly (naive 7.8, capped 5.4) assuming the two low clusters carried ~10% raw weight each; the audit didn't state those weights but the numbers only reconcile at that split, confirming the example. Everything else — the three dead gates appearing only in constants.ts plus two false comments, the zero-coverage→0.0→F path, the unweighted-mean confidence, the max(0.8,…) fact floor, and the log1p(60) magic number — was verified line-by-line against the source and the methodology spec (sections 2.2, 4.3, 6.1, 6.2).
Audit called it: Methodology fidelity
This part of the engine is supposed to do two jobs: (1) faithfully run the math your written rulebook (the "spec" — your own design document for how Drop Score is supposed to work) lays out for turning facts and opinions into a 0-10 score and a letter, and (2) enforce the rulebook's safety rules about WHEN there is enough evidence to let a grade stand on its own versus when it should be flagged as thin or held back. "Methodology fidelity" just means: does the code actually do what your own rulebook says it should do? The audit checked the real code against your spec, line by line.
The honest headline is split in two. The pure number-crunching is excellent — when your data is rich, the engine is a faithful, well-tested copy of your rulebook's math, down to the exact tunable numbers. But the parts of your rulebook that protect you when data is THIN are largely not wired in. Your spec names four "evidence gates" (minimum-evidence thresholds that decide whether a piece of data is allowed to move the grade). Only the weakest one is actually enforced in the code. The other three — the floor for "enough total opinions," the floor for "enough professional structured reviews," and the floor for "enough community votes" — exist as named constants and appear only in code comments saying they're handled "upstream," but no upstream code ever applies them. On top of that, the signature "Feels Worth It" fairness mechanic has lost its defining feature (pairing each value opinion with the price at the time it was said), and several anti-double-counting promises in the spec are aspirational comments rather than enforced rules. The net effect: the letter you get is right when reviews are plentiful, but it looks confident and precise exactly in the thin-data, long-tail situations your spec was most careful to protect.
The math you can see — weights, rounding, band lookups, the price-scaling on build length, the over-concentration guardrail — is genuinely faithful and well-tested; what is missing is the entire layer of "is there actually enough evidence here to trust this?" safety gates, which makes thin-data grades look far more confident than they should.
Each problem below is laid out the same way: what's happening, a concrete LEGO example, why it dents trust, and what the fix is.
Your rulebook defines a LOW_DATA_FLOOR of 5 — meaning a dimension fed only by opinions (not hard facts) needs at least 5 total opinions before it should count as genuinely 'scored.' The code defines that number but never uses it. The only check actually run is a weaker one, MIN_VOLUME_FOR_SENTIMENT of 3. So a dimension backed by just 3 or 4 opinions gets stamped 'scored' — as solid as one backed by 50 — instead of 'insufficient.' Because it reads as scored, it counts toward 'coverage' (the share of the grade that's actually backed by evidence) and is never recorded as a thin spot. ('Dimension' = one of the things you grade, like Build Fun or Display Appeal.)
Take a quiet little Botanicals set — say the Tiny Plants polybag. Almost nobody reviews it. Your mining pulls exactly 3 scraps of opinion about Build Fun, averaging an 8.5. The spec says: 3 is below the floor of 5, so Build Fun should be marked 'insufficient,' contribute nothing on its own, and the set should land in the data-light tier with a visible 'thin evidence' flag. What actually happens: 3 clears the only gate that's wired in (the 3-threshold), so Build Fun is stamped 'scored' at 8.5, counts as fully covered, and helps produce a confident-looking A- — built on three sentences from the internet.
This is the exact 'A- on 3 sources' failure your spec explicitly warns about. You'd publish a grade that looks fully backed and precise when it's really resting on a handful of comments. If a follower buys the set on the strength of that confident A- and it disappoints, the trust hit lands on you — the grade claimed certainty it never had.
In the dimension resolver, mark an opinion-only dimension 'insufficient' whenever its total opinion volume is below LOW_DATA_FLOOR (5). Keep the 3-threshold, if you want it at all, only as the separate, lower trigger for enriching with sentiment — not as the bar for calling something 'scored.'
Brickset publishes structured 1-to-5 sub-ratings (build, value, parts, playability). Your spec says these should only count once there are at least MIN_STRUCTURED_COUNT of them (8), because below that they skew toward older, adult, self-selected hobbyists and aren't representative — below 8 they should be discarded. The code that turns Brickset ratings into signals emits a signal the moment even 1 review carries a sub-rating. Its own comment literally says 'the MIN_STRUCTURED_COUNT gate is applied upstream, not here' — but no upstream code applies it. The gate exists as a number and a comment, and nowhere else.
Picture the UCS Razor Crest. On Brickset, 2 longtime AFOLs (adult fans) post early reviews and both rate Value for Money a stingy 2 out of 5. The code converts that to a value signal of 4 out of 10 (it multiplies the 1-5 rating by 2) and feeds it straight into the blend, nudging Feels Worth It — and through it, Smart Price — downward. Your spec says: only 2 reviews, that's below 8, throw them out entirely until more arrive. Instead, two grumpy early adopters quietly drag the set's value grade.
It means a tiny, unrepresentative cluster of early adult reviewers can swing a grade you publish, in precisely the direction your spec was trying to prevent. A set could read as a 'meh on value' B when the real, broader verdict is an A — and you'd have stood behind the wrong number.
Apply the MIN_STRUCTURED_COUNT gate before Brickset structured signals reach the blend — either in the Brickset adapter itself or in the assembly step — so fewer than 8 structured reviews contribute nothing to the computed score.
Your spec says community grades must be display-only below MIN_COMMUNITY_TO_MOVE (5) submissions — show them, but let them contribute 0 to the computed score until at least 5 people have weighed in. This is your anti-gaming protection. The community store instead emits one scoring signal per submission with no count check, and the assembly step merges those straight into the blend. So submission number 1 already moves the math.
A new set, the Orient Express, just launched. One person submits a community grade: Build Fun 10, Looks 10, Worth the Money 10, with the required written 'why.' That single vote enters the blend immediately and pulls the grade up. Now imagine a handful of friends — or a seller wanting the number to look good — each drop a glowing vote on day one. Five 10s and the set rockets. Your spec says: below 5 submissions, show them but contribute 0; only at 5-plus do they start moving the number, which blunts a tiny brigade. The code gives that protection away.
This is the exact anti-gaming failure your spec was designed to prevent. If a published Drop Score can be nudged by one or two motivated votes, the grade stops being something you can defend as independent — and a single coordinated push could make a mediocre set show an inflated grade with your name on it.
Gate community signals at MIN_COMMUNITY_TO_MOVE (5): below the threshold, emit them as display-only so they appear but contribute 0 to the score; at or above 5, let them enter the blend.
Your spec is emphatic that the Feels-Worth-It mechanic only works if every value opinion is paired with the street price at the moment it was said, and any opinion that can't be paired to a price is discarded. That's the fairness guarantee: a '$X is a steal!' said when the set was cheap can't distort today's value score after the price changed. The code does none of this. It simply nudges Smart Price toward the blended Feels-Worth-It value, capped at plus-or-minus 0.75. No opinion is matched to the price-at-the-time, nothing is discarded for being unpaired, and the only time input anywhere is a rough 'how many months old.'
The Millennium Falcon (75192). Two years ago, on a deep sale, dozens of people said 'an absolute steal at $480 — worth every penny.' Today it's back to $850 and harder to find. Your spec says: those 'steal' opinions were paired to the $480 price, so against today's $850 they should be reweighted or dropped — they were about a different deal. The code can't tell the difference. It blends all those old 'worth it!' votes in at face value and nudges today's Smart Price upward, making an $850 set look like a better value than it currently is.
It quietly breaks the central fairness promise of your headline mechanic. A set can show a flattering value score that's really an echo of a long-gone sale price — and 'Feels Worth It' is the feature you'd most want to be able to defend as honest, since it's the human-sentiment soul of the tool.
Either implement the pairing — attach the price-at-the-time to each value opinion, match it against the reconstructed sold-price history, and drop unpaired opinions before the nudge — or explicitly write into the spec that pairing is descoped, so the gap is honest rather than a silent broken promise.
Your spec has anti-double-counting rules so the same quality isn't rewarded twice. Rule 3 says generic 'element variety' (how many distinct kinds of pieces there are) should ONLY count toward Parts for Custom Building. But the What You Get score computes a 'variety ratio' from distinct elements over total parts and folds it in — and Parts for Custom Building has NO fact-based score at all. So variety is scored in What You Get and nowhere else: the exact inverse of the rule. A related rule (the Builder lens should up-weight only one part-value channel and hold Smart Price at its default) also isn't encoded.
A Pick-a-Brick-style parts pack, mostly a big mixed assortment for custom builders. Its strength is variety. Under your spec, that variety should light up Parts for Custom Building and stay out of What You Get. In the code, Parts for Custom Building produces no score, while What You Get quietly rewards the variety instead. So the one cluster that exists to celebrate a parts pack stays blank, and a different cluster takes the credit — the set is graded through the wrong lens, and the spec's promise that 'each quality is counted once, in its proper home' is silently violated.
When a quality is scored in the wrong cluster, a whole category of sets (parts packs, custom-builder sets) gets mis-graded in a structural way — and because it's a 'rule that isn't enforced,' nothing flags it. You could confidently publish a grade whose internal breakdown is simply pointing at the wrong strengths.
Make the anti-double-counting rules testable: add a test asserting generic variety contributes ONLY to Parts for Custom Building and novelty ONLY to Rare & New, then fix the inversion — give Parts for Custom Building its own fact score for variety and stop consuming distinct-element variety inside What You Get.
When a cluster has no data and drops out, its weight has to be redistributed to the survivors. Your spec says to prefer 'within-tier first': a dropped HIGH-importance cluster's weight should go to the other surviving HIGH clusters before any leaks down to MED clusters. The code redistributes strictly proportionally across ALL survivors with no awareness of HIGH versus MED tiers. It's faithful to the simpler 'proportional' clause but breaks the tier-preference clause.
Say the Display Appeal cluster (a HIGH-importance cluster) has no usable data and drops out on some set. Your spec says its weight should flow first to the other HIGH clusters. Instead, the code spreads it proportionally everywhere — so some of that weight leaks into MED clusters like 'hobbyist' or 'rare & new' that it wasn't supposed to reach first. The final number shifts slightly, and the importance hierarchy you designed gets blurred.
It's a quieter distortion than the gates, but it means a set with a missing major cluster gets weighted in a way your spec didn't intend — the things you decided matter most don't reliably hold their priority when data is incomplete, so the grade subtly drifts from your design.
Make the cap-and-redistribute step tier-aware: redistribute a dropped HIGH cluster's weight within surviving HIGH clusters before spilling any to MED — or amend the spec to drop the within-tier-first clause if you decide proportional is good enough.
Your spec has a lifecycle rule: opinions dated BEFORE the set was actually buyable (reveal-hype from photos and press) should be suppressed or discounted for value and build-experience dimensions, while Display Appeal and Fidelity may still use them (optionally discounted) — and the gap between 'reveal hype' and 'in-hand reaction' should be surfaced. None of this exists. Nothing in the pipeline carries an availability date or applies this scoped suppression. Reveal-hype reviews feed Build Fun and Feels-Worth-It with no filtering.
A hyped set like the Lord of the Rings Rivendell gets a flood of 'this looks INCREDIBLE, instant buy' reactions the week it's revealed — before anyone has touched it. Your spec says those should be discounted for Build Fun and value (you can't know a build is fun from a photo) but allowed for Display Appeal. The code can't tell reveal-day hype from a real in-hand review, so the pre-release excitement inflates Build Fun and Feels-Worth-It just as much as Display Appeal.
It means a set's grade can ride a wave of pre-release excitement that has nothing to do with the actual building or value experience — and you'd be publishing a Build Fun or value score partly built on hype from people who hadn't opened the box.
Either implement the lifecycle gate — carry an availability date, suppress/discount pre-availability sentiment for value and build-experience while allowing it (optionally discounted) for Display Appeal and Fidelity, and surface the reveal-vs-in-hand delta — or move it to an explicitly-deferred list in the spec.
Your spec describes the community 'worth the money' input as a thumbs (a simple yes/no, like a thumbs-up). But the data contract types it as a plain number, and the community store pushes it straight into the blend as if it were already a 0-to-10 score. A raw thumbs value — like 0 or 1 — entering a 0-to-10 blend is silently wrong. Brickset star ratings get scaled (times 2); there's no equivalent scaling defined for a thumbs.
A community member gives a thumbs-up on 'worth the money,' which arrives as the number 1. The engine treats that 1 as a value score of 1 out of 10 — a near-bottom score — and uses it to nudge Smart Price DOWN, when the person was actually saying 'yes, worth it.' One mis-scaled thumbs and the value signal points the wrong way entirely.
A signal that's supposed to push the grade up can push it down because nobody defined how a thumbs becomes a 0-to-10 number. That's a quiet data-corruption bug that can flip the direction of a value nudge on a published grade.
Define and apply a clear thumbs-to-0-10 mapping for worth-the-money (for example, thumbs-up maps to a defined high anchor, thumbs-down to a low one) before it enters the blend, the same way star ratings are scaled.
Any dimension defined by a hard fact gets a confidence (how sure the engine is, on a 0-to-1 scale) of at least 0.8, no matter what. But Smart Price has a fallback chain: it prefers real per-part market value, falls back to weight, and as a last resort uses raw piece count. Your spec says falling back to a weaker basis should LOWER confidence. The code holds it at the 0.8 floor regardless, so a Smart Price built on a crude piece-count guess reports the same confidence as one built on full part-value data.
Two sets. For the AT-AT (75313) you have real BrickLink part values, so Smart Price is well-grounded — 0.8-plus confidence is fair. For some obscure older set you have no part data and no weight, so Smart Price falls back to pure piece count — a much rougher guess. Your spec says that second one should report LOWER confidence to signal 'we're estimating.' Instead both report at least 0.8, so the rough piece-count estimate looks just as trustworthy as the solid one.
Confidence is the dial that tells you and your readers how much to trust a score. If it stays high while the underlying data quietly degrades to a piece-count guess, the tool is telling people 'we're sure' when it isn't — which is exactly the kind of false confidence that erodes trust when a grade turns out shaky.
Replace the flat 0.8 fact-confidence floor with a basis-aware confidence: full confidence for part-value, stepped-down for the weight and piece-count fallbacks, per the spec's rule that disclosed fallback to weaker bases lowers confidence.
Your spec says a licensing premium (the extra cost of a Star Wars or Marvel set because of the license) should be surfaced explicitly as a note, NOT silently held against the set's value score. The Smart Price code computes a value-per-dollar ratio and ranks it, but never computes or attaches any 'licensed premium' note. So a licensed set's naturally lower value ratio just drags its score down with no explanation — the exact silent penalty the spec forbids.
Compare a licensed Star Wars set and an unlicensed Creator set of the same piece count and price. The Star Wars set gives you fewer pieces per dollar because part of the price is the license. The code sees only the lower value ratio and quietly scores its Smart Price lower. Your spec says: surface a note like 'value reflects a licensing premium' so the reader understands WHY, rather than just seeing a dinged value score with no context.
It's the smallest of these issues, but it can make every licensed set look like a worse value than it fairly is, with no explanation — and licensed sets are a huge slice of what your audience cares about, so a quiet, unexplained value penalty across all of them chips at the grades' credibility.
Compute and carry a 'licensed premium' note on the result so the lower value ratio is explained rather than silently penalized — or, if you're deferring it, list it as explicitly deferred in the spec.
Every claim in this section checks out against the actual code. Verified directly: the three gates LOW_DATA_FLOOR (5), MIN_STRUCTURED_COUNT (8), and MIN_COMMUNITY_TO_MOVE (5) are defined in constants.ts:140-143 but appear ONLY in comments elsewhere (dimension.ts:20 and brickset.ts:58) and are never enforced; the only enforced opinion gate is MIN_VOLUME_FOR_SENTIMENT (3) at dimension.ts:47. assemble.ts mergeInputs concatenates community/brickset/mined signals with zero count-gating, confirming nothing is gated upstream. The Feels-Worth-It code (score.ts:51-57) is a plain capped nudge with no price-pairing or discarding. whatYouGetScore (factScores.ts:50) does compute varietyRatio from distinctElements/totalParts, and there is no partsForBuilding fact score in factScores.ts — confirming the inversion. capWeights (rollup.ts:47-79) renormalizes proportionally with no HIGH/MED tier logic. contracts.ts:64 types worthTheMoney as number and fakes.ts:49 pushes it in directly. dimension.ts:39 uses Math.max(0.8, sigConf) regardless of basis, while smartPrice.ts valueRatio has the partValue→weight→pieces fallback. No availability-date or licensed-premium handling exists anywhere in the pipeline or smartPrice.ts (grep returned nothing). One small nuance worth noting for precision: the audit's high-level summary lists 'MIN_VOLUME_FOR_SENTIMENT (3)' as the enforced gate and frames LOW_DATA_FLOOR as entirely unenforced — both true — but the dimension.ts:47 check ALSO returns 'insufficient' when blendSignals returns null (no signals at all), so a dimension with zero opinions is still correctly handled; the gap is specifically the 3-to-4 opinion 'thin but not empty' band, exactly as the audit states.
Audit called it: Discovery and set-matching
Before The Drop Score can grade a set, it has to find reviews of that exact set and be sure each review really is about it. This part of the engine scans your allowlist of trusted reviewer channels, looks at every video's title and description, and tries to decide "this video is about set 75355" versus "this video just mentions a Star Wars ship in passing." It tags each guess with a confidence level (high, medium, or low) so a later AI check can double-check the shaky ones before any review is used in a grade.
The matcher is genuinely well-built for a first pass, and the audit says so: a set's number in a video title is treated as the gold-standard match, names have to share at least two distinctive words to count (so a generic word like "Falcon" alone can't trigger a false match), and when two near-identical sets both fit, it refuses to guess and hands the decision to the AI. But the audit's core complaint is that the dials are set wrong for a catalog this big (~24,000 sets, many sharing names). The live pipeline accepts a match on a SINGLE weak clue — one bare number buried in a description, or one fuzzy name match — and writes it straight to the review pile, with the AI judge as the only thing standing between that weak guess and your published grade. Your own written spec explicitly says to require two agreeing clues before auto-accepting, and the code doesn't. On the flip side, the name-matching is too strict in a different way: it demands every word and treats "X-Wing" (with the hyphen) as a single glued-together word, so a video titled "X Wing" (no hyphen) silently fails to match, and dropping one word of a long official name kills the match entirely. So it both lets in too much junk AND misses real reviews — the worst of both.
The matching logic is thoughtfully engineered and its instincts are right — number-in-title is authoritative, names need real distinctive overlap, ambiguous twins are deferred not guessed — but for a 24,000-set catalog it lets single weak clues through (against your own two-clue spec rule) while a brittle name-matcher silently drops real reviews, so it leans too hard on the one AI judge as its safety net.
Each problem below is laid out the same way: what's happening, a concrete LEGO example, why it dents trust, and what the fix is.
Your written spec (the design doc, section 5.1) says: require two agreeing clues before auto-accepting that a review is about a set. The live code doesn't do that. A 'clue' here means one piece of evidence — a set number found somewhere, or a fuzzy name match. In the production matcher, a name match found in a title that fits exactly one set is accepted on its own at MEDIUM confidence, and a bare set number found only in a video's description is accepted on its own at LOW confidence. Both get written straight into the review pile (the 'review_index') with nothing backing them up except one later AI call. So a lone clue auto-accepts, which is exactly what the spec forbids.
Say a reviewer posts a 'February 2024 LEGO haul' video and, in the description, types the number 10497 (the Galaxy Explorer) just once among a list of ten sets — maybe it's even a typo or a set they only held up for two seconds. The matcher finds '10497', sees it's a real set number, and files a LOW-confidence claim that this video reviews the Galaxy Explorer. Nothing else in the video supports that. The only thing that can catch the mistake is one AI judge call. If that call slips (and across thousands of these it sometimes will), a two-second cameo becomes 'evidence' feeding the Galaxy Explorer's grade. What should happen: per your spec, a lone description-number with no second agreeing clue (like the set's name also appearing, or a Brick Insights confirmation) should NOT be auto-accepted as a real review — it should be held back or dropped, not queued as if it were solid.
This is the difference between a grade built on real reviews and a grade quietly padded with drive-by mentions. If a published Drop Score leans on a review that was never really about that set, the letter and the 0-10 number look authoritative but rest on sand — and you'd have no way of knowing which grades are affected.
Enforce the two-clue rule you already wrote down: don't accept a lone low-confidence description-number, or a lone fuzzy name match, unless a second clue agrees (the name also appears, the number is in the title, or Brick Insights independently links that video to that set).
When the engine pulls up the reviews it has on file for a set, it returns ALL of them regardless of confidence — high, medium, and low all come back mixed together (loadReviewRefs). Then the step that picks which reviews to actually analyze sorts them best-first but never throws the low-confidence ones out, and there's no limit on how many weak 'low' guesses a single set can accumulate. The audit estimates that across ~24,000 sets with roughly a quarter of set names overlapping each other, this produces thousands of bogus pairings — each one resting entirely on a single AI call to catch it.
Imagine three different sets all named 'Police Station' (this really happens in the catalog — there are multiple). A reviewer's 'best LEGO Police Stations ranked' video has all three numbers in its description. The matcher dutifully files a LOW-confidence claim linking that one video to all three Police Stations. Now multiply that by every haul, every ranking video, every 'my whole collection' tour. A single set can pick up a dozen of these flimsy links, and every last one gets shipped to the AI judge as if it were worth checking. What should happen: low-confidence guesses should be filtered out or capped — say, no more than a couple of 'low' guesses per set — so one chatty haul video can't inject a dozen junk pairings into the judge's workload and, eventually, into the grade.
Two ways this erodes trust. First, cost and reliability: flooding the AI judge with junk burns money and makes the judge the single point of failure for the whole catalog's accuracy. Second, every junk pairing the judge fails to reject is a fake review silently strengthening (or dragging down) a grade you publish, and you can't see it happening.
Filter the reviews-to-analyze list by confidence and cap how many LOW-confidence ones any single set can carry, so no one haul video can dump dozens of weak pairings into the judge queue.
When the matcher breaks a set name into words, it keeps hyphens stuck inside the word. So the set name 'X-wing Starfighter' becomes the two words 'x-wing' and 'starfighter'. It also requires EVERY word of the name to appear in the title. The problem: a video titled 'LEGO X Wing Starfighter review' breaks into 'x', 'wing', 'starfighter' — there is no glued 'x-wing' token, so the required word 'x-wing' is missing and the match fails. I confirmed this directly: both 'X Wing' and 'XWing' return nothing when matched against the set 'X-wing Starfighter'.
A reviewer titles their video 'LEGO X Wing Starfighter — is it worth it?' — no hyphen, which is incredibly common since people type it both ways. The set in your catalog is stored as 'X-wing Starfighter'. The matcher looks for the exact glued word 'x-wing' in the title, doesn't find it (the title has separate 'x' and 'wing'), and declares no match. A perfectly good, on-topic review is thrown away. What should happen: hyphens should be treated as spaces on both sides, so 'X-wing' and 'X Wing' and 'XWing' all line up and the review is found.
This makes coverage look thinner than it really is. A set might have ten great reviews on YouTube, but if reviewers wrote the name without a hyphen, the engine finds zero of them by name and may grade the set as 'data-light' or skip it — a confidently incomplete grade, when the data was there all along.
Normalize hyphens to spaces before matching, and switch from 'every word must appear' to 'most words must appear' (an N-of-M threshold). Add tests for punctuation and spelling variants.
Because the matcher demands that EVERY distinctive word of a set's name appear in the title, long licensed names are fragile. If a reviewer shortens the name by even one word — which everyone does — the match fails. I confirmed this: the set 'Luke Skywalker's X-wing Fighter' does not match a title 'Luke's X-wing Fighter build', because the title drops 'skywalker' and the matcher insists on it.
The set is officially 'Luke Skywalker's X-wing Fighter'. A reviewer titles their video 'Luke's X-wing Fighter — full build' (totally natural — nobody says the full name every time). The matcher's distinctive words for the set are 'luke', 'skywalker', 'x-wing', 'fighter', and it requires all four. The title has 'luke', 'x-wing', 'fighter' but not 'skywalker', so the match fails and the review is lost. What should happen: matching three of the four distinctive words should be plenty to confidently link this video to this set.
Long, licensed names (Star Wars, Harry Potter, Marvel) are exactly the sets people review the most, and they're the ones this hits hardest. You could end up with thin or missing grades on your highest-profile sets, which are the grades your audience scrutinizes most.
Use a 'most words match' threshold instead of 'all words match', so a dropped word or two doesn't sink an otherwise-clear match.
There are two matching functions. The one the bulk pipeline actually uses ('matchVideoToSets', plural) has the year-guard that ignores stray years like 2023. The other one ('matchVideoToSet', singular) instead has a rule that bails out if it sees any 5-or-more-digit number in the title, and has no year-guard at all. They disagree about how to handle numbers. The singular one is dead code — nothing in the live pipeline calls it — yet it's exported and covered by tests, so it looks like a shipping, trustworthy piece of the engine when it isn't.
Feed both matchers a title like 'LEGO X-wing review filmed in 2023'. The live matcher (plural) correctly ignores '2023' as a year and proceeds to name-match. The dead matcher (singular) has no year-guard, so '2023' isn't specially handled, and its number-rules behave differently. Anyone reading the tests, or reusing the singular function thinking it's production-grade, would inherit behavior the real pipeline abandoned. What should happen: there should be one matcher, with one agreed set of number-and-year rules, and the dead one should be deleted or clearly marked non-production.
Two matchers that disagree is a trap waiting to spring — a future change wired to the wrong one would silently grade sets by rules you thought you'd retired, and the passing tests would give false reassurance that everything's fine.
Merge the two matchers into one, reconcile their year and number rules, and delete or clearly label the unused one as not-for-production.
To avoid mistaking a year for a set number, the matcher ignores any 4-digit number starting with '19' or '20'. The catch: some real LEGO sets are literally numbered 2000, 2025, and so on, and they exist in the catalog. So those sets can never be matched by their number — the guard that protects against years also blinds the engine to them. I confirmed the guard treats both '2025' and '2000' as years.
There's a real classic set numbered 2025. A reviewer posts 'LEGO 2025 unboxing'. The matcher sees '2025', assumes it's a year, and refuses to treat it as a set number — so it finds nothing, even though '2025' IS the set. What should happen: instead of a blanket 'ignore anything that looks like a year', it should use context — e.g. require a 'set' or '#' prefix, or require the set's name to also appear — so a genuine year-numbered set can still be matched.
It's a small set of sets, but for those sets the engine is blind by number — they'd only ever get matched by name, if at all, quietly under-covering a slice of the catalog with no warning.
Disambiguate year-like numbers by context (a prefix, or the set name co-occurring) rather than dropping every 19xx/20xx number outright.
Many sets have variants (a regular release, a polybag, a promo) that share a base number but have slightly different names. The production catalog-loader ('loadCatalog') picks ONE name per base set by sorting set numbers alphabetically and taking the first. A different function ('catalogRowsToKnownSets') instead picks the variant with the most parts — usually the 'main' set with the proper name. The live sweep uses the first one, so a set can end up represented by an oddly-named polybag or promo variant instead of its real name.
Suppose base number 30000 has two variants: a 500-piece main set named 'Medieval Market Village' and a 12-piece promo polybag named 'Knight Minifigure Pack'. The 'most parts wins' function would correctly use 'Medieval Market Village'. But the production loader sorts by set number and might grab the polybag row first, so the matcher tries to match reviews against the name 'Knight Minifigure Pack'. Reviews of the actual Medieval Market Village would fail to name-match. What should happen: both code paths should agree and use the richest (most-parts) variant's name, so the set is always disambiguated against its real, recognizable name.
If a set is represented internally by the wrong (minor-variant) name, its real reviews won't name-match, and it'll look under-reviewed for a reason that has nothing to do with how popular it actually is — a silent, hard-to-trace gap in coverage.
Make loadCatalog choose the same canonical (most-parts) variant that catalogRowsToKnownSets does, so both paths agree on each set's name and year.
The check that rejects reviews published before a set existed only compares YEARS, not months. So a video from January 2023 passes for a set that didn't ship until December 2023. And if a set has no recorded year, the check is switched off entirely, offering no protection at all.
A set releases in December 2023. A reviewer's January 2023 video — eleven months before the set existed — happens to name-match it. The year-only check sees '2023 >= 2023' and lets it through, even though the set wasn't out when the video was made. Worse, if two near-identical twins differ only by a same-year reissue, this check can't tell them apart at all. What should happen: comparing the actual release date (year and month) would catch the eleven-month-early video and help separate same-year twins.
It's a narrow gap, but it lets through exactly the kind of 'this review predates the set' mismatch the check exists to stop, which can attach a wrong-set review to a grade in the tricky twin cases.
Compare full release dates (year and month) instead of just the year, and decide deliberately what to do when a set's date is missing rather than silently disabling the guard.
The name-matcher strips accent marks and requires at least two distinctive words. Any set whose distinctive name collapses to fewer than two usable words after that stripping is dropped from name-matching entirely — it simply can't be found by name. Short names and non-Latin-script (e.g. Cyrillic) names are the casualties. I confirmed a Cyrillic-only name reduces to zero usable words.
A short-named or foreign-market set — say one whose distinctive name is a single word, or written in Cyrillic — reduces to fewer than two matchable words. The matcher drops it from name-based discovery completely. It can still be found if its number shows up in a title, but a video that only says the name will never link to it. What should happen: such sets should at least have a fallback (single strong token plus year, or transliteration) rather than zero name-based discovery.
It's a small corner of the catalog, but those sets get systematically thinner coverage for a reason invisible in the final grade — nothing flags that they were never discoverable by name.
Add a fallback for short or non-Latin names (transliterate, or allow a single strong token plus a corroborating clue) so these sets aren't silently undiscoverable by name.
Brick Insights is a site that already links sets to their reviews — exactly the kind of independent second clue your spec wants. But the code stamps every link it harvests from there as 'medium' confidence no matter how good or weak the match is, and it fetches Brick Insights one set at a time on demand, with no caching and no backoff (no automatic slow-down when the site pushes back). So the 'free pre-matched set-to-review map' your spec describes is never swept in bulk, and the on-demand fetching is fragile at scale.
You grade 500 sets in a sweep. Brick Insights could have provided a solid second clue for many of them — the exact corroboration the two-clue rule needs. Instead, each set triggers its own uncached fetch, every harvested link is flat-rated 'medium' whether it's a perfect match or a loose one, and a burst of requests risks getting throttled with no backoff to recover. What should happen: Brick Insights pages should be cached and fetched politely with backoff, and used as the corroborating second clue — turning a fragile afterthought into the safety signal the spec intended.
This is the missing piece that could fix the biggest problem (the single-clue auto-accept). Leaving it fragile and uncached means the engine keeps leaning on the lone AI judge instead of the independent corroboration it was designed to have.
Cache Brick Insights pages, add backoff, and wire it in as the corroborating second clue the two-signal rule needs.
Every audit claim I could check held up against the actual code. Confirmed directly: (1) the single-signal auto-accept — youtube.ts line 159 tags a unique title-name match MEDIUM and line 151 tags a lone description-number LOW, both written to review_index, while the spec (docs/.../drop-score-data-sourcing-design.md section 5.1) does say verbatim 'Require two agreeing signals before auto-accept.' (2) loadReviewRefs (supabaseStore.ts 243-250) returns all tiers unfiltered, and selectRefsToMine (mine.ts 31-42) only ranks by confidence without dropping or capping LOW refs. (3) I ran the actual tokenizer: 'X Wing' and 'XWing' both fail to match 'X-wing Starfighter', and 'Luke's X-wing Fighter build' fails to match 'Luke Skywalker's X-wing Fighter' — both exactly as claimed. (4) The two matchers do diverge: matchVideoToSet (singular, youtube.ts 76-97) has a length>=5 block and no year-guard; matchVideoToSets (plural, 140-162) has isYear and no length block; and grep confirms the singular one is referenced ONLY in tests, never in the live pipeline. (5) isYear treats '2025' and '2000' as years, confirmed by running it. (6) loadCatalog (supabaseStore.ts 294, first-by-set_number-sort) and catalogRowsToKnownSets (catalog.ts 92, richest-by-num_parts) do pick different variants. (7) Brick Insights refs are hardcoded confidence:'medium' (brickInsights.ts 141) and harvested one fetch per set with no caching/backoff (158-161). One thing to flag as the audit's estimate rather than verified fact: the '24% name overlap' and 'thousands of spurious pairs' figures are the auditor's projection across the catalog — I did not independently recompute them, though the repo's own handoff doc corroborates the catalog size (27,080 sets / 24,644 bases) and the real existence of multiple identically-named sets (three 'X-Wing Starfighter' sets), which makes the overlap concern credible.
Audit called it: LLM judge and distillation
This is the part of The Drop Score that reads what real reviewers said about a set — their YouTube video transcripts, the top comments under those videos, and blog/forum write-ups — and squeezes each one down into eight numbers from 0 to 10 (build fun, instructions, sturdiness, looks, accuracy, playability, parts value, and worth-the-money). An AI (Claude) does the reading. It's smart enough to tell sarcasm from praise ("a steal at $0.10 a piece" is an insult, not a compliment) and to throw out a review that's actually about a different same-named set. Those eight numbers are what eventually feed into the letter grade — the original text is read once and never stored.
The design here is genuinely thoughtful, and the audit says so: the "is this review even about the right set?" check works, the sarcasm awareness is real, only the derived scores get kept (not the text), and one bad review can't crash the whole set. But the auditor's headline is "sound design, missing guardrails," and that's fair. The big problems are all about trust in the numbers: the AI's scores are never sanity-checked before they enter the grade, a score outside 0-10 wouldn't get caught, and the AI is never asked how sure it is — so a wild guess and a rock-solid read count exactly the same. There are also waste-and-noise issues: a loud brigade of angry comments reads like genuine consensus, the same review gets re-read (and re-billed) every single time you re-grade a set, and a momentary network hiccup silently drops a review instead of retrying. None of these are fatal, but several can quietly make a published grade look more confident, or more negative, than the evidence actually supports.
The AI-reads-reviews engine is well-conceived and handles the genuinely hard stuff (sarcasm, wrong-set detection) gracefully, but it ships the AI's raw output straight into your grades with no validation, no clamping, and no confidence — so the moment the AI hiccups, a wrong number flows through wearing a straight face.
Each problem below is laid out the same way: what's happening, a concrete LEGO example, why it dents trust, and what the fix is.
A 'validation gate' is a checkpoint that inspects the AI's output and refuses to let bad or suspicious numbers through. Your own spec (section 8) calls for one. The code doesn't have it. The function that turns the AI's reading into grade-feeding scores (distill.ts, line 36) takes whatever the AI returned and passes it straight on — no checking that the numbers make sense, that the evidence quotes actually exist, or that nothing looks off. Whatever the AI says, the grade believes.
Say you re-grade the Razor Crest. The AI reads a long, messy transcript and, on a bad run, returns build-fun as the text '9' instead of the number 9, plus an evidence quote that's blank. With no gate, the code happily accepts it. Downstream math may read the text '9' as zero (or choke), and now a beloved build is being dragged down by a glitch nobody caught. A validation gate should have looked at that output, seen 'this isn't a clean number with real evidence,' and dropped that one dimension — grading the Razor Crest from the reviews that came back clean, exactly like it does when a review fails to load.
This is the trust keystone. You're publishing letter grades as if a careful judge stood behind every number. Without a gate, the AI's worst moments — a malformed answer, a hallucinated quote, a fluke — flow into the grade with the same authority as its best work. A grade can look confident and be built partly on garbage, and you'd have no way to know which.
Add the validation gate the spec already describes: after the AI returns, check each dimension is a real number in range with non-empty evidence, and silently drop (not zero out) any dimension that fails — so a bad field is treated like a review that never came, rather than a fake data point.
'Clamping' means forcing a number back inside its allowed range — anything above 10 becomes 10, anything below 0 becomes 0. The eight scores are supposed to live on a 0-to-10 scale, but the code never clamps them before doing math with them (distill.ts, line 145, just hands the AI's raw numbers along). The AI is asked for 0-10, but nothing enforces it.
Imagine a glowing review of the Millennium Falcon UCS. The AI, over-enthusiastic, returns displayAppeal as 14 instead of capping at 10. Nothing trims it. That 14 then gets averaged in with other reviews' scores — so one review effectively votes one-and-a-half times on looks, quietly inflating the Falcon's display score above what any reviewer actually expressed. Clamping should have pulled that 14 down to 10 before it ever touched the average.
An out-of-range number doesn't just add one bad data point — it distorts every average it's mixed into, and it does so invisibly. A set could end up with a display or value score that no individual review supports, making the published grade subtly wrong in a way that's almost impossible to spot after the fact.
Clamp every score to 0-10 the instant it comes back from the AI, before any averaging or blending happens. It's a one-line guard that closes the hole.
'Confidence' here means a per-score measure of how strongly the review actually supported that number — was it stated outright, or half-inferred from a vague aside? The AI is only asked for a score and a short quote (distill.ts, lines 52-60); it's never asked for confidence. 'Weighting' means letting stronger evidence count for more. Without a confidence number, weighting is impossible — every score lands with identical force.
Two reviewers cover the Razor Crest. One spends three minutes raving in detail about how clever the build is — clearly a 9. The other mutters 'yeah, build was fine I guess' in passing — the AI also lands on a 7-ish, but it's basically a shrug. Right now both get blended in as equal-strength votes. The strong, detailed 9 and the offhand 7 pull on the grade with the same weight. With confidence captured, the detailed read should count for more and the throwaway aside for less, so the build-fun score reflects the reviewer who actually engaged.
It flattens the difference between 'the evidence is overwhelming' and 'the AI kind of guessed.' Sets where reviewers spoke clearly and sets where the AI scraped together a number from vague hints get presented with identical confidence. A grade built largely on shrugs looks just as authoritative as one built on detailed, passionate reviews — and you can't tell them apart.
Add a confidence field to what the AI returns for each dimension, then let the aggregation step lean on high-confidence scores more than low-confidence ones.
When the code pulls the top 30 comments off a YouTube video, each comment actually arrives with its like count attached. But the code keeps only the comment text and drops the likes (mine.ts, line 77 maps each comment down to just its words). A 'brigade' is a coordinated burst of similar angry comments. Without like counts, the AI can't tell one upvoted-500-times comment from thirty copies of the same complaint posted by a handful of people.
Take a polarizing set — say a $200 Botanicals centerpiece some fans felt was overpriced. Under one video, a small group floods 25 near-identical 'overpriced cash grab' comments, each with 2 likes, while the single most-liked comment ('worth every penny, gorgeous on a shelf' — 800 likes) sits there outvoted in raw count. Stripped of likes, the AI sees 25 angry voices versus 1 happy one and reads the room as 'people hate the value.' The worth-the-money score sinks. With likes kept, that 800-like endorsement would rightly outweigh the brigade, and the value score would land where the actual audience landed.
It lets a vocal minority hijack a grade. A set could be marked down on value or playability because a few determined commenters made noise, while the genuine majority — visible only through likes you discarded — gets ignored. The published grade then misrepresents what the community actually thinks.
Pass each comment's like count through to the AI (and weight comments by it), so a heavily-upvoted comment counts for more than a swarm of ignored ones.
A 'cache' is a saved copy of work you already did, so you don't redo it. There's no cache for a review's derived scores (mine.ts, line 74). So every time you re-grade a set, the system re-fetches each transcript and pays the AI to re-read it from the top — even though that exact review hasn't changed since last time.
You grade the Razor Crest today off five reviews. Tomorrow you tweak how the final letter grade is calculated and re-grade. Nothing about those five reviews changed overnight — but the system fetches all five transcripts again and pays Claude to re-read all five, just to land on the same eight numbers it already had. Multiply that across hundreds of sets and many grade-formula tweaks, and you're re-buying the same reading over and over. A cache would let it reuse yesterday's scores and only read genuinely new reviews.
This one's about cost and speed more than correctness, but it indirectly hurts trust: when re-grading is slow and expensive, you re-grade less often, so published grades drift out of date as new reviews land. A cache keeps grades cheap to refresh, which keeps them current.
Save each review's eight derived scores keyed to that review, and reuse them on re-grade unless the review itself is new or changed.
A 'retry' is automatically trying again after a failure; a 'timeout' is a deadline after which you stop waiting and give up. The loop that reads each review has neither (mine.ts, line 121). The safety net around it swallows any error quietly — good for not crashing, but it means a review that failed for a purely temporary reason (a one-second network blip) is discarded exactly like a review that's genuinely broken, with no second attempt and nothing logged.
You grade the Millennium Falcon off four strong reviews. As the system reads the third — the most detailed, highest-confidence one — YouTube's transcript service hiccups for a moment and the request fails. No retry. That review is silently dropped, and the Falcon gets graded off the remaining three. A simple 'try once more' would almost certainly have recovered the best review of the set, and the grade would rest on four voices instead of three — with no one ever knowing the strongest one nearly vanished.
Transient failures are random, so this quietly and unpredictably thins the evidence behind a grade. A set might be graded off three reviews instead of five purely because of bad luck with timing — and because nothing is logged, you can't even see that it happened or that the grade is thinner than it should be.
Add a short retry with a timeout around each review read, and log when a review is ultimately dropped so thin-evidence grades are visible rather than silent.
The code calls Claude Opus (the most capable, most expensive model) but runs it on 'low effort' — a budget setting on a premium model (distill.ts, line 135). It also doesn't use 'prompt caching,' a feature that stores the big fixed instructions once so you're not re-billed for re-sending the same several-hundred-word rulebook on every review.
Every review you distill re-sends that long sarcasm-aware instruction block from scratch and pays for it each time. Reading 500 reviews means paying to process those identical instructions 500 times over, on top of premium Opus pricing. A cheaper model with prompt caching could do this same straightforward extraction job at a fraction of the cost — the instructions get billed once and the per-review work is tiny.
Pure economics — it doesn't make any single grade wrong. But high running costs throttle how much you can grade and how often you can refresh, which over time leaves the catalog less complete and less current than it could be.
Either step down to a cheaper model for this extraction task or turn on prompt caching for the fixed instructions (ideally both), so reading lots of reviews stays affordable.
The 'does this review cover the right set?' decision is a plain yes/no (distill.ts, line 62) — true or false, nothing in between. When the AI is only 55% sure a borderline video is really about your set, that hesitation is thrown away; it's recorded as a confident 'yes' (or 'no') with no trace of the doubt.
A video titled 'Star Wars haul' spends two minutes on the new X-Wing among five other sets. The AI guesses 'yes, this covers it' but is genuinely on the fence. That shaky 'yes' is treated identically to a dedicated 20-minute X-Wing deep-dive that's an obvious 'yes.' Its scores pour into the grade at full strength. If the match confidence were kept, that borderline inclusion could count for less, or be flagged for a closer look, instead of carrying the same weight as a sure thing.
Low severity because the yes/no usually lands right — but on borderline videos, a coin-flip inclusion silently gets full voting power. A grade can lean on a review the system itself was only half-sure belonged, and that uncertainty never surfaces anywhere.
Have the AI report how confident it is that the review covers the target set, and let shaky matches count for less than clear ones.
For non-YouTube reviews, the code fetches a web page and pulls out its text. The only test before trusting that text is a length floor — it must be at least 200 characters (article.ts, line 38). There's no check that what came back is actually a review of the set rather than, say, a cookie notice, a paywall message, or an unrelated article.
The system follows a blog link for a review of a Botanicals set, but the page has been replaced with a 600-character 'Subscribe to read this article' wall plus navigation junk. It clears the 200-character bar, so it's handed to the AI as if it were the review. The AI tries to score a build it never actually read about — and whatever it invents flows into the grade. A real check would notice 'this doesn't read like a review of this set' and skip it, the way a failed fetch gets skipped.
Low severity because it only affects blog/forum sources and most pages are fine — but when it does misfire, the AI scores boilerplate or a paywall as though it were a genuine opinion. That's a fabricated data point dressed up as a real review, quietly nudging a grade off-true.
Beyond the length floor, add a light sanity check that the fetched text actually looks like a review of the target set before sending it to the AI — and lean on the validation gate above to drop scores derived from thin or off-topic text.
Every audit claim checks out against the real code. Confirmed specifically: (1) distill.ts line 36 passes the AI's parsed output straight through with no validation — no gate exists. (2) distill.ts line 145 (and the JSON.parse on line 147) returns raw scores with no clamping to 0-10. (3) The dimension schema at distill.ts lines 52-60 has only 'score' and 'evidence' — no confidence field. (4) mine.ts line 77 maps comments to '.text' only; I verified in youtube.ts that fetchComments returns a likeCount per comment (VideoComment interface, lines 174-176), so the like data genuinely exists and is discarded. (5) No derived-score cache anywhere in mine.ts. (6) The per-ref loop at mine.ts lines 121-128 has a bare try/catch with no retry and no timeout. (7) distill.ts line 135 uses model 'claude-opus-4-8' with effort 'low' and no caching. (8) coversSet is a plain boolean (distill.ts lines 66-69, schema). (9) article.ts line 43 gates only on text.length >= 200 (the audit cites line 38, the function start; the actual length check is line 43 — same function, trivial offset). One clarification, not a contradiction: the comment fetch limit is 30 (mine.ts line 77 passes 30 to fetchComments), which I used in the brigade example. All severities and the substance of every location match what the code does.
Audit called it: Smart Price & fact scoring
This is the part of the engine that answers "is this set worth the money?" It works out roughly what the set's loose parts would cost if you bought them piece by piece on the secondhand market (the "part-out value"), divides that by the price, and then ranks that value-for-money number against a group of comparable sets to produce a 0-10 Smart Price score. It also computes four simpler "fact" scores from plain numbers: how good the minifigure lineup is, how scarce the set is, how well it holds resale value, and how much you physically get in the box. Smart Price is the single heaviest money item in the whole grade, so getting it right matters a lot.
The plumbing here is genuinely well-engineered in the parts that do the arithmetic, but the headline promise of the feature does not actually run in real life. The methodology is emphatic that value should be measured against the "street price" — what people actually pay on the secondhand/discount market — and explicitly "not RRP" (the sticker price LEGO prints on the box). But the live data feeds never fill in a street price, so every real grade silently falls back to the box price — the exact thing the spec forbids — with no warning. On top of that, the "compare against similar sets" machinery is mostly a paper feature: those comparison groups only exist if someone has pre-priced a big library of sets first, so for almost every set Smart Price simply doesn't appear. And where it does appear, the comparison group can be as small as five sets, which is far too few to produce a trustworthy rank, and a set can accidentally be compared against itself. The four plain fact scores (figures, scarcity, resale, what-you-get), by contrast, are clean and do exactly what they claim.
The cost-estimation engine and the honesty checks around thin data are genuinely well-built, but the marquee "real market value vs. what people pay" promise is not delivered — the code quietly uses box price instead — and the comparison-group system is too sparse and too small to trust at the sizes it currently allows.
Each problem below is laid out the same way: what's happening, a concrete LEGO example, why it dents trust, and what the fix is.
The methodology says value should be measured against the 'street price' — the typical price people actually pay on the secondhand or discount market — and it underlines 'not RRP' (RRP is the recommended retail price, the official sticker price LEGO prints on the box; MSRP is the same thing). The code is written to use street price if it has one, and otherwise fall back to the box price. The problem: nothing in the live data pipeline ever fills in a street price. The one feed that supplies a price (Brickset) only hands over the official box price. So for every real-world grade, the 'street price OR box price' choice always lands on box price — the exact basis the spec forbids — and there is no flag, no warning, and no downgrade in confidence to tell anyone it happened.
Take the UCS Razor Crest, box price $600. Out in the real world it's heavily discounted and routinely sells around $450. Its parts come to a part-out value of about $540. The methodology WANTS: $540 ÷ $450 street price = 1.20 — a genuine bargain, because you're getting $540 of parts for the $450 people actually pay. What the code DOES: it never gets the $450, so it computes $540 ÷ $600 box price = 0.90 — which looks like merely average value. Same set, same parts, but the grade quietly understates the deal because it measured against a number the spec specifically told it not to use. The reverse happens too: a set that's hard to find and sells ABOVE box price would look like better value than it really is.
This is the most trust-damaging issue in the dimension. Seth would be publishing a Smart Price score under the banner of 'what the set is really worth vs. what people pay,' while the engine is actually only comparing against the box price — the one basis the methodology explicitly rejects. Every Smart Price grade is, right now, answering a different (and worse) question than the one it claims to answer, and nothing on the output admits it.
Either actually populate a street price in the live feeds — for example a median recent SOLD price for the whole set from BrickLink, or a discount/street price from a retailer or Brick Insights — so the feature does what it promises; OR change the methodology and all the public wording to honestly say Smart Price is measured against box price. The dangerous state is the current one: the spec promises one thing and the code silently does another.
To rank a set's value, the engine pulls a 'cohort' — a comparison group of similar sets — and sees where this set falls in the pack. But the query that builds that group grabs every comparable set without excluding the one being graded. So if the set you're grading has already been saved in the database, its OWN value number is sitting inside the group it's being measured against. Made worse: the ranking counts any set whose value is 'at or below' the target as below it — and a set is always 'at or below' itself. So every set gets counted as beating at least one comparable: itself.
Say you grade a 1,000-piece Star Wars set and the comparison group has 10 sets, including a saved copy of this very set. The set's true value-for-money would place it right in the middle — beating 5 of the 9 OTHER sets, a rank of about 0.56, which the curve turns into roughly a 6.3. But because its own copy is in the group AND counts as 'at or below' itself, it's now scored as beating 6 of 10, a rank of 0.60, nudging it up toward about 6.6. The effect is small in a big group but gets ugly in a small one — and the engine currently allows groups as small as five, where one self-vote is a fifth of the whole ranking.
It systematically tilts grades upward — every set looks slightly better than it is, and the smaller the comparison group, the bigger the thumb on the scale. A grade that should read 'middle of the pack' can read 'above average' purely because the set quietly voted for itself, which is exactly the kind of invisible bias that erodes trust in a published score.
Exclude the graded set's own row from its comparison group (filter it out by set number), and deliberately decide how ties are handled — for example counting a tied set as half-in, half-out — instead of the current rule where the best set always lands at a perfect 1.0.
The comparison group only contains sets that have ALREADY been priced and saved into the database. Smart Price only appears in a grade when such a group exists and has at least five members. Every group starts completely empty. So until someone goes and bulk-prices a big library of sets — broken out by theme and by size — Smart Price is dark, and the single heaviest money item in the grade (about 65% of the money score, and roughly 14% of the whole grade) just isn't there. The confident methodology language never warns about this 'cold start' gap.
You grade a brand-new 800-piece Botanicals set the week it launches. Nobody has pre-priced a library of other 800-piece Botanicals sets yet, so the comparison group is empty. Smart Price returns nothing, and the grade is computed without its biggest money component. The set might be a fantastic deal or a terrible one — the engine has no opinion, and the published grade leans entirely on the lighter factors instead, without making clear that the heavyweight money judgment is missing.
A grade that's quietly missing its most important money input can look complete and authoritative while actually being a partial answer. If Seth publishes grades before a big priced library exists, the most important 'is it worth it' signal is silently absent on most of them — and the reader has no way to tell a fully-backed grade from a hollow one.
Treat this as a real cold-start problem: batch-price a meaningful library of sets per theme-and-size group before relying on Smart Price, and/or clearly surface on the output when Smart Price had no comparison group so a grade without it is visibly distinguishable from one with it.
The minimum comparison-group size is set to five. Ranking something against only five others can only ever land on six possible positions (dead last, then 1-of-5, 2-of-5, and so on up to top). One new comparable set joining the group can jump a set's rank by a full step — and the curve that turns rank into a score is steep enough that one step is worth roughly a whole point of Smart Price. The output then presents a precise-looking 0-10 number whose underlying rank is statistically shaky. The methodology calls the method 'robust to outliers,' which is true with hundreds of sets but false with five.
A 500-piece Icons set is the 3rd-best value out of 5 — rank 0.60, which the curve turns into about 6.6, a solid C+. Then one more cheap-for-its-parts comparable gets added and now our set is 3rd-best out of 6 — rank 0.50, which maps to 6.0, a flat C. Nothing about the actual set changed; one new neighbor in the group knocked it down most of a letter grade. A score shown to one decimal place implies a precision the five-set rank simply doesn't have.
Publishing a confident 6.6 that would have been 6.0 if one more set existed makes the grade look more exact than the data supports. Two runs a week apart could disagree by a letter grade for reasons that have nothing to do with the set, which makes the number look unreliable when people notice.
Raise the minimum group size substantially (something like 15-25), and/or attach a confidence penalty that shrinks as the group shrinks, and at minimum show the group size on the output so a grade backed by five comparables is visibly different from one backed by two hundred.
When grading a set, the engine re-prices its parts live, using up-to-the-minute BrickLink prices. But every set in the comparison group is compared using whatever part value was saved the last time it was ingested, which could be months old. So the set being graded is measured with today's market, while its neighbors are frozen in the past. If part prices have drifted market-wide since those neighbors were saved, the comparison is unfair in a way nobody can see. There's no freshness check.
You grade a Modular building today; minifig and part prices have climbed about 15% across the board since last winter. The graded set's part-out value reflects the new, higher prices — say $310. But the comparison group was all priced last winter at the old, lower levels — around $270 each for similar sets. So the graded set looks unusually valuable purely because its numerator was updated and the group's wasn't. Measured fairly — everyone on today's prices — it would sit middle of the pack; measured as-is, it looks like a standout bargain and scores too high.
A grade can swing on WHEN each set happened to be priced rather than on the set itself. That's an invisible source of error that can make a perfectly ordinary set look like a deal (or a dud), and there's no way for a reader — or Seth — to know it happened.
Price the comparison group on the same basis and freshness as the graded set (re-price on read, or store a 'priced-as-of' date and refuse to use comparables that are stale relative to the live number).
The intended comparison is theme AND size — a Star Wars set against other Star Wars sets of similar size. But when there aren't enough same-theme sets, the engine drops the theme requirement and compares against ALL themes in that size bracket. Different lines have wildly different value-per-dollar patterns, so pooling them by size alone can be a category mistake. The methodology specifies theme-and-size precisely to avoid this; the fallback quietly throws the theme half away.
A 500-piece Botanicals set (mostly small, cheap, repeated petal and stem pieces — low part value for its size, because you're paying for the finished look) gets compared against a 500-piece licensed Star Wars set (loaded with pricey specialized and printed parts — high part value). Pooled together by size only, the Botanicals set looks like terrible value next to the part-dense Star Wars sets and scores low — not because it's a bad deal for a Botanicals set, but because it was ranked against a fundamentally different kind of product.
A set can get an unfairly harsh or generous money grade just because the only comparison available was a different category of LEGO. The reader sees a confident Smart Price number with no hint that it was computed against a mismatched crowd.
Make the cross-theme fallback visible on the output (a 'cross-theme comparison' flag), and consider only allowing it when the pooled themes are actually similar enough in value pattern to be comparable.
The rank is defined as 'fraction of the group at or below this set.' Because the top set is at-or-below itself, the best-value set in any group always lands at a perfect rank of 1.0, which maps to a 10.0 — no matter how slim its lead. And a set tied with the worst in the group doesn't score zero; it scores the fraction of ties. The curve was designed assuming a smooth spread of values, which doesn't hold when there are ties or a tiny group.
In a group of 8 sets, the best-value set beats the runner-up by a hair — pennies of part value per dollar. It still scores a flat 10.0, the same as a set that crushes its group. Two very different 'best deals' get the identical top grade, and the 'rip-off below 3 / bargain above 9' calibration the curve was tuned for drifts off because of how ties and small groups land.
A barely-best set wearing a perfect 10 overstates how exceptional it is, and the carefully-designed bargain/rip-off thresholds don't mean quite what they're supposed to. It's a smaller effect, but it makes the extremes of the scale less honest.
Use a balanced ranking rule (count tied sets as half, so the best set isn't automatically a perfect 1.0) so the score reflects how big the lead actually is.
When the comparison group is empty, the ranking function returns a neutral middle value that maps to a 6.0 (a C). In the normal pipeline this can't happen, because the five-set minimum blocks empty groups from reaching the scorer. But there's a secondary way to call the scorer directly that skips that gate — and through it, a set with literally no comparison data would be handed a confident 6.0. A neutral C is not the same as 'we don't have enough data'; it actively drags the money score toward mediocre instead of honestly abstaining.
Through the direct path, you score a set whose comparison group is empty. Instead of returning 'insufficient data,' it returns 6.0 — a clean C — as if it had weighed the set against peers and found it average. It weighed it against nothing. That fake-average then mixes into the money cluster and pulls the overall grade toward the middle.
A 6.0 looks like a real verdict, not a shrug. If that path is ever used, a set with no evidence behind its money score gets a confident-looking mediocre grade, which is more misleading than simply saying 'not enough data.'
Make the empty-group case return 'insufficient' and abstain (letting the rest of the grade re-balance) rather than emitting a confident 6.0.
The sourcing plan calls for a price-per-piece / Brick Insights sanity check as a cross-reference on Smart Price — a second, independent read on whether the price looks right. The Brick Insights module exists in the codebase, but nothing actually calls it for price-sanity; it's only used to find reviews. So the (already box-price-only) anchor has no outside check at all.
If a feed glitch left a set with a wildly wrong price — say a $90 set recorded as $9 — there's no second source whispering 'that price looks off.' Smart Price would just compute an absurdly good value and report it, with nothing to catch the bad input.
On its own it's minor, but it compounds the box-price problem: the price the whole feature hinges on is both the wrong basis AND unchecked by any independent source, so a bad or misleading price flows straight through to a published grade.
Wire the Brick Insights price-per-piece aggregate (or a retailer street price) in as the intended cross-check, so the price anchor has at least one independent reference.
The audit is accurate on every point I checked against the code. Confirmed specifically: (1) brickset.ts only sets msrp from set.LEGOCom.US.retailPrice (the official box price) at line 42 and never sets streetPrice; rebrickable.ts sets neither; grep shows the only real streetPrice values come from demo.ts, the test fixtures, and the Supabase cache round-trip — so smartPrice.ts:23 'facts.streetPrice ?? facts.msrp' does always resolve to box price on a live grade, exactly as claimed. (2) loadCohort (price.ts:74-81) has no self-exclusion filter, and percentile (smartPrice.ts:34) uses 'c <= value', confirming the self-inclusion-and-tie bias. (3) factInputs.ts:33 gates Smart Price on ctx.smartPriceCohort, and gradeSet.ts:127-131 only sets it at n>=5, confirming the cold-start gap. (4) MIN_COHORT_FOR_SMART_PRICE = 5 (constants.ts:145) and the anchors (constants.ts:50-56) are exactly as described — I verified the curve makes a 0.1-rank step worth close to a full point near the middle, supporting the medium severity. (5) The theme→size-only fallback is exactly gradeSet.ts:130's loadCohort(undefined, facts.pieces). (6) The empty-cohort 0.5→6.0 path is real (smartPrice.ts:32) and is indeed only reachable via the direct smartPriceScore/smartPriceCohort dependency path, not the gated main pipeline — matching the audit's 'dead code in the real pipeline but reachable' framing, which is why it's correctly rated low. Strengths also verified: bricklink.ts:100-101 defaults guide_type='sold'/new_or_used='N'; priceGuideToValue prefers qty_avg_price; the 0.8 coverage gate is at gradeSet.ts:117; resaleScore is null-guarded (returns null if currentValue or msrp missing); and Feels-Worth-It is a capped nudge at score.ts:54-57 with FEELS_WORTH_IT_CAP=0.75, not an additive sub-weight. No overstatements or understatements found.
Audit called it: Data and persistence
This is the part of the engine that saves things to a database (a Supabase database — basically a big online spreadsheet the tool can read and write) so it doesn't have to redo expensive work every time. When the tool grades a set, it mines reviews, computes per-category opinion signals, and looks up part prices. All of that is slow and costs API calls, so the engine stashes the results and reads them back on the next grade. This dimension of the audit asks one simple question: is that saved data trustworthy, and is it saved safely?
The save-and-reload layer is built carefully in a lot of ways — it never stores raw review text (just the distilled numbers), it labels every saved item consistently so it can be found again, and it cleverly works around a database limit when loading the big catalog. But it has two real holes, both rated high severity. First: nothing the tool saves ever expires. Once a set's opinion signals are cached, the engine treats them as good forever and quietly skips re-mining — so a re-grade months later silently reuses old data unless someone manually forces a fresh run. Prices have the same problem. Second: the function that updates a set's saved signals does a "delete the old, then add the new" in two separate steps with no safety net. If the delete silently fails, the new data gets added on top of the old, and the set ends up counted twice — which can warp its grade. Both of these can make a published grade look confident and current when it's actually stale or double-counted.
The storage layer is genuinely well-built for not losing or corrupting data on the happy path, but it has no concept of "this saved number is getting old" and one save routine that can silently double-count a set's signals — both of which can quietly make a published grade wrong.
Each problem below is laid out the same way: what's happening, a concrete LEGO example, why it dents trust, and what the fix is.
There is no expiry date on anything the engine saves. In database terms, there's no 'TTL' (time-to-live — a stamp that says 'this data is only good until X, then go refresh it') and no 'fetched_at' timestamp (a note of when the data was grabbed) on any saved row. The most visible victim is the opinion-signal cache. When the engine grades a set, the very first thing it does is check the database for previously-mined signals (gradeSet.ts line 89). If it finds any at all, it merges them in and flips a flag called fromCache to true, then skips the entire review-mining step — no new YouTube reviews discovered, no fresh distillation (gradeSet.ts lines 90-93). It does this no matter how old the cached signals are. The ONLY way to get fresh data is for someone to manually set a forceRemine flag (line 90). Part prices behave the same way: the price cache (supabaseStore.ts lines 304-349) has no timestamp, so a price grabbed a year ago is treated exactly like one grabbed this morning.
Say you graded the Razor Crest (75292) back when it first dropped. At the time, reviews were glowing — the early opinion signals land it at a 9.1, an A. The engine saves those signals. A year later it's retired, secondary-market reviews have soured a bit on the price-to-play ratio, and fresh reviews would pull the build-and-value signals down toward, say, 7.8 (a solid B). You re-grade it to publish an updated take. But because cached signals exist, the engine reads them, flips fromCache to true, and skips mining entirely (gradeSet.ts:89-93). It hands back 9.1 / A — the year-old verdict — and nothing on the result says 'this is stale.' What it SHOULD do: notice the signals are, say, 14 months old, decide that's past their freshness window, and re-mine for current reviews before grading.
This is a trust landmine. You'd publish an A on the Razor Crest in good faith, believing the engine just looked at current reviews — but it actually replayed a year-old verdict and told you nothing. Your audience treats the Drop Score as a live read on a set. If it's silently serving cached opinions from a different era, an old grade can directly contradict what reviewers are saying today, and you'd have no way to know it happened.
Add a 'fetched_at' timestamp to saved signals and prices, and make the read side age-aware: when the engine loads a cached set, check how old it is, and if it's past a freshness window you choose (say 90 days for signals), treat it as a miss and re-mine automatically — instead of needing a human to flip forceRemine by hand.
The function that updates a set's saved signals (saveDimensionSignals, supabaseStore.ts lines 147-168) works in two separate steps: first it deletes the set's existing signals, then it inserts the new ones. Two problems. First, the delete step (line 152) throws away its own error report — the code never even looks at whether the delete succeeded. Second, the two steps aren't wrapped in a 'transaction' (a transaction is a database all-or-nothing guarantee: either both steps happen, or neither does). So if the delete quietly fails but the insert succeeds (line 166), the old signals are still sitting there AND the new ones get added next to them. The set now has two full sets of opinion signals in the database, and the grading engine will read and count both. Unlike the engine's batch-save routine, which retries on a hiccup, this function has no retry and no safety net.
Take the UCS Millennium Falcon (75192). It's been graded before, so it has one set of cached signals — five reviews backing a build-quality score of 8.5. You trigger a re-mine to refresh it. saveDimensionSignals fires: step one tries to delete the old five-review signal set, but the database connection blips and the delete silently fails (the code ignores the failure, line 152). Step two then inserts the fresh signal set — another build-quality score, this time 7.9 from six reviews (line 166). Now the Falcon has BOTH in the database. Next time it's graded, the engine reads all of it: it sees eleven 'reviews' instead of six, and blends 8.5 and 7.9 together as if they were independent — inflating how confident and well-supported the grade looks. What it SHOULD do: guarantee the old signals are gone before the new ones land, so the Falcon always has exactly one current signal set.
This corrupts the two things a grade leans on most: the score itself and the confidence behind it. A double-counted set looks better-supported than it is — the engine thinks twice as many reviews agree, so it reports higher confidence and a blended score that no real set of reviews actually produced. You'd publish that number trusting it reflects real review volume, when it's partly a bookkeeping accident. And because the failure is silent, nothing flags the set as suspect — it just quietly reads as more authoritative than it earned.
Make the update atomic — either do the delete-and-insert as one all-or-nothing database operation (a transaction or a single server-side function), or switch to an 'upsert' that overwrites in place so there's never a moment with old and new data side by side. And stop ignoring the delete's error: if the delete fails, stop and raise it instead of charging ahead into the insert.
The audit's two weaknesses both check out exactly against the code. Confirmed: gradeSet.ts lines 89-93 read cached signals and, on any hit (absent forceRemine), skip mining entirely with no age check; no table in supabaseStore.ts carries a fetched_at/TTL column, and part prices (lines 304-349) are equally timeless. Confirmed: saveDimensionSignals (lines 147-168) does a delete (line 152) whose error is never captured, then a separate insert (line 166), with no transaction and no retry — so a failed delete plus a successful insert double-counts. One nuance on the second recommendation ('re-add the dropped FK or key all tables by base'): the repo has no SQL schema files (the Supabase schema lives remotely), so I can't see a foreign key in source to confirm one was dropped — that part rests on the auditor's view of the live database, not the code. The keying-by-base point is real and visible though: most tables key on the full set_number while set_catalog also carries base_number and loadCatalog (lines 284-302) dedupes by base_number, so the inconsistency the recommendation flags does exist in code. The strengths are accurate: no raw review text is stored (only distilled numbers, lines 153-164), keying is consistent between writes and reads, and loadCatalog genuinely paginates past the 1,000-row cap.
Audit called it: Confidence & coverage honesty
Every Drop Score ships with two little trust labels next to the letter grade. "Coverage" is breadth — out of all the things we COULD have graded about this set, how many did we actually have evidence for. "Confidence" is depth — for the things we did grade, how solid was the evidence (lots of reviews that agree = high; one stray review or reviewers who contradict each other = low). The whole point of these two badges is to stop a grade from looking authoritative when it's really resting on thin or lopsided data, so a B+ built on 200 reviews doesn't look identical to a B+ built on two.
The two badges are built correctly as separate ideas — breadth and depth really are measured by different math, and one safety valve (refusing to look confident when the whole grade leans on a single dominant source) is wired up properly. That foundation is good. But the honesty promise falls apart in the details. The single biggest problem: the rule the spec says should decide "we don't have enough data here" — needing at least 5 pieces of evidence — is never actually used. The code quietly uses a weaker bar of 3 instead, and a comment in the code even claims it's using the stricter rule when it isn't. On top of that, three other guard-rails the spec defines (a minimum for trusted Brickset ratings, a minimum before community votes count, and the requirement that source quality and recency feed into confidence) are declared as numbers but never applied anywhere. The worst symptom: one lone Amazon review can produce a perfect 1.000 confidence score — the exact opposite of what the spec promised that situation should show.
The two badges are genuinely well-separated in concept and the over-concentration safety valve works, but the headline "honesty" promise is broken — the main insufficient-data threshold and three other guard-rails are declared and never applied, so a grade can claim full coverage and even perfect confidence on data the spec says is too thin to trust.
Each problem below is laid out the same way: what's happening, a concrete LEGO example, why it dents trust, and what the fix is.
The spec names a primary threshold called LOW_DATA_FLOOR — a constant (a fixed number set once and reused) equal to 5. The rule: if a gradeable point has fewer than 5 total pieces of evidence behind it, mark it 'insufficient data' so it lowers the coverage badge. But the scoring code never checks against 5. Instead it checks against a weaker number, MIN_VOLUME_FOR_SENTIMENT, which equals 3. A search of the whole codebase confirms LOW_DATA_FLOOR appears only where it's first defined and in one comment — it's never wired into any decision.
Take a quiet set like a small Botanicals build. Suppose the 'build is fun' dimension has exactly 3 reviews behind it — that's it, 3. The spec says 3 is below the floor of 5, so this should be flagged insufficient and pull the coverage badge down toward 'data-light.' Instead the code sees 3, which clears its weaker bar of 3, so it scores the dimension, counts it as fully covered, and lets it feed the headline letter grade. The badge tells Seth's audience the set is well-covered when it was graded on a literal handful of reviews.
Coverage is the badge that's supposed to whisper 'go easy on this one, we didn't have much.' If the real floor is silently lowered from 5 to 3, sets graded on barely-there evidence get dressed up as fully covered. A grade Seth publishes looks more authoritative than the data underneath it can honestly support.
Actually check total evidence against LOW_DATA_FLOOR (5) when deciding whether an opinion-only dimension is 'scored' versus 'insufficient.' If the lower bar of 3 is meant to mean something different (like 'enough to bother analyzing'), keep it for that — but don't let it quietly replace the stricter 5 the spec requires.
Right above the function that resolves each dimension, there's a documentation comment stating that opinion signals are 'gated by LOW_DATA_FLOOR' (the floor of 5). But the actual code on the line below gates by the weaker 3. The comment describes behavior that isn't there.
Imagine you or an auditor open this file to confirm the engine is honest. The comment reassures you: 'yes, the 5-evidence floor is enforced here.' You move on, satisfied. But a set with 4 reviews on a dimension — which should be flagged insufficient — sailed through as fully scored, because the code is really using 3. The comment actively hid the bug from anyone checking.
This is a comment that says one thing while the code does another. It doesn't just fail to enforce the rule — it actively misleads anyone reviewing the engine into believing the safeguard is in place, which is how a flaw like this survives review and ends up shaping published grades.
Rewrite the comment to describe what the code actually does, and fix the code to match the spec. The comment and the behavior must tell the same story.
Confidence is volume (how much evidence) multiplied by agreement (how much the sources agree). Agreement is calculated as 1 divided by (1 plus the spread of the scores). With only one source there's no spread, so agreement comes out to a perfect 1.0. And 'volume' is just a number that can be set high. So a single review carrying volume 60 produces confidence = 1.0 (volume) times 1.0 (agreement) = a perfect 1.000.
Picture the Millennium Falcon graded off exactly one Amazon review that happens to carry a volume of 60. There's nothing to disagree with, so agreement is a flawless 1.0; the inflated volume maxes out the volume term; confidence lands at a perfect 1.000. The spec explicitly uses this exact scenario — all-Amazon, single source — as the textbook case that should read 'full coverage, LOW confidence.' The engine reports the opposite: maximum confidence, on one self-selected review.
This is the headline dishonesty the badges exist to prevent. A grade can wear a perfect-confidence badge while resting on a single, easily-gamed review. If Seth publishes that, his audience reads rock-solid certainty into a number that one anonymous reviewer could swing.
Stop one source from ever reaching perfect agreement. Discount confidence when every signal comes from the same source, or require at least two independent sources before agreement can climb past some cap. A lone review must never read 1.000.
Brickset reviews carry structured 1-to-5 sub-ratings the engine trusts highly. The spec says these should only count once there are at least MIN_STRUCTURED_COUNT (8) of them — below 8, they're meant to be thrown out as skewed by a small, self-selecting group of adult fans. The code that turns Brickset reviews into signals applies no such count check, and its own comment claims the gate 'is applied upstream, not here' — but the upstream code that calls it pipes the signals straight into the blend with no filter. Nobody applies the gate.
Say a newish set has just 2 Brickset reviews, both glowing 5-out-of-5s. The spec says discard them — 2 is far below 8, too few to trust. Instead they flow straight in, boost confidence, and can flip a dimension from 'not enough data' to 'scored and confident.' Two enthusiastic superfans just made the grade look both better and more certain than the evidence allows.
Brickset's audience skews toward devoted adult collectors, so tiny samples run rosy. Letting 2 reviews masquerade as solid evidence means a published grade can be both inflated and falsely confident — the precise bias the 8-review minimum was written to filter out.
Actually apply MIN_STRUCTURED_COUNT where the signals are assembled (in the pipeline, not just as a comment). Discard Brickset signals below 8 reviews, and add a test proving a 2-review set gains no coverage or confidence from them.
The spec says community submissions (votes from users) are display-only until there are at least MIN_COMMUNITY_TO_MOVE (5) of them — below 5 they should be shown as a count but contribute 0 to the actual computed score, both to protect against vote-stuffing and because a handful of votes isn't a real signal. The engine never applies this gate. Community votes enter with a volume of 1 each and get treated exactly like mined review signals, so a single one already moves a dimension's value and confidence.
Imagine someone submits one community vote on the Razor Crest's 'feels worth it,' rating it a 10. The spec says: show '1 vote' but let it move the score by zero until 5 votes exist. Instead that lone vote shifts the dimension's value upward and nudges confidence. One person — or one motivated fan rallying friends — can start steering the grade from the very first click.
This breaks both the 'display-only below 5' promise and the anti-brigading protection. A grade Seth publishes could be quietly nudged by a single submitter, which is exactly the kind of manipulation the gate was meant to block — and it makes the confidence badge complicit by treating that one vote as real evidence.
Enforce MIN_COMMUNITY_TO_MOVE where signals are merged: below 5 submissions, community votes contribute 0 to both value and confidence while still being surfaced as a display-only count. Add a test that a 1-submission set gains no coverage or confidence.
The spec lists four things that should shape confidence: source trust (a trustworthy outlet counts more), recency (older opinions count less), price-pairing, and a noise/sarcasm penalty. None of them touch the confidence number. Confidence is built from only two things: total volume and how much the scores agree. Source trust (STREAM_TRUST — Amazon rated 0.5, Brickset rated 1.0) and recency decay do exist, but they live in a different calculation that adjusts the VALUE, not the confidence.
Take two sets graded on the same number of reviews with the same level of agreement. Set A's evidence is entirely Amazon (trust 0.5, the lowest tier). Set B's is entirely Brickset (trust 1.0, the highest). The spec's own headline example says all-Amazon data should read lower-confidence. But because trust never enters the confidence math, both sets report the identical confidence badge. The system literally cannot tell its most-trusted source from its least-trusted one when stamping certainty.
Confidence is supposed to mean 'how much should you trust this.' If a wall of low-trust Amazon reviews looks exactly as certain as a wall of high-trust Brickset reviews, the badge is decorative. A grade can flaunt high confidence purely on the weakest, most gameable source — directly contradicting the spec's promise.
Fold source trust, recency decay, and a noise/sarcasm penalty into the confidence calculation so an all-Amazon dimension genuinely reads lower-confidence than an all-Brickset one. Right now those named inputs change the value but leave certainty untouched.
Confidence's volume term adds up a 'volume' field across signals. But each mined review is emitted as its own signal with volume 1. So 3 separate agreeing reviews total a volume of 3 — which produces the exact same volume factor as one single signal carrying volume 3. The math can't distinguish three independent voices from one voice claiming three votes.
Compare two scenarios for an X-Wing. In the first, three different YouTubers independently praise the build — three genuine, separate opinions. In the second, one source is logged as carrying three votes. Both produce an identical volume factor (about 0.337) and identical confidence. But three independent reviewers agreeing should be far more reassuring than one source asserting a count of three — and the engine treats them as equal.
The whole appeal of 'agreement' is that independent people converged on the same view. If independence is invisible, the agreement term means less than it appears to, and a confidence badge can be propped up by one source padding its own count rather than by a real chorus of reviewers.
Track the number of independent sources separately from raw volume, and feed source diversity into confidence so genuine independent agreement outweighs one source's self-reported count.
The cutoff that flips coverage from 'full' to 'data-light' is FULL_COVERAGE_MIN = 0.5 — meaning at least half of applicable dimensions must be scored. This number was invented directly in the scoring code. The spec talks about being 'above the threshold' but its table of tunable constants (the one place every other adjustable number lives) never defines this cutoff. And it's buried in the scoring file instead of the shared constants file.
A set that scores 4 of its 8 applicable dimensions hits exactly 0.5, so it's labeled 'full' coverage. A set that scores 3 of 8 lands at 0.375 and gets 'data-light.' That single dimension is the whole difference between a confident-looking badge and a cautious one — and the 0.5 boundary deciding it was never calibrated, justified, or written into the spec.
A boundary that flips the public coverage badge is load-bearing for how trustworthy a grade looks. If it's an unexamined guess hidden in code rather than a documented, agreed-on value, the badge's meaning can drift and nobody would notice — and Seth can't defend why 4-of-8 is 'full' but 3-of-8 isn't.
Move FULL_COVERAGE_MIN into the shared constants file alongside the other tunables and document it in the spec's constants table — or justify the 0.5 cutoff with real calibration. A boundary this important shouldn't be a magic number tucked away in scoring code.
Coverage and confidence are only computed across the six universal in-grade clusters. The holdValue badge (scarcity and resale potential) sits outside that set. But under the Investor lens, holdValue is weighted at 30% of the grade — the single largest factor. So the confidence number shown beside an Investor grade never reflects how well-evidenced that 30% chunk actually is.
Two Investor grades for a UCS-scale set. One has rich, well-sourced resale and scarcity data behind its holdValue. The other has almost nothing — a guess. Because holdValue is excluded from the confidence calculation, both can display the same confidence badge. The thinly-evidenced investor grade looks every bit as certain as the well-researched one, even though its single biggest input is essentially a shrug.
Investor grades are exactly where people put money on the line. If the confidence badge ignores the 30%-weighted cluster that the whole investor thesis rests on, an Investor grade can broadcast false certainty about the very thing an investor cares about most.
Compute and surface a separate confidence and coverage figure for the holdValue badge, and for any lens that weights it, so an Investor grade's reported certainty actually reflects the evidence behind its 30% cluster.
feelsWorthIt is mined from reviews and nudges the money-worth score (capped). But because it's treated as a nudge rather than a formally weighted sub-dimension, it's left out of both the coverage and confidence counts. The badges simply don't see it.
Suppose a polybag-sized set's only value-related evidence is strong 'feels worth it' sentiment from reviews, and it does shift the money-worth score upward. Yet the coverage and confidence badges for money-worth report zero contribution from it — the badges act as if no value evidence exists, even though that evidence quietly moved the score. The depth of real value evidence is hidden from the very labels meant to disclose it.
It's a smaller honesty gap, but it still means a badge can understate how much evidence actually shaped a score. The signal influences the grade while staying invisible to the labels that are supposed to tell the audience how well-evidenced that part is.
Let feelsWorthIt's evidence register in the money-worth coverage and confidence figures even though its effect on the score stays a capped nudge, so the badges reflect the value evidence that's actually present.
The audit is accurate against the real code; every substantive claim checks out. Two minor line-number drifts worth noting, neither of which changes the finding: (1) The 'volume = 1 per mined review' claim is attributed to distill.ts:40, but the actual emit with `volume: 1` is on line 44 of src/lib/pipeline/distill.ts (line 40 is the loop opening). The behavior described is exactly correct. (2) The Brickset 'applied upstream, not here' comment spans lines 56-59 of src/lib/sources/brickset.ts; the audit cites line 58, which is within that comment. I verified directly: LOW_DATA_FLOOR, MIN_STRUCTURED_COUNT, and MIN_COMMUNITY_TO_MOVE appear only at their declarations in constants.ts (lines 140, 142, 143) and are never referenced in any scoring decision; dimension.ts:47 gates on MIN_VOLUME_FOR_SENTIMENT (=3); the dimension.ts:20 comment does falsely claim LOW_DATA_FLOOR gating; signalConfidence (dimension.ts:6-15) uses only volume and variance with no source-trust input; STREAM_TRUST (amazon 0.5, brickset 1.0) lives in blend.ts and affects value not confidence; fakes.ts:44 pushes community signals at volume 1; and the investor lens weights holdValue at 30 while IN_GRADE_CLUSTERS excludes holdValue, confirming the confidence figure ignores it.
Audit called it: Test quality and coverage
Tests are little automated checks the code runs on itself: each one feeds the engine a known input and confirms it gets the expected answer, so a future change that quietly breaks the math gets caught before it reaches a published grade. This part of the audit asks two questions: are the tests good (do they actually try to catch mistakes, not just rubber-stamp the code), and do they cover enough of the engine (especially the risky parts that talk to the outside world — LEGO databases, YouTube, the price service, and the AI). The answer is: the tests that exist are genuinely good, but they only cover the calculator-like "pure" pieces. The big assembly-line pieces that fetch real data and run a real grade end-to-end are not tested at all.
The engine ships with 197 automated checks and every single one passes, and they're not lazy checks — they actively try to trip the code up, which is exactly what you want. But there's a sharp dividing line. The "pure" pieces (math, sorting, mapping one shape of data to another — the parts that, given the same input, always give the same output) are well covered. The "live" pieces — the ones that actually reach out to LEGO databases, YouTube, the price service, and the AI to grade a real set — have no tests at all. The two biggest assembly lines, the one that grades a set (gradeSet) and the one that prices a set's parts (priceInventory), have no test file whatsoever. The review-mining assembly line is "tested" only in the sense that three small helper functions inside it are checked; the actual mining engine, with all its branches and error handling, is never run by a test. Crucially, the engine was BUILT to be testable here — the live pieces accept their outside connections as swappable inputs, and the project already has fake stand-ins ready to use — so the gap is an unfinished job, not a design dead-end.
The tests that exist are honest and sharp and 197/197 pass, but they stop exactly where the risk starts: every part of the engine that fetches real data and produces a real grade is completely unverified.
Each problem below is laid out the same way: what's happening, a concrete LEGO example, why it dents trust, and what the fix is.
An orchestrator is the assembly-line code that calls everything else in order to produce one finished result. 'gradeSet' is the orchestrator that turns a set number into a full Drop Score: it fetches facts from LEGO databases, pulls in mined reviews, prices the parts, builds the comparison group, and runs the grade. 'priceInventory' is the live function that adds up what a set's parts are worth by calling the price service and the database. Neither has a test file at all. On top of that, the only tests near the review-mining engine ('mineSetReviews') exercise three small helper functions pulled out to the side — the actual mining engine, with its cache shortcut, its 'no indexed reviews, fall back to a live search' branch, and its 'skip a broken review instead of crashing' error handling, is never run by any test. So the entire path where the engine touches the network, the database, and the AI is unverified.
Take the Razor Crest (set 75331). Here's a bug that today's tests would sail right past. Inside gradeSet there's a rule: only trust the calculated part-out value if at least 80% of the set's parts got a price (the code checks partValueCoverage >= 0.8). Imagine a future tweak accidentally flips that to 'less than 80%.' Now the Razor Crest comes through with only 40% of its parts priced — a number that should be thrown away as too incomplete to trust — but the broken rule accepts it anyway. The half-finished price feeds into the value score, the Razor Crest's '8.4, A-' is really built on a quarter of the picture, and not one of the 197 tests fails, because no test ever runs gradeSet with a low-coverage set to check that the guardrail holds. The grade looks confident and is quietly wrong. What the tests SHOULD do: run gradeSet with a fake price source that returns 40% coverage and assert the engine ignored that value and fell back to its safer estimate — which is fully possible today, because gradeSet already accepts its price source as a swappable input.
This is the trust core of the whole tool. The pieces with zero tests are the exact pieces that produce the number and letter you publish. A change that corrupts a fetch, mis-wires the comparison group, double-counts mined reviews, or accepts junk-quality data can ship looking perfectly healthy — all 197 green checks — while every grade it produces is subtly off. Green tests would give you false confidence to hit publish on a wrong grade.
Write tests for gradeSet and mineSetReviews using the swappable-input hooks they already expose (GradeDeps and MineDeps) plus the in-memory fakes already in the project — feed them canned data and assert the finished score and the branch behavior (cache hit, search fallback, low-coverage rejection). Add a test for priceInventory the same way, and specifically test the failure paths, not just the happy path.
Distilling is the step where the AI reads one review and turns it into per-dimension scores. The very last line of that step takes the AI's reply and does JSON.parse on it — JSON.parse means 'read this text as structured data,' and it throws an error (crashes that step) if the text isn't perfectly formed. There's no safety net around it (no try/catch, which is the 'if this fails, handle it gracefully instead of crashing' wrapper). And because no test ever runs the live distiller, nothing in the suite would catch this. This sits inside the same untested live-mining path as the weakness above, which is why the audit groups them, but it's a distinct, concrete failure point worth calling out on its own.
You're grading the Millennium Falcon (75192) and the engine mines three YouTube reviews. For two of them the AI returns clean data. For the third, the AI hiccups and returns a reply with a stray character or a cut-off ending — not valid structured data. JSON.parse on that third reply throws. Right now that error does get caught one level up (mineSetReviews wraps each review in a 'skip the broken one' guard), so in practice that single review is silently dropped — but you lose a third of the Falcon's mined evidence with no signal that it happened, and you're relying on that outer guard staying exactly where it is. If a future refactor moves the parse or the guard, the same hiccup takes down the whole grade instead of one review. Either way, no test exists to tell you which behavior you actually have. What the tests SHOULD do: feed the distiller a deliberately broken AI reply and assert it returns 'nothing usable' cleanly instead of throwing — so a malformed answer can never escalate beyond the one review it came from.
AI replies are the one input you can never fully predict, so this is precisely where a safety net matters most. Without the wrapper and without a test pinning the behavior, the difference between 'lose one review quietly' and 'the Falcon's whole grade dies' rests on code staying perfectly arranged forever. A grade that silently drops a third of its evidence, or one that fails entirely on a set you've already promoted, both erode trust in different ways — and you'd have no early warning of either.
Wrap the distiller's JSON.parse in a try/catch so a malformed AI reply returns 'no usable result' instead of throwing, and add a test that feeds it a broken reply and confirms it degrades gracefully.
Mocking means swapping a real outside service (a LEGO database, YouTube, the price API, the AI) for a fake stand-in you control during a test — so you can rehearse 'what if YouTube is down?' or 'what if the database returns nothing?' without actually calling anyone. The audit found zero mocking anywhere in the test suite, and I confirmed it: a search for every common mocking tool came back empty. That means there is no test that rehearses an outage, a rate-limit, an empty result, or a timeout from any external service. Every failure path through the live code is unrehearsed.
Picture grading a brand-new set, say a 2026 Botanicals bouquet, the morning it's announced. The comparison group (the 'cohort' — the set of similar sets the engine measures this one against to judge its value) is loaded from the database, but for a day-one set there may be too few comparable sets to be meaningful. The code has a smart fallback for exactly this — if the comparison group is too thin, widen it. But no test ever feeds the engine a too-thin or empty comparison group to confirm the fallback fires correctly. So if that fallback breaks, the Botanicals set gets a value verdict computed against two random sets and presented with the same confidence as a set measured against fifty. What the tests SHOULD do: use a fake database that returns an empty comparison group and assert the engine widens the group or marks the value as 'not enough data' — instead of quietly grading on almost nothing.
The real world fails constantly — services go down, return nothing, or rate-limit you. If none of those scenarios is ever rehearsed, you only find out how the engine behaves under stress when it happens live, on a real grade, in front of your audience. The danger to trust isn't a visible crash — it's the quiet case where a service returns thin or empty data and the engine grades on it anyway, publishing a confident letter that's really a guess.
Introduce fake stand-ins for the external services (the project already has in-memory fakes to build on) and write tests that simulate outages, empty results, and rate-limits, asserting the engine degrades safely each time rather than producing an over-confident grade.
The audit checks out against the actual code, with two small clarifications. (1) The audit's summary calls priceInventory a 'live orchestrator'; in the code (price.ts:22-58) it's a live I/O function that prices parts, not a top-level assembly line like gradeSet — but the substance of the finding (untested, with a real network + database surface) is exactly correct: there is no price.test.ts. (2) On the unguarded JSON.parse at distill.ts:147: it is genuinely unwrapped as the audit/recommendation says, but in the current wiring its caller (mineSetReviews, mine.ts:122-128) wraps each review in a try/catch, so today a malformed AI reply drops that one review rather than crashing the whole grade. The fix is still correct and worth doing — the safety shouldn't depend on the caller's guard staying exactly where it is, and nothing tests either behavior. Everything else verified: 197/197 tests pass (I re-ran them), there is no gradeSet.test.ts and no price.test.ts, mine.test.ts and distill.test.ts cover only pure helpers, and a search for every common mocking tool returned zero matches, confirming 'no mocking.'
Audit called it: Cost, scale and robustness
This part of the engine isn't about how a single set is scored — it's about what happens when you point the tool at thousands of sets at once and let it run. It covers how the tool talks to the outside data sources it depends on (BrickLink for part prices, YouTube for reviews, a transcript service, your Supabase database), how much each set costs in real money to grade, whether a long run can survive a crash, and whether you'd even notice if something quietly went wrong. The audit calls this "Cost, scale and robustness" — basically, does the machine hold up under real-world load without burning money or silently producing junk.
The audit's verdict here is "mixed," and that's fair. The skeleton is genuinely well thought out: the part-price cache, the per-set fetch limits, and the database write logic are all the right shape for running cheaply at scale. But the engine has almost no defensive plumbing around the outside world. It never slows itself down or backs off when a data source pushes back, so a "too many requests" error and a "this doesn't exist" error look identical — both just vanish, and a set that got throttled looks merely thin instead of broken. It sends entire (sometimes 90-minute) video transcripts to the most expensive AI model with no length limit, which is the single biggest uncontrolled cost. There's no logging, no cost tracking, and no way to see when grades are coming out degraded. And a long catalog run that crashes restarts from zero. Five of the eight problems are high-severity, and they all share a theme: the tool can't tell the difference between "I succeeded" and "I quietly failed," which is exactly the difference that matters when you're publishing grades.
The cost-saving architecture (global price cache, fetch caps, careful database writes) is genuinely smart and matches the plan; what's missing is every safety net that tells you when a grade came out degraded instead of clean — no throttle handling, no cost cap on transcripts, no logging, no crash recovery.
Each problem below is laid out the same way: what's happening, a concrete LEGO example, why it dents trust, and what the fix is.
Every connection to an outside service (BrickLink for prices, YouTube for reviews, the transcript service) is a bare, fire-as-fast-as-possible request. There's no rate-limiting (deliberately spacing out requests so you don't overwhelm a service), no backoff (waiting a bit and trying again after a failure), and no honoring of the 'Retry-After' signal that services send to say 'wait this many seconds before asking again.' The written plan specifically required throttling — capping requests to no more than 60 a minute — and that throttling was never built.
Say you kick off a catalog run that part-prices a 700-piece Razor Crest (set 75292). BrickLink lets you make about 5,000 price calls a day. The tool fires all 700 lookups for that set back-to-back as fast as the network allows, then moves to the next set and does it again. A few hundred sets in, BrickLink starts replying 'too many requests, slow down.' The tool doesn't slow down — it just keeps hammering, every call now failing. What it SHOULD do is space the calls out, and the moment BrickLink says 'wait 30 seconds,' actually wait 30 seconds, then resume. Instead it wastes the rest of the day's quota hitting a wall.
When the tool gets throttled mid-run, the prices it couldn't fetch just don't make it into the grade. The Razor Crest's part-out value comes out based on a fraction of its pieces — or not at all — and the grade quietly shifts. You'd publish a 'value' verdict that was really shaped by a rate-limit error you never saw, not by the set's actual worth.
Route every outside request through one shared gatekeeper that spaces requests out, limits how many run at once, and pauses when a service says 'wait.' Build it once, and every data source benefits.
When a request fails, the code 'catches and skips' it — meaning it quietly swallows the error and moves on, with no record of what happened. The problem is that a '429' (the service saying 'you're being throttled, this would have worked') and a '404' ('there's genuinely no data here') both get swallowed the exact same way. So a set that was temporarily blocked looks identical to a set that simply has no data — it shows up as 'thin' (not much information) rather than 'degraded' (we hit an error and gave up).
Picture grading the Millennium Falcon (set 75192). Halfway through pricing its ~7,500 pieces, BrickLink throttles you. Every remaining price lookup throws a 'slow down' error — and the code silently skips each one. The set ends up priced on, say, 40% of its pieces. Because that's below the 80% coverage gate, the tool correctly declines to show a part-out value — but it records the set as simply 'not enough price data,' which reads as 'this set is poorly covered.' What SHOULD happen is the tool flags 'we got throttled here — come back and finish this one,' so the Falcon gets re-run later and ends up with a real, complete value. Instead it's permanently mislabeled as thin.
This is the trust-killer. A set that's missing a value because YOU got rate-limited looks the same as a set that genuinely lacks data. You can't tell which grades to re-run, and a published grade on a flagship set can be quietly wrong — built on partial data — while looking like a normal low-coverage result.
Surface throttle errors as their own distinct category, separate from 'no data found,' so degraded sets get flagged and re-run instead of silently mislabeled.
To turn a review into scores, the tool sends the video's full transcript to an AI model (this step is called 'distillation' — boiling a long review down to per-category signals). Two things compound here: the transcript is passed in raw with no cap or trimming, and the model it defaults to is 'claude-opus-4-8' — the premium-priced model. The written plan explicitly said to cap and trim long transcripts before sending them. That wasn't done. (For context, the premium Opus model costs about 5x more per word of input than the cheaper Haiku model — and input is exactly what a long transcript bloats.)
A 12-minute review of a small Botanicals polybag might be 2,000 words — cheap to process either way. But a 90-minute deep-dive on the Millennium Falcon can run 15,000–20,000 words. The tool sends that entire wall of text to the priciest model, costing roughly 12 cents for that one review. It does that for up to 3 reviews per set. Across a ~24,600-set catalog, that's the line item that balloons. What it SHOULD do is trim each transcript to a sane length and use the cheaper model for bulk work — cutting that same review from ~12 cents toward ~2 cents with no real loss in grade quality.
This doesn't make a single grade wrong, but it makes grading the whole catalog so expensive it may not be practical to keep current — and an out-of-date grade is its own kind of wrong. If cost forces you to grade fewer sets or refresh them less often, the published grades drift from reality.
Cap transcript length before sending, switch bulk distillation to the cheaper model, and add a running cost counter so the bill is visible as it grows.
There's no observability — no way to see what the tool is doing while it runs. Errors are swallowed by bare catch blocks, the only log lines are truncated to 50 characters, and there's no metrics, no run ledger (a record of what each run did), and no alerting. So when grades start coming out degraded during a big run, nothing tells you.
You run an overnight catalog ingest of 2,000 sets. Around set 600, BrickLink throttling kicks in and YouTube starts rejecting requests. Sets 600 through 2,000 all come out with shaky, partial data. In the morning the script prints '✅ 1,950/2,000 ingested' and you assume it went fine — because a set that errored on pricing still 'succeeds' overall, just with worse data. What SHOULD exist is a per-run summary: 'pricing degraded on 1,400 sets, 3,100 throttle errors, estimated cost $X' — so you'd immediately know the run was poisoned. Instead it looks like a clean success.
If you can't see degradation happening, you publish degraded grades believing they're solid. At scale, a quiet failure spreads to thousands of sets before anyone notices — and every one of those is a grade your audience trusts.
Add structured logging, per-run cost and health metrics, and a run ledger so a degraded run announces itself instead of masquerading as a success.
There's no resumability or cursor — no bookmark that remembers how far a run got. The review-discovery sweep collects everything into memory and only saves at the very end, and the batch grader walks its list with no checkpoint. So if anything crashes partway, the work done so far is lost and the next run re-pages everything from the beginning.
The review-discovery sweep is paging through dozens of reviewer channels across the whole ~24,600-set catalog. Three hours in, it crashes — network blip, machine restart, whatever. Because nothing was saved until the final step, all three hours of discovered review links evaporate, and you start over from set one. What SHOULD happen is the tool remembers the last video it processed per channel, so a restart picks up where it left off and you lose minutes, not hours.
It doesn't make any single grade wrong, but it makes keeping the whole catalog fresh painful and fragile. The harder it is to complete a full run, the more often grades go stale — and stale grades quietly stop matching reality.
Add a resumable cursor that records the last-seen video per channel and a run ledger, so a crashed sweep continues instead of restarting from zero.
When pricing a set's parts, the tool fetches prices strictly sequentially — one part, wait for the answer, next part, wait again — with no concurrency (running several lookups in parallel). For a big set that's hundreds of round-trips, each waiting on the last.
An X-Wing (set 75355) with several hundred distinct parts gets priced one lookup at a time. If each network round-trip takes a third of a second, that's a couple of minutes for one set. Now multiply that across the tens of thousands of less-popular sets in the catalog tail and it becomes impractically slow to ever part-cost them. What SHOULD happen is several price lookups run at once (carefully, under the same rate limit), turning minutes per set into seconds.
If most of the catalog is too slow to ever price, the Smart Price feature only really works for popular sets. The grades on everything else lean on rougher fallbacks, so the part-out-value signal is strong for flagships and weak for the long tail — an unevenness your audience wouldn't expect.
Run several price lookups in parallel under a shared concurrency limit, so big sets and the catalog tail become practical to price.
The plan called for an 'isolation spine' — running transcript scraping on its own separate accounts and addresses so that if that risky scraping ever got banned, it couldn't take down the safe, official data feeds with it. The code instead uses a paid managed transcript service (Supadata), which is actually the plan's own approved fallback. The catch: the code went straight to that paid service without reconciling it against the isolation plan, and without modeling or capping what each transcript actually costs. So there's a real per-transcript fee that's invisible and uncapped.
Across a full catalog mine, the tool fetches a transcript for every candidate review on every set — potentially tens of thousands of transcript fetches, each carrying a fee from the managed service. Because that cost is never tracked or capped, it doesn't show up anywhere until the bill arrives. What SHOULD happen is the transcript approach is explicitly reconciled with the plan and its per-fetch fee is budgeted and counted alongside the AI cost.
Like the transcript-length problem, this is a cost-and-sustainability risk rather than a wrong-number risk: an unbudgeted fee that could quietly make full-catalog grading too expensive to sustain, which pushes grades toward going stale.
Reconcile the managed-transcript approach with the written plan and budget its per-transcript fee so the full cost of grading is known up front.
The batch script checks that some credentials are present up front, but not all of them — the BrickLink, YouTube, transcript, and AI keys are pulled in with a shortcut that assumes they exist. If one is missing, an empty 'undefined' value gets passed into the request instead, and the failure only surfaces later as one of those swallowed catch-and-skip errors.
You start a priced run but forgot to set the BrickLink credentials in your environment. Instead of stopping immediately with 'BrickLink key missing,' the tool quietly sends blank credentials, every price request gets rejected, and each rejection is silently skipped. The run finishes 'successfully' with zero pricing on every set — and nothing told you why. What SHOULD happen is the script refuses to start and says exactly which key is missing.
It's a setup foot-gun, not a math bug, but it can produce a whole run of valueless grades that look complete. You'd only catch it by noticing every set mysteriously lacks a part-out value.
Validate every required key up front and stop with a clear message if any is missing, so a misconfigured run fails loudly instead of producing empty grades.
The audit's claims all check out against the actual code, with two small clarifications worth noting. (1) On the transcript 'isolation spine' weakness (supadata.ts): the code's use of a managed transcript service is not a rogue deviation — it's literally the plan's own approved 'managed-API escape hatch' (sourcing spec lines 99 and 153). The real, fair criticism is narrower than 'contradicts the spec': the code adopted that fallback without reconciling it against the isolation section and without budgeting or capping its per-transcript fee. I framed it that way. (2) The '24k-set tail' figure: the buildable-set crawl in the spec is ~13,200 sets, but the catalog the matcher runs against is '~24,600 base sets' (src/lib/sources/catalog.ts line 2), so the audit's 24k number is grounded in the real code. Everything else — the bare fetches with no throttling (bricklink.ts:107, youtube.ts:185, supadata.ts:50), the empty catch blocks that swallow 429s identically to 404s (price.ts:52-54, mine.ts:125-127, gradeSet.ts:122), the uncapped full transcript sent to the default model claude-opus-4-8 (distill.ts:130, 135), the 50-character truncated logging with no metrics (batch-ingest.ts:106), the no-cursor sequential loops (discover.ts:38-69, batch-ingest.ts:96-108), the strictly sequential pricing (price.ts:43-55), and the unvalidated non-null-asserted env keys (batch-ingest.ts:74-90) — matches the code exactly. The 80% coverage gate strength is real (gradeSet.ts:117). The 5x Opus-vs-Haiku input-price ratio is confirmed from the model catalog ($5 vs $1 per 1M input tokens).
If a word in this report tripped you up, it's defined here.
Plain-English rewrite of the 9-agent adversarial engine audit, generated 2026-06-18. Every finding was re-verified against the live source before translation. Companion docs: docs/2026-06-18-engine-audit.html (the original technical audit), docs/grading-walkthrough.html (the step-by-step grading procedure on the live Razor Crest), docs/how-it-works-plain.html (plain-English overview of the whole tool).