An A-grade deterministic math core dragged down by an evidence/honesty layer whose three named guardrails are dead constants, and a headline price feature that never reaches production data. It is honest and precise in the easy case, and confidently wrong in exactly the hard, low-data cases the methodology was written to protect. A strong skeleton with a hollow evidence layer — not yet trustworthy to ship grades to users.
If we fix one thing
Enforce the three evidence gates in the actual scoring path, with a test per gate so they can never silently die again. In resolveDimension, mark a signal-only dimension "insufficient" (not "scored") when total volume < LOW_DATA_FLOOR (5), and fix the doc-comment that lies about this; gate Brickset structured signals at MIN_STRUCTURED_COUNT (8) before the blend; gate community signals at MIN_COMMUNITY_TO_MOVE (5), display-only below it. It's the single most-corroborated finding (5 of 9 auditors), it restores the anti-brigading and anti-thin-data guarantees, it fixes the coverage over-reporting at the same time — and it's small, local, and immediately testable, unlike the street-price plumbing or the ops hardening.
The honest one-paragraph verdict
This is a well-architected deterministic scoring core wrapped around an evidence layer that does not yet enforce its own rules, sitting on a data layer whose marquee inputs aren't wired to live sources. The math you can see — cluster weights, sub-weights, the Feels-Worth-It cap, build-length price-scaling, letter bands with half-up rounding, the over-concentration water-fill, archetype N/A maps, the stream-trust blend — is a faithful, number-for-number, genuinely well-tested transcription of the spec (197 tests, adversarial in places). Hand it rich, clean data and it returns the right letter. But the spec was written precisely to protect the long-tail/low-data/adversarial regime, and that's exactly where the engine silently fails open: all three named evidence gates exist only as constants and comments; confidence is volume×agreement only, so one Amazon review with inflated volume reports 100% confidence; and the Smart Price "street price" anchor (~14% of the grade) is never populated by any live adapter, so every real grade quietly falls back to the MSRP basis the spec forbids — with no flag.
Strengths and risks, side by side
Biggest strengths
- The deterministic math is faithful and well-tested — cluster/sub-weights, the ±0.75 FWI cap, build-length price-scaling, the percentile anchors, and the letter bands all match the spec's §10 table number-for-number, with float-robust half-up rounding. Six of nine auditors praised this core.
- The over-concentration guardrail actually works — it water-fills weight to the 40% cap and forces the "data-light" tier when a grade rests on one dominant cluster.
- Coverage and confidence are real, distinct computations (breadth vs depth), and the N/A-vs-insufficient bookkeeping is correct — archetype-dropped clusters take no coverage penalty.
- The part-out costing + global part-price cache are the right design for BrickLink's daily limit: sold/new basis, qty-weighted averaging, fetch-only-the-missing, an 80% coverage gate. The non-Smart-Price fact scores are clean and spec-faithful.
- The discovery matcher is well-engineered for a first pass — number-in-title is authoritative with a year-collision guard, names need ≥2 distinctive tokens, pre-release reviews are rejected, twins are demoted to the judge rather than guessed. Derived-only storage is honored end to end.
Biggest risks
- The three named evidence gates are dead constants (confirmed in source). A 3-review dimension scores and reads "full coverage"; a 2-review Brickset signal moves the grade; a single community vote moves the score. Every anti-brigading/anti-thin-data promise is currently vapor. 5 of 9 auditors flagged this.
- Confidence is gameable and omits its own spec inputs — one signal → agreement 1.0, so a lone Amazon review reports 100% confidence (the exact dishonesty the badge was built to prevent). Source trust, recency, noise never enter the number.
- Smart Price's "street price" anchor is never populated — every real grade silently uses MSRP, the basis the spec forbids. Plus a cold-start cliff (no cohort = no score), cohort self-inclusion, and an n=5 staircase.
- Operational fragility — zero rate-limiting/backoff; catch-and-skip makes a 429 look like thin data; uncapped transcripts to Opus; no observability or resumability. The 24k-set sweep isn't yet practical or auditable.
- The live orchestrators are untested —
gradeSet/priceInventory have no test file; the 197 green tests cover the half that already works, giving false comfort where the bugs actually live.
Cross-cutting themes (patterns that recurred across audits)
- Faithful where deterministic, unfaithful where evidential. Pure-function math matches the spec; the moment a rule needs gating, pairing, or upstream separation, it degrades to an aspirational comment.
- Constants & comments are treated as if they were enforcement. The gates, FWI price-pairing, the §3.1 anti-double-counting rules, and the pre-availability gate are all "declared" but never run — and several comments actively misstate the behavior.
- The system fails OPEN, never closed. Missing data, throttling, empty cohorts, no street price, and sub-threshold evidence all resolve to a confident-looking number rather than an explicit "insufficient" — directly inverting the spec's core honesty principle.
- Honesty of output is the weakest property — despite being the stated goal. Coverage over-reports, confidence over-reports, Smart Price over-claims. The badges meant to convey trust are the least trustworthy components.
- Robustness, cost control & observability are uniformly absent at every external boundary (BrickLink, YouTube, Supadata, the LLM, Supabase) — the same first-pass-prototype maturity everywhere.
- Verification stops at the orchestration boundary. The pure core is tested adversarially; the live glue where every confirmed bug lives has no coverage.
The nine audits, in detail
Each auditor read the actual code + the methodology and sourcing specs, then returned strengths, weaknesses (with severity + file:line), and recommendations. All nine returned a verdict of mixed — strong bones, real gaps.
1 · Scoring engine math
mixed
The core 0–10→letter math is correct and tested, but the cap can rewrite the headline number (not just the tier), a no-data set computes a hard F instead of "data-light", and the three evidence gates are unenforced.
Weaknesses
highThree spec evidence gates are dead constants — MIN_COMMUNITY_TO_MOVE, LOW_DATA_FLOOR, and MIN_STRUCTURED_COUNT are never enforced, so 3 community grades, a 3-review point, and a 1–2-rating Brickset signal at trust 1.0 all move the grade.
src/lib/scoring/dimension.ts:43-51 · src/lib/sources/brickset.ts:61-72
mediumWhen the over-concentration cap binds it rewrites the headline score, not just the tier — a 3-survivor case computes a naive 7.8 but the cap yields 5.4 (≈ B+ → D+), undocumented.
src/lib/scoring/rollup.ts:91-105
mediumAn all-data-missing set computes score 0 → letter F. The spec wants a labelled "data-light estimate", not a confident F that reads as "bad set" rather than "no data".
src/lib/scoring/rollup.ts:99
mediumSpec refinements that gate/weight the blend are absent: price-pairing of value opinions (§1.3), the pre-availability lifecycle gate (§6.1), and outlier-dropping (§5.2). The blend treats every signal as already clean and paired.
design
lowFact dimensions get a hard 0.8 confidence floor even with zero corroboration; overall confidence is an unweighted mean; log1p(60) is a magic number absent from §10.
src/lib/scoring/dimension.ts:9,33-40
Top recommendations
→ Enforce the source-specific gates (with a test per gate). → Reconcile the insufficiency floor to LOW_DATA_FLOOR=5. → Special-case zero coverage to a withheld/null state, not F. → Compute the headline from uncapped weights and use the cap only for the tier.
2 · Methodology fidelity
mixed
The §3/§4 math is a faithful, number-for-number transcription — but fidelity collapses at the evidence gates, Feels-Worth-It loses its defining price-pairing property, and most §3.1 anti-double-counting rules are intentions, not enforced invariants (variety is scored in What-You-Get, the inverse of the spec).
Weaknesses (selected)
highLOW_DATA_FLOOR (5) — the spec's insufficient-data floor — is never enforced; the only gate is MIN_VOLUME_FOR_SENTIMENT (3), so a 3–4-review point is "scored" and counts as full coverage.
src/lib/scoring/dimension.ts:47
highFeels-Worth-It price-pairing is missing — §1.3 demands every value opinion be paired against the price at the time it was said (unpaired discarded). The code just nudges Smart Price toward the blended sentiment; the central fairness property is absent.
src/lib/scoring/score.ts:51-57
high§3.1 anti-double-counting is not enforced — whatYouGetScore uses element variety, which §3.1 says must live ONLY in Parts-for-Custom-Building, and partsForBuilding has no fact score at all: the inverse of the spec.
src/lib/facts/factScores.ts:44-56
mediumThe §4.3 "within-tier first" redistribution isn't implemented — a dropped HIGH cluster's weight leaks into MED clusters it shouldn't reach first.
src/lib/scoring/rollup.ts:47-79
mediumThe pre-availability lifecycle gate (§6.1) is absent — reveal-hype reviews feed Build Fun and Feels-Worth-It unfiltered.
design
mediumworthTheMoney "thumbs" has no defined 0–10 mapping — a raw 0/1 enters the blend as a 0–10 score, corrupting the Smart Price nudge.
src/lib/sources/contracts.ts:64
lowSmart Price "licensed premium" surfacing (§1.1) is unimplemented — a licensed set's lower ratio is silently penalized, the exact behavior the spec forbids.
src/lib/facts/smartPrice.ts:43-48
Top recommendations
→ Enforce the three gates in the scoring path. → Implement FWI price-pairing or explicitly descope it in the spec. → Add a test asserting variety contributes ONLY to partsForBuilding and novelty ONLY to rareNewParts. → Add tier-aware redistribution. → Define the thumbs→0–10 mapping.
3 · Discovery & set-matching
mixed
Well-engineered for a first pass, but the precision/recall posture is wrong for a 24k-set catalog with ~24% name overlap: the production path emits a LOW ref for every bare number and generic-name twin and hands ALL of them to the miner unfiltered, so the LLM judge is the sole backstop — violating the "two agreeing signals before auto-accept" rule. Recall is brittle too ("X Wing" never matches "X-Wing").
Weaknesses (selected)
highA single signal auto-accepts — a unique title-name match alone is MEDIUM and a lone description-number is LOW, both written to the index with only the judge as corroboration (spec §5.1 requires two agreeing signals).
src/lib/sources/youtube.ts:140-162
highloadReviewRefs returns every confidence tier unfiltered, so a haul turns each description number and every generic-name twin into a LOW ref — thousands of spurious pairs gated only by one Opus call each.
src/lib/sources/supabaseStore.ts:243-250
highThe name matcher requires EVERY token and glues hyphenated names into one token, so "X Wing Starfighter" never matches "X-Wing Starfighter" — verified empty.
src/lib/sources/youtube.ts:45-51
mediumTwo matchers with divergent guards coexist (matchVideoToSet vs matchVideoToSets); one is dead code in the bulk pipeline yet exported/tested as production.
src/lib/sources/youtube.ts:76-97
mediumThe year guard blinds the matcher to real year-numbered classic sets (2000, 2025 exist) — "LEGO 2025 unboxing" for set 2025 returns empty.
src/lib/sources/youtube.ts:149-151
mediumloadCatalog dedups by string sort while catalogRowsToKnownSets dedups by richest variant — the production sweep can disambiguate against the wrong (polybag/promo) name for a base.
src/lib/sources/supabaseStore.ts:287-298
Top recommendations
→ Require two agreeing signals before auto-accept; cap LOW refs per set. → Normalize hyphens and use an N-of-M token threshold instead of all-of-M. → Unify the two matchers. → Make loadCatalog pick the same canonical variant. → Emit per-confidence counts + judge-rejection rate so false-positive pressure is observable.
4 · LLM judge & distillation
mixed
Sound design (the coversSet gate, sarcasm flag, derived-only flow, per-ref try/catch, value-decay) — but missing the guardrails that make unbounded model output safe to feed into math.
Weaknesses
highNo validation gate (spec §8 demands one before scores count) — unvalidated scores enter grades.
distill.ts:36
highOutput is not clamped to 0–10 before the math — a model returning 11 or −3 corrupts the blend.
distill.ts:145
highPer-dimension confidence is missing from the schema, so confidence-weighting of distilled scores is impossible.
distill.ts:52
mediumComment like-counts are discarded — a brigade of low-quality comments reads as consensus.
mine.ts:77
mediumNo per-review derived-score cache — every re-grade re-distills, multiplying the uncapped transcript cost.
mine.ts:74
mediumNo retry/timeout — transient LLM errors are swallowed; Opus with no prompt caching is costly.
mine.ts:121 · distill.ts:135
Top recommendations
→ Build the validation gate; clamp output; add per-dimension confidence to the schema. → Add a per-review cache. → Use a cheaper model with caching for bulk distillation. → Pass like-counts; add retry + telemetry.
5 · Smart Price & fact scoring
mixed
The part-out costing and global cache are genuinely well-built; the fact scores are clean. But the headline value proposition isn't delivered in production: the "street price" anchor the whole story rests on is never populated by any live adapter, so every real grade silently uses MSRP — the basis the spec forbids. The cohort system is real but young (cold-start cliff, self-inclusion, n=5 staircase).
Weaknesses (selected)
highThe "street price" anchor is never populated in production — streetPrice is set only in demo/fixtures, so valueRatio's streetPrice ?? msrp ALWAYS resolves to MSRP, silently, for ~14% of the grade.
smartPrice.ts:23 · brickset.ts (never sets it)
highloadCohort doesn't exclude the set being graded from its own cohort — combined with the at-or-below percentile, a set is counted at-or-below itself, biasing its rank upward (worst at small n).
price.ts:74-81
highCohorts only exist where a corpus was pre-priced — every size bucket starts empty, so Smart Price (65% of Money & Worth) is dark for the overwhelming majority of sets. A cold-start cliff the spec's confident §1.1 language doesn't acknowledge.
design · gradeSet.ts:127-131
mediumMIN_COHORT=5 is far too low for a percentile — a 6-value staircase where one comparable swings the score ≈ a full point. The "precise 0–10" overstates a statistically meaningless rank.
constants.ts:145 · smartPrice.ts:32-36
mediumBasis/recency mismatch — the graded set is priced live this run while cohort rows use stale cached values; market drift shifts the numerator without the denominators.
price.ts:65-105
mediumThe theme→size-only fallback compares across value regimes (a 500-pc Botanical vs a 500-pc licensed set) — a category error the size bucket alone doesn't control for.
gradeSet.ts:128-131
Top recommendations
→ Implement the street-price anchor (BrickLink SET sold median, or a retailer street price) OR change the spec/copy to admit it's MSRP-anchored. → Add cohort self-exclusion + deliberate tie semantics. → Raise MIN_COHORT (15–25) and surface cohort n. → Flag cross-theme cohorts. → Treat the empty cohort as "insufficient", not 6.0.
6 · Data & persistence
mixed
The derived-only rule is honored (no raw text stored) and keying is consistent — but there's no freshness anywhere, and a non-atomic write can double-count a grade.
Weaknesses
highNo freshness/TTL anywhere — dimension_signals is a permanent cache hit (gradeSet skips mining forever); prices and scores are trusted indefinitely, so a re-grade never re-mines without manual --remine.
supabaseStore.ts · gradeSet.ts:89-93
highsaveDimensionSignals is a non-atomic delete-then-insert with the delete error discarded — a failed delete makes the insert append a second signal set, double-counting the grade. No transaction or retry.
supabaseStore.ts:147-168
Top recommendations
→ Add fetched_at and make reads age-aware so prices/signals expire. → Make saveDimensionSignals atomic and throw on the delete error; re-add the dropped FK or key all tables by base.
7 · Confidence & coverage honesty
mixed
The two badges ARE structurally distinct (breadth vs depth) and the over-concentration→data-light coupling works — the genuinely good part. But the honesty claim breaks on the details: the primary insufficient-data floor is never applied, two gates are declared-but-unapplied (with a comment that lies), and confidence is gameable — a lone source reports 100%.
Weaknesses (selected)
highLOW_DATA_FLOOR (5) is applied nowhere — a 3–4-review dimension is scored, marked "full coverage", and feeds the letter; the grade looks better-covered than the spec permits.
src/lib/scoring/dimension.ts:47
highThe doc-comment on resolveDimension claims signals are "gated by LOW_DATA_FLOOR", but the code gates by MIN_VOLUME_FOR_SENTIMENT — a documentation/behavior lie.
src/lib/scoring/dimension.ts:20
highA single noisy source can claim max confidence — variance of one signal is 0 → agreement 1.0, so a lone Amazon review (volume 60) reports confidence 1.000: the exact dishonesty §6 was written to prevent.
src/lib/scoring/dimension.ts:6-15
highSource trust, recency, price-pairing, and noise — all §6.1 confidence inputs — never enter the confidence number; an all-Amazon and an all-Brickset dimension with equal volume report identical confidence.
src/lib/scoring/dimension.ts:6-15
mediumFULL_COVERAGE_MIN=0.5 is a load-bearing magic number invented in the code, absent from the §10 tunables table.
src/lib/scoring/score.ts:33
mediumUnder the Investor lens, Hold-Its-Value is 30% of the grade but carries no confidence/coverage of its own — an investor grade can look as certain on thin data as on rich.
src/lib/scoring/score.ts:79
Top recommendations
→ Apply LOW_DATA_FLOOR as the real insufficient-data gate + fix the lying comment. → Stop a single source from claiming agreement 1.0 (require ≥2 independent sources). → Fold trust/recency/noise into confidence. → Actually apply the structured + community gates with tests. → Move FULL_COVERAGE_MIN into constants.
8 · Test quality & coverage
mixed
The tests that exist are strong and adversarial (the matcher and blend suites especially), and the 197 pass — but they only cover pure functions. With no mocking, every live orchestrator and I/O path is unverified, which is exactly where the confirmed bugs live.
Weaknesses
highThe live orchestrators gradeSet and priceInventory have no test file; mineSetReviews/distillReview are tested only via extracted pure helpers — so the cache/coverage/cohort branches, error paths, and the network/DB/LLM surface are unverified.
gradeSet.ts:71-150 · mine.ts:98-134 · price.ts:22-58
Top recommendations
→ Test gradeSet/mineSetReviews via the existing GradeDeps/MineDeps seams (they're already injectable). → Test failure paths. → Wrap distillReview's JSON.parse in try/catch. → Add per-adapter golden tests.
9 · Cost, scale & robustness
mixed
The discovery quota architecture and the part-price cache are the right cost designs, and writes are chunked with retry. But there's no rate-limiting anywhere, throttling is indistinguishable from missing data, the Opus transcript path is uncapped, and there's no observability or resumability — so the 24k-set tail isn't yet practical or auditable.
Weaknesses (selected)
highZero rate-limiting/backoff/Retry-After on any external API — all providers are raw fetch with no delay or concurrency cap.
bricklink.ts:107 · youtube.ts:184 · supadata.ts:50
highCatch-and-skip makes throttling indistinguishable from missing data — a 429 and a 404 both vanish from coverage, so a throttled set looks thin rather than degraded.
price.ts:52 · mine.ts:125 · gradeSet.ts:122
highYouTube transcripts go to Opus with no length cap — the largest uncontrolled per-set cost driver.
distill.ts:135 · mine.ts:75-80
highNo observability anywhere — bare catches, truncated logs, no metrics or run ledger, so degraded grades are invisible at scale.
batch-ingest.ts:105-107
mediumNo resumability/idempotent cursor — sweeps re-page everything and a batch crash restarts from zero; BrickLink pricing is strictly sequential.
discover.ts:38-69 · price.ts:43-55
Top recommendations
→ Add a shared rate-limited, concurrency-bounded HTTP client that surfaces 429s distinctly. → Cap transcript input + use a cheaper bulk model + a cost counter. → Add a run ledger + resumable cursor (last-seen videoId per channel). → Add structured logging + per-run health/cost metrics; validate env vars up front.
What to make of it
The headline is fair and worth sitting with: the parts of this engine you can write a unit test for are excellent; the parts that depend on messy real-world evidence are not yet enforced. That's a good problem to have — the skeleton is sound, the spec is right, and almost every finding is a specific, local, testable fix rather than a redesign. The single highest-leverage move (the evidence gates) is small and restores honesty exactly in the long-tail regime the whole methodology was built to protect. The two that follow — a real street-price anchor and basic operational hardening (rate limits, observability, capping the Opus path) — are what stand between "a strong prototype" and "trustworthy to ship grades to users."
Produced by a 9-agent adversarial audit on 2026-06-18. Companion docs: docs/grading-walkthrough.html (the step-by-step procedure on the live Razor Crest grade), docs/how-it-works.html (technical), docs/how-it-works-plain.html (plain-English overview).