A complete, code-grounded plan to close every problem the 9-agent audit found — sequenced into five waves, with each fix written as the precise change, the test to write first, the effort, and the risk. Eight design agents read the real source and the methodology spec to build it.
How to read this. The top half is the map: the principles every fix follows, the five-wave roadmap, the handful of calls that are genuinely yours to make, and the two things that live outside the code. The bottom half is the detail: all eight workstreams, every fix laid out the same way. You can decide the whole direction from the top half alone; the detail is there for when we build.
62
concrete fixes
5
waves
~2 wks
to restore honesty (waves 0–1)
~10–12 wks
full plan, focused
The principles every fix follows
Fail closed, never open.The audit's loudest pattern: when data is missing, thin, throttled, or unvalidated, the engine today returns a confident-looking number. Every fix bends the other way — when in doubt, the grade abstains ("insufficient / data-light / withheld"), it does not guess.
Test-first, always.Every fix starts with a failing test (that's how this codebase already works). The live "glue" code that talks to the internet and database — where every confirmed bug lives — gets tested through the dependency seams that already exist, so we can finally cover it.
One dashboard for every knob.Every threshold lives in <code>src/config/constants.ts</code>. No more magic numbers buried in logic, no more constants that are defined but never used. A gate that isn't wired up is treated as a bug.
Say what the code does.Several comments today describe behavior the code doesn't have. Part of every fix is making the words match the work — and adding a test so a guarantee can't silently rot again.
Derived data only.Unchanged and sacred: we store distilled numbers plus a source link, never raw review text. Nothing in this plan touches that.
The roadmap — five waves
Ordered by leverage and dependency. Each wave ends at a real, shippable milestone — you could stop after any wave and the engine is strictly more trustworthy than before it. Value is front-loaded: the cheapest wave restores the most trust.
0
Build the safety net enabler · ~3–4 days
Before changing how grades are computed, make it safe to change them. Stand up the in-memory test double for the database and the first real tests for the three "live" pieces that have none today — <code>gradeSet</code>, <code>mineSetReviews</code>, <code>priceInventory</code> — using the injection seams that already exist. And fix the one data-corruption bug right now: make the signal-save atomic so a half-failed write can never double-count a grade. After this, every later fix is test-driven at the level where the bugs actually live.
Ships: A test harness for the real machinery + the double-count bug closed.
workstreams: data (the atomic write + the fake-DB seam) · test scaffolding for gradeSet / mine / price
1
Restore honesty highest leverage · ~2 weeks
This is where the audit screamed loudest and the fixes are smallest. Turn the three evidence gates on for real (the #1 finding, 5 of 9 reviewers). Make confidence un-gameable — a lone review can no longer read 100%, an all-Amazon dimension reads lower than an all-Brickset one, and a fact built on a weak fallback stops claiming a high floor. And make the score itself fail closed: a no-data set returns a "withheld" state instead of a confident F, and the anti-lopsidedness cap stops secretly rewriting the headline number. Mostly small, local, immediately testable changes — and together they turn the letter and its badges from confidently-wrong into honestly-abstaining.
Ships: The grade and its two badges stop overstating on thin or gamed data.
Make the headline feature real the marquee fix · ~2 weeks
Smart Price is the "is this set worth the money?" story — and today it silently measures against the sticker price because no live source ever provides a real street price. Build the street-price anchor (a BrickLink "what sets actually sold for" lookup, cached) and label each grade honestly as street-anchored or MSRP-anchored. Then clean up the comparison group: exclude the set from its own cohort, raise the minimum size so a percentile means something, match price freshness, and abstain when the cohort is too thin instead of returning a confident "C." This wave also lands the data-freshness work (so prices and signals expire instead of being trusted forever) since Smart Price depends on it.
Ships: "Worth the money" measures real market value — or honestly says it couldn't.
workstreams: smartprice · data (freshness / TTL / keying)
3
Trust what feeds the math ~2–3 weeks
Now harden the inputs. The AI judge can currently hand the math an out-of-range or unvalidated score — clamp it, validate it, give every distilled opinion its own confidence that flows into the blend, cache it so re-grades don't re-pay, and (the big call) hold distilled "feelings" out of grades until a labeled accuracy check passes. And fix discovery's precision: stop auto-accepting a single weak match, restore recall on hyphenated and long licensed names, and emit counts so false matches become visible instead of silently leaning on one AI call each.
Ships: Bounded, validated, honestly-matched evidence going into every grade.
workstreams: judge · discovery
4
Harden for scale & finish fidelity ~2–3 weeks
Everything needed before running the full 24,000-set catalog: one shared web client that rate-limits, backs off, and tells a "we got throttled" error apart from "there's no data" (today they look identical); a cap and a cost counter on the expensive AI step; and a run ledger plus a resume cursor so a crashed batch doesn't restart from zero. Finish the remaining methodology gaps in the same wave — fix where "variety" and "novelty" are scored, map the thumbs input correctly, and either build or formally shelve the two big deferred mechanics (Feels-Worth-It price-pairing and the pre-release-hype filter).
Ships: A safe, observable, resumable batch run — and every spec rule honored or explicitly shelved.
workstreams: ops · fidelity (the remainder)
Your decisions — resolved
These forks weren't an engineer's to decide — they change scope, cost, or what the product promises. You made all ten. Here's where each landed, with the reasoning. (The smaller calls are also threaded into the affected workstreams below, marked "Updated by your decisions.")
1Smart Price: build a real street price, relabel as sticker-price, or both?
Right now it silently uses the sticker price (MSRP) — the audit's "most misleading thing" in that area. Building a real "what it actually sells for" anchor is the entire value proposition, but for brand-new or low-volume sets there isn't enough sales data to be reliable.
Your call <b>Hybrid.</b> We build the BrickLink street-price lookup (one cheap, cached call per grade) and use it wherever sales data is solid; where it isn't, we fall back to MSRP and <b>label the grade "MSRP-anchored"</b> so it's never silently one pretending to be the other.
2The AI judge: hold fan sentiment out of grades until an accuracy check passes, or keep it flowing with a badge?
The methodology says distilled review scores shouldn't count until they pass an accuracy check. That check doesn't exist yet, and building it needs a hand-labeled set of reviews to measure against — which is your time, not an engineer's.
Your call <b>Keep sentiment live, with a visible "unvalidated" badge.</b> Fan sentiment keeps flowing into grades as it does today, but every grade that leans on it is stamped "unvalidated" until the accuracy check passes — richer grades now, honest about the caveat. The badge comes off once you've built the labeled corpus (see the deep-dive below).
3Bulk distillation: keep using the top-tier (expensive) model, or use something cheaper at scale?
Running the full catalog through the most expensive model (Opus) is cost-prohibitive at scale; the cheapest model (Haiku) may read LEGO sarcasm less reliably — and sarcasm detection is the distiller's whole job.
Your call <b>Sonnet 4.6 as the default bulk model</b> (<code>claude-sonnet-4-6</code>) — the middle ground you picked: meaningfully cheaper than Opus, far better at nuance and sarcasm than Haiku. Opus stays available as an override for flagship/high-scrutiny sets, and we still validate Sonnet's sarcasm handling against your labeled corpus before trusting it at scale.
4One grade per set, or per variant? (e.g. does the polybag share the main set's grade?)
The data is inconsistent today — some tables key on the base number (75355), others on the variant (75355-1), and they don't line up. Per-base is the smaller fix; per-variant is richer but a larger re-keying job.
Your call <b>Per variant — with polybags explicitly deferred.</b> The grade unit is the variant (e.g. 75355-1, the main buildable set); we grade the main variant and shelve polybags/minor variants until much later. Deferring polybags removes most of the near-term pain (they're the worst offenders for name-collisions and cohort noise), so we get the richer per-variant path without paying full price for it now. This adds some re-keying scope versus per-base — flagged honestly in the Data workstream.
The six smaller tuning calls — also resolved
The redundant low gate (3 vs 5)Today two thresholds (3 and 5) overlap. <b>Decided: keep both and build the missing feature</b> — 5 stays the "enough to score a point" floor, and 3 becomes a real, newly-built trigger where light fan sentiment (3–4 mentions) may nudge a fact-based score's <i>confidence</i> (not the score). Spec-faithful.
The Smart Price cohort minimumA ranking over very few comparable sets is shaky. <b>Decided: keep the low floor (~5) and lean on a confidence ramp</b> — thin cohorts still score but are visibly low-confidence, and confidence scales up with how many comparables there are. Favors coverage now, with honesty carried by the label.
Feels-Worth-It price-pairingIts defining feature needs a history of past sale prices that doesn't exist yet. <b>Decided: shelve it on the record</b> — marked "deferred" in the spec with a test pinning current behavior; revisit once we're collecting price history. An honest "not yet" beats a misleading half-build.
The pre-release-hype filter<b>Decided: don't build it — you're keeping pre-release sentiment as legitimate signal.</b> Reveal/early reactions are part of how the community defines a set, so we won't discount them. Instead, grades become <i>living</i>: a <b>monthly re-grade CRON</b> re-assesses each set as opinions mature (see "Living grades" below). The spec's pre-availability-discount rule is intentionally retired.
The "display-only" community countWhen community votes are below the counting threshold, the spec wants them shown but not counted. <b>Decided: yes, surface it</b> — a small optional field on the grade output so the page can show "N community grades — not yet load-bearing." Additive, breaks nothing, and lets people see their votes building toward the threshold.
Throttle (rate-limit) behaviorWhen an outside service blocks us mid-grade, the result can be badly incomplete. <b>Decided: publish it, stamped "degraded," with an explicit breakdown of how incomplete it is</b> (e.g. "only 20% of parts priced, 0 reviews mined") — transparency over withholding. The engine still must detect a throttle and tell it apart from genuinely-missing data; the monthly re-grade then completes it.
The labeled review corpus — what it actually asks of you
You've seen "the corpus" referenced a few times, and two of your decisions depend on it (taking the "unvalidated" badge off grades, and trusting Sonnet at scale). Here's exactly what it is and what you'd do — in plain terms.
The problem it solves. The AI reads a review and outputs judgments — "build fun: 8," "this is sarcastic," "this review isn't even about this set." To trust those judgments, we have to know how often the AI is right. The only honest way to measure that is to compare the AI's answers against a batch of reviews where a human who knows LEGO has already written down the correct answer. That human-answered batch is the corpus.
Why it has to be you. Labeling needs LEGO judgment and brand taste — knowing that a reviewer saying "well, that's certainly a color choice" is throwing shade, not giving praise. That's you, not an engineer. The engineering half — the harness that runs the AI over your labels and scores it — is ours.
1
We hand you a batch. We pull a spread of real reviews — start with ~50, grow to ~100–150 — deliberately mixed: glowing, harsh, sarcastic, lukewarm, and a few that are about a different set, so the AI's "is this even the right set?" check gets tested too.
2
You skim each and jot a few quick answers. Per review: roughly what each dimension deserves (a number or a range is fine), is it sarcastic, and is it really about this set. ~2–4 minutes each. No writing, just judgments.
3
That's the whole ask: ~3–6 hours, once, for a solid first pass of 100 — easily split across a few sittings while you're around the sets you already know.
4
We grade the AI against you. We run it over the same reviews and compare. Clears the bar → the "unvalidated" badge comes off and Sonnet is cleared for the catalog. Falls short (usually on sarcasm) → we tune the prompt and re-check.
What the few hours buy you: it's not one-and-done busywork — it becomes the permanent accuracy backstop. Every time we change the model or the prompt afterward, we re-run it against your labels automatically. Invest once, and every future change is checked against your taste forever.
Want me to make this painless? I can scaffold a tiny labeling sheet that pulls the reviews and gives you the fields to fill — so it's a checklist, not a project.
Living grades — your call on hype, made into a feature
You decided that pre-release and reveal sentiment is legitimate signal, not noise to filter out — "these are some of the most crucial touch points that define the community's perception." That's a deliberate change to the original methodology (which wanted to discount pre-availability hype), and the plan now honors it: we don't filter early sentiment.
Instead, the concern that perception matures over time gets solved a better way — grades become living. A monthly re-grade CRON re-assesses every set, so a grade captured during reveal hype naturally settles as hands-on reviews and real sale prices accumulate. The data-freshness work already in this plan (a ~30-day score expiry) is exactly the trigger this needs — your monthly re-grade is that mechanism. It also quietly fixes the throttle case: a "degraded" grade simply completes itself on the next monthly pass.
Two follow-ups this creates: update the methodology spec to record the new stance (early sentiment is in, pre-availability discount is out), and add the monthly re-grade job as a small new workstream once the freshness columns land.
Two things that live outside the code
Database migrations (Supabase)
Several fixes need new columns (a <code>fetched_at</code> timestamp on prices/signals/scores, a <code>price_basis</code> on cohort rows) and ideally one stored procedure for a truly-atomic signal write. These live in the remote database, not in this repo, so they must be applied to Supabase in lockstep with the code. The interim JS guard closes the double-count bug today; the stored procedure is the durable follow-up.
A labeled review corpus
Two of your decisions hinge on this: taking the "unvalidated" badge <i>off</i> grades, and trusting Sonnet for the bulk catalog. Both need a set of hand-labeled reviews to measure the AI against — especially the sarcastic ones. It's a product input only you can give. There's a full deep-dive on exactly what it asks of you just below.
What to expect while we do this
Because of your calls, grades stay rich through this work rather than going quiet. The honesty shows up as labels and badges, not a coverage cliff: fan sentiment keeps flowing (stamped "unvalidated"), Smart Price keeps scoring thin cohorts (stamped low-confidence), and throttled runs still publish (stamped "degraded" with a breakdown). The one real shift is in Wave 1: a set resting on genuinely thin evidence — say 3 reviews — will correctly drop to "data-light" instead of wearing a confident letter. That's the audit's central problem being fixed, made visible. And once the monthly re-grade lands, even that is temporary: a set graded thin today automatically improves as reviews and sales data accumulate.
The eight workstreams, in detail
Jump to any one:
Wave 1Workstream 1 of 8
Enforce the three evidence gates
Done looks like: The three named minimum-evidence gates (LOW_DATA_FLOOR, MIN_STRUCTURED_COUNT, MIN_COMMUNITY_TO_MOVE) actually run in the scoring path, each is locked in by a dedicated test so it can never silently die again, community signals below their gate are display-only (0 to value AND confidence but a surfaced count), and every doc-comment describes what the code truly does.
Updated by your decisionsYour calls: the lower gate (3) is <b>kept and wired</b> — light sentiment will be built to nudge a fact's confidence, with 5 as the scoring floor. The below-threshold community count <b>is surfaced</b> on the grade ("N grades, not yet load-bearing").
The approach
The audit's #1 finding is true and I confirmed it directly: a grep shows LOW_DATA_FLOOR, MIN_STRUCTURED_COUNT, and MIN_COMMUNITY_TO_MOVE have zero usages anywhere outside constants.ts and tests; only MIN_VOLUME_FOR_SENTIMENT is wired up, and it sits in dimension.ts:47 occupying the slot the spec and the lying doc-comment both assign to LOW_DATA_FLOOR. The fix has a natural order. First, correct the in-place lie: swap dimension.ts to gate on LOW_DATA_FLOOR (the highest-leverage one-line change, since it also fixes coverage/confidence honesty downstream in score.ts for free, because score.ts already keys off state==='scored'). Second, gate the two source-specific streams where their signals are born/assembled — Brickset structured ratings need MIN_STRUCTURED_COUNT applied (today brickset.ts emits at >=1 review and its own comment admits the gate 'is applied upstream' though no upstream caller exists), and community signals need MIN_COMMUNITY_TO_MOVE applied as a display-only gate (contribute 0 to value and confidence, but surface the count). Because community signals must still surface a count while contributing nothing, this third gate needs a small type-surface addition to DropScore so the count has somewhere to live. Each gate gets a failing test first (TDD red), and pure-logic gates are unit-tested directly while the assembled/orchestrated paths use the existing GradeDeps seam. The MIN_VOLUME_FOR_SENTIMENT question (keep as a distinct lower enrichment trigger vs retire it) is a spec decision, not an engineering one, so it's in decisions with a recommendation rather than silently chosen.
Effort~3-4 days: 4 S (LOW_DATA_FLOOR gate+comment, Brickset gate+comment, fakes comment, coverage guard test + constant move) + 1 M (community display-only gate with the DropScore/DimensionInput type-surface addition and seam test), plus one deliberate fixture re-baseline pass.
Fixes5 · 4 small · 1 medium · 0 large
Depends onnone — this workstream is self-contained in the scoring/assembly/source-adapter layers and is the keystone the audit ranks #1. Note for coordination: any workstream that re-baselines src/__tests__/fixtures or changes DropScore consumers (UI/renderer) should sequence AFTER fix 3, since this workstream changes which fixture dimensions count as scored and adds an optional communityDisplayOnly field to DropScore.
The fixes, in order
small
1. Gate signal-only dimensions on LOW_DATA_FLOOR and stop the doc-comment lie
AddressesDead LOW_DATA_FLOOR + false doc-comment: src/lib/scoring/dimension.ts:47 (gate uses MIN_VOLUME_FOR_SENTIMENT=3) and dimension.ts:18-21 (comment claims 'gated by LOW_DATA_FLOOR'). Audits 0,1,6.
The changeIn resolveDimension (dimension.ts), change the insufficiency gate at line 47 from totalVolume < MIN_VOLUME_FOR_SENTIMENT to totalVolume < LOW_DATA_FLOOR, and update the import on line 1. Pending the decisions entry, either drop the MIN_VOLUME_FOR_SENTIMENT import entirely (retire) or keep it only as a separate enrichment trigger. Rewrite the resolveDimension doc-comment (lines 17-21) and the inline comment at lines 45-46 to state the real gate (LOW_DATA_FLOOR) — they currently lie. Do NOT touch signalConfidence. This change alone also corrects score.ts coverage/confidence honesty (lines 79-92) for free, since that code already keys off state==='scored' and confidence of scored dims.
Test firstIn src/lib/scoring/dimension.test.ts add: 'marks a signal-only dimension below LOW_DATA_FLOOR (total volume 4) as insufficient' — resolveDimension({id:'buildFun', signals:[{source:'youtube',score:8,volume:4}]}, noNa) expects state==='insufficient', value===null, confidence===0. Companion: 'scores at exactly LOW_DATA_FLOOR (volume 5)' expects state==='scored'. Both fail today (volume 4 currently returns 'scored' because the live gate is 3). Keep the existing volume-1 test green.
Watch outExisting dimension.test.ts line 20-25 uses volume 12 (still scored) and line 33-37 uses volume 1 (still insufficient) — both stay green. But fixtures in src/__tests__/fixtures and the score.test.ts flagship/obscure expectations may shift tier: any fixture dimension sitting at volume 3-4 will flip scored→insufficient, which can move a 'full' set to 'data-light' or change a cluster value. Re-run the full scoring suite and re-baseline fixtures intentionally (do not loosen assertions to hide a real tier change).
small
2. Apply MIN_STRUCTURED_COUNT to Brickset structured ratings
AddressesDead MIN_STRUCTURED_COUNT + lying comment: src/lib/sources/brickset.ts:56-60 ('gate is applied upstream, not here') and the emit at brickset.ts:61-72 (signal emitted at >=1 review). No upstream caller in gradeSet.ts applies it. Audits 0,1,6; data-sourcing spec line 113.
The changeIn bricksetReviewsToSignals (brickset.ts), discard the whole structured contribution when the review count is below the gate: per §2.2 'below that they're discarded' (not down-weighted, not shown low-confidence). Add an early if (reviews.length < MIN_STRUCTURED_COUNT) return {} after the empty check on line 63, importing MIN_STRUCTURED_COUNT from @/config/constants. Note the spec gates on review COUNT, and the per-sub-rating values.length can be lower than reviews.length when some reviews omit a sub-rating — gate on reviews.length to match the spec's '>=8 reviews' wording. Rewrite the lying comment at lines 56-60 to say the gate is applied HERE. This keeps the gate in the pure adapter (unit-testable, no live call) rather than scattering it into the orchestrator.
Test firstIn src/lib/sources/brickset.test.ts add: 'discards structured signals below MIN_STRUCTURED_COUNT' — build 7 BricksetReview objects each with rating.buildingExperience, expect bricksetReviewsToSignals(reviews) to return {} (no buildFun signal). Companion: '8 reviews produce signals' expects signals.buildFun to be defined. The 7-review test fails today (any single review emits a signal).
Watch outIf any existing brickset.test.ts case or fixture relies on a <8-review set producing signals, it flips to empty — update those expectations to reflect the intended discard. Confirm gradeSet.ts's bricksetReviews source stat (line 141, reviews.length) is unaffected; it reports raw count, not gated, which is correct for the sources breakdown.
medium
3. Make community signals display-only below MIN_COMMUNITY_TO_MOVE (0 to value AND confidence, but surface a count)
AddressesDead MIN_COMMUNITY_TO_MOVE + brigading hole: src/lib/sources/fakes.ts:42-45 (push at volume:1, no gate), assemble.ts mergeInputs:7-17 (concatenates ungated). A single community vote moves value and confidence. Audits 0,1,6; methodology §2.2/§4.5/§6.2.
The changeTwo parts. (a) GATING — the cleanest single chokepoint is mergeInputs (assemble.ts), which already groups signals by dimension and is pure/unit-tested: before attaching signals, partition each dimension's incoming signals into community vs non-community (by source==='community'); sum community volume; if the community sum < MIN_COMMUNITY_TO_MOVE, drop those community signals from the blend entirely (so they reach neither blendSignals nor signalConfidence → 0 to value AND confidence) but record the dropped count. Do NOT gate inside fakes.ts getSignals — the store should faithfully return submissions; gating is an engine concern so the live store and the fake share one code path. (b) SURFACING — add a field so the dropped count is not lost: extend DropScore (types.ts:139) with communityDisplayOnly?: { dimensionId: DimensionId; count: number }[] (or a per-dimension map), populated in score.ts from the dropped counts threaded through DimensionInput. Thread the count via a new optional field on DimensionInput (e.g. displayOnlyCommunityCount?: number) set by mergeInputs and read by computeDropScore. Honor the house rule: MIN_COMMUNITY_TO_MOVE already lives in constants.ts — import it, no new magic number.
Test firstUnit (pure): in src/lib/pipeline/assemble.test.ts add 'drops community signals below MIN_COMMUNITY_TO_MOVE but records the count' — mergeInputs(factInputs, {buildFun:[{source:'community',score:9,volume:1}] x4}) expects the resulting buildFun input to carry NO community signal in .signals and a displayOnlyCommunityCount of 4. Companion '5 community submissions are load-bearing' expects the signals to pass through. Seam test (fakes.test.ts or a score-level test): submit 4 community grades via InMemoryCommunityStore → assembleSetInputs → computeDropScore, assert the affected dimension's confidence and value are unchanged vs the no-community baseline, AND that score.communityDisplayOnly surfaces count 4. The 4-submission test fails today (1 submission already moves value+confidence).
Watch outTouches the DropScore type contract — any consumer/UI reading DropScore must tolerate the new optional field (it's additive/optional, so low blast radius, but grep the renderer). Mixing community + mined signals on the same dimension: ensure only the community portion is dropped, never the mined signals sharing that dimension. Edge case: community sum exactly == MIN_COMMUNITY_TO_MOVE must be load-bearing (use <, not <=). Volume semantics: each community submission enters at volume:1 (confirmed fakes.ts:44), so summed community volume == submission count, which is what the gate intends — but if a future community signal arrives with volume>1 this conflates submissions with volume; assert the volume:1 invariant in a test or gate on signal count rather than volume sum and document the choice.
Depends onShould follow fix 1 only to avoid two simultaneous re-baselines of the scoring fixtures; logically independent.
small
4. Stop fakes.ts from emitting at silent volume:1 in a way that hides the gate, and align its comment
Addressessrc/lib/sources/fakes.ts:42-50 (push helper emits volume:1 community signals with no annotation). Charter anchor: 'community signals pushed at volume 1, ~44'.
The changeKeep volume:1 (one submission == one opinion is correct), but once fix 3 lands, add a comment in fakes.ts getSignals clarifying that gating is the engine's job (mergeInputs), not the store's, so a future reader does not 're-add' a gate here and double-gate. No behavioral change in fakes.ts itself — this is a correctness-of-intent + comment fix that prevents the gate being re-buried. If the team prefers the gate visibly near the store, this is where a decision lands (see decisions).
Test firstCovered by the fix-3 seam test (community-below-gate is display-only end-to-end through the fake store); no separate red test strictly required, but keep fakes.test.ts line 38-44 ('returns it as signals') green to prove the store still faithfully returns raw submissions.
Watch outMinimal. Risk is purely that someone double-gates later; the comment mitigates it. If fix 3 instead chooses to gate at the store, this fix changes from a comment to the actual gate location — keep them consistent.
Depends onfix 3 (same gate; this is its companion comment/intent cleanup).
small
5. Lock coverage tiering to the corrected insufficiency gate
Addressessrc/lib/scoring/score.ts:79-92 coverage/confidence loop + the FULL_COVERAGE_MIN tier decision (line 90-91). Audit 1: 'Coverage tiering and confidence are computed against the wrong gate.'
The changeNo logic change needed in score.ts IF fix 1 lands — score.ts already computes coverage from state==='scored' and confidence from scored dims' confidence, so correcting the gate in dimension.ts automatically makes a volume-3/4 dimension count as an insufficient coverage gap rather than a scored point. The work here is a guard TEST that pins this wiring so a future regression in dimension.ts can't quietly inflate coverage again. Also move/confirm FULL_COVERAGE_MIN (currently a local const in score.ts:33) into constants.ts per the house rule that all tunables live in one file.
Test firstIn src/lib/scoring/score.test.ts add 'a dimension backed by sub-floor volume is an insufficient coverage gap, not a scored point' — construct inputs where one in-grade dimension has community/mined signals summing to volume 4; assert dimState(s, thatDim)==='insufficient' AND that it does not increment the scored count (e.g. coverageTier flips to 'data-light' or the cluster's scoredCount excludes it). Fails today (counts as scored).
Watch outRe-baselines the same fixtures as fix 1 (do them together). Moving FULL_COVERAGE_MIN into constants.ts is a mechanical refactor but touches the constants.test.ts surface — add it to any 'gates exist' assertion.
Depends onfix 1.
Calls to make in this workstream
Is MIN_VOLUME_FOR_SENTIMENT a real distinct threshold, or should it be retired in favor of LOW_DATA_FLOOR?
Options: (A) Retire it: delete the constant and its only use, gate everything on LOW_DATA_FLOOR=5. Simpler, one floor, removes the conflation the audit flagged. (B) Keep it as a genuinely distinct, LOWER enrichment trigger (3): a dimension is 'scored' only at LOW_DATA_FLOOR=5, but mined sentiment is *allowed to enrich/adjust* a fact-defined dimension's confidence at >=3. This matches §2.2's wording 'Mined sentiment enrichment fires only at MIN_VOLUME_FOR_SENTIMENT' as a separate clause from the LOW_DATA_FLOOR insufficiency clause.
Recommended: (B) Keep it, but only if you actually wire the enrichment path it names; otherwise (A). The spec clearly intends two different things (enrichment-fires vs point-is-scored), so retiring loses spec intent — but today there is NO enrichment path that consumes a 3-volume signal, so an unwired constant is just another dead gate. Recommend: keep the constant, and in fix 1 use it ONLY in the fact-defined branch (dimension.ts:33-40) to decide whether weak sentiment may nudge a fact dimension's confidence, while LOW_DATA_FLOOR gates the signal-only branch. If Seth doesn't want to build the enrichment nudge now, retire it (A) and add a spec note so §2.2 stops promising a behavior the code doesn't have.
Where should the community MIN_COMMUNITY_TO_MOVE gate physically live — in the engine (mergeInputs) or at the store boundary (fakes.ts / live CommunitySubmissionStore)?
Options: (A) Engine/mergeInputs: one code path for fake and live stores, gate is a pure-engine concern, easy to unit-test, store stays a faithful data source. (B) Store boundary: gate visible right where community data enters, but must be duplicated in every store implementation (fake + live Supabase) and risks drift.
Recommended: (A) Engine/mergeInputs. It keeps the gate in one pure, tested place that both the fake and the eventual live Supabase store inherit for free, and it matches the existing pattern where resolveDimension/score.ts own the scoring rules. The store's only job stays 'faithfully return submissions + enforce the written-why rule'.
How should the display-only community count be surfaced on the result?
Options: (A) Add an optional array/map field to DropScore (e.g. communityDisplayOnly: {dimensionId, count}[]). (B) Reuse the existing DimensionResult by adding an optional displayOnlyCommunityCount to each. (C) Fold it into the existing coverage note / confidence note only.
Recommended: (A) plus thread via an optional field on DimensionResult is cleanest for the UI string the spec wants ('N community grades, not yet load-bearing'). It's additive and optional so it won't break existing DropScore consumers. Avoid (C): a bare note loses the per-dimension count the spec explicitly calls for.
New constantsFULL_COVERAGE_MIN
SequencingFront-load fix 1 (LOW_DATA_FLOOR + doc-comment) — it is one line of logic, kills the highest-severity lie, and automatically fixes coverage/confidence honesty (fix 5 then becomes mostly a guard test). Re-baseline the scoring fixtures once, deliberately, right after fix 1+5 land together since they share the same fixture impact. Fix 2 (Brickset gate) is fully independent and pure — it can go in parallel or any time. Do fix 3 (community display-only) after fix 1 so there is a single intentional fixture re-baseline rather than two; fix 4 is its trailing comment cleanup. Quick wins to land first day: the two doc-comment corrections (dimension.ts and brickset.ts) and the Brickset gate, all S-sized and low-risk. The only M is the community display-only gate because it adds an optional field to the DropScore contract.
Wave 1Workstream 2 of 8
Make confidence honest and un-gameable
Done looks like: The confidence number (0–1) reflects the real strength of the evidence behind a value: a lone Amazon review never reads 1.000, an all-Amazon dimension reads lower-confidence than an all-Brickset one, a fact derived from a fallback/empty-cohort basis reads lower than one from real part-value, the holdValue badge carries its own confidence (so an Investor grade's reported certainty reflects its 30%-weighted cluster), and every tunable lives in constants.ts.
Updated by your decisionsYour call: Smart Price keeps a <b>low cohort floor (~5)</b> and leans on the <b>confidence ramp</b> here — thin cohorts score but read visibly low-confidence rather than being gated out.
The approach
All the damage lives in two pure functions — signalConfidence() and the fact-floor branch in resolveDimension() (dimension.ts), and the overall-confidence rollup in computeDropScore() (score.ts) — which makes this a tightly-scoped, fully unit-testable workstream with no live calls. The strategy is: (1) rebuild signalConfidence from volume×agreement into a product of four spec-§6.1 factors (volume, agreement, source-trust, recency) plus a source-diversity guard that prevents a single source from ever reaching agreement 1.0; (2) replace the unconditional 0.8 fact floor with a basis-aware floor so a fallback/thin-cohort fact reports lower confidence than a real-part-value one — which requires threading a small basis/fallback marker from smartPrice.valueRatio through factInputs into DimensionInput (the only cross-file plumbing); (3) make overall confidence coverage-aware and give the holdValue badge its own confidence; (4) move every new magic number (log1p denominator, FULL_COVERAGE_MIN, trust/diversity floors, fallback penalty) into constants.ts behind named exports. TDD throughout: each fix starts with a failing unit test asserting the exact dishonest number today, then the implementation flips it. Sequence puts the pure-math wins first (no plumbing), then the basis-threading, then the rollup/badge changes that depend on dimension confidence being correct.
Effort~1 week: 3 S (log1p extraction, FULL_COVERAGE_MIN move, diversity cap), 4 M (trust/recency fold, basis-aware fallback floor + plumbing, weighted overall confidence, holdValue badge confidence). No L items — the workstream is concentrated in two pure functions plus one thin cross-file basis marker; no new infrastructure or live calls.
Fixes7 · 3 small · 4 medium · 0 large
Depends onCoordinates with two siblings but is not blocked by them. (1) The 'evidence gates / fail-closed' workstream owns ENFORCING LOW_DATA_FLOOR as the insufficient threshold and DISCARDING sub-gate signals (MIN_STRUCTURED_COUNT for Brickset, MIN_COMMUNITY_TO_MOVE for community) before they reach the blend; my confidence math consumes whatever signals survive those gates, so honest confidence is strengthened by that work but doesn't require it to ship. (2) The 'Smart Price / data-sourcing' workstream owns the streetPrice-never-populated and empty-cohort-neutral-0.5 bugs; my Fix #4 only needs that workstream to surface a basis/fallback marker on the smartPrice fact (see decisions) — recommend they populate DimensionInput.basis so we don't both edit smartPrice.ts. If sequencing forces it, Fix #4 can self-detect fallback in factInputs as a stopgap.
The fixes, in order
small
1. Source-diversity guard: a lone source can't claim agreement 1.0
AddressesAudit 'Confidence & coverage honesty' high finding — lone Amazon review volume=60 → confidence 1.000 because variance of one signal is 0 → agreement 1.0 (src/lib/scoring/dimension.ts:6-15, esp :13). Charter goal: 'a lone Amazon review must NOT read 1.000'.
The changeIn signalConfidence (dimension.ts), after computing the variance-based agreement, gate it by the number of DISTINCT sources backing the signal set. Compute distinctSources = new Set(signals.map(s => s.source)).size. When distinctSources < CONFIDENCE_MIN_SOURCES_FOR_FULL_AGREEMENT (new constant = 2), cap agreement at SINGLE_SOURCE_AGREEMENT_CAP (new constant, e.g. 0.6) so one self-consistent source can never read as fully corroborated. Keep the 1/(1+variance) form for the multi-source case. This is the central un-gameable change; everything else stacks on it.
Test firstNew cases in dimension.test.ts: (a) resolveDimension({id:'buildFun', signals:[{source:'amazon',score:9,volume:60}]}) → confidence strictly < 1.0 and ≤ the new cap (assert toBeLessThan(0.9), and document expected ≈cap×volumeFactor×trust); (b) a 3-distinct-source agreeing set scores strictly higher confidence than the same total volume from ONE source — pins the diversity gate, not just the cap.
Watch outThe existing dimension.test.ts 'gives higher confidence with more agreeing evidence' (lines 39-56) already compares 1 vs 3 sources and must still pass — verify it does. Score.test.ts flagship/obscure assertions key off coverageTier and score ranges, not raw confidence, so they should be unaffected; run the full suite to confirm no fixture's confidence crosses a tier boundary indirectly.
medium
2. Fold source trust + recency into the confidence number
AddressesAudit high finding — STREAM_TRUST and recency live only in blend.ts and affect VALUE, never confidence, so all-Amazon (trust 0.5) and all-Brickset (trust 1.0) dimensions report identical confidence (dimension.ts:6-15). Spec §6.1 names source trust + recency as confidence inputs. Charter goal: 'an all-Amazon dimension must read lower-confidence than an all-Brickset one'.
The changeIn signalConfidence, add a trustFactor and recencyFactor computed from the SAME signal-weight primitives already in blend.ts (STREAM_TRUST[s.source]; exp(-ageMonths/HALFLIFE_MONTHS) when s.decays). Compute them as a volume-weighted mean across the signals (weight each signal's trust/recency by its own volume so a high-volume Brickset source dominates a 1-vote blog), then make final confidence = volumeFactor × agreement × trustFactor × recencyFactor. Reuse blend.ts's signalWeight decomposition rather than re-deriving constants, to keep one source of truth. Trust is already 0–1; recency is already 0–1; no extra normalization needed.
Test firstNew dimension.test.ts case: two dimensions with identical score/volume/agreement but source 'amazon' vs 'brickset' → brickset.confidence strictly greater (toBeGreaterThan). A second case: two identical decaying value signals at ageMonths 0 vs 48 → fresher reads higher confidence.
Watch outMultiplying four sub-1 factors compresses the confidence range downward — the obscure fixture and any test asserting an absolute confidence value could shift tier. Mitigate by NOT introducing a hard floor here; let fix #6 (coverage-aware overall) and the constants absorb calibration. Watch that fact-defined dimensions (which carry no opinion signals) are unaffected — they take the fact branch, not signalConfidence.
Depends onFix #1 (same function; land the diversity guard first so the trust/recency change is tested on the corrected agreement)
small
3. Promote log1p(60) volume-saturation denominator to a named constant
AddressesCharter + audit low finding — log1p(60) at dimension.ts:9 is a magic number absent from spec §10's constants table. House rule: all tunables live in constants.ts.
The changeAdd CONFIDENCE_VOLUME_SATURATION = 60 to constants.ts (with a comment tying it to §6.1 volume input and noting it's the count at which volumeFactor ≈ 0.99). In signalConfidence replace Math.log1p(60) with Math.log1p(CONFIDENCE_VOLUME_SATURATION). No behavior change — pure extraction.
Test firstExtend constants.test.ts (or a new confidence-constants test) to assert CONFIDENCE_VOLUME_SATURATION is exported and > 0; a dimension.test.ts characterization test pinning volumeFactor at a known volume (e.g. volume=60 → volumeFactor ≈ log1p(60)/log1p(60)=1.0 before other factors) guards against accidental denominator drift.
Watch outTrivial. Only risk is forgetting the import; tsc/test will catch it.
AddressesCharter + audit finding — fact dimensions get unconditional max(0.8, sigConf) regardless of how weak the basis was (dimension.ts:39). Spec §6.1: 'disclosed fallback to theme/sibling priors (down)'. Concretely smartPrice.valueRatio falls back partValue→weight→pieces (smartPrice.ts:22-29) and percentile() returns a neutral 0.5 on an empty cohort (smartPrice.ts:31-36), yet both still report 0.8. Charter goal: 'fallback facts must lower confidence, not hold a floor'.
The changeThree-part: (a) add an optional basis?: 'primary' | 'fallback' field to DimensionInput in types.ts; (b) in factInputs.ts buildFactInputs, when emitting smartPrice, tag basis='fallback' if valueRatio's basis !== 'partValue' OR the cohort length is below MIN_COHORT_FOR_SMART_PRICE (thread the cohort length in via FactContext, or have smartPriceScore return a {score, fallback} shape) — and likewise tag resale/scarcity 'fallback' when their inputs are derived from priors rather than real currentValue; (c) in resolveDimension's fact branch, replace Math.max(0.8, sigConf) with Math.max(input.basis === 'fallback' ? FACT_CONFIDENCE_FALLBACK_FLOOR : FACT_CONFIDENCE_FLOOR, sigConf), where FACT_CONFIDENCE_FLOOR=0.8 (existing behavior preserved for real facts) and FACT_CONFIDENCE_FALLBACK_FLOOR is a new lower constant (recommend 0.4). A fallback fact with no corroborating signals now reports 0.4, not 0.8.
Test firstdimension.test.ts: resolveDimension({id:'smartPrice', factValue:6, basis:'fallback'}) → confidence === FACT_CONFIDENCE_FALLBACK_FLOOR (and < the primary-basis case). factInputs.test.ts: a SetFacts with no partValueTotal/weight (pieces-only) and an empty cohort emits smartPrice with basis:'fallback'; one with partValueTotal + a ≥MIN_COHORT cohort emits basis:'primary' (or undefined).
Watch outThis is the only fix touching cross-file plumbing (types → factInputs → dimension). The neutral-0.5 empty-cohort percentile is a separate honesty bug arguably owned by the Smart-Price workstream; coordinate so we don't both edit smartPrice.ts. Keep this fix's edit confined to READING the basis, not changing the percentile math. Existing score.test.ts fixtures pass basis-free inputs, so the default must remain the 0.8 floor — assert that explicitly.
small
5. Move FULL_COVERAGE_MIN into constants.ts
AddressesCharter + audit medium finding — FULL_COVERAGE_MIN=0.5 is a hard-coded tier boundary invented in score.ts:33 with no entry in spec §10's table and not in the single-source-of-truth constants file. House rule: tunables live in constants.ts.
The changeAdd FULL_COVERAGE_MIN = 0.5 to constants.ts (comment: §6.2 coverage-tier cutoff, the scored/applicable fraction at/above which a grade is 'full' rather than 'data-light'). Delete the local const at score.ts:33 and import it. Pure move; the data-light branch at score.ts:91 is unchanged.
Test firstscore.test.ts already exercises the boundary indirectly (obscure → data-light). Add a constants assertion that FULL_COVERAGE_MIN is in (0,1]. No behavioral test needed beyond the existing data-light fixture continuing to pass.
Watch outTrivial extraction. Note for the spec-owner: §10's table should gain a row for this — flagged in decisions, since adding a calibrated cutoff to the spec is a product call, not an engineering one.
medium
6. Coverage-aware overall confidence (stop the unweighted mean over-reporting)
AddressesCharter + audit low finding — overall confidence is an unweighted mean over scored dimensions (score.ts:92), so a grade resting on few scored points can report the same confidence as a richly-covered one. Spec §6: confidence = depth of evidence behind the SCORED points; it should not look identical regardless of how many points were scored.
The changeIn computeDropScore, replace confidence = confSum/scored with a sub-weight-weighted mean of per-dimension confidence (weight each scored dimension's confidence by its SUB_WEIGHTS share within its cluster × the cluster's lens weight, mirroring how the score itself is rolled up), so a high-confidence figure on a low-weight dimension doesn't paper over thin evidence on a high-weight one. This makes confidence consistent with the weighting that produces the letter. Keep it 0 when scored===0. Do NOT additionally multiply by coverage — coverage is the separate badge (§6 keeps them distinct); the weighting alone fixes the over-report.
Test firstscore.test.ts: construct two inputs with identical per-dimension confidences but differing which dimension carries them (high-weight smartPrice well-evidenced vs low-weight instructions well-evidenced) → overall confidence differs, and the high-weight-evidenced one reads higher. Assert flagship's confidence stays within a sane band (>0, <1).
Watch outThe weighted mean needs the same lens-weight × sub-weight product the score uses; reuse clusterWeightsForLens + subWeightsOf rather than re-deriving, or extract a shared helper to avoid drift. Existing score.test.ts asserts no exact overall-confidence value, so low regression risk, but run all fixtures.
Depends onFixes #1, #2, #4 (per-dimension confidence must be honest before the rollup over it is meaningful)
medium
7. Give the holdValue badge its own confidence (Investor lens honesty)
AddressesAudit medium finding — coverage/confidence are computed only over IN_GRADE_CLUSTERS (score.ts:79), so the holdValue badge (scarcity, resale) carries no confidence, yet under the Investor lens holdValue is 30% of the grade (LENS_CLUSTER_WEIGHTS.investor). Charter anchor: 'holdValue carries no confidence though it's 30% under the investor lens'.
The changeCompute a holdValue confidence alongside the existing badge value: in computeDropScore, after rolling up clusters, aggregate the confidence of the scored scarcity+resale dimension results (weighted by holdValue's SUB_WEIGHTS) and attach it to holdValueBadge as a new confidence field (extend DropScore.holdValueBadge type in types.ts). Additionally, when lens==='investor' (or any lens where LENS_CLUSTER_WEIGHTS[lens].holdValue>0), include holdValue's dimensions in the overall-confidence weighting from fix #6 so the headline confidence beside an Investor grade reflects its 30%-weighted cluster. Gate purely on the lens weight being >0, not a hardcoded lens name.
Test firstscore.test.ts: investor-lens grade on flagship → holdValueBadge.confidence is defined and >0; a thin-data holdValue (scarcity scored, resale insufficient) reports lower badge confidence than a fully-scored one. Assert the universal grade's overall confidence is UNCHANGED by this fix (holdValue weight 0 under universal) so we don't regress the default.
Watch outTouches the DropScore public type (new optional field) — check no consumer destructures holdValueBadge expecting exactly {value,state}; make the field optional to stay backwards-compatible. The 'universal unchanged' assertion is the guardrail against accidentally financializing the default grade.
Depends onFix #6 (the overall-confidence weighting must exist before holdValue can be folded into it for the Investor lens)
Calls to make in this workstream
Where is a 'fallback fact' detected — in this workstream's factInputs, or surfaced by the Smart-Price/data workstream?
Options: (a) This workstream detects fallback locally in factInputs.ts by inspecting valueRatio's basis + cohort length; (b) the Smart-Price workstream (which already owns the streetPrice/empty-cohort honesty bugs) exposes a basis/quality marker on its fact output and we just consume it.
Recommended: (b) consume a marker the Smart-Price workstream exposes. valueRatio already returns a basis field that factInputs currently discards; the cleaner contract is for the fact layer to surface 'this was a fallback' and for confidence to react. Recommend we define the DimensionInput.basis field (this workstream) but coordinate so the Smart-Price workstream populates it for smartPrice — avoids two PRs editing smartPrice.ts and double-owning the empty-cohort neutral-0.5 bug.
What numeric value should FACT_CONFIDENCE_FALLBACK_FLOOR be, and should the §10 spec table gain rows for the new confidence tunables (saturation, single-source cap, fallback floor, FULL_COVERAGE_MIN)?
Options: Fallback floor anywhere from 0.3 (aggressive distrust) to 0.6 (mild). And: add these four constants to spec §10's table now, or ship them as code-only defaults pending calibration.
Recommended: Recommend FACT_CONFIDENCE_FALLBACK_FLOOR=0.4 (a fallback fact is meaningfully less trustworthy than a real part-value fact at 0.8, but still beats most thin sentiment). And yes — add all four to §10's table; they are load-bearing tunables and the house rule plus the audit both say a confidence/tier cutoff must not live only in code. The exact calibration is a product call once real sets are graded; ship sane defaults, table them, mark 'calibrate against real sets' like the other first-pass constants.
Should overall confidence be additionally multiplied by coverage, or kept strictly weight-aware-but-coverage-independent?
Options: (a) Keep confidence and coverage as two fully independent badges (spec §6 'two distinct badges'), fixing only the unweighted-mean bug; (b) also damp confidence when coverage is low, so a 3-of-9 grade can't report high confidence.
Recommended: (a). Spec §6 is explicit that they measure different things and gives a worked cross-case ('high confidence, low coverage') that REQUIRES them to vary independently. Multiplying confidence by coverage would collapse that distinction. Fix the weighting (fix #6) but leave the badges orthogonal; coverage already drives the data-light tier on its own.
New constantsCONFIDENCE_VOLUME_SATURATION, SINGLE_SOURCE_AGREEMENT_CAP, CONFIDENCE_MIN_SOURCES_FOR_FULL_AGREEMENT, FACT_CONFIDENCE_FLOOR, FACT_CONFIDENCE_FALLBACK_FLOOR, FULL_COVERAGE_MIN
SequencingFront-load the two zero-plumbing quick wins that de-risk everything else: Fix #3 (extract log1p(60) → CONFIDENCE_VOLUME_SATURATION) and Fix #5 (FULL_COVERAGE_MIN → constants) — both are pure extractions, both establish the constants-file entries the later fixes import. Then the un-gameable core in dependency order inside signalConfidence: Fix #1 (source-diversity agreement cap — the headline 'lone Amazon ≠ 1.000' fix) before Fix #2 (fold trust + recency), since both edit the same function and #2's tests should run against the corrected agreement. Next Fix #4 (basis-aware fact floor) — independent of the signalConfidence work but the one cross-file plumbing change, so isolate it in its own commit and coordinate with the Smart-Price workstream on who populates DimensionInput.basis. Finally the rollup/badge layer that depends on per-dimension confidence already being honest: Fix #6 (coverage/weight-aware overall confidence) then Fix #7 (holdValue badge confidence + Investor-lens fold), in that order. Every fix is red-test-first; run the full vitest suite after #1, #2, and #6 specifically since those have the highest chance of nudging a fixture across a tier boundary.
Wave 1Workstream 3 of 8
Close the methodology-fidelity gaps
Done looks like: Every rule the methodology spec defines is either honored by the math or explicitly descoped in the spec — no silent divergence. Concretely: the §3.1 variety/novelty home is correct (variety only in Parts-for-Custom-Building, novelty only in Rare & New), Feels-Worth-It either price-pairs value opinions or is descoped on the record, the over-concentration cap stops silently rewriting the headline number, a no-data set returns a withheld state instead of a confident F, and the thumbs/licensed-premium/lifecycle gaps are each either implemented or moved to an explicit "deferred" list with a test pinning the chosen behavior.
Updated by your decisionsYour calls: Feels-Worth-It pairing is <b>shelved on the record</b> (spec note + pinned test). The pre-release-hype filter is <b>not built</b> — you're keeping early sentiment as legitimate signal; grades are re-assessed monthly instead (see "Living grades").
The approach
This workstream is about restoring spec INTENT, so several items split into "implement now (cheap, code-only, no new data)" vs "descope honestly (needs data plumbing this engine doesn't have)". I front-load the two pure-logic, high-value, no-new-data fixes that the audit calls the worst: the §3.1 variety/novelty inversion (variety wrongly lives in What-You-Get, partsForBuilding has no fact at all) and the over-concentration cap distorting the headline number. Both are deterministic, unit-testable against the existing rollup/factScores tests, and ship behind constants. Next I add the no-data "withheld" state, which is a small score.ts change but touches the public DropScore type, so it sequences after the cap fix (they both live in the rollup→score path). Then I handle the four items that are really product/spec decisions because honoring them needs data the pipeline cannot yet supply: FWI price-pairing (needs a historical sold-price series), the pre-availability lifecycle gate (needs availability dates + per-review timestamps), the thumbs→0–10 mapping (needs the contract/store reshape), and the licensed-premium note (needs a licensed flag). For each I give both the implement-it and descope-it path with a recommendation, and where I recommend descope I still write a test that pins the documented behavior so the gap is an invariant, not a silent omission. The §4.3 within-tier-first redistribution is a self-contained capWeights rewrite I do last because it's the most intricate and lowest user-visible impact. Every threshold I introduce goes into src/config/constants.ts; nothing stores raw text; the FWI/lifecycle data hooks follow the existing adapter+contract seam so a live implementation can land later without re-touching the engine.
Effort~1.5–2 weeks: 3 S (cap-vs-headline, thumbs mapping, licensed note), 5 M (variety/novelty inversion, withheld state, FWI pairing-or-descope, lifecycle gate-or-descope, within-tier redistribution). Lower end (~1 week) if FWI pairing and the lifecycle gate are descoped (their M collapses to S: spec note + guard test) as recommended.
Fixes8 · 3 small · 5 medium · 0 large
Depends onEvidence-gates / fail-closed workstream (LOW_DATA_FLOOR, MIN_STRUCTURED_COUNT, MIN_COMMUNITY_TO_MOVE) — it edits the same resolveDimension/community-signal seams as my thumbs, FWI, and lifecycle items; sequence so we don't both rewrite resolveDimension at once. Smart-Price street-price / sold-price-series workstream — the real value of FWI price-pairing depends on it (which is why I recommend descoping pairing until that lands). Confidence & coverage workstream — co-owns score.ts:79-92 and the FULL_COVERAGE_MIN constant move; coordinate the withheld-state change with it.
The fixes, in order
medium
1. Fix the §3.1 variety/novelty inversion: move generic variety out of What-You-Get into Parts-for-Custom-Building
AddressesMethodology fidelity weakness §3.1 rule 3 — src/lib/facts/factScores.ts:44-56 (whatYouGetScore consumes inv.distinctElements/totalParts variety) and src/lib/facts/factInputs.ts:33-39 (partsForBuilding never emitted as a fact)
The changeTwo edits in concert. (1) In factScores.ts, strip the variety term from whatYouGetScore: delete varietyRatio/variety/the WHAT_YOU_GET_PRINTED_WEIGHT blend and make What-You-Get a function of printedShare (and the filler penalty) only — printed-vs-plain non-fig content density per §1.2. Retire WHAT_YOU_GET_VARIETY_ANCHORS and WHAT_YOU_GET_PRINTED_WEIGHT from this function. (2) Add a new pure fn partsForBuildingScore(inv: InventorySummary) in factScores.ts that maps the distinct-element variety ratio (distinctElements/totalParts) through a new PARTS_FOR_BUILDING_VARIETY_ANCHORS curve (seeded from the retired WHAT_YOU_GET_VARIETY_ANCHORS values), returning 0–10; null/0 when totalParts<=0. (3) In factInputs.ts buildFactInputs, when ctx.inventory is present, add('partsForBuilding', partsForBuildingScore(ctx.inventory)) alongside the whatYouGet add. Export partsForBuildingScore from facts/index.ts. Crucially this stays a FACT now feeding the hobbyist cluster; mined partsForBuilding signals still merge as enrichment via mergeInputs, so the parts-pack fixture (signal-supplied) keeps working.
Test firstFIRST write src/lib/facts/factScores.test.ts cases: (a) whatYouGetScore no longer changes when distinctElements varies but printedParts/totalParts/fillerRatio are held — two summaries identical except distinctElements yield identical What-You-Get (currently they differ — red); (b) partsForBuildingScore rises monotonically with distinctElements/totalParts and is 0 at totalParts=0. Then add the anti-double-counting invariant test in factInputs.test.ts: build inputs for an inventory and assert the whatYouGet factValue is independent of distinctElements AND a partsForBuilding factValue input is emitted — the §3.1 rule-3 invariant the audit asked to make checkable.
Watch outwhatYouGetScore calibration shifts (printed-only is a narrower base) so the flagship/parts-pack score-band assertions in score.test.ts may move; re-tune WHAT_YOU_GET_PRINTED_ANCHORS or the band expectations, don't loosen them blindly. The parts-pack fixture currently scores partsForBuilding via a forum signal — adding a fact value there means the fact+signal now blend; confirm the parts-pack test's 'scored' expectation still holds (it will: fact makes it scored regardless). SUB_WEIGHTS.hobbyist already splits playability/partsForBuilding so no rubric change needed.
small
2. Stop the over-concentration cap from rewriting the headline score; use it only to set the data-light tier
AddressesScoring engine math + Methodology fidelity — src/lib/scoring/rollup.ts:91-105 (capWeights output feeds the score) and src/lib/scoring/score.ts:68-91 (concentrationCapped → tier). Audit example: naive 7.8 becomes 5.4.
The changeIn rollupOverall, compute the headline score from the UNCAPPED proportional weights (the existing norm inside capWeights, i.e. raw/Σraw) and use capWeights' result ONLY to decide the capped boolean. Concretely: split capWeights into (a) the over-concentration DETECTION (does max normalized share exceed RENORM_CAP / are there too few survivors) returning just capped, and (b) keep the water-fill solely if a product decision says we still want capped weights for some other consumer — but rollupOverall's returned score must use proportional weights. Net: score = Σ(proportionalWeight_i · value_i); concentrationCapped unchanged in meaning (still forces data-light in score.ts:90-91). This matches §4.3's framing of the cap as a TIER guardrail, not a number distorter.
Test firstFIRST extend src/lib/scoring/rollup.test.ts 'flags concentration' case: assert BOTH concentrationCapped===true AND score equals the uncapped proportional average (for the audit's 3-survivor case — moneyWorth value 9 at raw weight 80, two others value 3 at small weights — assert score≈7.8, not 5.4). Currently it returns the capped 5.4 → red. Add a second case proving a non-concentrated set is unaffected.
Watch outscore.test.ts 'obscure long-tail … data-light' must still pass — it relies on concentrationCapped forcing data-light, which this preserves, but the obscure fixture's numeric score will rise to the uncapped value; if any assertion checks its score magnitude, update it. This is a deliberate, now-documented behavior change to the headline number for concentrated sets — flag it in the rollup doc-comment and the spec §4.3.
medium
3. Return an explicit withheld / no-data state instead of 0.0 → F when no in-grade cluster has a value
AddressesScoring engine math + Confidence/coverage — src/lib/scoring/rollup.ts:99 (rollupOverall returns score 0 for included.length===0) and src/lib/scoring/score.ts:108-117 (no zero-coverage special case → toLetter(0)=F)
The changeIn score.ts after the overall rollup, detect zero in-grade coverage (no IN_GRADE_CLUSTERS cluster has a non-null value, equivalently scored===0 across the in-grade loop at score.ts:79-88). When so, return a DropScore with score:null and a new top-level state: 'withheld' (or letter set to a sentinel and coverageTier forced 'data-light'), rather than score:0/letter:'F'. Add score: number | null and a gradeState: 'graded' | 'withheld' field to the DropScore type in types.ts; default 'graded'. rollupOverall keeps returning {score:0} internally but score.ts must not map that to F when coverage is zero. This honors principle 4 (every set graded, but a quiet set is data-light, never a confident failing letter).
Test firstFIRST add to score.test.ts: computeDropScore(facts, []) with an empty inputs list (or a facts object that yields no fact dimensions and no signals) asserts result.gradeState==='withheld', result.score===null, result.letter is NOT 'F', and coverageTier==='data-light'. Currently it returns score 0 / letter F → red. Keep an existing partial-data fixture asserting gradeState==='graded' so we don't over-trigger withheld.
Watch outDropScore.score becoming nullable ripples to every consumer (demo.ts, the app layer, any persistence). Grep all readers of .score/.letter and guard them; this is the main cost. Coordinate with the confidence/coverage workstream which also touches score.ts:79-92. Keep the change additive (new gradeState field) so existing 'graded' paths are untouched.
Depends onFix #2 (both edit the rollup→score headline path; land the cap fix first to avoid a merge conflict in rollupOverall)
medium
4. Decide + pin Feels-Worth-It price-pairing (implement the pairing seam, or descope in spec with a test)
AddressesMethodology fidelity (high) — src/lib/scoring/score.ts:51-57 (FWI is a plain capped nudge; no opinion is paired to street-price-at-time, none discarded). Spec §1.3/§6.1.
The changeIf implementing (recommended seam-only): extend OpinionSignal in types.ts with optional priceAtTime?: number; in score.ts step 2, before nudging, filter feelsWorthIt signals to those whose priceAtTime is present AND within tolerance of the reconstructed street price (a new dep: facts.streetPrice as the present anchor, since the historical sold-price series doesn't exist yet), discarding unpaired/mismatched opinions; only blend the survivors into the nudge. Gate the tolerance behind a new constant FWI_PRICE_PAIR_TOLERANCE. The distiller (distill.ts) and community store would later populate priceAtTime (community already collects pricePaid in CommunitySubmission). If descoping: add a spec note in §1.3 + §8 'deferred' stating FWI is an unpaired capped nudge until the historical sold-price series exists, and remove the 'unpaired discarded' claim from §1.3 so the doc stops promising a property the code lacks.
Test firstFIRST, for the implement path: a score.test.ts case with two feelsWorthIt signals — one with priceAtTime near facts.streetPrice, one far off (or absent) — asserts only the paired one moves Smart Price (the unpaired one is discarded; nudge equals the paired-only blend). For the descope path: a test asserting the documented behavior (all FWI signals nudge, capped at FEELS_WORTH_IT_CAP) PLUS a guard test that the spec file contains the deferred note — so the descope is an invariant.
Watch outThe 'present street price' proxy is NOT the spec's 'street price at the time it was said' — implementing pairing against today's price is a partial fix that could read as done when it isn't; be explicit in code+spec that the historical series is still deferred. streetPrice is unpopulated in live adapters today (separate Smart-Price-data workstream), so an implemented gate would discard ALL opinions in production until that lands — meaning implement-now mostly benefits tests/demo. That trade-off is why this is a decision, not a default.
medium
5. Decide + pin the pre-availability lifecycle gate (scoped to value/worth + build), or descope in spec
AddressesMethodology fidelity (medium) + Scoring engine math — absent everywhere; spec §6.1. Reveal-hype feeds Build Fun and Feels-Worth-It unfiltered.
The changeIf implementing: add optional availabilityDate?: string to SetFacts and optional saidAt?: string (or ageMonths already exists) to OpinionSignal; in a new pure helper applyLifecycleGate(signals, dim, availabilityDate) called from resolveDimension/blend assembly, suppress (drop) or discount (multiply weight by LIFECYCLE_PREAVAIL_DISCOUNT) signals dated before availability for the value/build dimensions ONLY (feelsWorthIt, buildFun, buildLength, instructions, integrityQC), while leaving displayAppeal/fidelityScale untouched per §6.1. The distiller would later emit saidAt. If descoping: spec §6.1 + §8 'deferred' note that the lifecycle gate is unimplemented (no per-review dates flow through the pipeline yet) and that pre-availability reveal sentiment is currently treated identically to in-hand sentiment.
Test firstFIRST, implement path: a unit test on applyLifecycleGate — a pre-availability youtube buildFun signal is dropped/discounted while a pre-availability displayAppeal signal is kept, given an availabilityDate. Descope path: a guard test asserting the spec deferred-list contains the lifecycle-gate entry, plus a behavioral test pinning that a pre-availability-dated buildFun signal currently DOES contribute (documenting the known gap).
Watch outNo availability date or per-review date currently flows from any adapter, so an implemented gate is dead in production until discovery/distill emit dates — same 'looks done but isn't' trap as FWI pairing. Touching blend/resolveDimension overlaps the evidence-gates workstream; coordinate so both don't rewrite resolveDimension simultaneously.
small
6. Define the thumbs→0–10 mapping for worthTheMoney (or descope the thumbs field)
AddressesMethodology fidelity (medium) — src/lib/sources/contracts.ts:64 (worthTheMoney typed number, no thumbs→0–10 mapping); community store pushes it straight in as a 0–10 score. Spec §5.1.
The changeIf implementing (recommended): change CommunitySubmission.worthTheMoney type to a thumbs union ('up'|'down' or boolean) in contracts.ts, and add a pure THUMBS_TO_SCORE mapping (constant in constants.ts, e.g. up→8, down→3 — values to calibrate) applied wherever community submissions become feelsWorthIt signals (the community store's getSignals). This stops a raw 0/1 entering the blend as a 0–10. Mirrors the Brickset star→score ×2 scaling that already exists. If descoping: keep number but document in contracts.ts + spec §5.1 that worthTheMoney is a pre-scaled 0–10 and rename the field/prompt so the 'thumbs' framing is dropped.
Test firstFIRST, a test on the community-store getSignals (or the mapping fn): a submission with worthTheMoney thumbs-up emits a feelsWorthIt signal at THUMBS_TO_SCORE.up, thumbs-down at .down — currently a raw value passes through unscaled → red once the type is a thumbs union. Assert no raw 0/1 ever reaches a signal score.
Watch outChanging the contract type touches the community submission API and any form/persistence binding; this is small but cross-layer. Coordinate with the evidence-gates workstream that also gates community signals (MIN_COMMUNITY_TO_MOVE) at the same seam.
small
7. Surface the licensed-premium note instead of silently penalizing licensed sets (or descope)
AddressesMethodology fidelity (low) — src/lib/facts/smartPrice.ts:43-48 (no licensed-premium note computed/carried). Spec §1.1 'licensing premium surfaced explicitly, not silently penalized'.
The changeIf implementing: add optional isLicensed?: boolean to SetFacts (populated later by an adapter from theme/IP metadata), and have smartPriceScore (or factInputs) emit a sibling note — e.g. a licensedPremium: boolean flag carried on the DimensionResult/DropScore — when isLicensed && the value ratio sits below cohort median. This is a SURFACED note, not a score change (Smart Price stays neutral to licensing, only annotated). If descoping: spec §1.1 + §8 note that the licensed-premium surfacing is deferred and Smart Price currently treats licensed sets identically (their lower part-value ratio lowers the score with no annotation).
Test firstFIRST, implement path: a test that smartPrice fact-build for an isLicensed set with a below-median ratio carries licensedPremium===true, and a non-licensed set does not — without changing the numeric Smart Price value (assert the 0–10 is identical to the unflagged case). Descope path: a guard test pinning the spec deferred note.
Watch outAdding a flag to DropScore/DimensionResult is additive but ripples to the app/output layer that renders the grade. Low risk to math (score is unchanged by design). isLicensed is unpopulated by live adapters today, so the note is dormant until an adapter sets it — acceptable for a surfaced annotation, unlike a gate.
medium
8. Add §4.3 within-tier-first redistribution to capWeights (or descope the within-tier clause)
AddressesMethodology fidelity (medium) — src/lib/scoring/rollup.ts:47-79 (capWeights renormalizes strictly proportionally with no HIGH/MED tier awareness). Spec §4.3 'renormalization prefers within-tier first'.
The changeIf implementing: pass each cluster's tier (HIGH/MED) into rollupOverall (a new CLUSTER_TIER map in constants.ts or rubric.ts derived from §3's table: moneyWorth/theBuild/finishedModel/minifigures=HIGH, hobbyist/rareNew=MED). When a cluster drops out, redistribute its weight first across surviving same-tier clusters proportionally, spilling to the other tier only if no same-tier survivor remains, THEN apply the RENORM_CAP water-fill. If descoping: amend spec §4.3 to drop the within-tier-first clause and keep strictly-proportional renormalization (which is what the code faithfully does today).
Test firstFIRST, implement path: a rollup.test.ts case where one HIGH cluster (e.g. minifigures) is null — assert its weight lands on surviving HIGH clusters (moneyWorth/theBuild/finishedModel) and does NOT raise the MED clusters' shares relative to the strictly-proportional baseline. Compare two redistribution outputs to make the tier preference observable. Descope path: a test pinning strictly-proportional behavior + a spec-note guard.
Watch outThis is the most intricate change to the water-fill; interacts with the over-concentration cap (do tier redistribution BEFORE the cap). Keep it last so the cap-vs-headline fix (#2) is already settled. Builder/investor lenses change tier composition (holdValue only under investor) — ensure CLUSTER_TIER covers holdValue or excludes it cleanly. Low user-visible impact, so descope is a legitimate call if effort is tight.
Depends onFix #2 (capWeights/rollupOverall are rewritten there; build tier-awareness on top of the settled cap-detection logic)
Calls to make in this workstream
Feels-Worth-It price-pairing: implement the pairing seam now, or explicitly descope it in the spec?
Options: (a) Implement a pairing gate against the PRESENT street price (partial — not the spec's historical at-the-time price), which in production discards all FWI opinions until streetPrice is plumbed; (b) descope in §1.3/§8 and remove the 'unpaired discarded' promise until the historical sold-price series exists.
Recommended: Descope now, with a pinned test. The spec's defining property is pairing against the street price AT THE TIME each opinion was said, which requires a historical sold-price series that does not exist anywhere in the pipeline — implementing against today's price would look done while silently doing something weaker. Honest descope + test beats a misleading partial. Revisit when the Smart-Price street-price/sold-series workstream lands.
Pre-availability lifecycle gate: implement, or descope?
Options: (a) Implement the scoped suppression now (touches blend/resolveDimension) but it stays dead until adapters emit per-review dates; (b) descope in §6.1/§8 and pin the current 'reveal sentiment treated as in-hand' behavior with a test.
Recommended: Descope now. No availability date or per-review timestamp flows through discovery/distill today, so an implemented gate is non-functional in production and overlaps the evidence-gates workstream's resolveDimension edits. Add the SetFacts.availabilityDate + OpinionSignal.saidAt fields as a typed seam (cheap, forward-looking) but defer the gate logic and document it.
worthTheMoney thumbs mapping: re-type as thumbs + map, or keep numeric and re-document?
Options: (a) Change the contract to a thumbs union + a THUMBS_TO_SCORE constant (cross-layer but small); (b) keep number, rename the field/prompt so it's an honest pre-scaled 0–10.
Recommended: Implement (a). It's an S-effort change that closes a real correctness bug (a raw 0/1 entering the blend as a 0–10 corrupts the Smart Price nudge), and it matches the existing Brickset star→×2 scaling pattern. Coordinate the edit with the community-gate (MIN_COMMUNITY_TO_MOVE) work since both touch the community signal seam.
§4.3 within-tier-first redistribution: implement tier-aware capWeights, or amend the spec to drop the clause?
Options: (a) Implement HIGH→MED tier preference in capWeights (M effort, intricate, interacts with the cap); (b) amend §4.3 to keep strictly-proportional renormalization, which the code already does faithfully.
Recommended: Implement (a) but schedule it LAST. It's a genuine spec property and the change is contained, but its user-visible impact is small and it's the riskiest water-fill edit — so it's also the most defensible to descope if the schedule tightens. Either way, pin the chosen behavior with a test.
Licensed-premium note: implement the surfaced flag, or descope?
Options: (a) Add isLicensed + a surfaced licensedPremium note (no score change); (b) descope in §1.1/§8.
Recommended: Implement (a) as an additive, score-neutral annotation — it's S-effort and directly satisfies the spec's 'surfaced explicitly, not silently penalized' wording. It stays dormant until an adapter populates isLicensed, which is acceptable for an annotation (unlike a gate).
New constantsPARTS_FOR_BUILDING_VARIETY_ANCHORS, FULL_COVERAGE_MIN, FWI_PRICE_PAIR_TOLERANCE, LIFECYCLE_PREAVAIL_DISCOUNT, THUMBS_TO_SCORE, CLUSTER_TIER
SequencingFront-load the two pure-logic, no-new-data wins the audit ranks worst and that need no product decision: (1) the §3.1 variety/novelty inversion and (2) the over-concentration cap no longer distorting the headline — both are deterministic, fully unit-testable today, and unblock honest scores. Then (3) the no-data 'withheld' state, which builds on the settled rollup→score path from #2. Pause for the product/spec decisions (FWI pairing, lifecycle gate, thumbs mapping, licensed note) — get Seth's calls, then implement the ones chosen 'implement' (thumbs and licensed note are the cheap, recommend-implement ones) and write the deferred-list spec notes + guard tests for the ones chosen 'descope' (FWI pairing and lifecycle gate are the recommend-descope ones). Do the §4.3 within-tier-first redistribution (8) LAST — it's the most intricate water-fill edit, lowest user-visible impact, and the easiest to descope if the schedule tightens. Quick wins to grab first day: #2 (S) and the thumbs mapping (S).
Wave 2Workstream 4 of 8
Make Smart Price measure real market value
Done looks like: Every live grade computes Smart Price against a real street-price anchor (BrickLink SET sold-median, MSRP only as an explicit, labeled fallback), percentiled against a clean, self-excluded, basis-and-recency-matched, sufficiently-large cohort — and when any of that is missing or thin, Smart Price returns an explicit insufficient/withheld state or a confidence-penalized score, never a silent confident MSRP-anchored C.
Updated by your decisionsYour call: <b>Hybrid anchor</b> — real BrickLink street price where sales data is solid, MSRP-labeled fallback elsewhere. Cohort floor stays ~5 with a confidence ramp (not raised to 15). Grading is <b>per variant</b> (polybags deferred).
The approach
The dimension is structurally faithful to spec §1.1 (part-value ÷ price → cohort percentile → curve) but its data plumbing betrays the promise: no live adapter ever sets facts.streetPrice, so valueRatio (smartPrice.ts:23) always falls to MSRP — the exact basis §1.1 forbids — and the cohort path has self-inclusion, tie, recency, size, and cross-theme defects, all silent. I sequence this as: (1) fail-closed plumbing first — make the price BASIS explicit and propagate it so the rest of the system can reason about and surface it; (2) build the real street-price anchor as a thin BrickLink SET sold-median adapter following the existing fetchPriceGuide pattern, populated in gradeSet and persisted; (3) fix the cohort math (self-exclusion, deliberate tie semantics, basis+recency match, raise MIN_COHORT, confidence-scale by n, label cross-theme fallback); (4) treat empty/thin cohort and missing-price as insufficient, not 0.5→6.0; (5) surface the licensed-premium note and wire Brick Insights PPP as a cross-check sanity flag. All tunables land in constants.ts; every fix starts with a failing test, pure logic unit-tested and gradeSet/priceInventory tested through the GradeDeps seam. The street-price-vs-relabel fork and the new-confidence-field fork are genuine product/architecture calls and go to decisions, not silently chosen.
Effort~2 weeks: 3 S (Fixes 1, 4, 7), 5 M (Fixes 2, 3, 5, 6, 9), 1 L (Fix 8). The L (cold-start bootstrap + batch corpus pricing) dominates and overlaps the data-ingest workstream; if that batch pricer is owned elsewhere, this workstream drops to ~1 week.
Fixes9 · 3 small · 5 medium · 1 large
Depends onDepends on the data-layer/ingest workstream for two things: (1) a schema addition on the sets cohort source — priced_at and price_basis (and ideally a street_price-as-of) columns — needed by Fix 5's recency/basis matching and Fix 3's caching; (2) the batch corpus-pricing pipeline Fix 8's cold-start bootstrap relies on (may already be the '--price batch-ingest' path referenced in price.ts:64). Fix 6's confidence-by-cohort and Fix 9's notes need the scoring-core/output-shape workstream to add DimensionInput.factConfidence and a DropScore notes/flags field (the engine-internals decisions above). Otherwise self-contained.
The fixes, in order
small
1. Make the value basis explicit and fail-closed (street vs msrp), no silent MSRP
Addresses'street price never populated → valueRatio always resolves to MSRP, no flag' (smartPrice.ts:23) + 'empty-cohort 0.5→6.0 is a confident C with no data' (smartPrice.ts:32)
The changeIn src/lib/facts/smartPrice.ts change valueRatio to return basis provenance for price too: extend its return to { ratio, basis, priceBasis: 'street' | 'msrp' } where priceBasis records whether facts.streetPrice was used or it fell back to msrp. Keep the value-basis fallback chain (partValue→weight→pieces) intact. This is the seam the rest of the workstream reads to fail closed: nothing else changes behavior yet, but the information stops being lost at line 23.
Test firstIn smartPrice.test.ts add: valueRatio({...base, partValueTotal:200, streetPrice:100, msrp:80}) returns priceBasis:'street'; valueRatio({...base, partValueTotal:200, msrp:100}) (no streetPrice) returns priceBasis:'msrp'. Red because the field does not exist today.
Watch outvalueRatio's return shape is consumed by loadCohort (price.ts:101) and smartPriceScore (smartPrice.ts:44) — both read .ratio/.basis only, so adding a field is non-breaking, but update the existing toEqual assertions in smartPrice.test.ts:10-25 which currently assert the exact object and will fail on the extra key.
medium
2. Add a BrickLink SET sold-median street-price adapter (thin client, live-pattern)
AddressesKEY DECISION + 'fetchPriceGuide defaults sold/new — could fetch a SET sold-median street anchor' (bricklink.ts:100-101); street anchor never built
The changeIn src/lib/sources/bricklink.ts add fetchSetStreetPrice(setNumber, creds): reuse the existing fetchPriceGuide('SET', setNumber, creds, { guideType:'sold', newOrUsed:'N' }) path (it already defaults to sold/new) and return its priceGuideToValue median, or null on miss/throttle. Keep it a thin client matching the existing OAuth1-signed pattern; no new HTTP plumbing. Export a typed contract.
Test firstUnit-test priceGuideToValue is already covered; add a test asserting fetchSetStreetPrice passes itemType 'SET', guide_type 'sold', new_or_used 'N' by spying on a injected fetch/fetchPriceGuide seam (mirror how bricklink signing helpers are unit-tested). Assert null is returned (not thrown) when the guide has no qty_avg/avg_price.
Watch outBrickLink SET sold guide can be sparse/zero for new or low-liquidity sets → must return null, never 0, so the caller fails closed to MSRP-labeled. Live-verify against the real API before trusting (house rule: live-verified adapters). BrickLink's ~5k/day cap: one SET call per grade is cheap vs the ~part-out budget, but cache it (next fix).
medium
3. Populate + persist facts.streetPrice in gradeSet, cached like part prices
Addresses'streetPrice set only in demo/fixtures; no live adapter assigns it' (brickset.ts:42 sets msrp only; rebrickable.ts sets neither)
The changeIn src/lib/pipeline/gradeSet.ts add an optional GradeDeps seam fetchStreetPrice?: (setNumber) => Promise<number|null> (parallel to priceParts) and, when present, set facts.streetPrice from it before assembleSetInputs. Wire the production impl to fetchSetStreetPrice with a Supabase street-price cache (store on the sets row's existing street_price column — supabaseStore.ts already round-trips it at :37/:60 — with a price-as-of timestamp). Do NOT touch bricksetSetToFacts/rebrickableSetToFacts; street price is a distinct sold-data source, not an API field they own.
Test firstAdd src/lib/pipeline/gradeSet.test.ts (new — the audit flags gradeSet has NO test) driving the GradeDeps seam: with fetchStreetPrice→200 and msrp 250, assert result.facts.streetPrice===200 and the smartPrice input was computed on 200 (priceBasis street); with fetchStreetPrice→null, assert it falls back to msrp AND the result carries an explicit msrp-anchored marker (next fix). Use the injectable seam, no real network.
Watch outThis is the highest-leverage behavior change in the workstream: existing end-to-end expectations that implicitly assumed MSRP will shift. Guard with the coverage/cache timestamp so a throttled street fetch fails closed to labeled-MSRP rather than blocking the grade. Persisting street price must reuse the existing sets-row store to avoid a schema migration where possible.
4. Self-exclude the graded set from its own cohort + deliberate tie semantics
Addresses'loadCohort no self-exclusion' (price.ts:74-81) + 'percentile c<=value tie-inclusion, best always =1.0 staircase' (smartPrice.ts:33-35)
The changeloadCohort (price.ts) takes the graded setNumber and adds .neq('set_number', setNumber) to the query. In smartPrice.ts replace the at-or-below percentile with a deliberate midrank: p = (countBelow + 0.5*countEqual) / n (excluding self), so a tie no longer reads as 'at-or-below' and the best set is not auto-pinned to 1.0→10.0. Keep the curve in constants.ts untouched.
Test firstsmartPrice.test.ts: a set whose ratio equals the single cohort value scores p=0.5 (was 1.0); a set above all 4 cohort members scores <1.0 via midrank, not exactly 1.0. loadCohort test (in gradeSet/price test via the Supabase query-builder seam): assert the .neq filter is applied with the graded set number.
Watch outChanging percentile semantics shifts every Smart Price score slightly — re-baseline the calibration anchors? No: the curve anchors (p50→6, p90→9) are defined on percentile, and midrank makes the realized mapping closer to the intended continuous CDF, so this is a fidelity improvement. Update smartPrice.test.ts cohort-median/bargain/rip-off expected values (they assume the old at-or-below def).
medium
5. Match cohort basis + recency to the live-priced graded set
Addresses'basis/recency mismatch: graded set priced live, cohort rows use stale cached part_value' (price.ts:65-105)
The changeloadCohort already keeps the cohort to basis==='partValue' (price.ts:102) — good. Add a freshness guard: store/read a priced_at timestamp on cohort rows and drop rows whose part value is older than COHORT_MAX_STALENESS_DAYS (new constant). When the graded set was re-priced live this run, only compare against cohort rows priced within the staleness window; if too few survive, treat as thin (→ insufficient, next fix). Also ensure the cohort ratio is computed on the SAME price basis (street vs msrp) as the graded set — only pool street-anchored cohort rows when the graded set is street-anchored.
Test firstprice/loadCohort test through the query seam: rows older than the staleness window are excluded; a graded set on street basis does not pool msrp-basis cohort rows. gradeSet.test.ts: when the surviving fresh same-basis cohort < MIN_COHORT, smartPrice resolves insufficient rather than scored.
Watch outRequires a priced_at/price_basis column on the sets cohort source — a schema addition (coordinate with the data-layer workstream). Until backfilled, every row looks stale → could dark out Smart Price entirely; mitigate with a migration default of 'unknown freshness treated as eligible' behind a feature flag, flipped on once backfill runs. This is the fix most coupled to the cold-start bootstrap decision.
Depends onFix 1, Fix 3
medium
6. Raise MIN_COHORT and scale Smart Price confidence by cohort n
Addresses'MIN_COHORT_FOR_SMART_PRICE=5 too low for a percentile; n=5 staircase swings ~1 full point; output over-precise' (constants.ts:145, smartPrice.ts:32-36)
The changeRaise MIN_COHORT_FOR_SMART_PRICE in constants.ts to a defensible floor (recommend 15–25; see decisions) and add SMART_PRICE_CONFIDENCE_BY_COHORT anchors (cohort n → 0–1 confidence multiplier). smartPriceScore returns not just a value but a cohort-size-derived confidence; thread that confidence into the dimension so a 5-comparable grade is distinguishable from a 200-comparable one. Surface cohort n in GradeResult.sources (e.g. smartPriceCohortN) so the output can show it.
Test firstsmartPrice.test.ts: score backed by n=MIN_COHORT carries lower confidence than n=200; constants change asserted by the gating test (a 10-member cohort now fails the gate if floor is 15). gradeSet.test.ts: result.sources.smartPriceCohortN reflects the cohort actually used.
Watch outresolveDimension (dimension.ts:33-39) currently hardcodes fact-defined confidence to max(0.8, sigConf) — there is NO channel to pass a per-fact confidence today. Threading cohort confidence requires either a new optional DimensionInput.factConfidence field or a Smart-Price-specific path; that is a real engine change touching the scoring core (see decisions). Raising MIN_COHORT worsens the cold-start cliff until bootstrap lands — sequence after Fix 8.
Depends onFix 4, Fix 7
small
7. Treat empty/thin/no-price cohort as insufficient, not a confident 6.0
Addresses'empty-cohort neutral 0.5→6.0 emits a confident C with no data' (smartPrice.ts:32, :43-48) + 'cohort omitted entirely → smartPrice silently dropped' (factInputs.ts:33)
The changesmartPriceScore returns null (→ insufficient) when the cohort is empty or below MIN_COHORT, instead of percentile()'s 0.5. In factInputs.ts:33 stop using truthiness of ctx.smartPriceCohort (an empty array [] is truthy and currently produces a 6.0); only emit the smartPrice input when a valid (≥MIN_COHORT) cohort produced a non-null score, otherwise leave it out so the engine marks it applicable-but-insufficient and renormalizes (spec §4.3). When price itself is unavailable (no street, no msrp), smartPrice is already null — keep it.
Test firstsmartPrice.test.ts: smartPriceScore(facts, []) returns null (was 6.0); smartPriceScore(facts, [1,2]) with MIN_COHORT>2 returns null. factInputs.test.ts: passing an empty cohort emits NO smartPrice input. gradeSet.test.ts: a set with no loadable cohort yields a DropScore where smartPrice state is insufficient/absent and moneyWorth renormalizes onto whatYouGet, not a phantom C.
Watch outThis removes a score that previously existed for thin-data sets, lowering coverage tier for many sets until cohorts fill — that is the CORRECT fail-closed behavior per the house rule, but it is a visible product change (more 'data-light' grades). The empty-cohort 0.5 path is reachable via the direct smartPriceCohort dep too, so fix both call sites.
Depends onFix 6 (MIN_COHORT)
large
8. Bootstrap cohorts + make cross-theme fallback explicit and labeled
Addresses'cold-start cliff: cohorts empty until a corpus is priced' + 'theme→size-only fallback mixes value regimes silently' (gradeSet.ts:127-131, :130)
The changeTwo parts. (a) Cross-theme: when gradeSet falls back from loadCohort(theme,…) to loadCohort(undefined,…) (gradeSet.ts:130), record a crossThemeCohort:true flag in GradeResult.sources and lower the Smart Price confidence (reuse Fix 6's confidence channel) so a pooled-theme percentile is never presented as theme-matched. Optionally gate the fallback behind a minimum cohort-homogeneity threshold (new constant). (b) Cold-start: add a one-time corpus pre-pricing step (a MineDeps/batch path) that part-prices and street-prices a seed set per theme×bucket so cohorts exist before live grading; until a bucket is seeded, Smart Price fails closed to insufficient (Fix 7), never to a fabricated number.
Test firstgradeSet.test.ts: when only the cross-theme cohort meets MIN_COHORT, result.sources.crossThemeCohort===true and smartPrice confidence is below the same-theme case. A bucket with zero priced corpus yields insufficient (not 6.0). Batch seed path tested through its dependency seam (no live network).
Watch outThe cold-start bootstrap is real infrastructure (batch pricing a corpus under BrickLink's 5k/day cap, persisting cohorts) and overlaps the data-ingest workstream — coordinate to avoid double-building the batch pricer. Cross-theme homogeneity gating risks darking out small themes; make the threshold a tunable and default it permissive.
Depends onFix 3, Fix 5, Fix 6, Fix 7
medium
9. Surface the licensed-premium note and wire Brick Insights PPP as a price sanity-check
Addresses'licensed premium not surfaced' (smartPrice.ts:43-48) + 'Brick Insights PPP aggregate exists, never used as price sanity-check' (brickInsights.ts)
The change(a) Per §1.1 ('licensing premium surfaced explicitly, not silently penalized'): when a set is licensed/exclusive and its value ratio sits low, attach a 'licensed premium' note to the Smart Price output (a string flag on GradeResult.sources / DropScore, not a score change). (b) Read the existing Brick Insights aggregate/PPP (parseBrickInsights already extracts it; brickInsights.ts) and, when present and reviewCount ≥ the spec's 2–3 trust floor, compare it to our computed value ratio; on a large divergence set a priceSanityFlag (does not alter the score, only flags for review) per sourcing-spec §47 ('Brick Insights PPP as sanity-check').
Test firstUnit test: a licensed set with a below-median ratio produces a 'licensed premium' note; an unlicensed one does not. Sanity-check test: a BI PPP that diverges >threshold from our ratio sets priceSanityFlag; reviewCount<trust-floor sets no flag (fail closed: thin BI data is not used). All via pure helpers + the gradeSet seam; never store BI raw text, only the derived number + source URL (derived-data rule).
Watch outBrick Insights coverage is uneven (sourcing spec §73) — most sets have no usable aggregate, so the sanity-check must be advisory and absent-by-default, never a gate. Notes/flags need a home: DropScore has no notes field today, so this needs a small additive output-shape change consumed downstream.
Depends onFix 1, Fix 3
Calls to make in this workstream
Build a real street-price anchor, or relabel Smart Price honestly as MSRP-anchored?
Options: (A) Build the BrickLink SET sold-median street anchor (Fixes 2–3) and populate facts.streetPrice live — delivers the marquee 'what people actually pay vs what the parts are worth' feature the spec §1.1 promises. (B) Relabel: change spec §1.1 + all output copy to say Smart Price is MSRP-anchored, and stop pretending streetPrice exists. (C) Hybrid: ship street price where the SET sold guide is liquid, explicitly label MSRP-anchored where it isn't.
Recommended: C (hybrid), built on A's adapter. The street anchor is cheap (one BrickLink SET call/grade, cached) and is the entire value proposition; but SET sold liquidity is thin for new/low-volume sets, so honesty demands an explicit per-grade 'street-anchored' vs 'MSRP-anchored (no street data)' label rather than silently mixing. Never the silent-MSRP status quo — that is, per the audit, 'the most misleading thing in this dimension.'
How high should MIN_COHORT_FOR_SMART_PRICE go, and how is thinness surfaced — gate, confidence penalty, or both?
Options: (A) Raise to ~15–25 hard gate (below it, Smart Price is insufficient). (B) Keep a low gate but scale confidence by cohort n so thin scores are visibly low-confidence. (C) Both: a hard floor AND a confidence ramp above it.
Recommended: C. A floor (recommend 15) stops statistically meaningless n=5 percentiles from being scored at all; a confidence ramp above it stops a 16-member cohort from looking as authoritative as a 200-member one. This pairs with the cold-start bootstrap (Fix 8) so the higher floor does not simply dark out Smart Price everywhere on day one.
How do we thread per-fact (cohort-size / cross-theme) confidence through the scoring core, which today hardcodes fact confidence to max(0.8, sigConf)?
Options: (A) Add an optional DimensionInput.factConfidence field read by resolveDimension (dimension.ts:33-39) — general, touches the engine core and its tests. (B) A Smart-Price-specific side channel (e.g. carry cohort n in GradeResult and adjust confidence post-hoc) — localized but special-cased. (C) Leave confidence as-is and only surface cohort n as informational metadata.
Recommended: A. A general DimensionInput.factConfidence is the cleanest and is reusable by other fact dimensions later; it is a small, well-bounded change to resolveDimension guarded by a focused test. B re-introduces the kind of special-casing the single-constants-file discipline exists to avoid; C leaves the over-confidence the audit specifically calls out.
Where do Smart Price's notes/flags (licensed-premium, cross-theme, price-sanity, anchor-basis) live in the output?
Options: (A) Add a notes/flags field to DropScore (engine output) consumed downstream. (B) Put them only on GradeResult.sources (orchestrator metadata), leaving the pure engine output unchanged. (C) Both — score-affecting context on DropScore, plumbing/provenance on GradeResult.sources.
Recommended: C. Basis label + cohort n + cross-theme + licensed-premium + price-sanity are provenance, not score, so most belong on GradeResult.sources; but the user-facing 'licensed premium' and 'data-light: street price unavailable' belong on DropScore so any consumer of the grade (not just gradeSet) sees them. Keep the pure engine free of network/string-formatting concerns.
New constantsSMART_PRICE_CONFIDENCE_BY_COHORT, COHORT_MAX_STALENESS_DAYS, BI_PPP_SANITY_DIVERGENCE, BI_MIN_REVIEW_COUNT_FOR_SANITY, CROSS_THEME_COHORT_MIN_HOMOGENEITY
SequencingFront-load the cheap, high-honesty seams first so later fixes have something to read. Order: Fix 1 (priceBasis seam, S) → Fix 4 (self-exclusion + midrank tie semantics, S — pure, immediate fidelity win) → Fix 7 (empty/thin → insufficient, S — the core fail-closed behavior) — these three are quick wins that stop the silent-confident-C immediately even before the street anchor exists. Then Fix 2 (SET street adapter, M) → Fix 3 (populate+persist streetPrice in gradeSet, M) deliver the marquee feature. Then the cohort-quality cluster: Fix 5 (basis+recency match, M) → Fix 6 (raise MIN_COHORT + confidence-by-n, M). Then Fix 8 (cold-start bootstrap + cross-theme label, L) — sequence LAST because raising MIN_COHORT (Fix 6) and insufficient-on-thin (Fix 7) widen the cold-start cliff that Fix 8 closes; landing Fix 8 before 6/7 ship to production prevents a window where most sets go dark. Fix 9 (licensed-premium note + BI sanity, M) can land any time after Fix 1/3 and is independent of the cohort cluster. Resolve the four decisions before Fix 2/6 — they gate the adapter scope and the engine-confidence change.
Wave 2Workstream 5 of 8
Data integrity and freshness
Done looks like: Stored prices, signals, and scores carry a fetched_at and expire via TTL (a stale set re-mines/re-prices instead of trusting a permanent cache hit); saveDimensionSignals is atomic and throws on a failed delete so a grade can never be double-counted; and base-vs-variant keying is reconciled so signals, refs, scores, and the cohort all line up on one canonical key.
Updated by your decisionsYour calls: canonical key is the <b>variant</b> (e.g. 75355-1), polybags deferred — a bit more re-keying than per-base, noted in the fixes. Score-freshness here powers the <b>monthly re-grade CRON</b> that makes grades living.
The approach
Three threads that share one new test seam. First, build a fake SupabaseClient query-builder (none exists today — the entire supabaseStore client layer is untested), because the atomic-write fix and every TTL read can only be exercised through it. Second, make saveDimensionSignals fail-closed: capture the delete error and throw before inserting, so a failed delete can never let the insert append a second signal set and double-count the dimension. Third, add fetched_at to the three cached tables (dimension_signals, part_prices, drop_scores) and make reads age-aware: loadDimensionSignals/loadPartPrices return only rows newer than a TTL, and the gradeSet cache-hit branch (currently permanent) treats an expired/empty cache as a miss so a re-grade re-mines on its own. Keying is reconciled by choosing the base number as the canonical signal/score/cohort key end-to-end, matching the base-keyed review_index that discovery already uses, and centralizing the base-extraction so scripts/grade-set.ts:57-59 stops mixing variant (signals) and base (refs) keys. TTLs and the canonical-key decision are product calls, so they go to constants.ts and to decisions respectively. Order: seam first, then the atomic fix (highest severity, smallest blast radius), then keying (it changes what the TTL reads key on), then freshness last.
Effort~1 week: 2 S (atomic guard, drop_scores stamp), 4 M (fake seam, keying helper, signals TTL, part_prices TTL). Add ~0.5 day coordination for the remote-schema migrations (fetched_at columns + optional replace_dimension_signals RPC) that live outside this repo.
Fixes6 · 2 small · 4 medium · 0 large
Depends onLoosely depends on the remote-schema/migrations track (the fetched_at columns and the optional atomic RPC live in Supabase, not in source). The drop_scores TTL only becomes actionable when a batch re-grade / staleness-trigger workstream consumes the isScoreStale signal. The cohort basis/recency mismatch (price.ts:65-105) overlaps the Smart Price / pricing workstream — this workstream only adds freshness + a stale-cohort flag, not the full re-pricing fix. Otherwise independent.
The fixes, in order
medium
1. Build a fake SupabaseClient query-builder test seam
AddressesEnabling infra for all three high-severity findings (supabaseStore.ts:147-168 atomic write; supabaseStore.ts + gradeSet.ts:89-93 freshness). The entire supabaseStore client layer is currently untested — supabaseStore.test.ts only covers pure row mappers, and there is no Supabase fake in src/lib/sources/fakes.ts.
The changeAdd src/lib/sources/fakeSupabase.ts exporting makeFakeSupabase(seed?) returning an object shaped like the subset of SupabaseClient the store uses: .from(table) yields a chainable builder supporting .select/.eq/.in/.gte/.lt/.not/.order/.range/.maybeSingle/.upsert/.insert/.delete, backed by per-table in-memory arrays. It must (a) record every call (table, op, args) for assertions, (b) let a test inject an error on a specific op (e.g. force delete to return { error }), and (c) honor eq/in/gte/lt filters so age-aware reads can be exercised. Cast to SupabaseClient at the boundary so store functions accept it unchanged.
Test firstsrc/lib/sources/fakeSupabase.test.ts: round-trip insert→select returns rows; .eq filters; an injected delete error surfaces as { error } from .delete(); call log records ops in order. This test defines the seam contract the later red tests depend on.
Watch outOver-fitting the fake to today's call shapes — if a store function uses an unmocked builder method it throws at runtime. Mitigate by implementing exactly the methods grep shows supabaseStore.ts/price.ts call, and throwing a clear 'unsupported op' error so gaps are obvious. This is a fake, not a real PostgREST emulation; keep it minimal.
small
2. Make saveDimensionSignals atomic / fail-closed on delete error
AddressessupabaseStore.ts:147-168 — delete-then-insert with the delete error discarded; a failed delete makes the insert append a second signal set, double-counting the grade (audit: high).
The changeIn saveDimensionSignals (supabaseStore.ts:147), capture the delete result: const { error: delErr } = await client.from('dimension_signals').delete().eq('set_number', setNumber); if (delErr) throw new Error(saveDimensionSignals delete: ${delErr.message}). Move the early if (rows.length === 0) return to AFTER the delete so an empty new set still clears the old rows (a set that lost all signals must not keep stale ones). Preferred end-state: replace delete+insert with a single client.rpc('replace_dimension_signals', {...}) Postgres function that deletes and inserts in one transaction (the only truly atomic option against PostgREST); if the RPC is descoped this round, the throw-on-delete-error guard is the fail-closed fallback. Apply the same capture-and-throw to any other delete-then-write in the store.
Test firstsrc/lib/sources/supabaseStore.client.test.ts using the fake: (1) inject a delete error → saveDimensionSignals rejects and the inserted-rows count stays 0 (no append); (2) happy path deletes existing rows for set then inserts the new set exactly once (assert one insert call, row count == new signals); (3) saving an empty DimensionSignals still issues the delete and leaves zero rows.
Watch outIf the RPC route is chosen, it needs a migration in the remote Supabase schema (out of this repo — see decisions); the JS-side throw-guard ships regardless and is independently testable. Watch that moving the empty-return after the delete doesn't break a caller that relied on a no-op for empty input — none found, but assert it.
Depends onFake SupabaseClient seam
medium
3. Reconcile base-vs-variant keying with a single canonical-key helper
Addressesscripts/grade-set.ts:57-59 (signals saved/loaded by variant '75355-1' but refs discovered by base '75355') and supabaseStore.ts:287-298 (loadCatalog dedupes by base_number); audit: medium — keys diverge across tables.
The changeAdd canonicalSetKey(setNumber) to a small util (e.g. src/lib/sources/setKey.ts) that strips the variant suffix to the base ('75355-1' → '75355'), replacing the ad-hoc sn.split('-')[0] at grade-set.ts:59. Route loadCachedSignals/saveCachedSignals (grade-set.ts:57-58) AND the upsertSet/upsertDropScore/getDropScore keys through canonicalSetKey so signals, refs, scores, and the cohort all key on the base — matching the base-keyed review_index discovery already uses. This is a deliberate base-as-canonical choice (see decisions). Document the keying contract in a comment in supabaseStore.ts next to loadCatalog. Note the related but separate loadCatalog/catalogRowsToKnownSets dedup mismatch (base sorts to first-by-string vs richest-by-num_parts) as a follow-up — out of scope here unless the canonical-key decision makes it trivial to align.
Test firstsrc/lib/sources/setKey.test.ts: '75355-1'→'75355', '75355'→'75355', a 7-digit base, and a no-dash id are idempotent. Then an orchestrator-level test (through GradeDeps' loadCachedSignals/saveCachedSignals seam): a grade run for '75355-1' saves signals under '75355' and a subsequent run for the sibling '75355-2' reads the same cached signals (proving the key is reconciled).
Watch outExisting rows in the remote DB are keyed by the OLD scheme (signals under variant); after this change those become orphaned and the set re-mines once. Acceptable (it self-heals on next grade) but call it out — there is no data migration in-repo. Verify no other reader keys dimension_signals by full variant before flipping.
Depends onFake SupabaseClient seam
medium
4. Add fetched_at + TTL to dimension_signals and make the cache hit expire
AddressesgradeSet.ts:89-93 (dimension_signals is a permanent cache hit → a re-grade never re-mines without --remine) and supabaseStore.ts (signals trusted forever); audit: high — no freshness anywhere.
The changeAdd fetched_at to the DimensionSignalRow type and the insert payload in saveDimensionSignals (set to new Date().toISOString() at write). Add loadDimensionSignals(client, key, opts?) supporting a maxAgeMonths/ttl filter via .gte('fetched_at', cutoffISO) where cutoff = now − SIGNAL_TTL_MONTHS; rows older than the TTL are simply not returned, so loadDimensionSignals yields {} → hasAnySignals(cached) is false → gradeSet.ts:90 takes the re-mine branch automatically. Keep --remine (forceRemine) as the manual override. TTL constant lives in constants.ts. The fetched_at COLUMN must exist in the remote schema (decisions / migration note).
Test firstIn supabaseStore.client.test.ts with the fake + a fixed clock (inject now via opts or a deps.now seam, never wall-clock): (1) a signal row written 2 months ago is returned when TTL=12mo; (2) a row written 18 months ago is filtered out → loadDimensionSignals returns {}; (3) saveDimensionSignals stamps fetched_at. Plus a gradeSet test through the mining seam: loadCachedSignals returning {} (expired) drives a re-mine (assert mineSetReviews path ran), while fresh cached signals still short-circuit.
Watch outRe-mining costs YouTube quota + LLM tokens (sourcing spec §5.6) — too short a TTL re-mines the whole catalog and blows the budget; spec §10 says look/build sentiment is timeless and only value sentiment decays, so a single blanket TTL is a blunt instrument (see decisions). Without the fetched_at column in the remote schema the filter no-ops or errors — gate behind the migration. Use an injected clock so the test isn't time-dependent.
Depends onFake SupabaseClient seam; canonical-key helper (so the TTL read keys on base)
AddressessupabaseStore.ts loadPartPrices/savePartPrices and price.ts:65-105 — prices trusted forever, and the graded set is re-priced live while the cohort reads whatever stale part_value_total was cached, distorting the percentile (audit: no freshness; cohort basis/recency mismatch, medium).
The changeAdd fetched_at to PartPriceRow + savePartPrices payload (stamped at write). Add an age filter to loadPartPrices (.gte('fetched_at', now − PART_PRICE_TTL_MONTHS)) so expired prices are treated as missing → priceInventory (price.ts:33-37) re-fetches them from BrickLink and re-caches, capped by maxFetches. For the cohort mismatch: at minimum document in loadCohort (price.ts:65) that cohort part values may be staler than the live-priced subject and add a guarded path to skip/flag the cohort when its rows are older than the subject's pricing run; full cohort-recompute-on-read is L and likely its own workstream — flag it, don't silently absorb it.
Test firstsupabaseStore.client.test.ts: a part price older than PART_PRICE_TTL_MONTHS is excluded from loadPartPrices' returned map → that part shows up in priceInventory's missing set (assert via the price.test seam with a fake BrickLink). Fresh prices are still served from cache (no BrickLink fetch). Cohort: a test asserting loadCohort flags/excludes rows beyond the freshness window rather than mixing them with a live-priced subject.
Watch outExpiring part prices forces BrickLink re-fetches against the ~5k/day cap (sourcing spec §6) — a too-aggressive TTL on a 700-part set can exhaust quota; TTL must be generous and is a tunable (decisions). The cohort fix is the riskiest to over-scope — keep this fix to freshness + a flag, not a full re-pricing engine.
Depends onFake SupabaseClient seam; constants for the TTL
small
6. Stamp fetched_at on drop_scores so a stale grade is re-gradeable
The changeAdd fetched_at to dropScoreToRow (supabaseStore.ts:71) stamped at write, and have getDropScore (supabaseStore.ts:131) select fetched_at and expose an isStale(maxAge) signal to callers (scripts/grade-set.ts reads it back at line 108) so a batch re-grade can choose to recompute expired scores. Score TTL constant in constants.ts. Does not change the grade math — only marks age. Keep the canonical base key here too (depends on the keying fix).
Test firstsupabaseStore.client.test.ts: upsertDropScore writes fetched_at; getDropScore returns it; a helper isScoreStale(row, ttl) is true past the window and false within it, using the injected clock.
Watch outNeeds the fetched_at column on drop_scores in the remote schema (migration note). Purely additive to the read path — low behavioral risk; the consumer (batch re-grade trigger) is a separate workstream, so this fix just exposes the signal.
Where does the canonical key live — base number ('75355') or full variant ('75355-1')? Today signals/scores/sets key on the variant while review_index and the catalog dedupe on base, so they don't join.
Options: (A) Base-as-canonical everywhere: signals, scores, cohort all keyed by base, matching discovery and loadCatalog's dedupe. (B) Variant-as-canonical: richer (distinguishes polybag vs main set) but requires re-keying review_index and the catalog, and loses the dedupe loadCatalog relies on.
Recommended: (A) base-as-canonical. It matches what discovery (grade-set.ts:59) and loadCatalog already do, is the smaller change, and the grade is conceptually per-set not per-variant. Note the cost: distinct variants of one base share a grade. If per-variant grading is a real product goal, that's a larger workstream, not this fix.
One blanket TTL, or per-signal-class TTLs? Methodology §10/§6.1 says value sentiment decays (HALFLIFE 12mo) but look/build sentiment is timeless.
Options: (A) Single SIGNAL_TTL_MONTHS for all dimension_signals — simple, but re-mines timeless look/build sentiment needlessly and re-spends quota. (B) Per-class TTL: short for value/worth dimensions, long/none for display/fidelity — honors the spec but needs the dimension→decays mapping (already on OpinionSignal.decays) plumbed into the read filter.
Recommended: Ship (A) as a single generous constant this round (unblocks the permanent-cache-hit bug), and file (B) as a fast-follow — the decays flag already exists on every signal, so per-class TTL is a clean later refinement, not a rewrite.
True atomicity for saveDimensionSignals requires a Postgres transaction (an RPC), which is a remote-schema migration outside this repo. Do we ship the RPC now or the JS throw-on-delete-error guard now?
Options: (A) Throw-on-delete-error guard only (pure JS, ships today, testable via the fake) — eliminates the double-count by aborting before the insert, but delete and insert are still two round-trips. (B) Add a replace_dimension_signals RPC (delete+insert in one txn) — truly atomic but needs a migration applied to remote Supabase.
Recommended: Ship (A) now (it fully closes the double-count path the audit describes) and queue (B) as the durable follow-up. Flag to Seth that all the fetched_at columns AND the RPC require remote-schema migrations that don't live in this repo — they must be applied to Supabase in lockstep with this code.
What TTL values? These are budget knobs, not engineering facts.
Options: Signals: ~12mo (aligns with the §10 value-sentiment HALFLIFE) vs longer to save quota. Part prices: long (weeks–months) given the 5k/day BrickLink cap. Scores: short (days–weeks) so re-grades pick up new reviews.
Recommended: Start SIGNAL_TTL_MONTHS=12, PART_PRICE_TTL_MONTHS=3, SCORE_TTL_DAYS=30 as constants.ts defaults and calibrate against real quota/cost once the batch re-grade trigger (separate workstream) exists. They're tunables, so cheap to change.
New constantsSIGNAL_TTL_MONTHS, PART_PRICE_TTL_MONTHS, SCORE_TTL_DAYS
SequencingFront-load the fake SupabaseClient seam (src/lib/sources/fakeSupabase.ts) — it is the prerequisite for every other red test in this workstream and the single highest-leverage piece, since the whole store client layer is untested today. Then the atomic saveDimensionSignals fix: highest severity, smallest blast radius, pure JS, and it's the one that can actively corrupt a grade right now (quick win). Then the canonical-key helper, because the TTL reads must key on whatever the reconciliation settles on (do it before freshness, not after). Then freshness in dependency order: dimension_signals TTL (unblocks the permanent cache-hit bug) → part_prices TTL → drop_scores stamping. Defer the per-signal-class TTL refinement and the full cohort-recompute-on-read; flag both. Coordinate every fetched_at column add and the optional RPC with a remote Supabase migration applied in lockstep — none of that schema lives in this repo, so code merged ahead of the migration will no-op or error.
Wave 3Workstream 6 of 8
Make the AI judge output safe to trust
Done looks like: Distilled LLM review scores can no longer enter a grade unvalidated, unbounded, or unweighted: every distilled score is schema-validated and clamped to 0–10, carries a per-dimension self-confidence that flows into the blend, fails closed (withheld, not zeroed) when the spec §8 accuracy gate hasn't passed; brigade-resistant like-counts shape comment weight; per-review derived scores are cached so re-grades don't re-distill; transient LLM/fetch failures retry with timeout and are distinguishable from "no data"; and bulk distillation runs on a cheaper, prompt-cached model.
Updated by your decisionsYour calls: sentiment <b>stays live with an "unvalidated" badge</b> (not switched off). Bulk model is <b>Sonnet 4.6</b> (Opus override for flagship), validated against your corpus before scale. The §8 work becomes the badge + the validation harness, not a hard off-switch.
The approach
The cluster splits into two layers. Layer A is pure logic with no external calls — the schema/clamp/validation gate in distill.ts, the new OpinionSignal.confidence field and its use in blend.ts, and the like-count → volume mapping. These are unit-testable today and are the trust-critical core, so they go first. Layer B is the live orchestrator hardening — per-review derived-score cache, retry/timeout, throttle-vs-missing distinction, and the cheaper cached model — tested through the existing MineDeps/GradeDeps seams with a fake Anthropic client (a stub messages.create). I sequence the fail-closed validation gate (spec §8) first because it is the single most-violated principle and gates everything downstream: until it's in, no other improvement matters because unvalidated scores still reach the math. Per-dimension confidence is next because it's a schema change that ripples into blend.ts and the dimension_signals persistence row, and several later fixes assume it exists. Caching, like-counts, retry, and the model swap are independent of each other and can be parallelized after the core lands. All tunables (model ids, retry counts, timeouts, like-weight curve, the validation-gate flag) move into src/config/constants.ts per the one-file rule. The derived-data+linkback rule is preserved throughout: nothing new stores raw text — the per-review cache keys on videoId/url and stores only the distilled DimEval numbers.
Effort~1.5 weeks: 3 S (clamp/validate, like-counts, model+cache), 4 M (confidence-weighting, per-review cache, retry/throttle, §8 gate enforcement). The hand-labelled validation corpus behind Fix 7 is additional product effort outside this estimate.
Fixes7 · 3 small · 4 medium · 0 large
Depends onLargely self-contained. Soft coupling to the scoring-engine workstream: this cluster adds OpinionSignal.confidence and feeds signalWeight (blend.ts), and the audit's separate finding that LOW_DATA_FLOOR/MIN_STRUCTURED_COUNT/MIN_COMMUNITY_TO_MOVE gates are unenforced lives in dimension.ts — the §8 'withheld' state (Fix 7) should be implemented consistently with however that workstream resolves the insufficient-data/coverage-tier handling, to avoid two divergent 'withheld' code paths. The per-review derived cache (Fix 4) should align with the data/persistence workstream's TTL/atomic-write fixes so it doesn't reintroduce the stale-forever cache bug. No hard ordering dependency on either.
The fixes, in order
small
1. Validate + clamp distilled scores at the distill boundary (fail closed)
AddressesNo validation gate before scores count — distill.ts:36 (distillationToSignals) and distill.ts:145-147 (raw JSON.parse, no clamp); spec §8 'HARD validation gate (not optional)'. Audit: 'Output not clamped to zero to ten before the math' (high) and 'No validation gate per spec section eight' (high).
The changeAdd a pure validateDistilled(parsed: unknown): DistilledScores | null in distill.ts that (a) confirms the parsed object has the required keys and types, (b) clamps every per-dimension score into [0,10] via a clampScore helper (null stays null), (c) returns null on any structural violation. Call it inside distillReview right after JSON.parse (don't depend on the caller's try/catch, per charter) AND have distillationToSignals clamp defensively too (score = clamp(evalResult.score) before emitting the OpinionSignal) so a bad score can never reach blendSignals even if a future caller bypasses distillReview. Wrap the JSON.parse itself in try/catch returning null. Out-of-range or non-numeric → drop that dimension (null), never coerce to a confident-looking number.
Test firstdistill.test.ts: distillationToSignals clamps score 14 → emits 10 and score -3 → emits 0; a parsed object with score 'high' (string) for buildFun → that dimension is omitted, others survive. New validateDistilled unit test: a JSON string missing coversSet → returns null; a well-formed object with one out-of-range score → returns the object with that score clamped. (Red first: assert clamp on 14 before implementing.)
Watch outdistillationToSignals is already tested; adding clamp must not change behavior for the in-range fixture (scores 5,7.5,8,9 stay identical). Watch that clamping null doesn't turn into 0 — null must pass through untouched, or every undiscussed dimension would wrongly score 0.
medium
2. Add a per-dimension confidence to the distillation schema and carry it into the blend
Addressesdistill.ts:52-60 (dimSchema has score+evidence but NO per-dimension confidence → confidence-weighting impossible); spec §8 'per-dimension self-confidence' + 'dimension score = confidence-weighted mean across reviews'. Audit: 'Per dimension confidence missing from schema; weighting impossible' (high).
The changeAdd confidence: { type:'number', 0–1 } to dimSchema and to the DimEval interface (distill.ts:7-10) and DISTILL_SYSTEM instructions. Add an optional confidence?: number field to OpinionSignal (types.ts:54-64). In distillationToSignals, set confidence on each emitted signal (clamped to [0,1], defaulting to a CONF_DEFAULT constant when absent so old cached rows still work). In blend.ts signalWeight, multiply the weight by (s.confidence ?? CONF_DEFAULT) so low-self-confidence reviews pull less — this realizes spec §8's confidence-weighted mean. Persist it: add a confidence column path in supabaseStore saveDimensionSignals/rowsToSignals (default when null).
Test firstblend.test.ts (mirror existing sig() helper + toBeCloseTo style): two equal-trust equal-volume signals scoring 8 (confidence 0.9) and 4 (confidence 0.1) blend strictly above 6 — the confident 8 dominates. distill.test.ts: a parsed dim with confidence 0.3 produces an OpinionSignal with confidence 0.3; a parsed dim with no confidence field → signal.confidence === CONF_DEFAULT. supabaseStore.test.ts: rowsToSignals round-trips confidence and defaults a null column to CONF_DEFAULT.
Watch outOpinionSignal is consumed in blend.ts, dimension.ts (signalConfidence), and supabaseStore row mappers — adding an optional field is backward-compatible, but signalWeight changing means existing blend.test expectations that assume confidence=1 must be checked (default CONF_DEFAULT must keep the current equal-weight tests passing — set default so a signal with no confidence behaves exactly as today). Decide whether the default is 1.0 (no behavior change for un-tagged signals) — recommended.
Depends onFix 1 (clamp helper reused for confidence [0,1])
small
3. Pass comment like-counts as a brigade-resistant sentiment weight
Addressesmine.ts:77 (comment like-counts discarded — fetchComments DOES return likeCount); sourcing spec §5.2 'Store like-counts as a sentiment weight'. Audit: 'Comment like counts discarded; brigade reads like consensus' (medium).
The changeStop mapping comments to bare strings at mine.ts:77. Keep the VideoComment {text, likeCount} objects and pass them through DistillParams (extend it to accept comments as {text, likeCount}[] or a parallel weights array). In the user-content assembly in distillReview, annotate each comment with its like-count (e.g. '- (▲42) text') so the LLM weights consensus by likes, and add a one-line instruction to DISTILL_SYSTEM: high-liked comments reflect more agreement, but the reviewer's body still outweighs comment noise. Do NOT let raw like-counts inflate volume directly — volume stays per-review =1; likes only shape the within-comment signal the LLM reads, consistent with the derived-only rule.
Test firstmine.test.ts via a fake distill seam (or a new distillReview unit test): assemble user content from comments [{text:'great', likeCount:50},{text:'meh', likeCount:0}] and assert the rendered prompt contains the like annotation for both, ordered/labelled. distillVideo: assert it forwards likeCount (spy/fake on distillReview captures comments-with-likes, not bare strings).
Watch outChanging DistillParams.comments shape touches distillRef (article path passes no comments — keep it optional/empty). Prompt growth is bounded (cap comments already 30). Must not let a single 9000-like brigade comment dominate — keep it as LLM-read context, not a numeric multiplier; consider capping the displayed like number.
medium
4. Per-review derived-score cache so re-grades never re-distill
Addressesmine.ts:74 (no per-review derived-score cache → every re-grade re-distills); sourcing spec §5.3/§10 'permanent derived-score cache keyed by videoId/review'. Audit: 'No per review derived score cache; every re grade re distills' (medium). Note: set-level MiningCache exists in gradeSet.ts but caches the whole set, not per-review.
The changeAdd an optional per-ref derived-score cache seam to MineDeps (mine.ts:44): loadDerived?(key) and saveDerived?(key, DistilledScores) where key = youtube:{videoId} or url:{href}. In distillRef/distillVideo, check loadDerived before fetching transcript/comments/LLM; on miss, distill then saveDerived. Store ONLY the DistilledScores numbers + the source key (derived-data rule), never transcript text. This is decoupled from the set-level cache: a video reused across two sets distills once.
Test firstmine.test.ts through MineDeps: a fake loadDerived returning a cached DistilledScores for videoId X makes mineSetReviews emit X's signals WITHOUT calling the fake anthropic.messages.create or fetchTranscript (assert those spies are never called). On a miss, saveDerived is called once with the distilled numbers and no raw text fields.
Watch outAdjacent audit (Data & persistence) flagged the existing set-cache has no TTL and a non-atomic save — do NOT replicate that. Give the new derived cache a freshness key or document that callers own TTL; keep saveDerived idempotent (upsert by key). Must not cache a withheld/failed distillation as if it were a real result (only cache on a validated, non-null DistilledScores).
Depends onFix 1 (only cache validated output)
medium
5. Retry + timeout on distill/fetch, and surface throttling as withheld (not missing)
Addressesmine.ts:121-128 (per-ref try/catch but no retry/timeout; bare catch swallows) and article.ts:43 (body trusted on length≥200 only). Spec fail-closed principle. Audit: 'No retry or timeout; transient errors swallowed' (medium); adjacent Cost/Scale audit: 'Catch-and-skip makes throttling indistinguishable from missing data' (high).
The changeAdd a small withRetry(fn, {retries, timeoutMs}) helper (new util) using AbortController for timeout and bounded exponential backoff, constants in config. Wrap the per-ref distill call in mineSetReviews. In the catch at mine.ts:125, classify the error: a rate-limit/throttle (HTTP 429 from youtube.ts ytGet, supadata non-404 !ok, or Anthropic RateLimitError) increments a stats.throttled counter and is recorded in MineResult.stats; a genuine 404/empty stays a clean skip. Surface stats.throttled up through gradeSet so a throttled set reads as withheld/insufficient, never as a confident thin grade. For article.ts:43, additionally require the extracted body to look like review prose (already gated by length≥200; tighten by also rejecting when the stripped text is mostly nav boilerplate — a min alpha-word ratio constant) so a thin/garbage page fails closed.
Test firstmine.test.ts via MineDeps: a fake anthropic.messages.create that throws a 429-shaped error twice then succeeds → withRetry yields the success and stats.distilled increments; one that always 429s → stats.throttled increments and that ref contributes no signal (and is NOT cached). A timeout test: messages.create that never resolves → AbortController fires, counted as throttled, not a silent skip. article.test.ts: a 250-char nav-only blob (low alpha-word ratio) → extractArticleText/fetchArticleText returns null.
Watch outAbortController wiring must thread into fetch calls in youtube.ts/supadata.ts/article.ts or be wrapped at the call site — wrap at the orchestrator to avoid touching every adapter. Retry must not amplify cost on a hard auth failure (only retry 429/5xx/timeout, never 4xx-auth — mirror error-code guidance). Changing the article gate risks dropping currently-passing thin reviews; calibrate the ratio constant against a couple of real fixtures before tightening.
6. Use a cheaper, prompt-cached model for bulk distillation
Addressesdistill.ts:130-135 (model claude-opus-4-8 at effort 'low', no prompt caching — costly for bulk); sourcing spec §5.6 (LLM tokens are the second cost) + §10 cache aggressively. Audit: 'Opus low effort, no prompt caching; costly' (medium) and adjacent Cost/Scale: transcripts sent uncapped to opus.
The changeMove the model id to constants (DISTILL_MODEL = 'claude-haiku-4-5' for bulk; keep an override path for a high-value head set). In distillReview: (a) switch default model to the constant; (b) add cache_control:{type:'ephemeral'} to the system block so the static DISTILL_SYSTEM+schema prefix is cached across the thousands of per-review calls; (c) drop output_config.effort when the model is Haiku (Haiku 4.5 does NOT support effort and will 400) — make effort conditional on an Opus-tier model; (d) cap transcript length to a MAX_TRANSCRIPT_CHARS constant before sending (the largest uncontrolled cost driver). Keep output_config.format json_schema (supported on Haiku 4.5).
Test firstdistill unit test with a fake anthropic.messages.create capturing the request: model === DISTILL_MODEL; system block carries cache_control ephemeral; effort is absent when model is Haiku; a 200k-char transcript is truncated to MAX_TRANSCRIPT_CHARS in the user content. (Red: assert model is the Haiku constant and effort is omitted before changing distillReview.)
Watch outTwo real gotchas verified against the claude-api skill: (1) Haiku 4.5 rejects output_config.effort (400) — the conditional is mandatory, not cosmetic. (2) The DISTILL_SYSTEM prefix is ~600 tokens, BELOW the 4096-token minimum cacheable prefix for Haiku/Opus 4.x — so cache_control will silently NOT cache (cache_creation_input_tokens=0) unless the stable prefix is padded past 4096 (e.g. fuller rubric/examples in the system prompt) or batched. Accuracy risk: Haiku may distill sarcasm worse than Opus — this is exactly why the §8 validation gate (Fix 7) must run on whichever model ships, and why model choice is a decision below, not a silent pick.
medium
7. Enforce the spec §8 validation gate: withhold distilled scores until accuracy clears the threshold
Addressesdistill.ts:36 — spec §8 'distilled scores do not enter any grade until accuracy clears a threshold'; the system must FAIL CLOSED. This is the charter's headline item and the most-violated principle.
The changeAdd a DISTILLATION_VALIDATED boolean (default FALSE) and DISTILLATION_MIN_ACCURACY threshold to constants. Thread a validated flag into the mining path (MineDeps or GradeOpts). When NOT validated, mineSetReviews still discovers + distills but tags the resulting signals as withheld — they contribute 0 to the computed score and are surfaced as an explicit 'distillation pending validation' coverage note (mirroring the §2.2 community display-only pattern), never silently dropped and never counted as confident. Provide a thin offline harness entry point (a function over a hand-labelled fixture set that computes per-dimension + sarcasm-flag accuracy and returns pass/fail vs the threshold) so flipping the flag is evidence-based; the harness corpus itself is a product input, not code.
Test firstmine/gradeSet test via the seam: with validated=false, a set whose ONLY signals are distilled returns coverageTier 'data-light' and those dimensions are state 'insufficient'/withheld (not 'scored'), and confidence does not rise from them. With validated=true, the same distilled signals score normally. Harness unit test: given a labelled fixture where the judge matches 9/10, accuracy computes correctly and compares to the threshold.
Watch outThis deliberately neuters the 'feelings half' until the gate passes — coordinate with the team so it's not perceived as a regression; it is the intended fail-closed posture. The full hand-labelled corpus (hundreds of reviews) is out of code scope; the fix delivers the enforcement mechanism + harness shell, not the labels. Make sure 'withheld' signals don't leak into the §6 confidence average (they must not count as scored).
Which model for bulk distillation, given Haiku is ~5x cheaper but may distill LEGO sarcasm less accurately than Opus?
Options: (a) Haiku 4.5 ($1/$5, 200K ctx) for the whole catalog; (b) Opus 4.8 ($5/$25) everywhere, accept the cost; (c) tiered — Haiku for the long tail, Opus 4.8 for the popular/high-value head where accuracy matters most.
Recommended: (c) tiered, defaulting to Haiku via the DISTILL_MODEL constant with an Opus override for head sets. But this choice is only safe AFTER the §8 validation gate (Fix 7) is enforcing — measure each model's sarcasm-flag accuracy on the labelled corpus and let the threshold, not intuition, decide. Do not ship Haiku to production grades before the gate passes on Haiku specifically.
Prompt caching needs a ≥4096-token stable prefix, but DISTILL_SYSTEM is ~600 tokens — pad the prompt to hit the cache floor, or skip caching?
Options: (a) Expand DISTILL_SYSTEM with a fuller rubric/few-shot examples so the cached prefix clears 4096 tokens (also likely improves accuracy); (b) leave the prompt lean and accept cache_control is a silent no-op (cache_creation=0) — rely on the cheaper model alone for cost; (c) use the Batches API (50% off) for bulk instead of per-call caching.
Recommended: (a) — expanding the system prompt with concrete sarcasm/coverage examples both unlocks caching AND is the kind of rubric detail that lifts distillation accuracy toward the §8 gate, so it pays twice. Verify with usage.cache_read_input_tokens>0 after. Consider (c) for the initial full-catalog backfill where latency doesn't matter.
Default for the new OpinionSignal.confidence when a signal has no LLM self-confidence (old cached rows, community/Brickset signals)?
Options: (a) default 1.0 — un-tagged signals behave exactly as today, zero behavior change; (b) default to a mid value (e.g. 0.6) — treats un-tagged signals as moderately trusted.
Recommended: (a) 1.0. It keeps every existing blend.test expectation green and confines the new weighting to genuinely LLM-distilled signals, which is the only place self-confidence is meaningful. Revisit only if calibration shows un-tagged sources are over-weighted.
Does enforcing the §8 gate (distilled scores withheld until validated) ship now, blocking the 'feelings half' of grades until a labelled corpus exists?
Options: (a) Ship the enforcement + harness shell now, flag default FALSE — grades fall back to factual/structured signals until the corpus is labelled and the gate flipped; (b) defer enforcement, keep distilled scores live but add a visible 'unvalidated' badge; (c) ship enforcement but flag default TRUE in dev so the pipeline is exercisable.
Recommended: (a) — this is the fail-closed posture the audit says is most violated; the engine should not present confident letters built on an unvalidated judge. Flag FALSE in production, TRUE in test fixtures/dev so the rest of the pipeline stays exercisable. Labelling the corpus is the unblock, and it's a product task, not an engineering one.
SequencingFront-load the two trust-critical pure-logic wins, then the fail-closed gate, then parallelize the rest. Order: (1) Fix 1 clamp+validate — smallest, highest leverage, unblocks safe caching. (2) Fix 2 per-dimension confidence — schema + blend + persistence ripple, needed by the gate. (3) Fix 7 §8 validation-gate enforcement — the headline fail-closed change; depends on 1 and 2. After the core lands, Fixes 3 (like-counts), 4 (per-review cache), 5 (retry/throttle), and 6 (cheap cached model) are mutually independent and can be done in any order or in parallel — but Fix 4 must precede the cache-related assertions in Fix 5 (don't cache throttled/failed refs), and Fix 6's accuracy risk is only acceptable once Fix 7 is enforcing. Quick wins to grab first: Fix 1 (S), Fix 3 (S), Fix 6 (S, minus the prompt-padding decision).
Wave 3Workstream 7 of 8
Fix review-discovery precision and recall
Done looks like: YouTube/Brick-Insights discovery produces refs whose confidence honestly reflects evidence — a unique title-name OR a description-number alone is no longer auto-accepted as mine-worthy (spec §5.1 two-agreeing-signals), recall on hyphenated/long licensed names is restored, both code paths disambiguate against one canonical catalog name, and every sweep emits per-confidence counts plus a judge-rejection rate so false-positive pressure is observable. The system fails closed: a thin/ambiguous match is withheld from mining rather than dumped on a single Opus judge call.
The approach
The cluster splits into three threads that share one new config block. Thread A (recall) fixes the tokenizer so hyphenated and long licensed names actually match — these are pure-logic edits to youtube.ts with the richest unit-test payoff and no downstream coupling, so they go first. Thread B (precision + fail-closed) is the spec-critical one: introduce a match-confidence concept (the existing high/medium/low tiers) backed by a "two agreeing signals" rule, then gate mining on it at the seam — loadReviewRefs/selectRefsToMine drop and cap LOW refs so one haul cannot inject dozens of spurious pairs gated only by the LLM judge. Thread C (consistency + observability) unifies the two divergent matchers, aligns the two catalog-dedup paths onto the richest-variant canonical, fixes the isYear blind spot, and adds the per-confidence + judge-rejection observability the charter demands. All magic numbers (tier ranks, the N-of-M token fraction, the per-set LOW cap, the cold-start search cap) move into src/config/constants.ts. Every fix is TDD: the pure matchers are unit-tested in youtube.test.ts/discover.test.ts; the mining gate is tested through the existing GradeDeps.mining.discoverRefs / MineDeps seam (gradeSet.test.ts / mine.test.ts) with fakes, never live. The guiding constraint throughout: when a match is ambiguous or a second signal is absent, withhold (return fewer/lower refs) rather than emit a confident-looking pairing.
Effort~1.5-2 weeks: 2 S (hyphen normalize, catalog-dedup align), 7 M (N-of-M, two-agreeing tiers, mining gate, matcher unify, year-context, observability, BI cache/backoff). No L items, but the two-agreeing-signals + mining-gate pair is the critical path and should get the most calibration/review time.
Fixes9 · 2 small · 7 medium · 0 large
Depends onCoordinates with the coverage/confidence workstream (the scoring-engine audit, index 0): when the mining gate withholds refs for a thin/ambiguous set, that set must downgrade to the data-light coverage tier and never present a confident letter — that fail-closed behavior is owned there, this workstream just stops feeding it spurious refs. Overlaps the cost/robustness workstream (audit index 4 / cost-scale) on rate-limiting + backoff: the Brick Insights cache/backoff should reuse a shared throttle util rather than grow its own, and the catch-and-skip-makes-throttling-look-like-missing-data finding applies to the mine path this workstream instruments. Otherwise self-contained.
The fixes, in order
small
1. Normalize hyphens/punctuation so X-Wing matches "X Wing"
AddressesDiscovery audit #3 (high) at src/lib/sources/youtube.ts:45-51 — nameTokens glues hyphens; 'X-Wing Starfighter' never matches a title 'X Wing Starfighter' (verified empty for both 'X Wing' and 'XWing')
The changeIn youtube.ts, change nameTokens (lines 45-51) to replace hyphens (and slashes/periods) with spaces BEFORE splitting, instead of keeping '-' in the character class — i.e. .replace(/[^a-z0-9\s]/g,' ') after first turning - into a space, so 'x-wing' tokenizes to ['x','wing']. Apply identically wherever the set name and the title are tokenized so the two sides use the same alphabet. Because 'x' and 'wing' are now separate tokens, revisit the t.length >= 2 filter in significantNameTokens (line 55): keep length>=2 but confirm distinctive 1-char tokens like the 'x' in 'X-Wing' are not the sole identifier (the >=2-distinct-token gate at line 61 still protects against generic single hits). Keep matchVideoToSet and the index path (buildSetIndex/significantNameTokens) on the SAME tokenizer so they cannot diverge.
Test firstAdd to youtube.test.ts (matchVideoToSets + matchVideoToSet blocks): a set {number:'75355',name:'X-Wing Starfighter',year:2023} must match titles 'LEGO X Wing Starfighter review', 'X-Wing Starfighter review', and (decision permitting) 'XWing Starfighter review' — all currently return [] / null. Assert the unique match is MEDIUM (or HIGH once the second-signal rule lands). Add a negative: 'wing nut factory tour' must NOT match (the 'x' is required, not just 'wing').
Watch outLoosening the alphabet can re-admit generic single-word hits; the >=2-distinct-significant-token gate (line 61) is the backstop and must be kept. Watch the existing passing test 'ignores a name without 2 DISTINCT distinctive words (Bricks Bricks Bricks)' — it must still pass.
medium
2. N-of-M token match for long licensed names
AddressesDiscovery audit #4 (medium) at src/lib/sources/youtube.ts:59-63 — nameMatches/subset require EVERY distinctive token, so "Luke Skywalker's X-wing Fighter" misses a title 'Luke's X-wing Fighter build' (drops 'skywalker'); long licensed names have depressed recall
The changeReplace the all-tokens-required predicate (nameMatches line 59-63 for matchVideoToSet, and the subset lambda line 155 for matchVideoToSets) with an N-of-M threshold: a name matches when at least ceil(sig.length * NAME_MATCH_TOKEN_FRACTION) of its distinctive tokens appear, with a floor of 2 matched tokens so short names still need full agreement. Pull NAME_MATCH_TOKEN_FRACTION (e.g. 0.7) and NAME_MATCH_MIN_TOKENS (2) from constants.ts. Keep the ambiguity rule intact: if the relaxed threshold makes two sets match, they stay LOW (the judge decides), preserving fail-closed behavior.
Test firstyoutube.test.ts: {number:'75301',name:"Luke Skywalker's X-wing Fighter",year:2021} matches title "Luke's X-wing Fighter build" (3 of 4 tokens) as a unique MEDIUM. Negative guard: a title with only 2 of 5 tokens of a long name does NOT match. Boundary: a 2-token name ('Galaxy Explorer') still requires BOTH tokens (floor), so 'Galaxy haul' does not match.
Watch outLowering the bar trades recall for precision — the biggest false-positive lever in the cluster. Mitigated because relaxed matches that collide land at LOW and are then dropped by the mining gate (next fix). The fraction is a product-tunable; see decisions.
3. Enforce two-agreeing-signals before a match is mine-worthy
AddressesDiscovery audit #1 (high) at src/lib/sources/youtube.ts:140-162 — a unique title-name match alone is MEDIUM and a lone description-number is LOW, both written to review_index with only the LLM judge as corroboration; spec §5.1 requires two agreeing signals before auto-accept
The changeKeep matchVideoToSets emitting all candidate tiers (it is the recall net), but make the SEMANTICS explicit and spec-aligned: HIGH = number-in-title (authoritative, one strong signal is enough per spec's 'number is authoritative'); promote a name+number agreement (e.g. unique title-name AND that set's number present in title or description) to HIGH as the canonical 'two agreeing signals'; demote a lone unique title-name from MEDIUM to a corroboration-pending tier and a lone description-number stays LOW. Centralize the tier ranks in constants.ts (MATCH_CONFIDENCE_RANK) so matchVideoToSets, selectRefsToMine and discover.ts stop each hardcoding their own {high:3,medium:2,low:1}. The 'two agreeing' enforcement itself lives at the mining gate (next fix), so discovery stays a complete index but mining only auto-accepts corroborated refs.
Test firstyoutube.test.ts: a title carrying BOTH a unique name and the matching set number → HIGH (new 'two agreeing' case). A lone unique title-name with no number → the corroboration-pending tier (assert it is NOT auto-mined downstream, see mine/grade tests). A lone description-number → LOW. Existing 'tags a UNIQUE name in the title MEDIUM' test is updated to the new tier name and a comment ties it to spec §5.1.
Watch outThis changes the public confidence tag semantics that discover.test.ts and youtube.test.ts assert on; those tests must be updated in lockstep (they are the spec executable). Risk of over-tightening recall — keep the lone-name ref in the index at the pending tier rather than dropping it at discovery, so the cold-start search fallback can still use it.
Depends onN-of-M token match (defines which name matches exist to corroborate)
medium
4. Gate + cap mining by confidence at the seam (fail closed)
AddressesDiscovery audit #2 (high) at src/lib/sources/supabaseStore.ts:243-250 (loadReviewRefs returns every tier unfiltered) and the charter note on src/lib/pipeline/mine.ts:31-42 (selectRefsToMine ranks but doesn't cap/drop LOW) — a 3-number haul yields 3 LOW refs, 3 same-named sets all kept, thousands of spurious pairs gated only by one Opus call each
The changeTwo-part gate. (1) In mine.ts selectRefsToMine (lines 31-42): after ranking, DROP refs below a minimum mine tier and CAP the number of LOW/uncorroborated refs per set, both from constants.ts (MINE_MIN_CONFIDENCE, MAX_LOW_REFS_PER_SET). A lone-name corroboration-pending ref or a bare description-number is not auto-mined unless it is corroborated or under the cap. (2) In gradeSet.ts the discoverRefs seam already injects loadReviewRefs (scripts/grade-set.ts:59, batch-ingest.ts:90); add an optional confidence filter parameter to loadReviewRefs (supabaseStore.ts:243-250) so the DB read can pre-filter tiers, defaulting to all-tiers for backward compat but called with the gate tier from the scripts. This is the fail-closed point: when discovery is thin/ambiguous, the miner withholds rather than dumping on the judge.
Test firstmine.test.ts (pure, no live calls): selectRefsToMine drops a lone LOW description-number ref when MINE_MIN_CONFIDENCE excludes it; with 5 LOW refs for one set and MAX_LOW_REFS_PER_SET=2, only 2 survive; HIGH/two-agreeing refs always survive and sort first (existing 'mines higher-confidence refs first' test extended). gradeSet.test.ts via GradeDeps.mining: a fake discoverRefs returning [1 HIGH, 4 LOW] for a set results in mineSetReviews being asked to distill only the gated subset (assert stats.found reflects the cap). supabaseStore.test.ts: rowToReviewRef/loadReviewRefs row-mapper test for the new tier filter (pure mapper test; the live query stays untested per house rules).
Watch outCapping LOW refs can starve genuinely-quiet sets of any signal — but that is the intended fail-closed posture (a withheld/data-light grade, methodology §6, beats a confident grade built on one uncorroborated haul mention). Keep the cold-start search.list fallback (mine.ts:103-117) reachable so a set with zero indexed corroborated refs can still try search. Coordinate with the coverage-tier workstream so a gated-thin set downgrades coverage rather than silently scoring.
Depends onEnforce two-agreeing-signals (defines the tiers this gate reads)
small
5. Align the two catalog-dedup paths onto one canonical variant
AddressesDiscovery audit #7 (medium) at src/lib/sources/supabaseStore.ts:287-298 — loadCatalog dedupes to the first row per base_number by string-sorted set_number, while catalogRowsToKnownSets (catalog.ts:88-94) dedupes to the richest variant by num_parts; the production sweep uses loadCatalog, so a base is disambiguated against the wrong (polybag/promo) name/year
The changeChange loadCatalog (supabaseStore.ts:284-302) to pick the richest variant per base_number to match catalogRowsToKnownSets: select num_parts alongside base_number,name,year, and when a base is already seen, replace it only if the new row has more parts (mirror catalog.ts:91-92 r.numParts > prev.numParts). Drop the reliance on .order('set_number') + first-wins. Keep pagination. Optionally extract a shared pickRichestPerBase helper used by both so they cannot drift again.
Test firstsupabaseStore.test.ts: a new pure helper test feeding catalog rows where base '75355' has variant '75355-1' (1949 parts, name 'X-Wing Starfighter') and '75355-2' (5 parts, name 'X-Wing promo polybag') in non-richest sort order asserts the resolved KnownSet is the 1949-part 'X-Wing Starfighter' with its year — i.e. the same result catalogRowsToKnownSets returns for the same rows (assert parity directly).
Watch outloadCatalog is currently DB-query-shaped (hard to unit-test live); extract the dedup into a pure function over rows so it is testable per house rules and the live .range() paging stays a thin untested shell. Adding num_parts to the select is a trivial schema-safe change (column already exists in set_catalog).
medium
6. Unify the two matchers; mark matchVideoToSet non-production
AddressesDiscovery audit #5 (medium) at src/lib/sources/youtube.ts:76-97 — matchVideoToSet uses a length>=5 block and no year guard; matchVideoToSets uses isYear and no length block; they disagree, and matchVideoToSet is dead in the bulk pipeline yet exported/tested as production
The changeMake matchVideoToSet (1:1) a thin wrapper over the index path: build a one-off SetIndex and return the single highest-confidence match from matchVideoToSets (or null on ambiguity), so the two cannot diverge on year/number guards. Reconcile guards onto the index path's behavior (isYear + year gating). Mark matchVideoToSet with an @deprecated JSDoc noting the bulk pipeline uses matchVideoToSets, and keep its tests as a thin parity layer rather than a second production spec.
Test firstyoutube.test.ts: a parity test asserting matchVideoToSet(video, catalog) === the single-best of matchVideoToSets(video, buildSetIndex(catalog)) across the existing matchVideoToSet cases (number-in-title, name-only, pre-release reject, ambiguous→null). The existing 'New Republic ... | 75460' case (a different set's number → null) must hold on BOTH paths.
Watch outmatchVideoToSet's current length>=5 'a specific other set's number → don't name-guess' behavior (line 84-85) must be preserved through the unification or a real precision guard is lost — fold it into the index path explicitly. Several existing matchVideoToSet tests encode subtly different behavior; expect to update assertions where the unified behavior is strictly better.
Depends onEnforce two-agreeing-signals (the unified tier semantics)
medium
7. Disambiguate year-like set numbers by context instead of blanket drop
AddressesDiscovery audit #6 (medium) at src/lib/sources/youtube.ts:149-151 — isYear blinds the matcher to real year-numbered classic sets (2000, 2025 exist in Rebrickable); 'LEGO 2025 unboxing' for set 2025 returns empty
The changeReplace the blanket isYear drop (youtube.ts:149) with context-aware acceptance: a 19xx/20xx token counts as a set number only when it is corroborated — either prefixed (e.g. 'set 2000', '#2000', 'LEGO 2000') or co-occurring with a name token of that known set in the title — otherwise it is treated as a year. Put the prefix patterns / behavior behind a small helper and keep the bare-year-rejects-by-default posture (precision-first). This dovetails with the two-agreeing-signals rule: a year-like number alone is never enough.
Test firstyoutube.test.ts: extend the existing 'does NOT match a bare 4-digit year' test — 'best LEGO sets of 2025!' still yields [] for set '2025', but 'LEGO set 2025 classic space unboxing' (prefixed) OR '2025 Galaxy Explorer review' (number + name token of set 2025) yields the year-numbered set. '75355 review filmed in 2025' still resolves only 75355 (the bare 2025 stays a year).
Watch outYear-numbered sets are a tiny, mostly-vintage slice (1970s-2000s); over-engineering here has low payoff and real false-positive risk (every 'sets of 2025' video). Keep the default a year-reject and only accept on explicit corroboration. Confirm the constants-driven prefix list is conservative.
AddressesCharter 'Add observability: per-confidence counts + judge-rejection rate'; Discovery audit recommendation 'Emit per-confidence counts and a periodic judge-rejection rate by tier'; cross-cutting with the no-observability finding (cost/robustness audit, batch-ingest.ts:105-107)
The changeDiscovery side already counts byConf in scripts/discover.ts:72-87 — promote that into a small reusable summarizer (e.g. summarizeRefs(refs) → per-tier counts + per-set fan-out) in discover.ts so it is unit-testable and reused by batch-ingest. Judge side: the LLM judge is distill.coversSet (distill.ts) invoked via mineSetReviews; thread a per-tier accept/reject tally out of the mine path (extend MineResult.stats with rejectedByTier / acceptedByTier counts based on the ref's confidence vs whether distill returned signals) and log it in scripts/grade-set.ts and batch-ingest.ts. This makes false-positive pressure (high LOW-tier rejection rate) visible without storing any raw text — counts only, honoring the derived-data rule.
Test firstdiscover.test.ts: summarizeRefs over a fixture of mixed-tier refs returns the correct per-tier counts and distinct-set count. mine.test.ts: mineSetReviews stats expose rejectedByTier — with a fake distillRef that returns null for LOW refs and signals for HIGH, assert stats.rejectedByTier.low === N and acceptedByTier.high === M (tested through the MineDeps seam with a stub distiller; no live Anthropic/YouTube calls).
Watch outThreading reject-by-tier through mineSetReviews touches the per-ref try/catch (mine.ts:121-128) — be careful not to double-count a fetch failure as a judge rejection (distinguish 'distill returned null' from 'fetch threw'). Keep it counts-only so nothing raw leaks into logs (derived-data rule).
Depends onEnforce two-agreeing-signals (tiers to bucket by); Gate + cap mining (the mine path being instrumented)
medium
9. Cache Brick Insights pages + backoff; stop hardcoding MEDIUM
AddressesDiscovery audit #10 (low) at src/lib/sources/brickInsights.ts:131-161 — refs hardcoded MEDIUM regardless of match quality, one on-demand fetch per set with no cache/backoff, so the spec's free crawl frontier (§4) is not swept in bulk and is fragile at scale
The changeIn brickInsights.ts: (1) keep BI refs as a legitimate corroborating signal but make the confidence honest — BI has already matched the set, so a BI YouTube ref whose video also passes matchVideoToSets is HIGH (two agreeing: BI map + our matcher), while an un-revalidated BI blog/forum link stays MEDIUM-as-corroboration, the constant moved to constants.ts (BRICK_INSIGHTS_BASE_CONFIDENCE). (2) Add a thin cache + polite backoff to fetchBrickInsightsHtml (brickInsights.ts:150-155) per spec §4 politeness (≤60 req/min, cache aggressively): an injectable fetch + a simple on-disk/Supabase HTML cache keyed by setNumber, and exponential backoff with Retry-After honoring, following the existing adapter pattern (thin client, typed contract). This lets BI serve as the bulk corroborating second signal the spec intends rather than a fragile per-set call.
Test firstbrickInsights.test.ts: brickInsightsRefs reads BRICK_INSIGHTS_BASE_CONFIDENCE from constants (assert the tier comes from config, not a literal); a BI YouTube ref that ALSO matches the index is tagged HIGH (corroborated), a bare blog link stays the configured corroboration tier. For the client: inject a fake fetch that returns 429 then 200 and assert one backoff/retry occurred and the cached body is reused on a second call (seam-tested, no live HTTP).
Watch outBackoff/caching is partly shared with the broader 'no rate-limiting anywhere' robustness finding (youtube.ts:184, bricklink.ts:107) — coordinate so BI doesn't grow a bespoke backoff that diverges from a future shared throttle util. The HTML cache must store derived/transient HTML within the processing window only; do not persist raw review prose (derived-data rule) — cache the page HTML transiently for re-parse, not the extracted text.
Depends onAlign catalog-dedup (BI corroboration needs the same canonical names); Enforce two-agreeing-signals
Calls to make in this workstream
What N-of-M token fraction (and floor) should a name match require? This trades recall on long licensed names against false positives on the 24k-set, 24%-name-overlap catalog.
Options: (a) 0.7 fraction with a 2-token floor (recommended) — recovers "Luke's X-wing Fighter" while keeping short names exact; (b) all-tokens (today) — highest precision, demonstrably misses real reviews; (c) 0.5 — aggressive recall, likely too many cross-set collisions that then lean entirely on the judge.
Recommended: (a) 0.7 + 2-token floor, with the relaxed-match collisions kept at LOW so the mining gate (not the matcher) absorbs the precision cost. Calibrate the fraction against a labeled sample of real reviewer-channel titles before raising recall further.
Where does 'fail closed' bite when discovery is thin — drop the ref at discovery, or keep it indexed but withhold it from mining?
Options: (a) Keep lone-name/lone-number refs in review_index but gate them out of mining via MINE_MIN_CONFIDENCE + MAX_LOW_REFS_PER_SET (recommended); (b) never index them at all.
Recommended: (a). Indexing-but-not-mining preserves the data for the cold-start search fallback and future re-evaluation while still failing closed at the point that costs money/credibility (the Opus judge + the grade). It also keeps the observability counts meaningful (you can measure how many LOW refs you're choosing not to mine).
How aggressively to cap LOW refs per set (MAX_LOW_REFS_PER_SET) and the minimum mine tier (MINE_MIN_CONFIDENCE)?
Options: (a) MINE_MIN_CONFIDENCE = corroboration-pending and MAX_LOW_REFS_PER_SET = 1-2 (recommended, strict fail-closed); (b) mine LOW but cap at ~3; (c) mine all (today).
Recommended: (a) Strict: only auto-mine HIGH/two-agreeing by default, allow a small capped number of pending/LOW only when a set has no corroborated ref (so quiet sets still get a shot), and let the judge-rejection-rate metric tell you whether to loosen. This directly answers the auditor's '3-number haul → 3 LOW refs' probe.
Should year-like set-number disambiguation (audit #6) be built now or descoped, given it affects only a tiny vintage slice?
Options: (a) Build the context-aware version (prefix or co-name corroboration); (b) descope and keep the blanket isYear drop, documenting the known miss for ~1970s-2000s year-numbered sets.
Recommended: Lean (a)-lite: implement only the cheap, high-precision corroboration (explicit prefix like 'set 2025'/'#2025' OR co-occurring name token) and keep bare-year rejection as default. It falls out almost free from the two-agreeing-signals machinery; full context disambiguation beyond that is not worth the false-positive risk on 'sets of 2025' videos.
New constantsMATCH_CONFIDENCE_RANK, NAME_MATCH_TOKEN_FRACTION, NAME_MATCH_MIN_TOKENS, MINE_MIN_CONFIDENCE, MAX_LOW_REFS_PER_SET, COLD_START_SEARCH_MAX_PER_DAY, BRICK_INSIGHTS_BASE_CONFIDENCE
SequencingFront-load the two pure-logic recall fixes (hyphen normalization, then N-of-M) — they are small, high-value, fully unit-testable in youtube.test.ts, and unblock everything that reasons about which names match. Then land the precision spine in order: (3) two-agreeing-signals tier semantics → (4) the mining gate/cap at the seam → these two together are the spec §5.1 fix and the single most important deliverable. In parallel and independent: (5) the catalog-dedup alignment (no dependency on the matcher changes) is a quick win that can go first if a second engineer is available. After the tier semantics exist, do (6) matcher unification, (7) year-like context disambiguation, and (8) observability (it needs the final tiers to bucket by). Finish with (9) Brick Insights caching/backoff, which depends on both the canonical catalog names and the tier semantics and overlaps the broader rate-limiting workstream. Quick wins to pull forward: hyphen normalization (S), catalog-dedup alignment (S).
Wave 4Workstream 8 of 8
Operational hardening for scale
Done looks like: The 24k-set sweep is auditable and survivable: every external call is rate-limited, backed-off, and concurrency-bounded; a throttled (429) set is recorded as an explicit "degraded/throttled" state distinct from genuinely-thin data (fail closed); transcripts are length-capped before the LLM with a per-run token/cost counter and a cheaper cached distill model; batch and discover runs emit a structured run-ledger (counts, costs, throttle events) and resume from a cursor after a crash; all env vars are validated up front so undefined never reaches an auth header.
Updated by your decisionsYour call: a throttled grade is <b>published with a "degraded" badge + an explicit incompleteness breakdown</b> (not held back). The engine must still detect a throttle and tell it apart from missing data; the monthly re-grade completes it later.
The approach
The spine of this workstream is one new shared HTTP client (src/lib/net/httpClient.ts) that all seven raw-fetch adapters route through: it adds bounded concurrency, token-bucket pacing, exponential backoff that honors Retry-After, and — critically — it throws a typed ThrottledError on 429/503 that is DISTINCT from a NotFoundError (404) and a generic HttpError. That typed distinction is what lets the pipeline fail closed: price.ts / mine.ts / gradeSet.ts currently catch-and-skip everything, so a throttled set silently reads "thin" and inflates coverage (audit: "a 429 and a 404 both vanish from coverage"). I change those three catch blocks to re-classify: 404/empty → skip (legit absence), Throttled/transient → record a degraded marker on the result so the grade is tagged throttled, never silently confident. Second pillar: a pure truncateForDistill() that caps+trims transcript chars before distill.ts builds the prompt, plus routing bulk distillation to a cheaper model with prompt caching and a CostMeter that accumulates input/output tokens per run. Third pillar: a RunLedger + resumable cursor for batch-ingest and discover so a crash re-pages from where it stopped, and a structured per-run health/cost summary replaces the truncated console.logs. Fourth: a single validateEnv() that replaces the three hand-rolled, inconsistent env checks in the scripts. Everything tunable (concurrency, base delay, max retries, transcript cap, model id, daily cost ceiling) goes into src/config/constants.ts. TDD throughout: the HTTP client, the error taxonomy, truncateForDistill, the ledger reducer, and validateEnv are all pure/seam-testable; the orchestrator fail-closed behavior is tested through the existing GradeDeps/MineDeps injection seams with fakes that throw ThrottledError.
Effort~2 weeks: 2 S (truncate, env), 8 M (http client, adapter rewiring, price fail-closed, mine fail-closed, gradeSet propagation, cost+cheaper-model, run ledger, cursor, bricklink concurrency). No L items — the work is many medium seams, not new infrastructure, because the injection points (GradeDeps/MineDeps/DistillOpts) and the constants file already exist.
Fixes11 · 2 small · 9 medium · 0 large
Depends onMostly independent (it adds a net-new src/lib/net/ layer and hardens boundaries). One soft coupling: the decision to SUPPRESS a confident score on throttle (vs only label it) belongs to the scoring/coverage workstream — this workstream guarantees the throttled state is detected, recorded, and surfaced; the scoring workstream decides whether coverageTier/letter is forced to Data-light/withheld on that signal. Also adds new entries to src/config/constants.ts, so coordinate constant additions with whoever else is editing that file to avoid merge churn.
The fixes, in order
medium
1. Shared HTTP client with error taxonomy (ThrottledError vs NotFoundError)
AddressesCost/scale audit 'Zero rate-limiting/backoff/Retry-After on any external API' (bricklink.ts:107, youtube.ts:185, supadata.ts:50); also the catch-and-skip indistinguishability root cause (price.ts:52, mine.ts:125, gradeSet.ts:122)
The changeCreate src/lib/net/httpClient.ts exporting (a) a typed error set — class ThrottledError (429/503, carries retryAfterMs), class NotFoundError (404), class HttpError (other non-ok) — and (b) async function httpFetch(url, init, opts?) that wraps global fetch, maps status→typed error, and retries ThrottledError/network errors with exponential backoff honoring the Retry-After header, up to HTTP_MAX_RETRIES. Concurrency/pacing live in a small token-bucket gate (acquire()/release()) created per host. This is a pure adapter over fetch — the existing 'thin client, live-verified' pattern. Do NOT yet rewire callers in this fix (keep it isolated and green).
Test firstsrc/lib/net/httpClient.test.ts: inject a fake fetch. (1) 404 → rejects with NotFoundError; (2) 429 with Retry-After: 1 → fake fetch called twice (retried once after the header delay, with a fake clock), succeeds; (3) 429 every time → after HTTP_MAX_RETRIES rejects with ThrottledError, retryAfterMs populated; (4) 200 → returns the response, fetch called once; (5) 500 → HttpError, not ThrottledError. Use vi.useFakeTimers() so no real waiting.
Watch outBackoff timing must be injectable (pass a sleep/now seam) or tests will be slow/flaky. Token bucket must not deadlock if a call throws before release — wrap in try/finally. Keep the public signature close to fetch so adapter rewiring is mechanical.
medium
2. Route all seven adapters through httpFetch
AddressesCost/scale audit rate-limiting finding across every provider (bricklink.ts:107, youtube.ts:185, supadata.ts:50, plus brickset.ts:81, rebrickable.ts:83, article.ts:40, brickInsights.ts:151 found in the same fetch-site sweep)
The changeReplace the raw await fetch(...) + if(!res.ok) throw new Error(...) in each adapter's thin-client function (bricksetCall, rbCall, ytGet, fetchPriceGuide, fetchTranscript, fetchArticleText, fetchBrickInsightsHtml) with await httpFetch(...). Preserve each adapter's existing semantics: supadata must still translate NotFoundError→null (its 404=no-transcript contract, supadata.ts:53), and rebrickable's fetchRebrickableSet keeps its try/catch→null. Pass the per-host concurrency gate from constants.
Test firstExtend the adapters' existing live-verified pattern with seam tests where feasible: supadata.test.ts — fetchTranscript returns null on NotFoundError but PROPAGATES ThrottledError (asserts the throttle is no longer swallowed as 'no transcript'). For adapters without a fetch seam today, add a minimal injected-fetch parameter or module-level fetch override consistent with how article.test.ts already exercises the client.
Watch outBrickset returns 200 with a JSON status:'fail' body (brickset.ts:96) — that is NOT an HTTP error and must stay a null/empty result, not a throttle. YouTube's 403 quotaExceeded should map to ThrottledError-like backoff, not a hard fail — decide its status mapping explicitly. Don't change supadata's 404→null contract or you break the 'skip video gracefully' path.
Depends onShared HTTP client with error taxonomy
medium
3. Fail closed on throttle in priceInventory
AddressesCost/scale audit 'Catch-and-skip makes throttling indistinguishable from missing data' (price.ts:52)
The changeIn priceInventory (price.ts:43-55) the per-part try/catch currently swallows everything. Re-classify the caught error: NotFoundError/null → skip the part (legit, unchanged); ThrottledError → stop fetching further parts this run AND surface it. Add a throttled: boolean (and degradedReason?) to PartValueResult (partValue.ts:9) so the caller knows coverage is artificially low because we backed off, not because parts are unpriceable. sumPartValue stays pure; priceInventory sets the flag.
Test firstprice.test.ts (new orchestrator test using a fake fetchPriceGuide injected via the existing creds/client seam, or by stubbing the bricklink module): given 3 parts where the 2nd throws ThrottledError, result.throttled===true and pricedQty reflects only parts fetched before the stop; given a 2nd part that throws NotFoundError, result.throttled===false and the run continues to part 3.
Watch outPartValueResult is consumed in gradeSet.ts:115-120 and printed in scripts — adding a field is additive but the coverage gate (>=0.8) must now ALSO refuse to trust partValueTotal when throttled even if coverage happens to look ok. Don't regress the existing 0.8 coverage behavior for the non-throttled path.
The changeIn mineSetReviews (mine.ts:121-128) the per-ref try/catch hides throttles. Re-classify in the catch: a per-ref NotFoundError/null → skip that ref (unchanged); a ThrottledError → increment a throttledRefs counter and stop mining further refs this set. Extend MineResult.stats (mine.ts:57) with throttled: boolean / throttledRefs: number so gradeSet can mark the grade degraded rather than 'no reviews found'.
Test firstmine.test.ts: add a fake distillRef path (inject MineDeps whose fetchTranscript/distill throws). Two selected refs, the 2nd throws ThrottledError → stats.throttled===true, stats.distilled===1, and mining stops (3rd ref not attempted). A ref throwing NotFoundError → stats.throttled===false and the next ref IS attempted.
Watch outdistillRef calls fetchTranscript, fetchComments (Promise.all in distillVideo), fetchArticleText, and the Anthropic client — any of them can throw. Make sure a comments-only ThrottledError vs a transcript NotFound are classified independently so we don't over-trip the throttle stop. MineDeps is the documented seam — use it, don't reach into modules.
Depends onShared HTTP client with error taxonomy
medium
5. Propagate degraded/throttled state onto the grade in gradeSet
AddressesCost/scale audit 'a throttled set looks thin rather than degraded' (gradeSet.ts:122); methodology §4.3/§6 fail-closed principle (insufficient-data must lower coverage/confidence, never present a confident letter)
The changegradeSet.ts already returns a sources block (gradeSet.ts:54-67). Add throttled: boolean and degradedSources: string[] to GradeResult.sources, fed by mined.stats.throttled and the priced result.throttled, plus the inventory try/catch (gradeSet.ts:110-124) which must now distinguish ThrottledError from a real fetch miss. When throttled, the result is explicitly marked so the caller/UI shows 'grade withheld/degraded — data was throttled, retry' instead of a confident data-light F. Do not change the scoring math here; this is a truthful-labeling fix at the orchestrator boundary.
Test firstThrough GradeDeps seam: a fake priceParts that returns {throttled:true} and a fake mining whose mineSetReviews reports throttled → GradeResult.sources.throttled===true and 'bricklink'/'youtube' in degradedSources. Control case (no throttle) → throttled===false and degradedSources empty. Assert the existing happy-path gradeSet tests still pass unchanged.
Watch outGradeResult.sources is consumed by both scripts' console output — extend the print to surface the degraded flag. The deeper question of whether the SCORE/coverageTier itself should change on throttle is a methodology call (see decisions); this fix only guarantees the truth is RECORDED and surfaced.
Depends onFail closed on throttle in priceInventory; Fail closed on throttle in mineSetReviews
small
6. Cap and trim transcripts before distillation
AddressesCost/scale audit 'YouTube transcripts sent to the model with no length cap — largest uncontrolled per-set cost' (distill.ts:130-135, mine.ts:75-80); sourcing spec §11 'cap/trim long transcripts for LLM'
The changeAdd a pure exported truncateForDistill(text, maxChars) to distill.ts (or a new distill helpers file) that trims a transcript to DISTILL_TRANSCRIPT_MAX_CHARS, cutting on a sentence/word boundary and keeping head+tail (reviewers state verdicts at both ends) rather than a blind head slice. Call it in distillReview (distill.ts:123) on params.transcript before building userContent, and apply the same cap to the joined article text path in mine.ts. No raw text is stored — this is purely the transient processing window, honoring the derived-data rule.
Test firstdistill.test.ts (pure): truncateForDistill on a 200k-char string returns <= DISTILL_TRANSCRIPT_MAX_CHARS, ends on a boundary (no mid-word cut), and includes both the first and last sentence markers; a short string passes through unchanged. Plus a distillReview test with an injected fake Anthropic client asserting the userContent it receives is within the cap.
Watch outOver-aggressive trimming can drop the verdict and skew scores — head+tail strategy mitigates, but the cap value is a calibration decision (put it in constants, default generously, e.g. ~24k chars ≈ a long review). Keep coversSet reliability: ensure the set number/name still appears in the kept window.
medium
7. Route bulk distillation to a cheaper cached model with a cost meter
AddressesLLM-judge audit 'Opus low effort, no prompt caching; costly' (distill.ts:135); cost/scale audit transcript cost (distill.ts:135); sourcing spec §5.6 (LLM distillation tokens are the second-largest budgeted cost)
The changeIn distill.ts: (1) make the model id come from constants (DISTILL_MODEL_BULK) instead of the hardcoded 'claude-opus-4-8' default (distill.ts:135), so the 24k sweep runs a cheaper model while single-set grading can override via DistillOpts.model; (2) add prompt caching on the large static system block (DISTILL_SYSTEM) via cache_control so the ~1.5k-token instructions aren't re-billed per set; (3) return token usage from distillReview (or accept a CostMeter callback) and accumulate input/output tokens into a per-run CostMeter (new src/lib/net/costMeter.ts — pure accumulator). Wire a DAILY_DISTILL_COST_CEILING check so a runaway sweep self-halts.
Test firstcostMeter.test.ts (pure): add({input,output}) accumulates; estimatedCostUSD uses per-1M rates from constants; overCeiling()===true past DAILY_DISTILL_COST_CEILING. distill.test.ts: distillReview with a fake client whose response carries usage → the meter records those tokens, and the request it built sets the bulk model id + cache_control on the system block. Single-set override (opts.model) still wins.
Watch outModel id, effort level, and the exact caching field are Anthropic-SDK-version specific — verify the SDK shape via the claude-api skill before asserting field names (the code already uses output_config.effort 'low' and a json_schema). Don't let prompt caching change the JSON-schema output contract. The cheaper-model choice has a quality tradeoff (see decisions).
Depends onCap and trim transcripts before distillation
AddressesCost/scale audit 'Missing env vars inject undefined into auth and fail only as swallowed catch-and-skip' (batch-ingest.ts:74-90, the BRICKLINK_*! non-null assertions)
The changeCreate src/config/env.ts exporting validateEnv(required: string[], env?): Record<string,string> that reads process.env (or an injected map), collects ALL missing names, and throws one clear aggregated error listing every missing var (never returns undefined). Replace the three hand-rolled, divergent checks in batch-ingest.ts:38-45, discover.ts:25-35, and grade-set.ts:25-71 — including the dangerous process.env.BRICKLINK_CONSUMER_KEY! non-null assertions (batch-ingest.ts:74-77) — with validateEnv calls gated by the flags actually in use (--price needs BRICKLINK_*, --mine needs YOUTUBE/TRANSCRIPT/ANTHROPIC).
Test firstenv.test.ts (pure): with a stubbed env object, validateEnv(['A','B'],env) where both set → returns {A,B}; where both missing → throws once, message contains BOTH 'A' and 'B'; where one missing → message names exactly the missing one.
Watch outScripts use top-level await + process.exit today; keep that UX (catch the aggregated error, print it, exit 1) so behavior for an operator is unchanged except clearer. Make validateEnv take the env map as an optional arg to stay pure/testable.
medium
9. Run ledger + structured per-run health/cost summary
AddressesCost/scale audit 'No observability anywhere: bare catches and truncated logs, no metrics/run ledger/alerting' (batch-ingest.ts:105-107)
The changeAdd src/lib/net/runLedger.ts: a pure reducer that accumulates per-set outcomes (graded / skipped / FAILED / throttled), provider call counts, throttle events, and CostMeter totals into a RunSummary, plus a formatSummary() for a structured end-of-run block. In batch-ingest.ts (the loop at 96-108) and discover.ts, record each set's outcome into the ledger instead of (only) the truncated e.message.slice(0,50) line, and print the RunSummary at the end (graded N, throttled M, failed K, est cost $X, provider calls). Optionally persist the summary as one JSON line to a run_log file/table for auditability.
Test firstrunLedger.test.ts (pure): feeding a sequence of outcomes yields correct tallies; a throttled outcome increments throttled AND is reflected in the summary as a distinct bucket from failed; formatSummary includes cost + throttle counts. The orchestrator wiring is covered by a thin batch test using injected deps that returns a throttled grade and asserting the summary counts it as degraded, not success.
Watch outKeep the ledger pure and the file/DB write a thin seam so tests don't touch disk/Supabase. Don't double-count when a set is retried. The 'alerting' part of the audit is out of scope for v1 — a structured summary + machine-readable run_log line is the honest first step (note in decisions if Seth wants real alerting).
Depends onCost meter (from the distillation fix) for the cost totals; Propagate degraded state in gradeSet for the throttled tally
medium
10. Resumable cursor for batch-ingest and discover
AddressesCost/scale audit 'No resumability or idempotent cursor: sweeps re-page everything and a batch crash restarts from zero' (discover.ts:38-69, batch-ingest.ts:96-108)
The changeAdd a small cursor abstraction: for batch-ingest, persist the last successfully-ingested set number (to a checkpoint file or a run_state row) after each iteration and skip already-done sets on restart; for discover.ts's allowlist sweep, checkpoint per-channel progress (which channelIds completed) so a crash resumes mid-allowlist instead of re-paging from the first channel. Make the cursor store injectable (a tiny interface with load()/save()) so it can be an in-memory fake in tests and a file/Supabase row in prod.
Test firstA pure cursor-reducer/store test (src/lib/net/cursor.test.ts): given a checkpoint of {done:[A,B]}, the resumable iterator over [A,B,C,D] yields only [C,D]; saving after C advances the checkpoint; an empty checkpoint yields all. For discover, a channel-cursor test: completed channels are skipped on resume.
Watch outIdempotency: upserts (upsertSet/upsertDropScore/saveReviewRefs) are already idempotent, so re-running a partially-done set is safe — the cursor is an efficiency+auditability win, not a correctness gate. Write the checkpoint AFTER the durable upsert so a crash mid-set re-does that one set (safe) rather than skipping it. Validate the checkpoint on load so a corrupt file can't silently skip the whole catalog.
Depends onRun ledger (shares the run-state persistence seam)
medium
11. Bounded-concurrency BrickLink part pricing
AddressesCost/scale audit 'BrickLink pricing is strictly sequential with no concurrency, making the 24k-set tail impractical' (price.ts:43-55)
The changeReplace the strictly-sequential for-loop in priceInventory (price.ts:43-55) with a bounded-concurrency map over the missing parts using the HTTP client's concurrency gate, capped at BRICKLINK_CONCURRENCY (small, e.g. 2-4, to respect the ~5k/day BrickLink limit per sourcing spec §6). The maxFetches cap and the global part-price cache stay; only the in-flight parallelism changes. The ThrottledError fail-closed behavior from the earlier fix must compose with concurrency (first throttle cancels the remaining in-flight batch).
Test firstprice.test.ts: with an injected fetchPriceGuide that records concurrent in-flight count, assert it never exceeds BRICKLINK_CONCURRENCY; assert total fetches still respect maxFetches; assert a ThrottledError mid-batch sets result.throttled and stops scheduling new fetches.
Watch outConcurrency + a global daily BrickLink budget is easy to over-shoot — keep the concurrency tiny and the token-bucket pacing authoritative. Ordering of fresh writes is no longer deterministic; savePartPrices is a set upsert so order is irrelevant, but confirm no test asserts call order. Do this LAST so the fail-closed semantics are already in place.
Depends onFail closed on throttle in priceInventory
Calls to make in this workstream
On a confirmed throttle (429 storm) for a set, should the SCORE/coverage tier itself be suppressed, or do we only LABEL the grade as degraded and still show the number?
Options: (a) Engineering-only: record sources.throttled and surface a 'degraded — retry' badge, but compute and store the number as today. (b) Methodology change: when a primary source (mining or pricing) was throttled, force the set into Data-light/withheld and do NOT publish a confident letter until a clean re-run (true fail-closed per methodology §4.3/§6, principle 4).
Recommended: (b) for the publish boundary, (a) for storage. The spec's load-bearing principle is fail-closed: a throttled set must not wear a confident letter. Implement the label first (cheap, this workstream), then gate publishing on it. Full re-scoring suppression touches the scoring workstream, so coordinate.
Which cheaper model do we route bulk distillation to, and what quality bar must it clear before the 24k sweep trusts it?
Options: (a) Keep claude-opus-4-8 everywhere (highest quality, highest cost — current state, blocks scale on budget). (b) Route bulk to a cheaper Claude tier for the sweep, keep opus for single-set/flagship grading via the existing DistillOpts.model override, and validate the cheaper model on a labeled sample (sarcasm cases especially) before trusting it. (c) Two-pass: cheap model first, escalate to opus only when self-confidence is low.
Recommended: (b) with a validation gate. The sweep cannot run on opus at 24k sets × multiple reviews within the §5.6 budget; the DistillOpts.model seam already exists for the override. Gate it on a sarcasm-aware eval sample — the distiller's whole value is sarcasm detection. Verify the exact model id/pricing against the claude-api skill before committing the constant.
Where does run-state (cursor + run-ledger) live — local checkpoint file, or a Supabase run_state/run_log table?
Options: (a) Local JSON checkpoint file (zero infra, but not durable across machines and invisible to any dashboard). (b) Supabase table(s) (durable, auditable, queryable, but adds a migration and write path). (c) Both behind the injectable store interface.
Recommended: (c): ship the injectable store interface now with a file-backed default (keeps tests/dev hermetic), and add a Supabase-backed implementation when the sweep moves off a single machine. The interface is the decision; the backend is a swap.
What initial values for the new operational tunables (concurrency, base backoff, max retries, transcript cap, daily cost ceiling)?
Options: Pick conservative first-pass numbers now vs. block on calibration. The spec gives anchors: Brick Insights ≤60 req/min (§4), BrickLink ~5k/day (§6).
Recommended: Ship conservative defaults in constants.ts and calibrate later (the file's own header says 'numbers are calibration defaults; the structure is the decision'). Suggested seeds: BRICKLINK_CONCURRENCY 3, HTTP_MAX_RETRIES 4, HTTP_BACKOFF_BASE_MS 500, DISTILL_TRANSCRIPT_MAX_CHARS ~24000, generous DAILY_DISTILL_COST_CEILING. Tune against a real partial sweep.
SequencingFront-load the two foundations that unblock the most: (1) the shared HTTP client + error taxonomy, and in parallel (2) the two zero-dependency quick wins — truncateForDistill (S) and validateEnv (S) — since they need nothing and visibly cut cost/risk immediately. Then route adapters through httpFetch. Then the fail-closed trio (priceInventory → mineSetReviews → gradeSet propagation), which is the correctness heart of the workstream and depends only on the typed errors. Then the cost meter + cheaper cached distill model (needs truncation in place). Then observability: run ledger (needs the cost meter + the degraded flags to have something to tally) and the resumable cursor (shares the run-state seam). Do bounded-concurrency BrickLink pricing LAST so the fail-closed throttle semantics already exist for the parallel batch to compose with. Every step starts red: write the failing test, then the code. Re-run the full 197-test suite after each fix to confirm no engine regression.
Start here
The first pull request
If you greenlight the direction, the smallest high-value starting point — provable in a day, and the keystone the whole plan rests on:
PR #1 — Turn on the LOW_DATA_FLOOR gate (+ stop the comment that lies)
Write the failing test first: a dimension backed by 4 reviews must come back insufficient, not scored (it passes as "scored" today).
Change the one gate in dimension.ts to use LOW_DATA_FLOOR (5), and rewrite the doc-comment that currently claims it already does.
Re-run the full scoring suite and deliberately re-baseline any fixtures that shift tier — confirming each change is a real honesty correction, not a regression to paper over.
It's a one-line behavior change that also fixes coverage over-reporting for free, and it's the keystone of Wave 1. From there: the other two gates, then the confidence fixes.
Appendix
Every new knob, in one place
All tunables this plan adds to src/config/constants.ts — kept here as the single dashboard, exactly as the house rule requires. 46 in total.
ARTICLE_MIN_ALPHA_WORD_RATIO
Minimum share of real words in extracted article text before it's trusted (tightens article.ts:43's length-only floor — fail closed on nav/boilerplate).
BI_MIN_REVIEW_COUNT_FOR_SANITY
Minimum Brick Insights reviewCount (spec's 2–3 trust floor) before its PPP aggregate is trusted as a sanity-check; below it, no flag (fail closed on thin BI data).
BI_PPP_SANITY_DIVERGENCE
Relative-divergence threshold between our computed value ratio and the Brick Insights PPP aggregate above which a non-score-affecting priceSanityFlag is raised.
BRICK_INSIGHTS_BASE_CONFIDENCE
Replaces the hardcoded MEDIUM literal in brickInsightsRefs (brickInsights.ts:141) so the BI corroboration tier is configurable and can be promoted to HIGH when our matcher independently agrees.
BRICKLINK_CONCURRENCY
BrickLink-specific in-flight cap for part pricing (tiny, to respect the ~5k/day limit, sourcing §6).
CLUSTER_TIER
HIGH/MED tier per cluster (from §3's rubric table) to drive within-tier-first redistribution in capWeights (§4.3). Belongs in rubric.ts/constants.ts as structural config.
COHORT_MAX_STALENESS_DAYS
Max age of a cohort row's cached part value (relative to the live-priced graded set) before it is dropped from the cohort, closing the numerator/denominator recency drift.
COLD_START_SEARCH_MAX_PER_DAY
Daily cap on search.list cold-start fallback calls (spec §5.1 'small daily cap'); currently the search fallback in mine.ts:103-117 has no global cap. (Tunable; wire enforcement in the batch path.)
COMMENT_LIKE_DISPLAY_CAP
Cap on the like-count surfaced to the LLM so one brigaded comment can't read as overwhelming consensus (spec §5.2 anti-brigade).
CONFIDENCE_MIN_SOURCES_FOR_FULL_AGREEMENT
Distinct-source count required before the uncapped 1/(1+variance) agreement applies. Default 2.
CONFIDENCE_VOLUME_SATURATION
Replaces the magic log1p(60) denominator in signalConfidence; the review count at which volumeFactor saturates to ~1 (§6.1 volume input). Default 60.
CROSS_THEME_COHORT_MIN_HOMOGENEITY
Optional gate on the theme→size-only fallback: minimum cohort homogeneity required before pooling all themes in a size bucket, preventing display-line vs licensed value-regime mixing. Default permissive.
DAILY_DISTILL_COST_CEILING_USD
Self-halt threshold for a runaway sweep; CostMeter.overCeiling() trips the run ledger.
DISTILL_CONF_DEFAULT
Default OpinionSignal.confidence (recommend 1.0) for signals lacking an LLM self-confidence, so existing behavior is unchanged.
DISTILL_EFFORT
Effort level applied ONLY when the model is Opus-tier (Haiku 4.5 rejects output_config.effort with a 400); omitted for Haiku.
DISTILL_HEAD_MODEL
Optional higher-accuracy model ('claude-opus-4-8') for popular/high-value head sets under the tiered strategy.
DISTILL_MODEL
Bulk-distillation model id (default 'claude-haiku-4-5'); single source of truth replacing the hardcoded 'claude-opus-4-8' at distill.ts:135.
DISTILL_MODEL_BULK
Cheaper model id used for the bulk 24k sweep (single-set grading overrides via DistillOpts.model).
Per-1M-token rates the CostMeter uses to estimate run cost (verify against the claude-api skill).
DISTILL_RETRY_ATTEMPTS
Bounded retry count for distill/fetch transient (429/5xx/timeout) failures.
DISTILL_RETRY_BASE_MS
Exponential-backoff base delay between retries.
DISTILL_TIMEOUT_MS
AbortController timeout per distill/fetch call.
DISTILL_TRANSCRIPT_MAX_CHARS
Hard cap on transcript chars sent to the distiller (cap-and-trim; sourcing §11).
DISTILLATION_MIN_ACCURACY
Per-dimension + sarcasm-flag accuracy threshold the offline harness must clear before DISTILLATION_VALIDATED may be flipped true.
DISTILLATION_VALIDATED
Master fail-closed flag (default FALSE): when false, distilled scores are withheld from the computed grade per spec §8.
FACT_CONFIDENCE_FALLBACK_FLOOR
Lower floor for a fact derived from a fallback basis (weight/pieces instead of part-value) or an empty/thin cohort, per §6.1 'disclosed fallback → down'. Recommend 0.4.
FACT_CONFIDENCE_FLOOR
Confidence floor for a primary-basis fact-defined dimension (preserves today's behavior). Default 0.8; extracts the hardcoded 0.8 at dimension.ts:39.
FULL_COVERAGE_MIN
Move the currently-local const in score.ts:33 (0.5) into constants.ts so the coverage-tier threshold is a single-source tunable like every other gate (house rule: all tunables in one file).
FWI_PRICE_PAIR_TOLERANCE
Only added if FWI price-pairing is implemented: max relative gap between an opinion's priceAtTime and the anchor street price for the opinion to count (unpaired/out-of-tolerance discarded).
HTTP_BACKOFF_BASE_MS
Base delay for exponential backoff (doubled per attempt; overridden by Retry-After when present).
HTTP_HOST_CONCURRENCY
Default per-host max in-flight requests for the shared client's token-bucket gate.
HTTP_MAX_RETRIES
Max backoff retries for a throttled/transient external call before surfacing ThrottledError.
LIFECYCLE_PREAVAIL_DISCOUNT
Only if the lifecycle gate is implemented: weight multiplier (0 = suppress, <1 = discount) applied to pre-availability signals on value/build dimensions per §6.1.
MATCH_CONFIDENCE_RANK
Single source of truth for the high/medium/low ordinal ranks, replacing the three separate hardcoded {high:3,medium:2,low:1} maps in matchVideoToSets (youtube.ts:141), selectRefsToMine (mine.ts:32), and any sort in discover.ts.
MAX_LOW_REFS_PER_SET
Cap on uncorroborated/LOW refs mined per set so a single haul/roundup cannot inject dozens of spurious pairs into the judge queue.
MAX_TRANSCRIPT_CHARS
Cap transcript length sent to the LLM — the largest uncontrolled per-set cost driver (spec §11 cap-and-trim).
MINE_MIN_CONFIDENCE
Lowest discovery-confidence tier the miner will auto-accept; enforces the spec §5.1 two-agreeing-signals rule at the mining seam so lone uncorroborated refs are withheld (fail closed).
NAME_MATCH_MIN_TOKENS
Floor on matched distinctive tokens (e.g. 2) so short names still require full agreement even under the N-of-M rule.
NAME_MATCH_TOKEN_FRACTION
Fraction of a set name's distinctive tokens that must appear in a title for a name match (N-of-M), e.g. 0.7 — restores recall on long licensed names without requiring every token.
PART_PRICE_TTL_MONTHS
Age beyond which a cached part_prices row is treated as missing by loadPartPrices, forcing a capped BrickLink re-fetch in priceInventory. Default 3 — generous to respect the ~5k/day BrickLink cap (sourcing §6).
PARTS_FOR_BUILDING_VARIETY_ANCHORS
Distinct-element variety ratio → 0–10 curve for the new partsForBuildingScore (seeded from the retired WHAT_YOU_GET_VARIETY_ANCHORS values). Makes Parts-for-Custom-Building the sole home of generic variety per §3.1 rule 3.
SCORE_TTL_DAYS
Age beyond which a stored drop_scores row is considered stale, letting a batch re-grade recompute it. Default 30.
SIGNAL_TTL_MONTHS
Age beyond which cached dimension_signals are treated as expired, so loadDimensionSignals returns {} and gradeSet re-mines automatically (replaces the permanent cache hit at gradeSet.ts:89-93). Default 12, aligning with the §10 value-sentiment HALFLIFE.
SINGLE_SOURCE_AGREEMENT_CAP
Max agreement a single-source signal set may claim, so a lone source's variance-0 'self-agreement' can't read 1.0 (the lone-Amazon-1.000 bug). Recommend 0.6.
SMART_PRICE_CONFIDENCE_BY_COHORT
Piecewise anchors mapping cohort size n → 0–1 confidence multiplier, so a Smart Price backed by few comparables is presented at lower confidence than one backed by many.
THUMBS_TO_SCORE
Maps community worthTheMoney thumbs (up/down) to a 0–10 feelsWorthIt score, the thumbs analogue of Brickset's star→×2 scaling (§5.1).
Comprehensive remediation plan for The Drop Score engine, generated 2026-06-18 from a code-grounded 8-agent design pass over the 9-agent audit. Companion docs: docs/2026-06-18-engine-audit-plain.html (the audit in plain English), docs/2026-06-18-engine-audit.html (the technical audit), docs/grading-walkthrough.html (the live grading procedure).