A clarity briefing on whether the calculator drifted into over-engineering — and what the best version actually looks like. Ten independent reviewers. No flattery.
Your instinct is half-right, and the half you've got wrong is the good news. It does feel too math-heavy — but the math itself is mostly justified (almost every piece traces to a real fix you signed off on). The thing that actually went sideways is simpler: you asked one little letter grade to carry a job that's really three separate questions, and collapsing them into a single number is what forced all the gnarly machinery into existence.
The feature you think you lost — "weight it by what I care about" — was never lost. It's fully built, tested, and database-ready in your engine right now. It's just switched off in the live grader. Turning it on is a small wiring change, not a rebuild.
And the fix for "too math-heavy" is not to gut the engine (that would turn you back into the price-per-piece tool you exist to replace). It's to change what you show: lead with a letter and one human sentence, keep the math backstage, and let people see a couple of honest reads instead of one over-cooked number.
You asked six things. Here they are, each with a blunt verdict first and the reasoning under it. No softening.
Yes & noToo heavy on the surface; mostly justified underneath.
By the numbers it's genuinely heavy: 61 hand-set constants, ~22 transformations from a raw input to a letter, six weighted categories, archetype re-balancing, a four-layer price engine. But when the auditors traced each piece, almost all of it maps to a real failure it prevents (judging a flower set on its missing minifigs; a sarcastic "what a steal 🙄" review; a set with no reviews faking a confident grade). The math is the right kind of tool. What's not well-suited is the final output shape — see Q3.
Mostly — one wrong oneRight philosophy. One wrong question is poisoning the rest.
The right questions you're asking — "is it a good deal AND is it actually good?", "what do real people think?", "how sure are we?" — are excellent and rare. The one wrong question is "what's the single number for this set?" That question forces a hidden trade ("how many dollars of overpricing equal how much building fun?") that no two buyers answer the same way. Almost all the machinery you're reacting to exists only to defend that one forced number. (Full reasoning in §5.)
NoYour logic is sound. Your output format is the weak link.
"Value AND joy, from real humans, honest about certainty, never price-per-brick" — that's not flawed, it's your moat. The market check (§3) confirms no competitor even attempts to fuse value and joy; they all keep them in separate boxes. Your thinking is ahead of the field. The flaw isn't in the logic — it's that you're cramming sound logic into a one-letter container that can't hold it.
The idea, yes. The calculation, no.And you should never try to explain the calculation on air.
You can pitch the idea in 10 seconds, beautifully. You cannot honestly explain how a specific grade is computed without reaching for arithmetic — and you shouldn't. The fix is to decide, on purpose, that the calculation is the "under the hood" part you never narrate. (Your actual script is in §6.)
Not lostIt's built, tested, and DB-ready. It's just turned off.
Three auditors found it independently in the code. Five "lenses" (Display, Builder, Parent, Investor, + Universal) are fully implemented, proven by passing tests, with database columns waiting for them. The live grader simply runs the "Universal" weighting only and never calls the others. This is the biggest relief in the whole audit — §2.
YesShow a few honest reads, not one over-cooked number. Keep the engine; change the face.
The better way isn't more math or a teardown — it's three legible reads (Worth it? · Will you love it? · Will it hold value?) with one simple Value ↔ Joy slider, riding on the engine you already have. That single move answers Q2, Q3, Q4, Q5 and revives the feature from Q5 — all at once. The full plan is §7.
You said: “At one point I feel like it was talked about having a filtering that, based on what you care about, weights the score differently — a more accurate grade.” You remembered correctly. It's real, and it's already in your engine.
The system is called "lenses." There are five: universal, display, builder, parent, investor. Each one re-weights the six categories to a different priority — e.g. the Builder lens cranks up parts & build value and dials down "finished look"; the Investor lens brings the "hold its value" badge into the grade at a heavy 30%.
It is fully written, covered by passing tests that prove the score actually changes per lens, and the database is already keyed to store one grade per lens. The only missing piece: the live grader (gradeSet.ts) calls the scorer without asking for a lens, so it defaults to Universal every time. The other four are computed nowhere but the demo.
So the honest status is "built and stranded," not "lost." Your old notes even say why: "ship the universal grade first; add lenses later." You parked it on purpose — and then the accuracy and sourcing work pulled focus, and it stayed parked.
You called it a "more accurate" grade. One auditor pushed back hard on that word, and they're right: re-weighting doesn't make a grade more accurate — it makes it more relevant to you. The Universal grade isn't wrong for a builder; it's just answering a slightly different question than the builder is asking. That distinction matters, because it tells you how to ship this: not as "the real grade, finally," but as "the same honest data, tuned to what you care about." Sell it as relevance, not accuracy, or you'll over-promise.
The agents worked blind to each other. When independent reviewers converge, that's the strongest signal in the whole exercise. Five things came up again and again.
The math is firewalled inside the code; a user never has to see it. The reason it feels out of control is that the current output (a CLI dump full of insider labels like data-light, split verdict, coverage vs. consensus) puts the engine block on the dashboard. The market auditor put it cleanly: "The fusion is the product; the math should be backstage." Every successful competitor shows one or two numbers and hides everything else.
Several agents went further: the reason there's so much machinery is that one blended letter has to be defended against every edge case. Split the output into a few honest reads and whole categories of complexity simply evaporate (the re-balancing, the "no value anchor → withhold" tangle, half the confidence damping). You're not under-built; you're over-reconciling.
A grade built from two reviews that flat-out disagreed (one said 3/10, one said 9/10) still prints a confident "6.0 / C" — with the uncertainty shoved into a side badge most people ignore. The one-decimal score claims a precision the data often can't back. (More in §6's sibling point and §7.)
Both real-world audits graded outputs C− for the same reason: the tool is great on well-reviewed sets and shaky on everything else. That's a coverage gap (no reviews on the long tail) plus a value-anchoring gap (licensed sets, unpriced sets) — and you cannot compute your way out of missing data. The proposed next step (a 15-constant hedonic price regression) is, in one auditor's words, "complexity metastasizing to patch the failures of complexity."
Only about three lines are genuinely inert (a reserved constant, a flag nothing populates, an always-empty set). The "out of control" feeling is cognitive load from paused/visible-but-dormant features, not code bloat. Good news: that's cheap to quiet down.
You asked me not to just agree with you. So I assigned one agent to defend the complexity as hard as possible, and another to argue for tearing it down. Here's the real disagreement, not a strawman.
They agree on more than they disagree. Both say: the math goes backstage, the surface gets simple, and you stop adding more math. The only real split is the headline — one blended letter (Keep) vs. a few separate reads (Cut).
And that split dissolves once you notice the resolution is the same move as your lens feature. Keep the rich engine (the defenders win on "don't gut it"). But surface it as a few honest reads with a Value↔Joy slider (the simplifiers win on "stop forcing one number"). The slider is the personalization you remember. Everyone's right; nobody has to lose.
If you take one idea from this whole dispatch, take this: one number is doing a job that three numbers should do — and that single decision is the source of nearly everything that feels "out of control."
To roll "value" and "joy" into one letter, the engine has to secretly decide an exchange rate: how many dollars of overpricing cancel out how much building fun? Your weights (22 / 19 / 19 / 16 / 12 / 12) are that decision — made once, for everyone, invisibly. But a bargain-hunter and a Star Wars superfan are asking opposite questions. One letter lies a little to both.
You already pulled "Hold Its Value" out of the main grade into its own side badge — because resale corrupts the buy-or-skip signal. That was exactly the right instinct. It's also proof the single number doesn't hold: if resale deserves to stand on its own, so does value, and so does joy. Your own value-calc v2 notes have even started splitting "VALUE vs JOY" inside the engine. Your design is already cracking the monolith — you just haven't done it on purpose yet.
Here's what the buyer is actually asking when they look at a set — almost never "what's the one true score":
| The buyer's real question | Who asks it | What answers it |
|---|---|---|
| Is it worth the money? | Almost everyone — often the only question | Smart Price (part-out value vs. street price) |
| Will I love building / displaying it? | The hobbyist, the self-buyer | Mined reviews → build & display sentiment |
| Will it hold its value? | The investor, the saver | Resale / scarcity (already a clean side read) |
| What's the single blended 0–10? | Nobody. | An exchange rate the engine invents and then spends huge effort defending |
The better questions, in plain buyer language: "Worth the money?" · "Will you love it?" · "Will it hold value?" Show those three as their own reads. Lead with one as the headline (Worth It is the most universal and the most defensible). Then let the user nudge a single Value ↔ Joy slider to weight it to themselves. That's the personalization, the simplification, and the honesty — in one stroke.
Hides that it's a great build that's overpriced. The badge that would tell you so gets ignored. To produce this one cell, the engine re-balances, caps, damps, and withholds.
Verdict: you can nail the pitch today; you cannot explain the calculation — and that's fine, because you never should. Here are the words.
“It's a report card for LEGO sets — one grade, A+ to F, for whether a set is both a good deal and genuinely good, built from real reviews instead of just counting pieces. And it tells you how confident it is.”
That's true, complete, and contains zero math. The longer 90-second version the auditor wrote holds up just as well — restaurant-critic analogy, "judging a meal by price-per-calorie," real reviewers, "when we don't have much we say so." It works because it describes what the score means, never how it's computed.
Where a plain explanation genuinely breaks — the things you should never try to narrate live:
Almost everything above is fine to leave backstage. The one exception is Smart Price — it's the only mechanism where even the intent needs two sentences ("an absolute part-value ratio blended with a percentile rank weighted by cohort size"). Your most important, most-defended feature is also your least explainable. If you can't say your value story in one breath, that's the thing worth genuinely simplifying — not the explanation of it, the design of it.
Ordered by leverage. The first three are mostly presentation work on an engine that already exists — high impact, low risk. None of this is a rebuild.
gradeSet.ts to run them and the DB to store them. Ship it as a simple slider/toggle, not a five-persona menu. Sell it as "tuned to what you care about," not "more accurate."data-light, split verdict, coverage vs. consensus on the surface. Replace with one plain line: "Based on 12 reviews" / "Early read — only 2 reviews so far." The six categories live one tap down, not on the front door.The risk now is over-correcting. You asked for honesty, so: do not let "too math-heavy" become "rip the guts out." These are the parts that earn their keep — protect them.
| Keep this | Why it's load-bearing |
|---|---|
| Keep Fusing value and joy | Your entire moat. Confirmed: no competitor does this — they all keep value and quality in separate boxes. This is the gap you fill. |
| Keep Part-out value (not price-per-piece) | The founding promise, in code. Even 2025-era AI tools still just divide price by pieces. Yours is more honest about what "value" means. |
| Keep Real-reviewer sentiment (the AI judge) | Subjective things (fun, looks) can only be scored honestly by reading what real people said. Sarcasm-aware. This is "real human opinion at scale" working. |
| Keep Honesty about certainty | The instinct to withhold rather than fake a grade is right and rare. Just say it in plain language and let it widen the number, not hide in a badge. |
| Keep Archetypes (the one complex thing worth it) | Not docking a botanical set for "no minifigs" is perceivable and correct. Keep the classification — it's the one piece of heavy logic users would agree with. |
| Keep The A+→F letter as the headline | No LEGO tool uses letter grades; it's the most instantly legible verdict a casual buyer can read. Make it one of your reads, not a blend of all of them. |
| Watch The "Feels Worth It" ±0.75 nudge | Even the defender couldn't fully justify its weight. Worth a "does this actually change any grade enough to keep it?" test. |
| Rethink The single blended 0–10 | The one genuinely wrong-shaped output. Demote it: show it crisp only when data is strong; otherwise a band. See §5 & §7. |
Every recommendation here is reasoning from your code, your docs, and the market — not from real users. Multiple auditors flagged the same gap: there is no app yet and no user research. The single-number-vs-three-reads question has never actually been put in front of a LEGO buyer.
So treat "three reads + a slider" as the strongest hypothesis, not gospel. The cheapest possible next move — before another line of scoring math — is to mock both versions (the one-letter card and the three-read card) and show them to ten real buyers. Your own instinct to pause the value-calc v2 build over a validation worry was exactly the right reflex. Apply it here too.
Bottom line: you didn't build the wrong thing — you built a genuinely great engine and then asked it to speak through a too-small mouth. Don't shrink the engine. Widen the mouth: a few honest reads, one slider, plain words, and the confidence to leave the math under the hood where it belongs.