No jargon. Just the real story of how a LEGO set gets its grade, told with everyday comparisons — and explained slowly enough that none of it feels like magic.
The Drop Score is a tool that gives any LEGO set a single, honest grade — like a report card, but for a toy. The grade is a letter (A+ through F) with a more precise number underneath it (say, 8.2 out of 10), the way a movie might get "Certified Fresh" plus a 91%.
A really good restaurant critic. A critic doesn't judge a meal by one thing — they weigh the food, the value, the room, the service, and they tell you how the place compares to others in its class. The Drop Score is that critic for LEGO sets, except instead of one person's opinion, it pools thousands of real opinions and a pile of hard facts, then writes the verdict.
The point isn't just to spit out a number. It's to do it fairly, at scale (every set, not just the famous ones), and honestly (it tells you when it's guessing).
For years, LEGO fans have judged whether a set is "worth it" with one crude shortcut: price per piece. Divide the price by the number of pieces; if it's under about ten cents a piece, people call it a good deal.
It's like judging a restaurant by price per calorie. By that logic a bucket of fast-food fries beats a beautiful steak dinner — more calories per dollar! But nobody actually thinks the fries are the better meal.
Price-per-piece does the same thing to LEGO. A set crammed with hundreds of tiny identical 1×1 tiles looks like a "bargain," while a set with big, complex, genuinely valuable pieces (a giant curved windscreen, rare printed parts, sought-after minifigures) looks "expensive" — even though the second set is obviously the better buy. The shortcut counts pieces; it has no idea what those pieces are worth or whether the set is any fun.
So The Drop Score throws that shortcut out and asks the real questions: Is it a fair deal? Is it fun to build? Does it look great when it's done? Are the minifigures good? — and a few more. Then it combines the answers.
A grade is only as good as what it's built from. We pull from three different kinds of information, because each one knows things the others don't.
| Source | What it is | The everyday version |
|---|---|---|
| ① Hard facts | Numbers you can look up: price, piece count, weight, how much the individual parts are worth, how many minifigures, whether the parts are brand-new molds. | The lab results. |
| ② What reviewers said | Real opinions already published online — mostly YouTube reviews and their comment sections — read and scored by an AI. | The expert witnesses. |
| ③ What fans tell us | Grades that visitors to the site leave themselves (no account needed, but they have to say why). | The crowd / the jury. |
A doctor making a diagnosis. They don't rely on one thing. They look at your bloodwork (the hard facts), they listen to your description of the symptoms and maybe a specialist's second opinion (the reviews), and they factor in how you actually feel day-to-day (the crowd). Any one alone can mislead; together they're trustworthy.
Say five people review the same set. They won't agree, and they shouldn't all count equally. We weigh each opinion by three common-sense factors:
Deciding where to eat based on advice at a dinner party. You lean toward the friend who's eaten there a dozen times over the one who went once. You trust the chef at the table a little more than the guy who "thinks he drove past it." And if someone raves about a place but their last visit was five years ago, you mentally discount it — the menu's probably changed. You're not ignoring anyone; you're just weighting them sensibly. That's exactly what the math does.
One subtle thing: each extra review matters less than the one before it.
Reviews on Amazon. Going from 1 review to 10 changes your trust a lot. Going from 1,000 to 1,010 changes basically nothing — you were already convinced. So the formula gives fast credit for the first few opinions and then flattens out, instead of letting a set with 5,000 reviews bulldoze a set with 50.
We grade six big categories, and a few smaller points live inside each one. Think of it as a report card where some subjects matter more than others.
| Category | How much it counts | What it's really asking |
|---|---|---|
| 💰 Money & Worth | 22% | Is this a fair deal — judged smartly, not by counting pieces? |
| 🔧 The Build | 19% | Is it fun to put together, or a boring slog? |
| 🏆 The Finished Model | 19% | Does the thing look great when you're done? |
| 🧑🚀 Minifigures | 16% | Are the little people good? (Often the whole reason people buy.) |
| 🛠️ For the Hobbyist | 12% | Fun to play with, and good as a parts donor for custom builds? |
| ✨ Rare & New Parts | 12% | Does it have exciting new pieces or colors? |
| 📈 Hold Its Value | side note | Will it be worth more later? (Shown on the side — see below.) |
This used to be a fake number — the same middling score for every single set, because we didn't have the data to do it properly. It got rebuilt from two things we can measure:
Keeping those two separate is the whole point: a Star Wars set can have cheap parts (low floor) but strong collector demand (high outlook), and a plain parts pack can be the reverse.
Each percentage is how heavily that category pulls on the final grade. They add up to 100. The exact splits are adjustable dials, not laws of nature.
This is the heart of fixing the price-per-piece problem, so it's worth going slow.
Instead of counting pieces, we add up what the individual parts are actually worth on the open market, then compare that to what people really pay for the set (not the sticker price — the street price, including the discounts everyone actually gets).
Appraising a used car. A smart buyer doesn't judge by the asking price alone — they ask "what are this car's actual parts and condition worth?" A set whose pieces are collectively worth a lot, sold cheap, is a genuine bargain. A set of cheap common parts sold at a premium is a rip-off — no matter how many pieces are in the box.
A grocery cart. Two carts each have 100 items. One is 100 packets of instant noodles; the other is 100 cuts of steak. "Price per item" says they're comparable. Obviously they're not. Smart Price weighs what's in the cart, not how many things are in it.
Then comes the clever bit: we don't grade that in a vacuum. We compare the set to others of its own kind — same theme, same size class.
Grading a car's gas mileage. 30 mpg is fantastic for a big truck and mediocre for a tiny hatchback. You only learn anything by comparing it to its own class. So a 2,000-piece Star Wars set is judged against other big Star Wars sets — "is this a good deal for what it is?" — landing it somewhere from "clear bargain" to "clear rip-off."
The grade tells you whether a set is priced fairly in one letter. On the set's page, True Value shows you the receipts behind that — laid out so you can see where the money goes.
The itemized bill at a restaurant instead of just the total. It shows the anchor price (what people actually pay), what the parts are worth, and the ratio between them as a plain line: "the parts are worth about 80% of the price." Then it tags the set Bargain, Fair, or Overpriced, and breaks the value into plain bricks, printed & detailed pieces, functional pieces, the minifigures, and any brand/licensing premium.
The hard price math is one thing; whether fans feel it's worth the money is another. We let the crowd's feeling adjust the price score — but only a little.
A thermostat with a limit on it. The crowd can nudge the temperature up or down a couple of degrees if everyone agrees the room feels off — but they can't crank it to 100. The vibe can fine-tune the hard math; it can't overrule it. (Concretely, the nudge is capped at about ¾ of a point on the 10-scale.)
This also stops double-counting: the "feels worth it" opinion tweaks the price score rather than getting counted a second time as its own separate thing.
Lots of sets don't have full information — maybe nobody reviewed them, maybe a category doesn't apply. We handle that two ways.
If a category doesn't apply (like "playability" for a flower bouquet meant only for display), we don't give it a zero — that would be unfair. We just drop it and let the other categories count for proportionally more.
A test where the teacher throws out a bad question. You're not marked wrong on it — it's removed, and the remaining questions are simply worth a little more each. The flower set isn't punished for not being a playset; it's judged on the things that actually matter for a flower set.
If the grade ends up resting on too little, we refuse to act overconfident — but we still give you a real letter.
Election night. The networks won't call a winner when 2% of precincts have reported — but they still show you the count so far. A Drop Score works the same way: a thin set still gets a real grade (never a blank dash), it's just labeled an "early read" (or "too quiet" when it's really thin) so you know to hold it loosely.
Here's where a lot of the magic (and the real cleverness) lives. There's a goldmine of opinion sitting in YouTube reviews — but it's people talking, not numbers. Something has to turn the talking into scores. That something is an AI (Claude).
Hiring a sharp, tireless intern who watches every LEGO review video, reads the comments, and fills out the exact same scorecard for each one — "build fun: 7, looks: 9, worth the money: 5, and here's the quote that proves it." It never gets bored, never plays favorites, and uses identical criteria every time. That's the AI's job: read one review, output consistent numbers.
Why a real AI and not a simpler word-counter? Because LEGO reviewers are relentlessly sarcastic, and sarcasm is a trap for dumb software.
A reviewer says: "Oh GREAT, another $0.10-a-piece masterpiece." A keyword robot sees the word "great" and marks it positive — completely backwards. A person hears the eye-roll instantly. We specifically tell the AI to listen for that tone, and it does: on a deliberately sarcastic test review, it correctly scored the "riveting" build a 3 out of 10, gave the genuinely-praised looks a 9, and flagged that it caught sarcasm.
The AI is also told to leave a category blank if the review didn't actually mention it (no making things up), to back every score with a real quote, and to trust the reviewer's overall take over noisy comments. We then need a few reviews agreeing before that category counts — we don't trust a single voice.
Not booking a restaurant off one Yelp review. You wait until a handful of people say the same thing before you believe it. We wait for at least three.
Every grade comes with two honesty signals, and they mean different things:
We used to call that second one "confidence," but the word made it sound like the tool was unsure of its own math. It isn't. A low number there almost always means one real thing: the reviewers disagreed — or barely anyone weighed in. So we name it for what it actually is.
A weather forecast. "We have data for the whole week" is coverage. "But the models strongly disagree about Thursday" is consensus. You can have lots of data the experts argue over, or a little data they all nod at — and an honest forecaster tells you which.
So instead of a confusing "34% confidence," a grade wears a plain badge that says what's actually going on:
"Locked in" is a pile of agreeing evidence; "Split verdict" means we read plenty but the room was divided (that's the X-Wing); "Early read" and "Too quiet" mean there just wasn't much to go on yet. Those last two retire the old jargon label "data-light."
Different people want different things from a set. A parent cares about playability; a collector cares about how it looks on a shelf; a custom builder cares about the parts. So besides the main "everyone" grade, you can flip a lens that re-weights the categories toward what you care about.
The filters on a hotel or restaurant app. Same places, same data — but "good for kids" and "romantic date night" surface very different rankings. The set didn't change; what you're optimizing for did.
There are four optional lenses — Display fan, Builder, Parent, and Investor — and the Investor lens is the one place the resale stat is actually allowed to move the grade.
Before we grade anything, we figure out what kind of set it is, so we don't judge it by the wrong yardstick.
Not docking a documentary for lacking a plot twist. You judge a film by what it's trying to be. A giant tub of loose bricks shouldn't be marked down for "no impressive display model" — building a display model isn't its job; being a great pile of useful parts is. A flower bouquet shouldn't lose points for "no minifigures." We detect these cases (figureless sets, pure display pieces, bulk parts tubs, tiny bags sold for one rare minifigure) and quietly switch off the categories that don't apply.
Let's walk one through (numbers here are illustrative, to show the flow). Say we grade a mid-size Ideas set with strong reviews:
| Step | What happens | Result |
|---|---|---|
| 1. Identify it | Look up the facts: ~2,100 pieces, $150, 3 minifigures, not a display-only line. | normal set |
| 2. Smart Price | Its parts are worth a lot vs. its street price; compared to similar Ideas sets, it's a good deal. | 7.5 |
| 3. What You Get | Lots of printed parts and variety, no filler padding. | 7.0 |
| 4. Build (from 5 reviews) | Reviewers call the build clever and satisfying. | 8.4 |
| 5. Finished look (reviews) | Widely praised as gorgeous and accurate. | 8.8 |
| 6. Minifigures | Three solid, fitting figures. | 7.0 |
| 7. Combine, weighted | Each category × its importance, added up. | ≈ 7.9 |
| 8. Letter + honesty | 7.9 → B; five agreeing reviews → high confidence, full coverage. | B (7.9) |
Now the real example we actually ran — the UCS X-Wing (#75355). With review mining on, six real reviews got read by the AI, and the set scored a C− (5.7) — a "Split verdict": full coverage, but only 34% consensus. Why so middling? Because the real reviewers were tougher than fan instinct — the build drew "repetitive" complaints, and the value genuinely lags (Star Wars sets carry a licensing tax). The reviewers also disagreed with each other, which is exactly why consensus came out low. The grade isn't being mean; it's being honest.
Nothing here is made up. Each piece of information has a real, checkable home:
| What we need | Where it comes from |
|---|---|
| Price, pieces, minifig count, official review ratings | Brickset (a LEGO database with an open data feed) |
| The full parts list, colors, new-part detection | Rebrickable (catalogs every set's exact contents) |
| What each part is worth + resale prices | BrickLink (the LEGO marketplace's price guide) |
| Which reviews exist + a sanity-check score | Brick Insights (a review aggregator — like Metacritic for LEGO) |
| The actual review transcripts | Supadata (fetches a video's captions for us) |
| Finding reviews + their comments | YouTube's official data feed |
| Reading the reviews into scores | Claude (the AI judge) |
The main way is a hand-picked list of 48 trusted LEGO reviewers (JANG, MandRproductions, RacingBrick, Tiago Catarino, and 44 more). A background sweep quietly checks each of their channels for new uploads, works out which set each video is about, and files it in a "review index." So by the time we grade a set, its reviews are already sitting on the shelf, waiting — no expensive searching needed. A plain YouTube search is only a last resort, for a brand-new set nobody on the list has covered yet.
A librarian who shelves new arrivals as they come in, instead of ransacking the whole building every time someone asks for a book. The shelving (the sweep) happens quietly in the background; the lookup (grading) is then instant and nearly free.
And once the AI has read a set's reviews, we save the scores — so re-grading that set later reuses them instead of re-watching (and re-paying for) every video.
It's harder than it sounds — and easy to get wrong. A title like "LEGO 75355 review" is easy; the number's right there. But plenty of reviewers just say "the new UCS X-Wing," a single "haul" video rattles through ten sets at once, and — the real trap — different sets share words (there are many "X-Wing" and "Millennium Falcon" sets across the years). So name-matching has to be careful: it reads the title only (not the description, where round-ups list everything), needs every distinctive word of the name, ignores a title that names a different set's number, and refuses to guess when two sets could fit. When we first tried it loosely on the X-Wing it wrongly swept in 10 unrelated videos — review round-ups and a different set, the New Republic X-Wing (75460) — so we added those guards, and now it keeps only the real ones. We also fold in a review aggregator's (Brick Insights) already-matched links as a second, pre-sorted source.
For the X-Wing, the reviews that fed the grade came from the trusted-reviewer sweep — for example this UCS X-Wing review and this one.
The grading recipe itself (which categories, how heavily they count) wasn't invented either — it was built by studying how respected reviewers and the LEGO community actually judge sets, and the case against price-per-piece. That research is written up in the project's design specs.
Three principles keep the whole thing trustworthy:
Rotten Tomatoes. It shows a score and links you to the full review — it doesn't photocopy the critic's article onto its own site. We do the same: we store the numbers we calculated and a link back to the original video, but never the transcript or the comments themselves. That keeps us clean on copyright and respects people's privacy — we're not hoarding anyone's words, just the takeaway.
Most of the data is free. The places where the meter actually runs:
A kitchen that preps ingredients in bulk. You don't re-chop the onions for every order — you prep once and reuse. Caching is exactly that: do the expensive work once, reuse the result forever.
✅ Done and live today
All of it — the brain and the face. The full site is live at thebrickdrop.com: nearly 5,800 sets graded, per-set pages, search, browse-by-release, side-by-side compare, saved sets, a written "Brick Drop Take" in the brand's voice on every set, a downloadable share card, and the True Value price breakdown. Finding reviews runs off the trusted-reviewer sweep and reuses saved scores, so grading is fast and cheap. It's covered by around 800 automated tests.
🔧 Built, but could be smarter
🚧 Not built yet
Want the precise, technical version with formulas and exact numbers? See docs/how-it-works.html next to this file.