No card beats a wrong card
Debugging an LLM that recommended houses to a customer asking about style
A customer asked about "phong cách indie" — indie style. The bot's text answer was reasonable: it talked about tables, chairs, bookshelves. Then, underneath that perfectly sensible paragraph, it rendered a product card for a prefabricated house. Words about furniture; a card selling a building. And the worst detail for anyone who has to fix it: it happened sometimes. Ask the same thing twice and you might get furniture talk with no rogue card at all.
When an LLM product does something absurd, the reflex is to open the prompt and start pleading with the model. I want to argue the opposite. Across fourteen production incidents on this project, the root cause was almost never the model — it was the deterministic code around the model: search, context, ordering, display. The durable fix is to let the LLM do the expressive work and check its output with code. This is the story of one bug that taught me that, layer by layer.
01Reproduce before you reason
My first instinct was to theorize about the model. I resisted it and tried to reproduce instead. The first four ways I phrased the question did not reproduce the bug at all — clean furniture answers, no house. That's the moment an intermittent bug tries to convince you it's random and unknowable.
It isn't. An intermittent bug almost always has a deterministic core. The trick is to stop staring at the flaky surface and go hunting for the thing that is wrong every single time. Here, that thing was the product list being handed to the LLM: it was always wrong. The house was always in the candidate set. The card only surfaced when the model happened to choose a particular kind of lead-in phrasing — so the visible symptom flickered while the underlying fault sat there, rock-steady, on every request.
If the symptom is "sometimes," find the layer where it's "always." The flicker is downstream; the bug is upstream and boringly consistent.
02A display bug is usually several bugs stacked
I had seen this shape once before, on a sibling incident: a customer asked for "nội thất phòng ngủ cho cặp đôi" (bedroom furniture for a couple) and got three dining-table cards priced at 0 ₫. One wrong display, and when I pulled it apart it was not one bug — it was a chain. Putting both incidents together, here is the anatomy, seven layers deep:
- A short category name swallowed a long query. A broad parent category ("Nội thất" — two words) matched a six-word request, because only a subset of keywords had to line up. Breadth won over intent.
- Ordering did the rest. Candidates came back in alphabetical order, so one product family floated to the top every time — not because it was relevant, but because of its name.
- The relevance gate wasn't on this path. The check that would have caught it only ran on a secondary branch, not the one the query actually took.
- The prompt insisted on showing items. A rule pushed the model to present products, so it dutifully surfaced the bad candidates instead of saying "nothing fits."
- A display slip made it grotesque. A price of 0 rendered as "0 ₫" instead of "contact us," turning a wrong result into an absurd one.
- Search couldn't tell a thing from a description. It didn't distinguish words that name a product from words that describe a style — and some product names literally contain style words, so they matched "phong cách" by pure coincidence.
- Loosening ran the wrong way. When a query found nothing, the code relaxed it by dropping the last word — which, for "phong cách indie," meant dropping "indie," the only word that carried any distinguishing meaning.
A single visible error is often five errors stacked. Fix the top one and the bug doesn't die — it changes shape. You have to walk it down to the first layer.
03The layer that surprised me most: loosening in the wrong direction
Layer seven is worth sitting with. The intention behind "drop a word when nothing matches" is good — be forgiving, don't dead-end the customer. But dropping the last word assumes the last word is the least important. In "nhà phong cách bắc âu" maybe it is. In "phong cách indie," the last word is the whole point. The relaxation strategy was quietly throwing away the single token that distinguished one request from another, and then matching on the leftovers — which is how "style" quietly became "house."
04My own wrong assumption
I have to be honest about this one, because it's the most useful part. In an earlier version, I had exempted one search path from the relevance gate — the path that matched several words inside a product's name. My reasoning: if a query matches lots of words in a product, it's surely relevant. That felt obviously true. It was wrong.
The number of matched words does not measure relevance. The kind of matched word does. Three coincidental matches on filler and style tokens tell you nothing; one match on the actual product noun tells you everything.
That assumption is exactly the kind that hides bugs — the "obviously correct" belief nobody re-examines. The fix wasn't cleverer; it was removing the exemption I'd granted out of misplaced confidence.
05The model that replaced the guesswork: anchor + modifier
The durable fix was to stop treating every word as equal and give search a grammar. A query has an anchor — the thing being asked for — and modifiers that shape it. Search resolves the anchor first; modifiers only ever refine. Three symmetric cases show how it behaves:
"phong cách indie" → no anchor (only a descriptor)
→ DON'T guess. Ask one question + show
product-group buttons.
"nhà phong cách bắc âu" → anchor = "nhà" (house), modifier = style
→ return the right houses.
(viewing a sofa) "phong cách → inherit the anchor from context (sofa)
indie" → return sofas in that style.
A descriptor with no anchor is not a failure to match — it's a signal to ask, not to guess. That single reframing turned a whole class of "confident nonsense" into "one clarifying question."
06The backstop: no card beats a wrong card
Even with better search, I don't trust any single layer. So there's a gate after the LLM whose entire job is one rule: a card the bot shows must belong to the product group the bot is talking about. If the words say "furniture" and the card is a house, the card is suppressed. The customer gets the text answer and no card — which is always better than a confidently wrong recommendation.
No card beats a wrong card. Silence costs you a little; a wrong recommendation costs you the customer's trust.
07A Vietnamese trap worth its own section
Vietnamese made one layer especially sharp. To help customers who type without diacritics, search can match accent-blind —
so someone typing ban an still finds "bàn ăn." Useful. But accent-blindness is a double-edged sword: strip the
marks and different words collide.
- "bán" (to sell, a verb) collapses onto "Bàn" (table) — so "do you sell water filters?" can surface a tea table.
- "đen" (black) collides with "đèn" (lamp).
- "đa năng" (multi-purpose) collides with "Đà Nẵng" (a city).
The lesson: accent-insensitivity is a feature for the person typing without marks and a bug for the person who typed them correctly. The two need to be handled as two deliberate modes, not one blurry default. Domain language is dictionary work, not model work — an LLM understands the slang, but it can't filter a database.
08Verifying the fix (and finding a bug I wasn't looking for)
A fix you can't re-run isn't finished, so all of this is pinned by a self-grading regression set. The first full run came back 23/24 — and the one failing case had nothing to do with houses or style. It exposed an old, unrelated gap: a topic named by a single word ("sofa," "lavabo") wasn't being treated as a strong-enough topic for a follow-up question to inherit. It had been latent for a long time; a wide-enough test finally dragged it into the light.
A good test suite finds the bugs you weren't looking for. Don't write tests only to confirm the thing you just fixed.
09Three questions before you touch the prompt
When an LLM product misbehaves, I now run through these before opening the system prompt:
- Was the model even given the right data? A perfect prompt over a wrong candidate list produces confident nonsense.
- Is anything checking the model's output? If nothing verifies what comes back, the prompt is your only line of defence — and prompts only lower the frequency of errors. Code takes them to zero.
- Which layer is "always wrong"? If the symptom flickers, the fault is deterministic somewhere upstream. Find that layer before you touch anything downstream.
The prompt is almost never the bug. It's the last, most visible thing in a chain of quiet, deterministic decisions — and those decisions are where reliability is actually won.
On the "I" in this piece, and how it was built: the system was developed under an AI-coding-agent orchestration model. A "wrong assumption of mine" means an approach that was proposed, that I approved, and that the project later proved wrong — and I think owning that makes the story more honest, not less. Every change was verified on production.
Written by Nguyễn Trung Tâm — AI Engineer & AI-native builder.
More work: nguyentrungtam-portfolio.pages.dev · GitHub:
nguyentrungtamwork-hue
© 2026 Nguyễn Trung Tâm. All rights reserved.