← Nguyễn Trung Tâm · Portfolio

A friendly fallback is a silencer on a fire alarm

How a polite chatbot hid its own outage — and the discipline of making failure visible again

By Nguyễn Trung Tâm AI Engineer ~9 min read Live: tanphatsolutions.vn
ObservabilityLLM ReliabilityFallback Telemetry Truth-computed TestsProduction

For a stretch of time, the sales chatbot on our store was unfailingly polite. Whatever you asked — a price, a product, a comparison — it replied with some version of "please give our hotline a call and we'll help you right away." Courteous. On-brand. Completely broken. That single reply was not an answer; it was the sound a system makes when it has stopped working and nobody wired up an alarm. This is a write-up of three friendly veneers — a fallback reply, a confident claim, a green test — each of which made the product look fine while it was quietly failing, and what it took to make failure visible again.

01The reply that was too polite

The first veneer was the "call our hotline" message. It looked like a deliberate design choice — a graceful hand-off to humans. It wasn't. It was a hard-coded fallback the chat widget shows whenever the response coming back has no real answer in it, and it was firing on every request.

The root cause sat one layer below the chatbot entirely. Each message the widget sent carried a short-lived platform verification token — the kind a web platform embeds into a page at load time and expires after a day or so. The page never refreshed it. So a tab left open overnight, or a page served from a long-lived cache, would send an expired token, and the platform rejected the request before it ever reached the chatbot's code. The irony that stung: the chat endpoint is public — it never needed that token in the first place.

The bug wasn't in the AI, or even in my code. It was in a security handshake the request had to survive to reach my code — and the friendly fallback made that handshake failing look like ordinary customer service.

Two things about this one still shape how I work. First, I made myself prove the blast radius: I sent a request with no login cookie at all and a deliberately stale token, and it was still rejected — confirming that even anonymous, first-time visitors were hitting this, not just logged-in edge cases. Second, and more uncomfortable, I had to write down what I couldn't know: how many real customers this had turned away. I can't measure it. The failure happened upstream of any logging I controlled. An honest postmortem names the number you don't have as clearly as the ones you do.

02The principle: put a signal behind every fallback

Removing the token from public requests was the easy fix (with a check that the other endpoints still reject stale tokens — I fixed a bug, I didn't loosen security). The fix that mattered was structural:

A friendly fallback is a silencer bolted onto a fire alarm. Every time it fires, it makes a failure sound like a feature. So every fallback must carry telemetry — or it converts an outage into silence.

Now, every time the widget has to fall back to a canned reply, it reports it to the dashboard — the failure type, the error code, the page it happened on. The server does the same whenever it has to say "the system is busy." I was careful about what this isn't, and wrote the limitation down next to the feature: the signal travels through the website itself, so if the whole site is down, the telemetry is silent too. This is not uptime monitoring. It's something narrower and, for this failure class, more useful: a count of every time a customer got a polite non-answer instead of a real one.

03The second veneer: confident nonsense

The next veneer didn't dodge the question — it answered it, fluently and wrongly. Asked for the cheapest option in a category, the bot said "our most affordable is model X." Model X ranked 7th of 84 by price. It wasn't lying so much as overreaching: it had been handed a small sample of three to five products and had spoken for the entire catalog, with total confidence and no idea the list it saw was only a sample.

There were really three faults stacked here, and only one of them was the model. The retrieval had handed the LLM a partial list; a "no superlatives" brand rule lived only as an instruction in the prompt, with nothing checking the output; and the phrasing was seductive enough that no one questioned it. The durable fix moved the truth out of the model:

An LLM must never infer a whole-catalog fact from a sample. Anything that claims to be the cheapest, the only, or all of them has to be computed by a query. The model gets to phrase the truth; it doesn't get to decide it.

The wording we settled on is deliberately modest — "the lowest listed price in this group is …" — a fact the code can stand behind, not a marketing superlative the model felt like reaching for.

04The third veneer: "the test passed"

Here's the humbling part. When I built that ranking engine, my own fix shipped three bugs in a row — and I want to talk about them, because the way they were caught is the actual lesson.

Every one of those was caught by tests that compute their own answer key. For price-ranking questions there is no hand-written expected result. The test loads the entire catalog — all 1,330 products — through the store's public API, applies the group-scope rules itself, computes the true cheapest and dearest, and compares that to the bot's cards, both the top card's price and the order of the cards.

A test with a hand-written answer only checks what you already thought of. A test that computes its own answer from the source of truth catches the thing you didn't — including the bug you just wrote, because the person writing the expected answer is the same person who made the mistake.

That second bug — ranking across all toilets instead of the smart ones — is exactly the kind a hand-typed expected value would have blessed, because I'd have typed the wrong expectation with the same blind spot that produced the code. The catalog didn't share my blind spot.

05The flaky veneer — and the test that took down the office

One rule kept passing twice and failing the third time: a "these items are quote-only, don't show a card" behaviour that lived only in the prompt. Whenever the model happened to open with a certain phrasing, the safety net injected a card anyway. Intermittent by construction.

A test that's green twice and red once isn't flaky luck — it means a branch depends on how the AI phrased itself. Find that branch and replace it with deterministic code.

So the quote-only rule became a code-level block on both paths that can emit a card, and the flakiness went away. But the same investigation surfaced a second, more embarrassing signal: a burst of my regression runs made everyone sharing the office network get "the system is busy," because the message quota was counted per IP and my tests were spending the real users' allowance. The thing that caught it was the fallback telemetry from section 02 — my own monitoring reported an outage I had personally caused. The fix: test requests now carry a secret key (compared in constant time) that exempts them from the quota and keeps them out of the business stats.

Test infrastructure must never spend real users' resources. If your QA can take down the office, it's part of production whether you meant it to be or not.

06A note on the treadmill I stopped running on

Two smaller cases taught the same meta-lesson from a different angle. Vietnamese search kept colliding on words: the verb "bán" (to sell) matching the noun "Bàn" (table); the verb "tìm" (to look for) getting treated as a product word and finding nothing. Each time, the tempting fix is to add one more word to a filler-word list. I did that, twice.

When you're patching the same list for the third time, change the mechanism — don't patch it a fourth. I moved from a hand-maintained list to a vocabulary derived from the catalog itself, so filler words fall away on their own instead of being chased one by one.

07The through-line

Every one of these started as something that looked fine. A courteous fallback. A confident recommendation. A passing test. A rule that worked yesterday. The failure mode of an AI product isn't usually a crash — it's a smooth surface with a hole behind it. So the job, over and over, is the same two moves:

None of this makes the chatbot sound smarter. It makes it honest — which, for a system that talks to your customers all day, is the more valuable trait.

On how this was built: the system was developed under an AI-coding-agent orchestration model — I designed it, set the guardrails, made the calls, and found most of these failures by testing the live site by hand; a coding agent (Claude Code) wrote code to that charter, and every change was verified on production before it counted as done. The "three bugs in my own fix" were mine to own — and the tests were built to catch exactly that.

Written by Nguyễn Trung Tâm — AI Engineer & AI-native builder.
More work: nguyentrungtam-portfolio.pages.dev · GitHub: nguyentrungtamwork-hue

© 2026 Nguyễn Trung Tâm. All rights reserved.