Teaching a production chatbot to learn from its own failures
A 48-hour self-improvement loop you can actually leave running
A customer typed "tủ âm bàn trong tầm giá từ 10 củ đổ lại" — a cabinet, budget around ten million or under. The chatbot read "từ 10 củ" as a floor and started showing the expensive end. The customer didn't complain. They just left. Nobody on our side knew it happened, because knowing would have meant someone sitting down to read the chat logs, and nobody reads the logs.
That single silent failure is the whole problem in miniature. So I stopped asking "how do we catch mistakes like this?" and started asking a harder question: how does a mistake like this get detected and fixed without a human in the room? This is a write-up of the answer — a 48-hour loop that turns each answer the bot got wrong today into an answer it gets right the day after tomorrow. The interesting part isn't a smarter model. It's two much less glamorous things: structured error signals and permission boundaries.
01"Occasionally read the logs" is not a strategy
Most commercial chatbots are trained once, at launch, and then stand still. Real customers don't stand still — they ask in thousands of ways nobody scripted: money slang ("dưới 10 củ", "tầm 5 chai"), regional names for the same object, requests phrased by room instead of by product name, and things the shop doesn't even sell. Every one of those is a customer quietly walking away.
The industry's default answer is "someone will look at the logs sometimes." In practice, no one does — not because people are lazy, but because reading logs is unbounded, unrewarding work with no obvious stopping point. A strategy that depends on a human doing that forever is not a strategy. So the design goal became: learning from failure should happen automatically, on a schedule, safely — and the owner should only ever have to read one short report.
02Pillar one — the bot must know why it failed
A learning loop is only as good as the error signal feeding it. The tempting move is to use an LLM to grade the LLM: ask a model "did you answer that well?" It's expensive, slow, and unstable. Instead, the bot diagnoses itself with cheap rules, right inside the chat turn, tagging every "stuck" moment with a reason code from more than ten categories. The codes separate failures by cause, not by symptom:
- Search — no products found, or the wrong products found.
- AI — the model promised results and forgot to attach them, or over-promised.
- Policy — a question about shipping or warranty the FAQ doesn't cover yet.
- Budget — nothing in stock at the price the customer asked for.
- By design — "stuck" on purpose (e.g. items that are quoted, not sold off the shelf) and only tracked.
Why this matters: the reason code decides the cure. A search failure is fixed with a dictionary entry; a policy gap is fixed with an FAQ; a genuine model slip is fixed with a prompt or code change. Skip the diagnosis and the loop is just guessing — trying random remedies and hoping.
And here's the part I earned the hard way: a wrong diagnosis is worse than none. At one point the "stuck questions" board was pairing a flagged failure with the wrong message entirely — it read the last line of the session instead of the line that actually failed. The self-learning loop then read that same wrong field and would have gone off to reproduce the wrong question. Your observability layer has bugs too, and its bugs are twice as dangerous because they steer every fix that comes after. The instrument needs calibrating before you trust its readings.
03Pillar two — separate the "training knobs" from the code
Everything about a store that changes over time — the other names people use for a product, filler words to ignore, style descriptors, items that ship only from the workshop, extra FAQ answers — is pulled out of the code and turned into knobs: configuration data, editable instantly, no deploy. The code holds only the mechanism; the knobs hold the knowledge.
One concrete example makes the boundary real. A human added a regional alias — "bệ xí" as another word for "bồn cầu" (toilet) — through a knob, and it took effect immediately, with no deployment. That's the whole point: the thing that changes is data, and data can be handed to a machine to change safely. The consequence is the core safety property of the entire loop: an agent can improve the chatbot without ever touching code.
04Pillar three — give the agent power in tiers, always less than it could take
Letting an AI agent change the behaviour of a system that is selling to real customers is a serious thing. So the agent's authority is tiered, and deliberately narrower than what it's technically capable of:
| Type of change | Who may do it |
|---|---|
| Add a safe knob (synonym, filler word, FAQ…) | Agent does it itself — only once the owner has switched autonomous mode on |
| Change a prompt or code | Agent may only propose, with a root-cause analysis; a human decides |
| Delete or edit an existing knob | No one in the automated loop may. The agent only ever adds. |
The first couple of cycles run in propose-only mode no matter what: the owner reads what the agent thinks is wrong and judges whether it's right before handing over the keys to fix things itself. Trust is built in steps — exactly the way you onboard a new hire on probation. You don't give someone commit access on day one because they seem clever; you watch how they reason first.
05What one cycle actually does
Failures accumulate continuously as customers chat. On a schedule, a gate checks "has enough time passed?" and, if so, a cloud agent wakes up with permission to call exactly one narrow API and nothing else:
That verify-then-keep discipline is what makes the loop trustworthy. "Fixed" is not a claim the agent is allowed to make from reasoning alone — it's a status the system only grants after re-asking the live bot and seeing a different, better answer. A change it can't verify gets rolled back automatically and demoted to a human-reviewed proposal.
06The safety model — why I dared switch it on in production
Seven layers stand between "an agent had an idea" and "a live store changed." In plain language:
- The API is closed by default. It opens only when the owner personally configures a server-side secret; the key comparison is constant-time (so timing can't leak it), calls are rate-limited, and every call is written to an audit log.
- No admin surface. The agent can't reach the admin panel, can't log in, and can't read customer data beyond the stuck questions themselves.
- Customer chat is data, not commands. The agent is told explicitly to ignore any instruction sitting inside a customer's message — that's prompt-injection defence through the chat channel itself.
- Business checks live on the server, which does not trust the agent. If the agent proposes treating a word as filler, the server tries it first and refuses if that word is currently finding real products. The agent couldn't break search even if it wanted to.
- A ceiling on how much can change per cycle, add-only, never delete, and every change carries a before/after log and a one-tap restore.
- No fake data. The agent's test sessions are marked so they never pollute analytics, and the agent is forbidden from putting a phone number into chat (no fake leads into the CRM).
- A call budget. The number of trial questions per cycle is capped, and stays under the site's own rate limits.
The one I'm proudest of is the fourth. A safety model that asks the agent to behave is not a safety model. This one assumes the agent might be wrong and puts the veto in deterministic server code that the agent can't argue with.
07The first day in the real world
The first real cycle on production ran in 54 seconds — full circuit: check, pick up work, report. But the day I'm fondest of is the one where it didn't work. The cloud environment blocked its outbound connection. The agent stopped in 13 seconds and reported the exact cause — no guessing, no trying to route around the block, no creative improvisation. It failed closed and told me why.
Stopping at the right moment is a feature, not a shortfall. An agent that keeps going when it hits something abnormal is the one that quietly corrupts your data. I would rather have an agent that halts and hands me a clear reason than one that's determined to be helpful.
08What I'd take to the next system
- Structured error signals matter more than a smarter model. One correct reason code is worth more than ten LLM calls spent on "self-reflection."
- Separate mechanism from knowledge. Whatever changes over time has to be data — only then can you safely hand it to a machine.
- Design an agent's authority like a probationary employee's: narrow, capped, big moves reviewed, everything logged and reversible.
- An agent must verify against reality, and must know how to roll itself back. "Done" only counts when re-asking changes the answer.
- Stopping at the right time is a feature. When an agent hits something abnormal, it should halt and report — not invent.
None of this required a bigger model. It required treating the loop like a system with failure modes of its own, and being honest that the most dangerous bug is the one in the thing you use to watch for bugs.
On how this was built: the system was developed under an AI-coding-agent orchestration model — I designed the loop, set the permission boundaries, decided the reason-code taxonomy, and accepted or rejected every change; the coding agent wrote code to that charter, and every change was verified on production before it counted as done.
Written by Nguyễn Trung Tâm — AI Engineer & AI-native builder.
More work: nguyentrungtam-portfolio.pages.dev · GitHub:
nguyentrungtamwork-hue
© 2026 Nguyễn Trung Tâm. All rights reserved.