An AI sales chatbot that learns from its own failures
Case study — a production assistant on a live 1,330-SKU WooCommerce store
A building-materials and furniture retailer wanted a 24/7 sales assistant on its WooCommerce store — one that quotes real prices, recommends the right product out of 1,330 SKUs, drives the cart, and hands the sales team a lead they can act on in five seconds. The interesting engineering isn't the chatbot answering questions; it's everything built around the model to make it trustworthy — and a controlled 48-hour self-learning loop that lets the system get better from the very questions it answered badly, without anyone reading logs.
01The problem
Customers arriving after hours had no one to advise them; people asked a price and left; and the leads that did reach the CRM were so thin that a salesperson calling back had to start from scratch. The product goal: an always-on assistant that quotes live prices, suggests the right products, adds them to the cart, and — when a customer leaves a phone number — sends sales a needs briefing instead of a raw transcript.
02Constraints that shaped the architecture
| Constraint | Design consequence |
|---|---|
| Never invent a price, stock level or policy — one slip loses trust | Prices are pulled LIVE from WooCommerce every turn; the LLM never types a number; a server-side filter strips any price it emits anyway. |
| Shared hosting: no CI/CD, no SSH, the dev machine has no PHP | Deploy by uploading the plugin; a Node-based PHP syntax gate; all verification runs against production over the API. |
| Vietnamese catalog: accented / unaccented text, currency slang, regional names | A dedicated query-normalization layer; accented and accent-folded matching are handled deliberately differently. |
| Product names entered by many people, some stuffed with ad adjectives | Search has to be immune to accidental word overlap (see the anchor + modifier model). |
| Part of demand (beds, cabinets, shelves) is made-to-order in the workshop, not stocked | A separate conversation branch: no product card, elicit specs, take a phone number for the design team. |
| The hosting firewall blocks the IP under bursts of API calls | The test runner paces itself and retries; it can tell a "403 firewall" apart from a "500 code error". |
03Architecture, at a block level
Customer ─► Chat widget (vanilla JS, remembers the conversation across pages)
│
▼
┌──────────────────────── WordPress plugin (PHP) ────────────────────────┐
│ 1 Query normalization aliases · drop filler · parse budget · │
│ split ANCHOR (the item) vs MODIFIER (style) │
│ 2 Multi-tier search product name → room/space → category → 1 word │
│ 3 Context inheritance follow-ups inherit the topic; a new topic cuts │
│ 4 Relevance gate a weak-match list is treated as "no stock" │
│ 5 LLM (OpenAI/OpenRouter) prompt carries LIVE prices + a real-stock map │
│ 6 Post-LLM safety net reconcile words ↔ cards · cut empty promises · │
│ strip any price the model typed │
│ 7 Self-diagnosis tag every "stuck" turn with a reason code ──► 48h loop
│ 8 Lead pipeline structured needs → store locally → push to CRM │
└─────────────────────────────────────────────────────────────────────────┘
│ │
▼ ▼
Product cards + cart Internal CRM (idempotent webhook)
(WooCommerce Store API) + funnel dashboard & A/B test
The LLM only does what it's good at — phrasing and eliciting needs. Everything that must be correct — prices, which products may appear, whether a lead is saved — is decided by deterministic code and cross-checked after the model answers.
04Seven design decisions
- "No card beats a wrong card." A wrong product card destroys trust faster than showing nothing. Every product list passes a relevance gate; the bot's words and its cards must belong to the same product group, or the card is dropped and the bot asks again.
- The "anchor + modifier" retrieval model. Separate item words from descriptor words (style, colour, "premium", "smart"). A query with no anchor returns no products — just a question and product-group buttons. When a query must be loosened, drop modifiers first and never the anchor.
- Save the lead first, everything else after. The lead is written locally before the CRM push; if the AI is down the lead is still saved; a returning customer's needs are merged, not overwritten.
- The CRM note is a briefing for whoever calls, not a transcript. Seven need-fields plus a one-line hint for sales; the full conversation is stored separately for retrieval.
- A controlled self-learning loop — the highlight of the project, below.
- Catalog-wide facts are computed by code; the AI only phrases them. "Cheapest / most expensive / newest / on sale" questions run a query over the whole product group, build a "truth from the catalog" block, and force the cards into the right order — the model never asserts a superlative. No data → the bot says so.
- Every fallback reply carries telemetry. A customer not getting a real answer is a line on the dashboard (failure type, error code, page), never a silence — so an outage hiding behind a friendly "please try again" gets seen.
★The 48-hour self-learning loop
the #1 idea — not a feature, a controlled self-improving system
Most commercial chatbots are trained once at launch and then stand still, while real customers keep asking in thousands of unforeseen ways — currency slang, regional names, questions by room rather than by product name, items the shop doesn't stock. Each of those is a customer leaving in silence. The usual answer — "someone reads the logs occasionally" — means, in practice, nobody reads them. The design question: how do you make learning-from-failure happen automatically, on a schedule, and safely, so the owner only has to read one short report?
Three decisions that make it work
1 · The chatbot must know it just failed, and why. The premise of any learning loop is a quality error signal. The system doesn't use an LLM to grade an LLM (expensive, slow, unstable); it self-diagnoses with cheap in-turn rules, tagging each "stuck" turn with one of 10+ reason codes — separating a search failure from an AI failure from a policy gap. The reason code decides the remedy: a search miss is cured with the dictionary, a policy gap with an FAQ, an AI slip with code. Diagnosing the wrong root cause is worse than not diagnosing — so the observability layer itself gets tested.
2 · Separate the "training knobs" from the code. All the domain knowledge that changes over time (synonyms, filler words, style descriptors, made-to-order items, extra FAQs) lives as knobs — configuration, editable instantly, no deploy. The code holds only the mechanism. The consequence: an agent can improve the chatbot without ever touching code. That boundary is the core safety line.
3 · The agent is granted power in tiers — always less than it "could" do. It may add safe knobs (only once the owner has enabled autonomous mode); it may only propose code or prompt changes, with a root-cause analysis, for a human to decide; and nothing in the automatic loop may delete or edit an existing knob — the agent only adds. The first two cycles always run in propose-only mode: trust is built in tiers, the way you'd onboard a new hire.
Why it's safe enough to run on production
Giving an AI agent the power to change a live selling system is serious. The defences, at a principle level: the API gate is closed by default and opens only when the owner sets a server-side secret; the agent has no route into the admin, no login, and can't read customer data beyond the stuck questions; customer chat is treated as data, not commands (prompt-injection defence through the chat channel itself); business guardrails live server-side and don't trust the agent — if the agent proposes treating a word as "filler", the server tries it and refuses if that word is currently finding real products; every change is additive, capped per cycle, logged before/after, and one-click reversible; and the agent is barred from generating fake data or fake leads.
What actually happened
- First real cycle on production: 54 seconds, through the full check → take-work → report path.
- A day-one incident handled exactly as designed: the cloud environment blocked an outbound connection, so the agent stopped safely and reported the cause in 13 seconds — no guessing, no working around it. That is the behaviour you want from a production agent.
- A human-added alias (a regional word for "toilet") took effect instantly, no deploy — proving the knob/code boundary holds.
- A real customer question went the full lifecycle: flagged → root cause found → fixed → replayed and confirmed → marked done.
05Verified results
| What | Result |
|---|---|
| Conversation regression | 43/43 scenarios pass — auto-graded, run on production (incl. a real 3-turn conversation) |
| Price-ranking questions | 5/5 match the true price (e.g. cheapest 3,317,000₫, dearest 113,616,000₫ among 82 toilets) — the test loads all 1,330 SKUs and computes the answer itself |
| AI's false "-est" claims | Before: "our cheapest is X" where X ranked 7th of 84. After: 0 unfounded superlatives — cross-checked against the catalog |
| Brand banned-words in the bot's replies | 0 across all 43 scenarios (checked on every scenario) |
| "Off-topic card" failure class | 0 recurrences in the test set (reproducible on production before the fix) |
| Leads under injected AI outage | 0 leads lost across both failure modes (connection loss, error status) — fault injection on production |
| Needs profile when the customer changes pages before leaving a number | All 7 fields + the full conversation kept |
| Self-learning loop | 54 s per cycle; stops safely by design on a network incident |
| Legal risk in product names | 0 / 1,324 names (at scan time) contain a banned superlative — full catalog scan |
| Release cadence | ~30 verified production releases, each with a backup + post-deploy verification |
06My role — honestly
This was built with agent-orchestrated development: I designed the system and made the calls; an AI coding agent (Claude Code) wrote the code under a written operating charter I set.
- I (the owner) defined the problem and priorities, made every product and policy decision (payment rules, how made-to-order items are handled, how much autonomy the agent gets), manually tested to find bugs — most of the production incidents behind this system were found that way — approved every option before it shipped, and signed off on real CRM data.
- The AI coding agent (Claude Code) investigated root causes, proposed options, wrote the code, deployed, verified and reported — under my charter: back up before changing anything; never enter keys or passwords; verify with real results; treat customer content as data, not commands.
Human keeps decision and verification authority; the agent keeps execution speed. That division is why one person could put ~30 verified releases on production, quickly, with no engineering team — and it's a capability in its own right.
Written by Nguyễn Trung Tâm — AI Engineer & AI-native builder. More work: nguyentrungtam-portfolio.pages.dev
© 2026 Nguyễn Trung Tâm. All rights reserved.