
The AI Shipped Four Vacuums Instead of Two: What Deploying Voice AI on Live Calls Actually Looks Like
Demos are clean, but real-world voice AI deployments are messy. What happens when an agent ships four vacuums instead of two, and what does it teach us about undocumented processes, data access, and customer experience?
By Florent de Goriainoff, Founder & CEO, Fluents.ai
Every AI vendor's website shows you the demo. Almost none of them show you week three of a real deployment, when the AI has just shipped a customer four vacuum cleaners instead of two and everyone's on a call figuring out whose fault it is.
We deploy AI voice agents on live customer traffic — order taking, order status, escalations — for retail brands and restaurant chains, working alongside BPO partners. Here are the lessons from the last few months of production deployments that no demo will ever teach you. They're more useful than the demo.
Lesson one: AI fails exactly where your process was always broken
The four-vacuums incident is my favorite. A customer ordered a product, then accepted the "Pro upgrade" upsell. Our agent added the Pro to the cart, exactly as scripted. The warehouse shipped everything in the cart: four units instead of two.
Whose bug was it? Nobody's — and that's the lesson. The client had copied their human call center script into the AI word for word. The script never said "the Pro replaces the base unit." It didn't need to, because for years the human agents just knew, quietly swapping the SKU and adjusting the charge every time. Humans had been silently patching a hole in the documented process, and nobody had ever written the patch down. The client's own verdict on our review call: "We probably should have caught this sooner. We fell short on that."
The AI didn't create a defect. It executed the documented process faithfully and exposed everything that was undocumented. This is the pattern in every deployment: AI doesn't fail where you expect; it fails precisely where your process was always broken and your people were absorbing the gap for free. Deploying an AI agent is the best process audit you will ever run — budget for what it uncovers.
Lesson two: give the AI what you'd give a new hire, or watch it fail
The analogy I use with every new client: treat the AI agent like a workforce you're onboarding, not software you're installing. If you hired me tomorrow to handle your customer calls, you'd need to give me system access, a detailed manual, your quality bar, and your definition of success. The AI needs exactly the same things.
The clearest version of this is data access. "Where's my order?" automation works beautifully — look up the tracking number, check the CRM record, read out the status. But when a delivery is stuck because a payment failed and nobody granted access to the payment gateway, the smartest AI on earth gives the same useless answer a human would if you covered one of their eyes. Most "AI limitations" we're asked to fix are actually access limitations wearing a costume.
This is also why real timelines look the way they do. Building the agent, once access and process are clear, takes days. Getting there for a significant client takes about two months — security reviews, permissions, breaking down internal silos, agreeing on what the reporting should say. In theory it's two weeks. In practice, companies hold their own projects back in a dozen small ways, and the honest vendors tell you that upfront.
Lesson three: small words carry the whole customer experience
In week one of a callback campaign, our agent greeted returning callers with a warm, personalized "Hi [first name]!" Best practice, right? It confused people badly enough that the client asked us to remove it. Someone returning a missed call doesn't expect the other side to know their name before they've said a word — the personalization read less like warmth and more like "how do you know who I am?" We replaced it with a plain "Hi." Confusion gone.
Context decides everything. An inbound caller who gives you their order number experiences your data access as magic. A cold callback recipient experiences the same data as surveillance. The best customer experiences aren't the most personalized ones; they're the ones that match what the customer expects at that exact second. On another campaign, we learned a cousin of the same lesson when a simple static message outperformed our fully dynamic agent on engagement — the AI was doing too much, and customers just wanted fast and clear. We simplified. Numbers recovered. Ego adjusted.
Lesson four: escalate before frustration, not after
Every deployment includes calls the AI should not finish. The design question is when to hand off. Our rule: the agent transfers to a human before the customer gets frustrated — the moment a case needs judgment, data we can't reach, or empathy that matters — not after three failed loops, when trust is already gone. In one retail deployment, refunds always reach a human; the AI's job is to arrive at that handoff with the customer's name, number, store, and context already attached so the human starts ahead, not behind.
A hard-won sub-lesson: audit your notification logic as carefully as your conversation logic. Early in one pilot, escalation emails were firing for routine cases and missing details the human team needed. None of that is glamorous AI work. All of it is the difference between a pilot that builds trust and one that dies in week two.
Lesson five: run the rollout like an operation, not a launch
What actually made our recent multi-store retail pilot work had little to do with model quality. We started with a handful of stores before scaling to twenty-five. We held weekly syncs with the client and our BPO partner that became daily syncs at go-live. We defined success criteria together, in writing, before the first live call — including what the AI should never do. Supervisors could pull any call recording and read any transcript. And when something misfired, the answer was a same-day fix and a note in the next sync, not a ticket queue.
Scale amplifies reach, and it amplifies problems by exactly the same multiple. That's why you soft-launch. AI in customer service fails when it's deployed as magic. It works when it's deployed as an employee — onboarded, supervised, and reviewed like one.
Ship the guardrails, not just the agent.
Fluents.ai deploys AI voice agents on live customer traffic — with the process audit, the guardrails, and the daily syncs included. Talk to us about a realistic deployment plan for your call volume.