← Back to all articlesWe Ran an AI Call Center Head-to-Head Against Humans. Here's What the Numbers Actually Said.

We Ran an AI Call Center Head-to-Head Against Humans. Here's What the Numbers Actually Said.

What happens when you run an AI voice agent platform head-to-head against a traditional human call center under the same brand, commercial, and number? A look at conversion rates, cost structures, and audited carrier metrics.

By Florent de Goriainoff, Founder & CEO, Fluents.ai

Everyone selling AI voice agents has a slide that says AI beats human call centers. I run an AI voice platform, and I'm going to tell you something more useful: what happened when one of our clients put us head-to-head against the human call center they'd used for years — same brand, same TV commercials, same phone number — and we all sat down to compare the reports.

Spoiler: both sides learned more about the human call center than about the AI.

The setup

The client is a direct-response retail brand. Customers see a TV spot, call a number, and place an order. For years, those calls went to a traditional outsourced call center billing on a per-order basis. For the past month, a portion of the traffic has gone to our AI agents instead, billed per minute.

When the first comparison report landed, here's roughly how it looked. The human call center showed a 44% gross conversion rate and 66% net. The AI showed 14% gross and 43% net. The human center reported a 3% abandonment rate; the AI reported 28%. On cost, human telemarketing worked out to roughly $66 per order. The AI came in around $2.50 to $3.00. And on upsells, the AI roughly doubled the human take rate on several offers, producing a higher average order value.

If you only read the conversion column, the humans won. If you only read the cost column, the AI won by a factor of twenty. Both readings are wrong, and the reason why is the real story.

The dirty little secret of call center metrics

Take that abandonment gap: 28% versus 3%. Was the AI hanging up on nine times more customers? No. The two numbers were measuring different universes.

Our abandonment number comes straight from the carrier: every call that connects and ends within 30 seconds is counted, no exceptions. The incumbent's number came from their internal call center platform — and as an industry veteran on the call put it, the dirty little secret of call centers is that internal systems vastly under-report total calls. Very short calls get lost in the handoff between the phone carrier and the call center's software and simply never appear in any report. Fewer recorded calls means a smaller denominator, which means every percentage looks better.

Nobody was lying. The client had simply been reading reports from a single source for years without ever auditing where the numbers came from. Their own words on our call: "We've been using this reporting for years and years. We never really had a chance to question it. We just took it at face value."

That sentence should keep every CX and operations leader up at night. Not because your vendor is cheating you — probably nobody is — but because a metric you've never audited isn't a metric. It's a habit.

"Net conversion" isn't a number. It's an opinion.

The deeper we went, the more definitional gaps we found. What counts as a "customer service call" that gets excluded from the sales denominator? The client couldn't say how their call center classified it. Meanwhile, we were counting callers with zero purchase intent — someone phoning about a completely different product they'd seen on TV — as missed conversions against the AI.

Human agents classify calls subjectively, one disposition at a time, under time pressure. Our AI classifies calls by analyzing the full transcript after the fact. Neither approach is dishonest, but they will never produce the same numbers, and pretending they're comparable is how companies make bad decisions with confident spreadsheets.

Even "abandoned" had two definitions inside the same incumbent report.

So who actually won?

Honest answer: the comparison isn't finished, and anyone who declares victory on round one is selling you something. There is a real conversion gap — humans still close more of the callers they get, and closing that gap is script optimization work measured in weeks, not days. There is also a real cost gap — at $3 versus $66 per order, the AI can be significantly worse at converting and still be the better economic engine for a large share of the traffic.

The pragmatic playbook that's emerging looks like this: AI takes the high-volume, routine work — order taking, order status, after-hours coverage — where its cost structure is unbeatable and its consistency is an asset. Humans take the calls where their conversion edge genuinely pays for itself. The interesting management question is no longer "AI or humans?" It's "what's the right routing between them?" — and you cannot answer that question until both sides of the comparison are measured with the same ruler.

What to do before you run your own bake-off

If you're benchmarking an AI agent against your existing operation — and in 2026, nearly everyone is — settle three things before you look at a single result. Agree on where call counts come from, and prefer carrier-level data, because it has no incentive to flatter anyone. Write down the exact formula behind every rate you'll compare, especially what gets excluded from denominators. And ask your incumbent vendor the questions you've never asked: what's in "total inbound calls"? What's classified as customer service? What are the two definitions of abandoned doing in the same report?

You'll learn something regardless of what you decide about AI. Our client did — and most of what they learned was about the vendor they'd trusted for years, not about us.


Fluents.ai builds AI voice agents for customer service, order taking, and outbound engagement — deployed alongside human teams, measured with honest numbers. If you want to run a benchmark like this on your own call volume, get in touch.