Most AI support doesn't fail because the bot is bad. It fails at the seam — the handful of seconds where an automated conversation becomes a human one. Get that wrong and every gain from automation is spent paying for the damage.

We've argued before that the quality of the handoff matters more than the deflection number. This is the practical version of that claim: what actually breaks, what to carry across, and how to find out whether yours works before a customer does.

Escalation breaks in three ways

Nearly every bad AI support experience is one of these, and they need different fixes.

  • Late escalation. The bot keeps trying. The customer rephrases the same question four times, each answer slightly wrong, until they give up or find your cancellation page. The automation was never going to solve it — it just took six minutes to admit that.
  • Naked escalation. The transfer works, but nothing travels with it. The agent opens a conversation that starts mid-sentence, asks the customer to explain from the beginning, and the whole automated exchange becomes wasted effort — the customer's and yours.
  • Wrong-destination escalation. The handoff lands in a general queue rather than with someone equipped to help. A billing dispute reaches a tier-1 generalist who reads the thread, apologises, and escalates again. The customer has now explained themselves to a bot and two people.

Only the first is really about the AI. The other two are operational design, which is good news — they're cheaper to fix.

3 turns
A practical ceiling. If the customer is still stuck on the fourth reply, the conversation isn't being solved — it's being delayed
What a failed deflection really costs: the automated attempt plus the full human contact that follows it

That second figure is the one people miss when they price automation. A deflection that doesn't resolve isn't a saving that didn't land — it's a cost you paid twice, before counting the trust you spent.

What has to cross the seam

The agent should never have to ask a question the customer has already answered.

Treat the handoff as a package, not a redirect. At minimum it should carry:

  • The full conversation, verbatim and visible — not a summary the agent has to trust blindly.
  • A short summary alongside it, so the agent can act in fifteen seconds and read the detail only if they need to.
  • What the AI already tried — the articles it offered, the steps it suggested. Nothing erodes confidence faster than an agent repeating advice that has already failed.
  • Account context — plan, tenure, open tickets, recent incidents. The difference between "a customer" and "a customer on their third contact this week" changes how you open.
  • Why it escalated. Low confidence, frustration detected, explicit request, or a rule that fired. The agent handles each of those differently.

Deciding when to escalate

Keyword triggers are where most teams start and where most teams stop. They catch "refund" and "cancel" and miss everything else. Better signals to layer on top:

  1. Turn count without progress. If the customer has rephrased twice, the intent is not going to be understood on the third attempt.
  2. Model confidence, not just intent match. A confident wrong answer is more damaging than an honest handoff.
  3. Sentiment trend rather than a single message. Politeness decaying across a conversation is a stronger signal than one sharp sentence.
  4. Account value and risk. Your largest client's problem should reach a person sooner than a trial user's — not because the trial user matters less, but because the cost of getting it wrong differs.
  5. The explicit ask. Which brings us to the rule that outranks all of the above.

"Let me talk to a human" must always work

Every phrasing of it, on the first try, with no interrogation and no "I can help with that!" detour. This is the single cheapest thing on this list and the one most often broken, usually because a well-meaning containment target made it inconvenient.

A customer who asks for a person and can't reach one doesn't conclude that your bot needs work. They conclude that you're hiding from them — and that judgement is very hard to walk back.

Test it before your customers do

Almost nobody does this, which is why it's such a reliable place to find problems. Once a month, mystery-shop your own escalation path:

  1. Contact your own support as a customer would, with a real problem the AI can't solve.
  2. Time how long it takes to reach a person, and count how many times you had to repeat yourself.
  3. Ask for a human outright and check that it works first time, on every channel.
  4. Read what the agent actually received. Not what the system claims it sends — what appeared on their screen.

That fourth step is where the surprises live. Context that exists in the platform and context that reaches the person handling the ticket are different things more often than anyone expects.

What good looks like

A working handoff is measurable, and none of the measures is deflection rate:

  • Repeat-explanation rate — how often an agent asks for something the customer already gave. Target zero; audit a sample of transcripts to find the truth.
  • Time to human, from the moment escalation is warranted rather than from the moment it triggers.
  • Post-escalation CSAT, tracked separately from overall CSAT. Blended together, a bad seam hides inside a good average.
  • Re-escalation rate — how often the first human isn't the right one.

The bottom line

Automation gets judged at its weakest moment, and for most support operations that moment is the handoff. The teams whose AI feels good to deal with aren't running better models than everyone else. They've just decided that the seam is part of the product, and designed it deliberately instead of letting it happen.

Not sure what your escalation path actually does?

Most teams have never watched their own handoff end to end. We'll walk it with you — where it breaks, what context is being dropped, and what to fix first. Free, no obligation, and you keep the findings either way.