Why the threshold matters more than anything else in setup
A confidence threshold set too low lets the assistant answer questions it's actually unsure about, which means it sometimes guesses โ and a guessed refund window or an improvised billing rule is the kind of mistake that shows up in a chargeback or a support escalation days later, not immediately. Set too high, and it hands off questions it could plainly have answered, which just relocates the workload back onto your team without reducing it.
There's no single correct number, because the right threshold depends on what a wrong answer costs in that specific topic โ which is why per-topic thresholds matter more than one global setting.
Set thresholds per topic, not one number for everything
Billing, refunds and anything involving money or a legal commitment should sit at a stricter threshold than something like opening hours or how to reset a password โ being wrong about the second is a minor annoyance; being wrong about the first is a real cost. Start conservative on anything with financial or contractual weight, and loosen it later once you've reviewed enough transcripts to trust the pattern.
The opposite mistake is just as common: setting every topic to the same cautious threshold ends up routing simple, low-stakes questions to a person for no reason, which is exactly the workload the assistant was meant to take off the queue.
Never-automate topics skip the threshold entirely
Some topics shouldn't be judged on confidence at all โ they should go straight to a person regardless of how sure the assistant is. Cancellations tied to a retention conversation, anything involving a safety complaint, or a legal or compliance question are the usual candidates: not because the assistant would necessarily get them wrong, but because the right response depends on judgment a threshold can't capture.
Marking a topic never-automate is a stronger guarantee than a high threshold โ a confident wrong classification can still slip past a threshold; a never-automate rule can't.
Handoff is the part that decides whether escalation feels like a failure
What happens when the assistant hands off matters as much as when it decides to. A handoff that arrives with the full transcript, the customer record, and a note on what the assistant already tried means the person picking it up isn't starting from zero โ and the customer isn't asked to repeat themselves, which is the single fastest way to make someone angrier than the original problem did.
Escalation should also be available on request, not just triggered by a low confidence score โ a customer who explicitly asks for a person should get one immediately, whatever the assistant's confidence in its own answer.
Route the handoff, don't just release it
A handoff that lands in a generic queue is better than no handoff, but routing it to the right person or channel based on topic โ billing questions to whoever handles billing, technical issues to the team that can actually fix them โ cuts the time between escalation and a useful reply. That routing is worth setting up once thresholds are stable, rather than in the first week, since it depends on already knowing which topics tend to escalate.
Review and adjust, don't set once and leave it
Thresholds and never-automate lists are worth revisiting after the first few weeks of real transcripts, not just at setup. A topic that seemed risky enough to route to a person every time might turn out to have a narrow, consistent answer the assistant handles well โ and a topic that seemed safe might turn out to have more edge cases than expected. The resolution report is the place to see which topics are escalating and why; see the guide on reading it for what to look for.
