What is Autonomous Resolution Rate?
ObservabilityThe share of incoming requests a system resolves end to end without a person taking over. It separates a useful deployment from an expensive demo, and it is the easiest metric in this field to inflate.
Why It Matters
When a deployment is described by a number, this is usually the number. It answers the question the business actually asked: how much of this work no longer needs a person. That makes it the most consequential metric in the category and the one most likely to be quietly redefined on the way to a slide.
It is also the metric that has to survive contact with the customer. A system can resolve a request in the sense that the conversation ended, while the person resolved nothing and went elsewhere. The difference between those two outcomes is the entire argument for measuring outcomes rather than volume.
How It Works
The denominator. Requests that arrive in the channel, or requests that the system was eligible to handle? Both are defensible, and the number moves substantially depending on which you choose. Whichever it is, state it.
What counts as resolved. Not “the conversation closed”. Resolved means the customer’s request was satisfied: the order was traced, the refund issued, the policy question answered from the current source, and nobody had to reopen it.
The window. A request resolved on Monday that returns on Thursday was not resolved. The measurement window has to be long enough to catch the return, which is usually days rather than minutes.
The escalation path. Requests handed to a person are not failures, and hiding them flatters the number. Track them separately: how many, at what point, and whether the handoff was timely.
Cost per resolution. The rate on its own says nothing about whether the deployment pays. Pair it with the cost per interaction against the human baseline it replaced.
Where It Breaks
Deflection counted as resolution. The most common inflation, and the reason the metric has a credibility problem. A customer who gave up was deflected, not served.
Measuring the channel. Someone who asks the assistant, fails, and calls the contact centre appears as two separate events, one resolved and one new. The measurement has to follow the person across channels.
Volume as the win condition. A system optimised to close conversations quickly will close them quickly, including the ones that needed more time. The optimisation target is the thing being measured, so choose it deliberately.
A rate without the tail. An aggregate can hide a category where the system performs badly and a category where it performs well. The rate belongs alongside its breakdown by request type.
Claims that cannot be reproduced. If the number came from a vendor dashboard rather than the system’s own records, it is a claim rather than a measurement. Our published commerce figure is a client-reported outcome from an anonymized deployment, which is why it is labelled that way rather than presented as a benchmark.
How Flytebit Handles It
We define the metric before the build: the denominator, what counts as resolved, the window that catches returns, and how escalations are counted. Resolution is confirmed against the system of record rather than inferred from a closed conversation, and the rate is reported with its breakdown by request type and its cost per resolution alongside. The operating model behind it is our LLMOps work, and the deployment we publish figures for is on the Financial Services page.
More info
- 2026 state of AI in retail Reported containment bands by workflow, and which support categories they come from.