AI Chatbot Resolution Rates Explained

Updated September 14, 2026 · AI-assisted guide. Sources and illustrative examples are identified below.

There is no single resolution rate you should assume for an AI chatbot. The result depends on which questions it receives, what information it can use, and how you define success. Compare equivalent measures, then test the assistant against your own workload.

Editorial correction: This page previously presented 80% as a broadly achievable share of questions handled instantly. The page did not provide evidence for that claim. It has been replaced with the measurement framework below.

Choose a definition before looking at a percentage

A customer leaving a chat does not tell you whether the problem was solved. They might have found an answer, given up, called the office, or opened a ticket elsewhere. Treat the outcome as unknown unless your measurement process supports a stronger conclusion.

The following are working definitions for a small-business pilot. Reporting products may use different definitions, so document the formula beside every result you share.

MeasureWorking definitionWhat it does not prove
Bot participationEligible conversations in which the bot gave a substantive responseThat the response was correct
ContainmentConversations with no staff handoff recorded in that channelThat the customer completed the task
Verified resolutionA completed task or customer-confirmed answer, checked against a defined repeat-contact windowThat every unreviewed conversation succeeded
Ticket reductionChange in comparable ticket volume across all relevant channelsThat the chatbot caused the entire change
Answer qualityA reviewed answer meeting your accuracy, relevance, and source requirementsThat a subsequent action completed

How the denominator changes the result

Consider an illustrative month with 500 eligible customer conversations. A bot participates in 300. Of those, 180 end without a recorded staff handoff, while 120 meet your stricter verified-resolution definition. These are invented counts for explaining the calculation.

Containment among bot conversations is 180 ÷ 300 = 60%. Verified resolution among bot conversations is 120 ÷ 300 = 40%. Verified resolution across all eligible conversations is 120 ÷ 500 = 24%. All three calculations can be correct, yet they describe different things.

Report the counts with the rates: “120 verified resolutions from 300 bot conversations, out of 500 eligible conversations.” Keep the 60 contained conversations that lack verified outcomes in their own group. Do not quietly count them as successes.

Build a useful scorecard

Pick a consistent reporting period and record the channel, topic, total eligible volume, bot participation, verified resolution, staff handoffs, unknown outcomes, and repeat contacts. Add the total operating cost and staff review time so a rising resolution rate does not hide rising effort.

Review representative conversations, including failed and abandoned ones. For each answer, ask whether it used an approved source, answered the actual question, made an unsupported promise, and offered a useful route to a person. Track recurring errors by topic so the next content fix addresses a real problem.

Use the same test set before and after prompt, model, or knowledge changes. Include ordinary questions, ambiguous wording, outdated policies, missing information, and requests outside the assistant’s scope. This is a recommended evaluation method, not a guarantee that a test pass eliminates future errors.

Separate information from actions

Explaining a returns policy is different from authorizing a return. Describing appointment options is different from booking a time. In AI Engine, function calling uses functions executed by the surrounding application; generating a conversational reply does not itself perform the requested business action. See AI Engine’s function-calling documentation.

For an information-only assistant, a successful answer might include the correct policy and a working contact link. For a booking assistant, require evidence from the booking system before recording completion. If the connection fails, the response should say the booking did not complete and provide the next step.

Investigate changes before claiming a win

Ticket volume can change when sales volume, opening hours, website traffic, product mix, or support channels change. Compare similar periods and keep a note of launches, outages, and promotions. A seasonal Wisconsin retailer should be especially careful about comparing a quiet month with a holiday rush.

At low volume, report raw counts and individual problems alongside percentages. Two extra successes can move a small sample sharply. Extend observation when you need more evidence, and avoid treating an early result as a stable forecast.

When should a chatbot hand off?

Route requests when an approved source is missing, a policy conflicts, a customer asks for a person, or the task exceeds the assistant’s permissions. Make the handoff useful: explain the next step and include a short, accurate summary if the workflow supports one. A justified escalation can be a good customer outcome even though it lowers containment.

If you are planning a pilot, review AI Curdy’s implementation options and define the measurement plan before choosing a target percentage.

Keep exploring

Have a question this guide should answer? Send it to AI Curdy, or share this guide with the person responsible for the project.

Similar Posts