What a 96% AI Support Claim Means
An AI customer support agent is software that reads a customer’s request, retrieves the relevant account or policy information, replies or takes an approved action, and hands the case to a person when it should not decide alone. When you compare vendors, judge verified customer outcomes, the quality of that handoff and the full cost before the headline no-escalation rate. The TSA example below shows why.
Before the Transportation Security Administration deployed its AI customer support agent, TSA Ace, it built two less glamorous foundations. In Salesforce’s account of the rollout, a Customer Service Managers App joined traveler inquiries across email, phone and social channels; a TSA Cares App routed special-assistance requests. Ace now handles about 100,000 routine traveler conversations a month, and Salesforce reports that 96% of routine inquiries need no human escalation. Read the September 2026 case study.
The sequence matters more than the headline percentage. If you’re buying an AI support agent, compare independently verified outcomes, the handoff to a person and full cost before you compare no-escalation rates. The Salesforce report gives packing liquids and navigating a checkpoint with a medical device as examples of routine questions. Those questions may need different context and escalation rules. The reported 96% is a routine no-escalation measure, not an independent audit of correctness or satisfaction across every contact.
Before choosing Salesforce Agentforce Service, Zendesk AI Agents, Intercom Fin or a custom build, decide which requests the agent may handle, what counts as a successful outcome, and what happens when it is wrong.
What is an AI customer support agent, exactly?
An AI customer service agent reads the customer’s request, retrieves relevant information, replies or takes an approved action, and hands the case to a person when needed. Unlike a fixed FAQ bot, it may use account or order context and several steps of reasoning. That added ability also creates new failure modes: stale knowledge, wrong account data, an action outside its authority, or a handoff that forces the customer to repeat everything.
The important distinction is between answering and resolving. “Here is the refund policy” is an answer. A refund successfully issued under the right policy, with the customer’s consent and an auditable record, is an outcome. In a regulated or high-value case, escalating early may be the best outcome of all.
The 96% question: what did the number count?
Take three impressive-sounding claims from recent vendor-backed reports:
| Claim | What the public source actually describes | What it does not establish |
|---|---|---|
| TSA Ace: 96% | Routine traveler inquiries completed without a human escalation; the same release reports roughly 100,000 routine conversations monthly. Source | An independently audited correctness or satisfaction rate across every traveler contact. |
| Mobileye: 98% | A reported overall success rate for an internal engineering-support system built with Amazon Bedrock AgentCore. The case also says responses fell from hours to about one minute and describes more than 100 tickets monthly. Source | Performance for consumer-facing customer service, or proof that the same setup will work with your data. |
| Fin “resolution” | A defined billable outcome in Intercom’s pricing rules; under those rules, a user leaving after an answer may initially count as a resolution, with a later return to the same conversation changing the charge. Procedure handoff and disqualification can also be billable outcomes. Source | Your own measure of a durable customer fix, unless you reconcile billing events with reopenings and quality review. |
These are not comparable benchmarks. TSA is a public-facing service; Mobileye’s example is internal support; Fin’s outcome definition is a commercial event. Each can teach you something, but only if you keep its context attached.
For this guide, containment or no-escalation means a request stayed out of the human queue. A vendor-counted outcome follows that platform’s published rules. A verified resolution means the customer’s task was completed under criteria you set, with reopenings checked over a defined period. Vendor definitions vary; write down your own before comparing dashboards.
For your pilot, record four separate numbers: eligible requests, agent-handled requests, vendor-counted outcomes, and independently verified customer outcomes after a defined follow-up period. Add reopenings and human escalations. Otherwise the dashboard can improve while the queue quietly gets worse.
How should you choose an AI customer support agent?
Choose the operating model before the logo.
| Route | Named example | Best starting point | The question to put in the demo |
|---|---|---|---|
| Native service platform | Salesforce Agentforce Service or Zendesk AI Agents | Your cases, knowledge and human agents already live in that platform. | Can the agent use the same case history and permissions as your team, and hand off with its reasoning and actions visible? |
| Add-on to an existing helpdesk | Intercom Fin, including its supported external-helpdesk route | You want to keep the current ticketing system while testing automated answers. | Which outcome types are billed, what is the minimum commitment, and what happens when a case reopens? Intercom pricing FAQ |
| Custom agent runtime | Amazon Bedrock AgentCore as one possible building block | Your policies, data access and permitted actions do not fit a product’s standard workflow. | Who will own identity, tool permissions, observability, evaluation, incident response and integration maintenance? |

This is not a ranking. Salesforce and Zendesk have their own product and plan variations; Fin’s external-helpdesk route has separate commercial terms; AgentCore is infrastructure, not a finished helpdesk. Ask for a demonstration using your redacted cases, not a polished sample with a perfect knowledge base.
Run a case audit before a vendor pilot
Start with a deliberately small discovery sample, say 50 recent cases across your main channels. Fifty is enough to expose categories and edge cases, not enough to forecast a company-wide return. Include the awkward cases that sales demos avoid:
- Separate policy lookups from account-specific requests, transactions, complaints, sensitive data and emergencies.
- Mark the source of truth for each case: knowledge article, CRM record, order system, human judgment or unavailable information.
- Mark what the agent may do: answer only, draft for approval, update a record, or take a customer-facing action.
- Identify what must trigger handoff. Examples: the policy is ambiguous, the account identity is uncertain, the customer asks for a person, or the model cannot cite the right source.
- Record the current baseline: time to useful answer, reopen rate, customer effort, escalation reason and any quality score you already trust.
If half the sample depends on tribal knowledge or conflicting policy documents, fix that first. An agent will not turn contradictory instructions into a reliable service operation. Our agent-evaluation guide is useful for designing test cases; AI agent security practices covers the authority boundary behind any action-taking design.
Five questions to ask in a live demo
1. What happens when the knowledge source changes? Ask the vendor to update a policy, then show when the agent starts using it, what source the answer cites, and whether an old answer can be traced afterward.
2. Can it see the right customer context, and only that context? A product that can fetch an order status must prove account identity and permission boundaries. Run a negative test with two similar customers.
3. Does handoff carry the whole conversation? Ask the human agent to show the prior messages, retrieved facts, attempted actions and reason for escalation. A transfer that restarts the interview simply moves the work.
4. What is the billing event? Zendesk describes an automated-resolution allowance and paid tiers; Intercom publishes specific Fin outcome definitions. Salesforce has multiple service and Agentforce commercial routes. There is no safe single “price per agent” to compare across them. Get the current contract definition, exclusions, minimums, seats and overages in writing. See our broader AI agent pricing guide for the cost categories beyond the model call.
5. What can an auditor reproduce? Save the prompt or policy version, retrieved source, tool call, customer-visible answer, handoff and eventual case result. If the platform gives you only a “resolved” counter, you cannot investigate what improved.
The deceptively cheap $0.99 outcome
Intercom’s published Fin outcome guide lists $0.99 for several outcome types in its described chat/email-with-Intercom arrangement. It also lists $9.99 for sales qualification; an external-helpdesk setup has separate terms. That makes an excellent example of why one unit price is not a budget.
Imagine 1,000 incoming conversations. You judge 30% eligible for the pilot, or 300. If 60% of those produce a $0.99 billable outcome, that is 180 × $0.99 = $178.20 in that one usage component. This is an illustrative calculation, not a Fin quote or expected performance. It excludes any required subscription or commitment, implementation, integrations, human review, escalations and other billable outcome types. It also says nothing about how many customers received a durable fix. For a custom build, add runtime, model, logging, storage and operational ownership as well as initial engineering; AWS lists these cost components for AgentCore.
Your economic unit should be cost per independently verified useful outcome, with quality and customer effort beside it. This is more work than copying the vendor’s headline. It is also the number that survives a finance meeting.
A pilot that can actually change your decision
Give each shortlisted route the same bounded case set and success rules. Start in a non-customer-facing environment with redacted historical cases. Require correct answers with traceable sources, safe refusal on excluded cases, a complete handoff and a cost record. Have service managers review a blind sample. Only then expose a controlled slice of live traffic, with a human fallback and a clear stop condition for harmful answers or data leakage.
At the end, compare the routes on verified outcomes, reopenings, human minutes saved or added, customer effort, cost and operational burden. A vendor may win on immediate setup and lose on integration flexibility; a custom build may handle a distinctive workflow and lose on maintenance. Those are useful trade-offs, not disappointing results.
If your volume is low or your cases mostly ask for a stable FAQ, a cleaner help center, better search or a deterministic workflow may beat an AI agent. If your case data is fragmented, start with that integration problem. CRM automation is a separate, often simpler route for internal handoffs and record updates.
The right question is not “Which vendor has the highest resolution number?” It is “Which approach gives our customers the right answer or action, at a cost and risk we can explain?” If you want to scope that test against your own case data, Powercode Group’s AI and data team can help define the pilot and the integration boundary before you buy or build.
Sources and editorial notes
Primary sources are linked beside each changing claim. Vendor and joint customer case studies are identified as reported results, not independent verification. Product features, outcome definitions and prices were checked on 23 September 2026 and may change. The cited Mobileye deployment is internal engineering support; it is included only as a contrast in use case and measurement.