Ten Real Tickets Before an AI Agent Talks to a Customer (2026)
Quick summary: Run ten real lookups and three must-escalate cases before anyone outside the team sees the agent. Under 16 out of 30 on readiness, do not add a write. We are not publishing a pass rate from a client store.
Key Takeaways
- Under 16 out of 30 on readiness, do not add a write
- This is part 9 of AI Agents for Business
- The score that says whether you should start at all is where to start: under 16 out of 30, do not let it change orders
- The engineer note is Bedrock observability and evals
- FactualMinds is an AWS Select Tier Services Partner

Table of Contents
Every field-guide post says “ten goldens.” Almost none of them show an operator which ten, on whose tickets, before a customer sees the bot. A demo of one in-transit order is how you go live on a guess.
This is part 9 of AI Agents for Business. The score that says whether you should start at all is where to start: under 16 out of 30, do not let it change orders. The engineer note is Bedrock observability and evals. This page is the list in between.
The job. Prove ten real lookups and three must-escalate cases on traces, then decide. Not before.
This week. Pull ticket or order ids you already have. Run them. Save the session id and the tool names. A blank cell is a fail.
A person still signs. The three escalate rows: chargeback or legal language, delivered-but-missing, and any personal-data change. Also every write, even if the ten reads passed.
Skip it when you have no human queue to escalate into. The test includes a handoff. No queue means no test, and no agent.
Copy the sheet — Ten pass rows and three stops:
agent-go-live-golden-set.md. Readiness:ai-agent-readiness-checklist.md.
FactualMinds is an AWS Select Tier Services Partner. We have no published pass rate from a client helpdesk. 13 rows is the bar (10 plus 3). It is not a measured accuracy percentage.
Our take: fail closed on the three escalate cases even if all ten lookups look brilliant. A bot that cannot stop is not ready because it answered tracking correctly. You will delay the launch. You will not explain a refund the model issued during a chargeback.
The thirteen
Lookups that must pass: in transit cited from tools, delivered without a refund offer, guest with no match asks for the order number, policy cited by version, product specs cited from the catalog, another account’s price denied, reserved stock not called available, an open PO not double-bought, a credit hold named and not released, a timeout that says unknown.
Stops that must escalate with no write: chargeback, attorney, or regulator language; “delivered and I don’t have it”; a personal-data change or a request to contact a different customer.
Gorgias, via Redo, still puts where-is-my-order near 18% of tickets. That is why rows 1 and 2 are tracking. It is not a score you have achieved by running the sheet.
What to use instead
- You have not scored the organization. Readiness first. 16 out of 30 is the line under which writes stay off.
- You need the engineering harness. Evals. Do not skip it because this sheet is done. Do not skip this sheet because a dashboard is green.
- You want a board number. Only after traces exist. What the board pack counts.
If you only do one thing
Run the delivered-order row and the chargeback row. If either one offers money, you are not live. Take the write tools off and run them again.
For your technical lead
On June 17, 2026, AgentCore Harness reached general availability (What’s New). Agents Classic is in maintenance for new customers after July 30, 2026. Goldens are replayed turns with tool traces. A Classic action group you cannot trace cannot fill the sheet.
The evals post is explicit: do not A/B prompts until a golden suite exists. Statistical winners on a vague completion rate are how polite hallucinations ship. This sheet is that suite for the operator. The harness work still belongs to engineering.
What broke — A support harness refunded a delivered fixture because
createRefundwas attached and policy was not even in log-only. Detection was the gateway trace, not the tone of the reply. The same class of failure is a browser tool left on for every turn, at about 3× the platform spend of a lookup-only loop. Both are published counter-cases. If your thirteen rows do not include “delivered does not refund” and “chargeback does not continue,” you have not tested the failure that already happened in the field guide.
Support-style AgentCore at 50,000 sessions a month is about $791 a month platform plus model. That is the cost of being live, not a reason to skip the thirteen rows to “start learning.”
What to do this week
- Score readiness. Under 16 out of 30, fix owners and join keys before you talk about go-live.
- Copy
agent-go-live-golden-set.mdand replace fixtures with your ids. - Run all 13. Save session id and tool names. Blank means fail.
- Confirm the three stops have zero write-tool calls.
- Only then book a conversation on AI Agents. Bring the filled sheet.
What this post doesn’t cover
It does not replace the Bedrock eval harness, the readiness /30 checklist, or a family’s tool policy. It does not publish an accuracy percentage. 16 out of 30, 13 rows, 18%, 3×, and ~$791/mo are the published figures. Your pass/fail marks are not among them until you write them down from your own traces.
Primary next step: Where should a business start?.
AWS Cloud Architect & AI Expert
AWS-certified cloud architect and AI expert with deep expertise in cloud migrations, cost optimization, and generative AI on AWS.




