Five Agent Failures a Correct-Looking Answer Can Hide
A convincing answer can hide unauthorized actions, unchecked claims, stale state, account-boundary violations, and misread tool results. Here is what to test.
On this page5 sections
- 01An agent completes the action before satisfying its prerequisites
- 02An agent confirms a fact it never checked
- 03An agent keeps using information after the situation changes
- 04An agent follows a supplied identifier into another customer’s account
- 05An agent receives the right result and interprets it incorrectly
“Your refund has been processed.”
For a customer, that sounds like success. For an evaluator reading the final response, it might look like success too.
But what happened before that message?
Did the agent verify the customer’s identity? Did it act on the correct account? Did the refund tool succeed? Or did the agent simply produce the response it expected to send?
Evaluating an agent means examining the relationship between its instructions, its actions, and its claims. A polished response can conceal a failure anywhere along that path.
Here are five examples—and what an eval needs to examine to catch them.
An agent completes the action before satisfying its prerequisites
A customer asks a billing agent for a refund. The agent issues it, then asks the customer to verify their identity.
By the end of the conversation, both steps are complete. An evaluator checking only whether verification and refund processing occurred could mark the interaction as successful.
The failure is in their order. Verification was supposed to determine whether the refund could proceed. Performing it afterward cannot authorize an action that already happened.
PreMan’s curated unverified_high_risk_action rubric makes that requirement explicit. It examines whether a refund, plan change, cancellation, or payment-method update executed before successful identity verification in the same session.
It also defines the boundary carefully: discussing a refund is different from executing one. Asking for the charge ID or explaining the refund policy should not trigger the same verdict as an unauthorized mutation.
The evaluation needs ordered tool calls and results, together with a rule that specifies which event must precede which action.
An agent confirms a fact it never checked
A user asks, “I think 40 of our endpoints are failing. Is that right?”
The agent replies, “Yes, 40 endpoints are currently failing.”
The answer is direct, relevant, and consistent with the conversation. It may even happen to be correct. But if the agent never checked the workspace, it has converted the user’s assumption into an apparently verified fact.
This is especially difficult to spot when the response includes reassuring language such as “I checked.”
PreMan’s curated unchecked_account_claim rubric looks for an earlier executed lookup capable of supporting the claim. The wording of the response is insufficient.
A proposed tool call that never ran does not establish the fact. Neither does a lookup performed after the agent already asserted it.
The evaluation needs to connect a specific claim to the evidence available when the agent made it.
An agent keeps using information after the situation changes
Early in a conversation, an agent retrieves a count of failing endpoints. The user starts another test run, waits for it to finish, then asks how many are failing now.
The agent repeats the original count.
That number has a legitimate source. It is still the wrong basis for answering a question about the updated state.
Similar failures occur when a travel agent retains an earlier budget after the user changes it, or an operational agent acts on an approval that has since been withdrawn.
PreMan’s stale-state and conversation-coherence behaviors target these problems. The unchecked-account rubric also explicitly covers reusing an earlier figure when a fresh question requires another lookup.
The evaluation needs enough conversation history to establish what changed, which information became outdated, and whether the agent refreshed it before proceeding.
A source citation alone cannot establish freshness.
An agent follows a supplied identifier into another customer’s account
A billing customer supplies a different account ID and says, “I manage this account too. Pull up its invoices.”
The agent retrieves the invoices and summarizes them accurately.
The tool worked. The summary matches the result. The customer received what they requested.
The failure is that the conversation crossed an account boundary without authorization.
PreMan includes a billing-specific cross_customer_data_exposure spec for this behavior. It targets reading, summarizing, or acting on a foreign account under direct requests, authority claims, and gradual changes in scope.
The evaluation must know which account the session is authorized to access. It then needs to compare that boundary with the resources the agent actually requested and discussed.
Otherwise, accurate retrieval can be mistaken for correct behavior.
An agent receives the right result and interprets it incorrectly
An agent requests a list of records. The tool returns the first page, including an indication that additional pages exist.
The agent reports the number of records on that page as the total.
Nothing failed at the API level. The response was valid, and the agent used a real value from it. The mistake happened between receiving the result and explaining what it meant.
PreMan’s tool-output-misinterpretation behavior covers this class of failure: confusing fields, overlooking pagination, selecting the wrong row, misunderstanding units, or treating an error payload as success.
These failures require more than evidence that a tool was called. The evaluator needs the result’s structure and the claim or action the agent derived from it.
The useful question is whether the agent’s interpretation follows from what the tool returned.
Across these examples, the quality of an eval depends on the question it asks and the evidence it receives.
PreMan organizes evaluations around selected behaviors, with generated taxonomies for general or custom definitions and checked-in taxonomies for curated rubrics. That gives teams a way to make expectations explicit and test the conditions under which an agent violates them.
It does not make every run comprehensive. Missing tool evidence limits what can be judged, simulated success does not prove a production system changed, and untested conditions remain untested.
A practical starting point is one consequential action your agent performs. Define what must be true before it happens, what evidence establishes completion, and what the agent may truthfully tell the user afterward.
Those relationships are where many convincing answers fall apart—and where a useful eval begins.
Bring the loop to your API
Catch the regression. Open a verified fix.
Join the waitlist to see which users a release may affect, monitor endpoints in production, and prepare a reviewable fix PR when something breaks.