All field notes

How to Test AI Agents That Call Your APIs

An AI agent that calls your APIs is two products glued together. Both can fail. Here is the three-layer test plan that catches bugs at the seam.

On this page4 sections
  1. 01Three failure modes worth testing
  2. 02A practical test plan
  3. 03What to log
  4. 04Don't skip the boring layer

An AI agent that calls your APIs is two products glued together: the model's reasoning and your endpoint's behavior. Both can fail. Testing each in isolation isn't enough; the bugs live at the seam.

Three failure modes worth testing

1. The model picks the wrong tool. You have get_user and search_users. The agent calls search_users for a single known ID, gets 50 results, and confuses itself. This is a description bug, not an endpoint bug.

2. The model fills in bad arguments. The agent calls create_invoice with amount: "twenty dollars" instead of amount: 2000. Your endpoint returns a 400. The agent retries. The user sees a five-second hang.

3. The endpoint returns something the model can't reason about. Your endpoint returns a 50KB JSON blob with 80 fields. The model truncates, drops the relevant field, and hallucinates a reply.

A practical test plan

Layer 1: tool-level tests. For each tool, run the same harness you'd run on any API. Valid inputs, invalid inputs, auth boundaries, response shape. These are the cheapest to run and catch the most bugs.

Layer 2: tool-selection tests. Write 20 to 50 user prompts. For each, record which tool you'd expect the model to pick. Run them through your model with your tools registered. Diff actual vs expected. The fixes here are usually description rewrites, not code changes.

Layer 3: end-to-end agent tests. Run a realistic multi-turn conversation that should require two or three tool calls. Score the final answer for correctness. These are slow and expensive; keep the suite small (5 to 15 scenarios) and run on PR merge, not on every commit.

What to log

  • Every tool call's name, arguments, and response
  • Token count per turn
  • Latency per tool call
  • The model's chosen tool versus the expected one

Don't skip the boring layer

Most teams jump to layer 3 because it's the most fun. Then they spend a week debugging an agent failure that was a layer-1 bug all along (a 500 from a bad fixture). Build in order.

PreMan handles all three layers in one workspace. Endpoint tests run as request collections. Tool-selection tests run as agent simulations with your registered tools. End-to-end tests run as full conversations with replay.

→ Build your agent test suite in PreMan

Bring the loop to your API

Catch the regression. Open a verified fix.

Join the waitlist to see which users a release may affect, monitor endpoints in production, and prepare a reviewable fix PR when something breaks.