All field notes

How to Write API Descriptions That LLMs Actually Understand

The single highest-leverage thing you can change to make AI agents pick the right tool. With before/after examples and a checklist you can ship today.

On this page7 sections
  1. 01What the model sees
  2. 02The four jobs of a good description
  3. 03Before and after
  4. 04Parameter descriptions matter as much as the tool description
  5. 05Six patterns I see go wrong
  6. 06A checklist you can ship today
  7. 07How to know if it's working

The model isn't reading your code. It's reading three short strings: the tool name, the description, and each parameter's description. Those strings are the entire surface the model uses to decide whether to call your tool, and what to pass.

Most teams underrate this. They spend a week tuning prompts and ten minutes copy-pasting JSDoc into their tool descriptions. Then they wonder why the agent is broken.

Here's how to write descriptions that actually work, with the patterns I see go wrong most often.

What the model sees

For a typical OpenAI or Anthropic tool definition, the relevant fields are:

{
  "name": "search_orders",
  "description": "Find recent orders by customer email.",
  "parameters": {
    "type": "object",
    "properties": {
      "email": { "type": "string", "description": "The customer email." },
      "limit": { "type": "integer", "description": "Max results." }
    },
    "required": ["email"]
  }
}

That's it. The model never sees your function body, your test cases, your README, or your Slack channel. If a fact about your tool isn't in those strings, the model will guess.

The four jobs of a good description

A description has to answer four questions, in order:

  1. What does the tool do? (One verb phrase.)
  2. When should the model use it? (The trigger condition.)
  3. What does it return? (Shape and rough size.)
  4. What should the model not use it for? (The nearest-neighbor confusable.)

Skip any of these and the model fills in the gap with a plausible guess. Plausible is rarely correct.

Before and after

Before. The kind of description that ships when nobody is paying attention.

description: "Search orders."

The model knows there is a tool called search_orders that does something with orders. That's about it. If you also have list_orders, get_order, and find_invoice, expect chaos.

After. Same tool, descriptions that actually steer the model.

description: |
  Find recent orders for a single customer, looked up by email.
  Use when the user asks about one customer's purchase history
  ("what did Alice buy last month?", "did this customer order
  anything this year?"). Returns up to 20 orders, newest first,
  with id, total, status, and date.

  Do NOT use this for:
  - Listing all orders across customers (use list_orders).
  - Looking up a single specific order ID (use get_order).
  - Invoices or refunds (use search_invoices, search_refunds).

That description is doing real work. It triggers on a clear user intent. It rules out three confusable cousins. It tells the model what the response will look like, so the model can plan its next step.

Parameter descriptions matter as much as the tool description

A common pattern: a great tool description, then email: "The email." for every parameter. The model passes whatever string looks email-shaped, often the wrong one.

A better parameter description gives the model a rule, an example, and a constraint:

email:
  description: |
    The customer's primary email address. Must be a valid RFC 5322
    email. If the user gave you a name instead of an email, do not
    call this tool; ask for the email or use search_customers first.
  examples: ["alice@example.com"]

That last sentence saves you from a class of failures where the model invents an email from a name.

Six patterns I see go wrong

1. Restating the function name as the description. get_user: "Gets a user." This is zero information. Delete and rewrite.

2. Engineering-internal vocabulary. "Returns the user record from the auth-service shard." The model has no idea what your shards are. Describe behavior in user-facing terms.

3. Optional parameters that look required. If a parameter's description ends with "if applicable," the model will try to fill it every time. Write "Optional. Only pass if the user explicitly asks for X."

4. Tools that overlap. search_customers and find_customers will be picked roughly at random. Pick one name and delete the other, or write descriptions that explicitly partition the space.

5. Walls of disclaimers. The model reads top-down and weighs early text more. Don't bury the trigger condition under three paragraphs of legal warnings.

6. Lying about the response. If your description says "returns up to 20" but the endpoint can return 2,000, you'll blow the model's context window and it will hallucinate a summary. Make the description match the contract.

A checklist you can ship today

Before you call your tool definitions done, run each one through this:

  • Description starts with a verb that tells the model what happens
  • At least one example user query that should trigger this tool
  • At least one near-confusable tool ruled out
  • Response shape and size described in plain text
  • Every required parameter has an example value
  • Optional parameters say "Optional" as the first word
  • No code-internal terms a user would never use

How to know if it's working

The honest answer is: run the tool with realistic prompts and watch what the model picks. PreMan lets you register a set of tool definitions and a list of prompts, then runs each prompt through the model and shows you which tool got selected. When you see a wrong selection, you don't change the prompt; you change the description and rerun. It's the fastest way to turn vague "the agent feels weird" reports into specific, fixable description bugs.

→ A/B test your tool descriptions in PreMan

Bring the loop to your API

Catch the regression. Open a verified fix.

Join the waitlist to see which users a release may affect, monitor endpoints in production, and prepare a reviewable fix PR when something breaks.