All field notes

How to Audit an MCP Server: The Eight Checks That Matter

A practical, hands-on audit you can run against any MCP server in under an hour. The eight checks that catch most production bugs and the patterns that keep showing up across the ecosystem.

On this page11 sections
  1. 01Before you start
  2. 021. The handshake
  3. 032. Discovery
  4. 043. Schema validity
  5. 054. Description quality
  6. 065. Happy-path execution
  7. 076. Bad-input handling
  8. 087. Auth boundaries
  9. 098. Latency and stability
  10. 10Patterns that show up over and over
  11. 11Run the audit yourself

The MCP ecosystem is young. Most servers work most of the time. The interesting question isn't "is this server up?" The better question is "is this server going to behave well when a real model is driving it?"

This is the audit we'd run against any MCP server before pointing a production agent at it. Eight checks, in order. Each one targets a class of bug that shows up regularly and tends to be invisible until you're already in production.

Before you start

You need:

  • An MCP client that lets you see raw protocol traffic (PreMan, Inspector, or your own harness)
  • Whatever credentials the server requires
  • About 45 minutes for a full pass

Don't do this in your head while staring at the source. Run it as a checklist with notes. The whole point is to catch things you'd miss by reading code.

1. The handshake

Send initialize. Read the response carefully.

  • Does the server return a sensible serverInfo.name and version?
  • Does it advertise the right capabilities for what it actually exposes? A server that says it has prompts but doesn't implement prompts/list is broken.
  • Is the protocol version compatible with your client?

Handshake failures are the cheapest bug to fix and the most embarrassing to ship.

2. Discovery

Run tools/list, plus resources/list and prompts/list if those are advertised.

  • Are tool names well-formed? [a-zA-Z0-9_-]+, no spaces.
  • Are there duplicates or near-duplicates that will confuse a model? search_user and find_user next to each other almost guarantees mis-selection.
  • Does the count match what you expect? A list that returns nothing usually means an auth bug or a registration bug.

3. Schema validity

For every tool, run inputSchema through a JSON Schema validator. This is the single most-skipped check and one of the highest-leverage.

What you're looking for:

  • Schemas that aren't valid JSON Schema at all (typos, wrong meta-schema)
  • Schemas that disagree with what the tool actually accepts (drift)
  • Required fields that the description treats as optional, or vice versa
  • Type mismatches (schema says number, tool expects a numeric string)

When schema disagrees with reality, the model constructs technically-valid calls that the server rejects. Users see a flaky agent. The fix is almost always five minutes once you find it.

4. Description quality

Open every tool's description and grade it against four criteria:

  • What does it do? Should be a verb phrase in the first sentence.
  • When should the model use it? A trigger condition, ideally with an example user query.
  • What does it return? Shape and rough size.
  • What should it not be used for? At least one near-confusable tool ruled out.

A description under 30 characters almost always fails this. So does a description that just restates the tool name. Rewriting these is the highest-ROI change you can make to a working MCP server.

5. Happy-path execution

For each tool, synthesize valid inputs from the schema and call it once.

  • Status: success
  • Response shape matches what the schema and description claimed
  • Response size: under a few KB for typical tools, under 50KB for "list" tools with limit

Watch for "list" tools that take no limit and return everything. Those are silent killers. They'll fit in your test session and explode in production when the user has 10,000 records.

6. Bad-input handling

Send malformed inputs on purpose:

  • Missing required fields → expect 4xx with a structured error
  • Wrong types → expect 4xx, not a 500 or a hang
  • Empty strings → expect either accept-and-return-empty or 4xx, never silent garbage
  • Extremely long strings (1MB+) → server should reject early, not OOM

A server that crashes on bad input takes the whole MCP connection with it. The client has to reconnect. The user sees a flicker and a confused model.

7. Auth boundaries

For hosted servers requiring auth, run each tool with:

  • A valid token from user A
  • An expired token
  • No token at all
  • A valid token from user A, but arguments referencing user B's data

You're looking for clean 401, 403, and "not found" responses. Anything that returns user B's data when called by user A is a tenant-isolation bug, and those have ended companies.

8. Latency and stability

Run 20 calls per tool back-to-back. Note P50 and P95.

  • P95 under 1 second: fine
  • P95 between 1 and 5 seconds: livable but worth profiling
  • P95 over 10 seconds: effectively broken inside a chat interface; the model will time out

Then run 1,000 sequential calls against the busiest tool. Watch for memory leaks, dropped connections, or response-time drift. A server whose latency climbs steadily over a long sequential run almost always has a leak somewhere.

Patterns that show up over and over

After running this audit against enough servers, the same handful of issues keep appearing. If you only have time to fix three things on your own server, fix these:

  1. Tool descriptions that are too short. "Searches data" is not a description. Add the trigger, the return shape, and the near-neighbors to rule out.
  2. List tools without a limit parameter. Add one. Default to a small number. Cap server-side regardless of what the client passes.
  3. Schemas that drift from reality. Validate them in CI. Keep them as the source of truth, generate Zod/Pydantic from the schema rather than the other way around.

These three account for the majority of "the agent is being weird" reports.

Run the audit yourself

PreMan ships a saved version of this audit you can point at any MCP server. It runs the eight checks, scores each tool, and tracks scores over time so you can see whether you're improving on each release.

→ Run the audit suite against your MCP server in PreMan

Bring the loop to your API

Catch the regression. Open a verified fix.

Join the waitlist to see which users a release may affect, monitor endpoints in production, and prepare a reviewable fix PR when something breaks.