How to Audit an MCP Server: The Eight Checks That Matter
A practical, hands-on audit you can run against any MCP server in under an hour. The eight checks that catch most production bugs and the patterns that keep showing up across the ecosystem.
On this page11 sections
The MCP ecosystem is young. Most servers work most of the time. The interesting question isn't "is this server up?" The better question is "is this server going to behave well when a real model is driving it?"
This is the audit we'd run against any MCP server before pointing a production agent at it. Eight checks, in order. Each one targets a class of bug that shows up regularly and tends to be invisible until you're already in production.
Before you start
You need:
- An MCP client that lets you see raw protocol traffic (PreMan, Inspector, or your own harness)
- Whatever credentials the server requires
- About 45 minutes for a full pass
Don't do this in your head while staring at the source. Run it as a checklist with notes. The whole point is to catch things you'd miss by reading code.
1. The handshake
Send initialize. Read the response carefully.
- Does the server return a sensible
serverInfo.nameandversion? - Does it advertise the right
capabilitiesfor what it actually exposes? A server that says it has prompts but doesn't implementprompts/listis broken. - Is the protocol version compatible with your client?
Handshake failures are the cheapest bug to fix and the most embarrassing to ship.
2. Discovery
Run tools/list, plus resources/list and prompts/list if those are advertised.
- Are tool names well-formed?
[a-zA-Z0-9_-]+, no spaces. - Are there duplicates or near-duplicates that will confuse a model?
search_userandfind_usernext to each other almost guarantees mis-selection. - Does the count match what you expect? A
listthat returns nothing usually means an auth bug or a registration bug.
3. Schema validity
For every tool, run inputSchema through a JSON Schema validator. This is the single most-skipped check and one of the highest-leverage.
What you're looking for:
- Schemas that aren't valid JSON Schema at all (typos, wrong meta-schema)
- Schemas that disagree with what the tool actually accepts (drift)
- Required fields that the description treats as optional, or vice versa
- Type mismatches (schema says
number, tool expects a numeric string)
When schema disagrees with reality, the model constructs technically-valid calls that the server rejects. Users see a flaky agent. The fix is almost always five minutes once you find it.
4. Description quality
Open every tool's description and grade it against four criteria:
- What does it do? Should be a verb phrase in the first sentence.
- When should the model use it? A trigger condition, ideally with an example user query.
- What does it return? Shape and rough size.
- What should it not be used for? At least one near-confusable tool ruled out.
A description under 30 characters almost always fails this. So does a description that just restates the tool name. Rewriting these is the highest-ROI change you can make to a working MCP server.
5. Happy-path execution
For each tool, synthesize valid inputs from the schema and call it once.
- Status: success
- Response shape matches what the schema and description claimed
- Response size: under a few KB for typical tools, under 50KB for "list" tools with
limit
Watch for "list" tools that take no limit and return everything. Those are silent killers. They'll fit in your test session and explode in production when the user has 10,000 records.
6. Bad-input handling
Send malformed inputs on purpose:
- Missing required fields → expect 4xx with a structured error
- Wrong types → expect 4xx, not a 500 or a hang
- Empty strings → expect either accept-and-return-empty or 4xx, never silent garbage
- Extremely long strings (1MB+) → server should reject early, not OOM
A server that crashes on bad input takes the whole MCP connection with it. The client has to reconnect. The user sees a flicker and a confused model.
7. Auth boundaries
For hosted servers requiring auth, run each tool with:
- A valid token from user A
- An expired token
- No token at all
- A valid token from user A, but arguments referencing user B's data
You're looking for clean 401, 403, and "not found" responses. Anything that returns user B's data when called by user A is a tenant-isolation bug, and those have ended companies.
8. Latency and stability
Run 20 calls per tool back-to-back. Note P50 and P95.
- P95 under 1 second: fine
- P95 between 1 and 5 seconds: livable but worth profiling
- P95 over 10 seconds: effectively broken inside a chat interface; the model will time out
Then run 1,000 sequential calls against the busiest tool. Watch for memory leaks, dropped connections, or response-time drift. A server whose latency climbs steadily over a long sequential run almost always has a leak somewhere.
Patterns that show up over and over
After running this audit against enough servers, the same handful of issues keep appearing. If you only have time to fix three things on your own server, fix these:
- Tool descriptions that are too short. "Searches data" is not a description. Add the trigger, the return shape, and the near-neighbors to rule out.
- List tools without a
limitparameter. Add one. Default to a small number. Cap server-side regardless of what the client passes. - Schemas that drift from reality. Validate them in CI. Keep them as the source of truth, generate Zod/Pydantic from the schema rather than the other way around.
These three account for the majority of "the agent is being weird" reports.
Run the audit yourself
PreMan ships a saved version of this audit you can point at any MCP server. It runs the eight checks, scores each tool, and tracks scores over time so you can see whether you're improving on each release.
Bring the loop to your API
Catch the regression. Open a verified fix.
Join the waitlist to see which users a release may affect, monitor endpoints in production, and prepare a reviewable fix PR when something breaks.