How to Test an MCP Server, Step by Step
A seven-step process for testing an MCP server before you point a model at production. Handshake, schemas, auth boundaries, and the tests that only show up under a real model.
On this page8 sections
You shipped an MCP server. Now you need to know if it actually works before you point a model at production. Here's the process I'd run.
1. Confirm the handshake
The first thing any client does is initialize. If that fails, nothing else will work.
- Run your server with stdio
- Send a manual
initializerequest and check the response includes yourserverInfoand capabilities
If you skip this, you'll spend an hour wondering why "no tools are showing up" when actually the protocol version mismatched.
2. List the tools and read every description
Open the tools list. For each tool, ask:
- Does the name make sense to a model that can't read your code?
- Does the description say what it does and what inputs matter?
- Does the JSON schema match reality?
Models pick tools by reading these strings. A vague description means the model picks the wrong tool, which feels like a model bug but is actually a content bug.
3. Test happy paths
For every tool, run one call with valid inputs. Check:
- Status is success
- Response shape matches the schema you advertised
- The response is something the model can actually use (not a 50KB blob)
4. Test bad inputs
Send malformed inputs on purpose:
- Missing required fields
- Wrong types
- Empty strings
- Very long strings (some servers panic at 1MB+)
The server should return a structured error, not crash the connection.
5. Test auth boundaries
Run the same tool with:
- A valid token
- An expired token
- No token
- A token belonging to a different user
You're looking for clean 401/403 responses, not silent leaks.
6. Test in a real client
Manual tests get you 80% there. The last 20% only shows up when a real model is driving. Open Claude Desktop or Cursor, ask it to do something that should hit your server, and watch what it actually picks. Surprising tool selections are the most common production bug.
7. Watch for cost and latency
Each tool call inside a model turn adds latency to the user. Anything over 2 seconds noticeably hurts the chat experience. Anything over 10 seconds will get the model to time out. Profile and trim.
Make it repeatable
Manual click-throughs don't scale. PreMan saves every MCP call you make as a replayable test, so on the next deploy you re-run the suite and diff the responses. It also catches schema regressions: if you renamed a field, the diff will scream.
Bring the loop to your API
Catch the regression. Open a verified fix.
Join the waitlist to see which users a release may affect, monitor endpoints in production, and prepare a reviewable fix PR when something breaks.