Skip to content

Agents (MCP)

wirebench mcp serves one project to any MCP client: a coding agent can import a WSDL or an OpenAPI document, list the operations, generate a sample request, send a saved request, validate and query a response, and compare two responses from History. No model runs inside Wirebench; the agent is yours, and the server only answers its tool calls. Nothing is sent anywhere except the requests you or the agent ask send to make.

Most MCP clients start a local server as a command. Give it the project folder as an absolute path:

{
"mcpServers": {
"wirebench": {
"command": "npx",
"args": ["-y", "@wirebench/cli", "mcp", "--project", "/path/to/project", "--allow-send", "-e", "<environment>"]
}
}
}

Replace <environment> with the name of an environment your project defines; send may then use only that one (a comma-separated list names several, and leaving -e out allows any). The client starts the process, so it chooses the working directory, and it ends the server by closing its input. A relative path in a tool call (an import source, a file to validate) resolves against that directory: tell the agent to use absolute paths.

wirebench mcp --http 8931 --project /path/to/project serves Streamable HTTP on http://127.0.0.1:8931/mcp, on that loopback address only. Set WIREBENCH_MCP_TOKEN before starting it, or copy the token it prints once to stderr, and give the client the URL and the header. A token you set must be at least 16 characters with no spaces; a shorter one stops the server from starting.

{
"mcpServers": {
"wirebench": {
"url": "http://127.0.0.1:8931/mcp",
"headers": { "Authorization": "Bearer <token>" }
}
}
}

A request without the token gets 401, and one from a web page on another origin, or with a Host other than 127.0.0.1 or localhost on that port, gets 403. Each client session has its own server with the same gates. One process holds at most 64 live sessions: past that, a new client’s initialize closes the session idle longest (one with no GET stream open first), and that client must initialize again. A request that is not a real initialize closes nobody’s session. All clients share the one token, so treat it as one identity. Any path other than /mcp gets 404, and --http 80 is refused (a port from 1 to 65535 otherwise).

Tool Does Needs
operations Lists SOAP operations and REST endpoints, with the references the other tools take and the saved requests send takes. none
generate Builds a sample SOAP envelope or REST JSON body for one operation. Saves nothing. none
send Sends one saved SOAP or REST request as wirebench run does, and records it in History. --allow-send
validate Checks a message against the WSDL schema, or a REST response against the OpenAPI response schema. none
query Runs XPath on XML or JSONPath on JSON. none
history_list Lists recent sends, the app’s and the agent’s. none
history_diff Diffs two responses semantically, ignoring the paths you name. none
import Adds a WSDL or OpenAPI document to the project. --allow-write

Every tool is always listed. One whose flag is off answers with an error that names the flag, so the agent can tell you what to allow. -e local,staging limits the environments send may use, and with -e a send that would use no environment is refused.

Both flags apply to the contract tools below as well. Each one needs --allow-send, and -e limits the environments it may use in the same way, refusing a call that would use no environment.

validate and query read a History entry (historyId), a file (file) or the text itself (text); pass exactly one. A file is any regular file of 16 MiB or less that the server process can read; any file in the History folder in use is refused, so read History through historyId. send also takes a body that replaces the saved body for that one send. It is sent as written, so a placeholder such as ${…} in it is refused; the saved request’s own body still expands placeholders as usual.

import reads any local path the server can read, or fetches any http(s) URL that its network can reach, so allow it only for a project and a machine you are content to have the agent work on. The desktop is stricter (it imports from project roots and files you pick); the server is not.

A few limits keep a result small: query returns at most 64 Ki characters per result and 256 Ki in all, history_diff at most 256 Ki characters of changes, and both say truncated when they cut. send cuts a response body at 256 Ki characters and sets bodyTruncated. history_list returns 20 entries unless asked for more, up to 200.

REST is covered in part: validate checks responses only, using the status from the History entry (or the status you pass, else 200), and generate gives a REST operation’s OpenAPI path template. A REST body validate could not compare with a schema (none declared, over 1 MiB, or out of time) comes back valid with checked: false, and its contract says why; valid is false only when problems were found. Both SOAP and REST list at most 50 problems and say truncated when there were more. gRPC and WebSocket requests are listed in the project but send refuses them with unsupported-kind.

Pass baseline: true to send to compare the response with the golden response saved beside the request in the Snapshot tab. The comparison is by meaning and honours the golden’s ignore rules, as wirebench run --baseline does in CI.

The result gains a baseline field:

status Meaning
matched The response matches; ignored counts the differences the ignore rules hid.
differs changes lists each kind (changed, added, removed), path, expected and actual, the first 100; truncated says there were more. The send’s outcome is failed.
missing No baseline saved for this request. The send is judged on its own assertions.
unreadable, too-large The golden could not be read, or a body is over 2 MB. The send errors (exit 3 from wirebench send).
unsupported A WebSocket request: there is no golden to compare.

A body override can be combined with it: the saved request’s golden still applies, so you can change a request and check that the response did not change. Secret values the send resolved are masked across the whole result, changes included.

An agent calls it like any tool:

{ "item": "Pets/pets/List pets", "environment": "local", "baseline": true }

and, when the response has drifted, reads (each value is JSON-encoded, so a string keeps its quotes):

{
"outcome": "failed",
"baseline": {
"status": "differs",
"format": "json",
"ignored": 0,
"changes": [{ "kind": "changed", "path": "/0/name", "expected": "\"Rex\"", "actual": "\"Fido\"" }]
}
}

Besides the tools above, every imported operation is a tool of its own: each SOAP operation of an interface and each endpoint of an OpenAPI-backed API. The agent calls it with typed JSON arguments and gets the response back as JSON, so it can use a legacy SOAP service without writing XML. Each call needs --allow-send, goes through the interface’s or API’s endpoint, auth and secrets, and lands in History.

A tool is named after its container and operation, in snake_case: the Add operation of the CalculatorService interface is calculator_service_add. operations shows each operation’s tool name, and wirebench call <name> --schema prints its arguments’ schema.

An agent calls it like any tool:

{ "name": "calculator_service_add", "arguments": { "environment": "local", "a": 2, "b": 3 } }

and reads:

{
"tool": "calculator_service_add",
"operation": "CalculatorService/Add",
"kind": "soap",
"status": 200,
"statusText": "OK",
"ok": true,
"durationMs": 41,
"headers": { "content-type": "text/xml; charset=utf-8" },
"result": { "result": 5 },
"notes": [],
"historyId": "01K6…"
}

A SOAP fault is a normal result with ok: false and a fault; arguments that do not fit the schema, or a body the XSD refuses, are refused before anything is sent.

A server serves at most 128 of these tools. A larger project refuses to start until --tools names the interfaces and APIs to serve (--tools none serves only the tools above). The server watches the project, so after the agent imports a WSDL the new operations appear as tools without a restart.

A send from an agent lands in the same History as the app’s, tagged mcp, and an open History panel shows it at once. history_list and history_diff read it back, so an agent can compare today’s response with yesterday’s. A send that got no response is not written. If the History file is busy, the send’s result still comes back, without a historyId, together with a warning. An agent’s send never trims History below what the file already holds, so it cannot drop entries you keep under a larger History cap than the default; the app’s own cap applies again on its next write.

By default the server uses the desktop’s History folder. Point --history-dir at another folder to keep an agent’s sends apart from the app’s, or to reach the folder of a development build.

The same capabilities are terminal verbs (wirebench send, validate and the rest). They exit 0 on success, 1 when a send assertion failed or validate found the message invalid (a REST body it could not check exits 0 and prints not checked: <reason>), 2 for a refused or wrong call, and 3 for everything else.

See the CLI reference for every flag, and the security notes for what the server allows.