MCP tool surface
Thirteen tools over streamable HTTP. The split between them is not a naming convention — it is the field that decides whether a call can happen while you are getting coffee.
Read-only tools run autonomously. Anything that writes or destroys is gated. Investigation should never need a click; remediation always should.
| Tool | Risk | Gated | What it does |
|---|---|---|---|
get_incident | @read-only | — | The incident record, its severity and its notes. |
list_incidents | @read-only | — | Everything currently open on the estate. |
get_service_health | @read-only | — | Live deployment id, replica counts, and the named checks. |
list_recent_deployments | @read-only | — | Deployment history for a service, newest first. |
get_deployment | @read-only | — | One deployment: version, author, changed files. |
get_deployment_diff | @read-only | — | The unified diff. Where mechanism comes from. |
export_metrics_csv | @read-only | — | Raw golden-signal samples. Deliberately no analysis. |
preview_remediation | @read-only | — | What a destructive call would change, computed without doing it. Free to call. |
post_incident_note | @write | yes | Append a finding to the incident. Mutates shared state, so it asks. |
record_finding | @write | yes | The conclusion as structure: claims paired with sources, confidence, what was ruled out. |
audit_finding | @write | yes | A second reviewer scores the evidence, not the conclusion. Identity is self-declared. |
rollback_deployment | @destructive | yes | Redeploy the previous version. The action the whole product is arranged around. |
restart_service | @destructive | yes | Roll the pods. Cheap-looking, still production. |
Why export_metrics_csv returns nothing useful
It returns raw samples and no analysis, and that is the single most deliberate decision in the tool surface. A tool that answered “p95 is up 3.7x” would let the agent report a number it never derived, and the sandbox would become decoration. Returning samples forces the change point, the settled means and the ratio to be computed — which is the part an operator can actually check.
How a tool is registered
export const rollbackDeployment = defineTool({
name: 'rollback_deployment',
risk: 'destructive', // required — no overload without it
description: 'Redeploy the previously live version of a service.',
inputSchema: { deployment_id: z.string() },
handler: async ({ deployment_id }) => { ... },
});
// annotations are derived, never hand-written:
// read → { readOnlyHint: true }
// write → { readOnlyHint: false }
// destructive → { readOnlyHint: false, destructiveHint: true }riskis required. There is no overload without it, so “forgot the annotations” is not a state this codebase can reach.- Annotations are derived from
risk, never written by hand, so they cannot disagree with it. registry.test.tsre-checks both against the harness’s own predicates, and against the agent spec’s approval list.
The eleventh tool
There is one more tool in the registry that is not in the table above: rollback_deployment_unsafe, a deliberately unannotated twin of the real thing. It exists so the conformance suite can drive the exact bypass this project is built around and observe what the harness does, rather than asserting what it would do.
npm run dev:mcp:labVerify the surface yourself
curl -s -X POST http://localhost:8940/mcp -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'npm test --workspace @sentinel-agent/mcp-server