The approval gate
This is the part worth reading closely, because it is the part most likely to be built wrong — and when it is built wrong, everything still looks correct.
sentinel-agent’s whole claim is that irreversible actions pause for a human. That claim rests entirely on a mechanism it does not own: the harness derives approval from the annotations an MCP server publishes, and dispatches or pauses accordingly.
How the harness decides
Three predicates, evaluated against the annotations that come back from tools/list:
// trueforge-core/src/core/mcp/toolSelectors.ts
function isReadOnly(a?: ToolAnnotations) {
return a?.readOnlyHint === true;
}
function isWrite(a?: ToolAnnotations) {
return a?.readOnlyHint === false && a.destructiveHint !== true;
}
function isDestructive(a?: ToolAnnotations) {
return a?.destructiveHint === true;
}An agent’s require_approval_for_tools defaults to ["@write", "@destructive"]. Tools matching those tags pause. Everything else runs autonomously.
The failure mode this project is built around
So a rollback_deployment that forgot its annotations fires straight at production, silently. And nothing in review looks wrong: the tool is correct, the agent config is correct, the policy is correct. The gate just never triggers. There is no error, no warning, and no log line that says a call skipped approval — the successful path and the unsafe path are byte-identical from the outside.
The shipped bring-your-own-mcp cookbook example publishes zero annotations. Copying it as a template — which is exactly what you do when you are starting out — is how you inherit this.
Three layers, so it cannot happen here
- 1
Structural — you cannot register an unclassified tool
Every tool is created through
defineTool, which takesrisk: 'read' | 'write' | 'destructive'as a required field and derives the annotations from it. There is no code path that registers a tool without them, because there is no way to call the function without saying what kind of tool it is.typescript// risk is required — there is no overload without it defineTool({ name: 'rollback_deployment', risk: 'destructive', // ← derives the annotations below // readOnlyHint: false, destructiveHint: true ... }); - 2
Tested — against the harness's predicates, not our labels
registry.test.tsasserts that every tool is annotated, that its annotations match its declared risk, and that every production-mutating tool is named in the agent spec’s approval list — using TrueForge’s own selector functions rather than a local reimplementation of them. Add a destructive tool without classifying it and CI fails before review does. - 3
Belt and braces — literal names as well as tags
The agent spec names
rollback_deploymentandrestart_serviceliterally inrequire_approval_for_tools, alongside the tags. A literal name matches unconditionally, so the gate holds even if an SDK version drops annotations somewhere in transit.
curl -s -X POST http://localhost:8940/mcp -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'How approval actually resolves
There is no approval endpoint in TrueForge, and no approval id. This surprises people, so here is the whole flow.
1. The harness emits an event that tells you almost nothing
{
"type": "tool.approval_required",
"thread_id": "thr_9f2a",
"tool_calls": [
{ "id": "call_71c", "source_event_id": "evt_4410" }
]
}
// ↑ no tool name. no arguments. nothing to render.To render “Approve rollback of dpl-4c21?” a client must keep a Map<eventId, event> of everything it has seen, follow source_event_id back to the originating model.message, and find the matching tool call inside it. That index is mandatory, not an optimisation — without it there is nothing to put on the screen but an id.
2. Resolution is a new turn
POST /api/v1/sessions/{session_id}/turns
{
"input": [{
"type": "user.tool_approval",
"thread_id": "thr_9f2a",
"tool_call_id": "call_71c",
"approval": { "status": "allow" } // or "deny"
}]
}- One item per pending call.
- Approval items must never be mixed with user messages in the same turn — the harness returns 422.
- On a cold reload,
GET /turns/{turn_id}→state.required_actionsrecovers anything still pending, so a refresh mid-investigation does not strand the run.
The gate protects a path, not a tool
This is the consequence that a code review caught and that is worth internalising: approval is enforced by the harness. The MCP server knows nothing about it. So anything that reaches the MCP server directly never encounters the gate, because there is nothing there to encounter.
| Route | Passes the gate? | Control |
|---|---|---|
| agent → harness → MCP server | Yes — this is the gate | Annotations + policy |
| anything on the LAN → MCP server | No — bypasses it entirely | Binds 127.0.0.1; optional OPS_MCP_TOKEN bearer auth |
| a browser on another site → UI proxy | No — would approve on your behalf | Sec-Fetch-Site check |
| local curl → UI proxy | No — sends no Sec-Fetch-Site at all | Operator token x-sentinel-operator, fails closed |
Binding the MCP server to all interfaces therefore did not weaken the safety model — it offered a way around it entirely. That finding, and the five others from the same review, are written up in the repository’s Qodo section.