The safety model

The approval gate

This is the part worth reading closely, because it is the part most likely to be built wrong — and when it is built wrong, everything still looks correct.

sentinel-agent’s whole claim is that irreversible actions pause for a human. That claim rests entirely on a mechanism it does not own: the harness derives approval from the annotations an MCP server publishes, and dispatches or pauses accordingly.

How the harness decides

Three predicates, evaluated against the annotations that come back from tools/list:

typescript
// trueforge-core/src/core/mcp/toolSelectors.ts
function isReadOnly(a?: ToolAnnotations) {
  return a?.readOnlyHint === true;
}
function isWrite(a?: ToolAnnotations) {
  return a?.readOnlyHint === false && a.destructiveHint !== true;
}
function isDestructive(a?: ToolAnnotations) {
  return a?.destructiveHint === true;
}

An agent’s require_approval_for_tools defaults to ["@write", "@destructive"]. Tools matching those tags pause. Everything else runs autonomously.

The failure mode this project is built around

How a missing annotation disables the approval gateANNOTATEDrollback_deploymentMCP tooldestructiveHint: trueannotations@destructivederived tagin policy listrequire_approval_for_toolsPAUSEShuman decidesUNANNOTATEDrollback_deploymentidentical tool— none published —annotationsmatches nothingnot even @read-onlynot in policy listnothing to matchEXECUTESno prompt at allSame tool. Same production system. The difference is four lines of metadata.

So a rollback_deployment that forgot its annotations fires straight at production, silently. And nothing in review looks wrong: the tool is correct, the agent config is correct, the policy is correct. The gate just never triggers. There is no error, no warning, and no log line that says a call skipped approval — the successful path and the unsafe path are byte-identical from the outside.

The shipped bring-your-own-mcp cookbook example publishes zero annotations. Copying it as a template — which is exactly what you do when you are starting out — is how you inherit this.

Three layers, so it cannot happen here

  1. 1

    Structural — you cannot register an unclassified tool

    Every tool is created through defineTool, which takes risk: 'read' | 'write' | 'destructive' as a required field and derives the annotations from it. There is no code path that registers a tool without them, because there is no way to call the function without saying what kind of tool it is.

    typescript
    // risk is required — there is no overload without it
    defineTool({
      name: 'rollback_deployment',
      risk: 'destructive',        // ← derives the annotations below
      // readOnlyHint: false, destructiveHint: true
      ...
    });
  2. 2

    Tested — against the harness's predicates, not our labels

    registry.test.ts asserts that every tool is annotated, that its annotations match its declared risk, and that every production-mutating tool is named in the agent spec’s approval list — using TrueForge’s own selector functions rather than a local reimplementation of them. Add a destructive tool without classifying it and CI fails before review does.

  3. 3

    Belt and braces — literal names as well as tags

    The agent spec names rollback_deployment and restart_service literally in require_approval_for_tools, alongside the tags. A literal name matches unconditionally, so the gate holds even if an SDK version drops annotations somewhere in transit.

see the annotations yourself
curl -s -X POST http://localhost:8940/mcp -H 'Content-Type: application/json' -H 'Accept: application/json, text/event-stream' -d '{"jsonrpc":"2.0","id":1,"method":"tools/list","params":{}}'

How approval actually resolves

There is no approval endpoint in TrueForge, and no approval id. This surprises people, so here is the whole flow.

1. The harness emits an event that tells you almost nothing

json
{
  "type": "tool.approval_required",
  "thread_id": "thr_9f2a",
  "tool_calls": [
    { "id": "call_71c", "source_event_id": "evt_4410" }
  ]
}
// ↑ no tool name. no arguments. nothing to render.

To render “Approve rollback of dpl-4c21?” a client must keep a Map<eventId, event> of everything it has seen, follow source_event_id back to the originating model.message, and find the matching tool call inside it. That index is mandatory, not an optimisation — without it there is nothing to put on the screen but an id.

2. Resolution is a new turn

http
POST /api/v1/sessions/{session_id}/turns
{
  "input": [{
    "type": "user.tool_approval",
    "thread_id": "thr_9f2a",
    "tool_call_id": "call_71c",
    "approval": { "status": "allow" }   // or "deny"
  }]
}
  • One item per pending call.
  • Approval items must never be mixed with user messages in the same turn — the harness returns 422.
  • On a cold reload, GET /turns/{turn_id} → state.required_actions recovers anything still pending, so a refresh mid-investigation does not strand the run.

The gate protects a path, not a tool

This is the consequence that a code review caught and that is worth internalising: approval is enforced by the harness. The MCP server knows nothing about it. So anything that reaches the MCP server directly never encounters the gate, because there is nothing there to encounter.

RoutePasses the gate?Control
agent → harness → MCP serverYes — this is the gateAnnotations + policy
anything on the LAN → MCP serverNo — bypasses it entirelyBinds 127.0.0.1; optional OPS_MCP_TOKEN bearer auth
a browser on another site → UI proxyNo — would approve on your behalfSec-Fetch-Site check
local curl → UI proxyNo — sends no Sec-Fetch-Site at allOperator token x-sentinel-operator, fails closed

Binding the MCP server to all interfaces therefore did not weaken the safety model — it offered a way around it entirely. That finding, and the five others from the same review, are written up in the repository’s Qodo section.