Get started

About sentinel-agent

An agent that takes a production incident from alert to prepared fix without asking you anything — and then stops dead before the one action that matters, and asks.

When checkout latency triples, an engineer opens five tabs. Dashboards, to see the shape of it. The deploy log, to see what changed. GitHub, to read the diff. A terminal, to work out whether the change is big enough to explain it. And then a decision — roll back or keep digging — taken under time pressure on partial evidence.

The investigation is mechanical. The decision is not. Almost every attempt to automate this goes wrong in one of two directions: the tool only reports, and leaves you exactly where you started, or it acts on its own, and now a language model’s inference is wired straight into your production control plane.

sentinel-agent does the mechanical part completely, and stops at the decision. That split is the product: investigation is automated, execution is authorised.

7
read-only tools that run with no prompt at all
3
production-mutating tools, every one gated
3.70×
the regression, computed in a sandbox from raw samples
149
tests, including the ones that guard the gate

What it actually does

Given one incident id, the agent reaches real systems over MCP, delegates three parallel lines of investigation to subagents, writes and runs diagnostic Python in an isolated sandbox to compute the magnitudes rather than estimate them, correlates the evidence into a root cause with a stated mechanism and a confidence number — and then pauses, holding the remediation until a human approves it.

  • It reads. Incident, service health, deployment history, diffs, raw golden-signal samples. All seven of those tools are read-only, so none of them ever interrupts you.
  • It computes. export_metrics_csv returns samples and no analysis, on purpose. The change point, the settled means and the ratio all have to be derived — which is what makes the sandbox load-bearing rather than decorative.
  • It argues. Before a gated call it must state action, target, evidence, mechanism, expected effect, risk, reversibility and confidence. A thin case is a failure, because a reasonable approver will decline it.
  • It stops. Not by choice — by construction. The pause is enforced by the harness, in a place the agent cannot reach.

How it works

Four processes, one credential boundary. The UI is a view over harness events and holds no agent logic; the harness holds every key; the MCP server is reachable only from loopback; the sandbox gets its tool calls bridged back out so no credential ever enters it.

sentinel-agent system mapHTTP + SSEMCPcreate_sub_agentexecsentinel-agent UINext.js 16 · React 19 · 127.0.0.1:3000, loopback onlyA view over harness events. Holds no agent logic./tf/[...path] route handlerAttaches TRUEFORGE_TOKEN server-side · streams SSERefuses cross-origin and untokened mutationsTRUEFORGE HARNESS · localhost:8790Agent loop · tool routing · approval gating · subagent delegationSandbox orchestration · session persistence · context managementsentinel-ops MCP :89407 read-only · 1 write GATED2 destructive GATEDSubagents (dynamic)perf / deploy / codeisolated contexts, conclusions onlySandbox · Python 3.13pandas, requests, pydanticno credential ever enters itEvery credential lives inside the harness box.The UI holds none. The MCP server holds none. The sandbox holds none — its tool calls are bridged back out.

Why a harness is load-bearing

Remove TrueForge and this project does not degrade — it stops existing. Six things it carries that would otherwise have to be built from scratch, and gotten right under time pressure:

CapabilityWhat it carries
MCP tool routingReaching the ops estate at all — discovery, schemas, dispatch.
Approval gatingThe entire safety model, enforced in the harness where the agent cannot bypass it.
Sandbox orchestrationIsolated Python on demand, with tool calls bridged back so no credential enters it.
Subagent delegationThree investigation lines in parallel, isolated contexts, conclusions only.
Session persistenceSurviving a page reload mid-investigation.
Context managementCompaction and large-response offloading, so 61 samples plus four diffs still fit.

The agent loop itself — plan, call tools, observe, decide, pause, resume — is the harness’s. sentinel-agent contributes the domain, the safety classification, the methodology, and the view.

Security posture

  • No credential reaches this repo, the sandbox, or the UI. All three live in the harness.
  • The MCP server binds 127.0.0.1 by default and supports bearer auth, because the gate protects a path, not a tool — anything reaching the MCP server directly never encounters it.
  • The UI proxy refuses cross-origin mutations and requires an operator token for every state-changing method. It fails closed when unconfigured.
  • The estate is simulated, so no real system is reachable from this repo. The protocol traffic against it is not simulated.

Try it in two commands

ops MCP server on 127.0.0.1:8940
npm install && npm run dev:mcp
checks all five prerequisites before you waste a run on a 422
npm run doctor

Where to go next