Architecture
Four processes and one rule: every credential lives in the harness, and nothing else is trusted to hold one.
The pieces
| Component | Runs on | Holds |
|---|---|---|
| sentinel-agent UI Next.js 16 · React 19 | 127.0.0.1:3000 | No agent logic and no credentials. A view over harness events: timeline, subagent threads, evidence, approval gate, audit trail. |
| /tf route handler server-side proxy | same process | The harness token, server-side only. Streams SSE through; refuses cross-origin and untokened mutations. |
| TrueForge harness | localhost:8790 | Every credential. Agent loop, tool routing, approval gating, subagents, sandbox orchestration, session persistence, context management. |
| sentinel-ops MCP server | 127.0.0.1:8940 | Thirteen tools and a simulated estate. No credentials. Optional bearer token for its own protection, not for anyone else’s. |
Credential boundaries
- The UI never sees the harness token — it is attached server-side by the route handler, so a browser devtools tab has nothing to steal.
- The sandbox never sees any key. Its tool calls are bridged out to the harness and executed there.
- The one external credential this repo ever handles is a model provider key, read once by
npm run provisionfrom.envand handed to the harness. Never read again. OPS_LAB_TOKENandOPS_MCP_TOKENare secrets you generate yourself withopenssl rand -hex 24, not issued by anything.
Trust model
Two consequences, both found by code review rather than by design, and both now closed:
| Was | Meant | Now |
|---|---|---|
MCP server bound 0.0.0.0, /mcp unauthenticated | rollback_deployment reachable from the LAN, never passing through the harness | Binds 127.0.0.1 (OPS_MCP_HOST to override); optional bearer auth with a constant-time compare; insecure posture logged at error; /estate CORS narrowed from * to known origins |
| Proxy attached the server-held token for any caller | Anything able to reach :3000 could approve a production rollback | Sec-Fetch-Site refuses cross-origin browser requests, and an operator token (x-sentinel-operator) is required for every state-changing method — closing local curl callers, which send no Sec-Fetch-Site at all. Fails closed when unconfigured |
The second one took two rounds. The first attempt added the origin check and documented caller authentication as out of scope — and the reviewer did not mark it resolved, correctly: an origin check is not authentication, and the guard explicitly allowed non-browser callers, so a local curl could still submit an approval. The operator token is the actual fix.
Event flow
agent decides → model.message { toolCalls: [{ id: call_71c, ... }] }
harness sees @destructive → tool.approval_required { tool_calls: [{ id, source_event_id }] }
turn ends → state.required_actions populated
UI joins on source_event_id → renders "rollback_deployment(dpl-4c21)" + the evidence brief
you decide → POST /turns { type: 'user.tool_approval', approval: { status } }
harness dispatches or not → tool.response, or the agent continues without it
estate records → /estate/audit (independent of the event stream)The join on source_event_id is the non-obvious part — the approval event carries no tool name and no arguments, so a client without an event index has nothing to show but an id. The gate page has the detail.
Deliberate omissions
- No database. Session state lives in the harness; the estate is in-process with its own audit log. Nothing here is worth persisting past a restart.
- No auth system. One operator token, checked in constant time, failing closed. A login screen would be a bigger surface for no gain on a loopback-only tool.
- No client-side agent logic. Every decision is the harness’s. If the UI could decide anything, the UI would be part of the safety model.