How it works

Honest limitations

Everything on this page is a thing the project does not do. It is a page rather than a footnote because a safety claim with an undisclosed gap in it is worse than no claim.

The rule applied throughout: where something is unverified, it says so. Two of the four Gate Prover verdicts are deliberately not a pass for the same reason.

Subagent role names are prompt convention

The harness has no way to declare named subagents. AgentSpec has no subagent field, and create_sub_agent takes a name the model invents and a brief the model writes. The three roles are specified in the instructions and in the skill, and the model follows them — but the fan-out is not guaranteed and the names are not enforced.

context →

The estate is simulated

Real MCP protocol traffic, fixture data. No real system is reachable from this repo, which is also why /estate/audit exists — so the agent’s account of what it did can be cross-checked against what actually changed.

context →

The conformance report is a partial run

reports/gate-conformance.json currently contains one probe, P4, with "complete": false. The P1 and P3 "gate held" verdicts and the P2 bypass come from earlier runs written up in PR #4. Reproducing all four in one fresh report is open work.

context →

P4 has never actually been exercised

The sandbox-bridge route reports route_not_exercised because the model never provisioned a sandbox and called the tool through the bridge. That verdict means untested. It does not mean safe, and it is deliberately not rendered as a pass.

context →

No compaction or sandbox-command events exist

TrueForge does not emit them, so the console cannot surface either directly. You can see that a sandbox turn happened; you cannot see the code it ran or its stdout from the UI.

context →

Skills load only from public repository URLs

github.com and gitlab.com only. There is no private-repo credential field, so a skill in a private repository cannot be registered.

context →

Model behaviour is not deterministic

The fixtures are byte-identical on every boot; the investigation path is not. Runs vary in tool order, subagent count, and occasional detours. What does not vary is the evidence required before a gated call.

context →

Things that used to be listed here and no longer are

Known open work

ItemBlocked on
Reproduce the P2 bypass in a fresh conformance reportRecreating the sentinel-ops-unsafe connector with the current lab token
Exercise P4 for realGetting the model to provision a sandbox and call the tool through the bridge, rather than reporting route_not_exercised
Demo videoNot started
Point the tool surface at a real observability APIA credentials change, not an architecture change
Multi-incident triage — rank concurrent incidents by blast radiusNot started
Post-incident report generation from the session evidence graphNot started

What is not on this page

  • The safety model. Thirteen of thirteen tools carry annotations on the wire, five are gated three ways, and the tests assert it against the harness’s own predicates.
  • The credential boundary. No key reaches this repo, the sandbox, or the UI. npm audit reports zero vulnerabilities.
  • The arithmetic. The 3.70x figure is computed from 61 raw samples in a sandbox, and you can reproduce it from a clean clone.

If you find something that belongs on this page and is not on it, that is a bug in the docs rather than a difference of opinion. Back to the start →