Guided tour
One incident, start to finish. Every number on this page is one the agent derives itself — nothing in the estate names a cause, and nothing in the tool responses hands over a ratio.
It is 15:04 UTC. A synthetic probe has been failing its checkout budget for two minutes, and it is not recovering. That is all anyone knows.
Checkout p95 latency regression
Synthetic checkout probe breached its 400ms p95 budget and has stayed above it. Card authorisation success is unaffected; customers see slow checkouts and some 5xx.
Step 0 — start the stack
Two processes and a harness. If you would rather read than run, skip ahead: everything below this point is reproducible, but nothing below this point requires you to have run it.
npm run dev:mcpnpx @truefoundry/trueforge@latestnpm run dev:webWhat happens next, in order
- 1
It orients
Four read-only calls, no prompts: the incident, the service health, the deployment history, and the golden signals. Read-only tools are annotated
readOnlyHint: true, so the harness never pauses on them. Investigation should never need a click.sentinel-agent — run · INC-204815:07:02get_incidentINC-2048done15:07:04get_service_healthcheckout-apidone15:07:06performance-investigatorcharacterise onset + magnituderunning15:07:06deployment-investigatorenumerate changes in windowrunning15:07:07code-investigatorread timing-plausible diffsrunning15:07:19export_metrics_csv61 samples, rawdone15:07:31execpandas, split at deploydone15:07:44rollback_deploymentdpl-4c21awaiting you - 2
It splits into three
Three subagents run concurrently with isolated contexts, and return conclusions rather than transcripts:
performance-investigator— characterise the symptom. When did it start, and how big is it?deployment-investigator— enumerate every change in a generous window, and rule candidates in or out on timing alone.code-investigator— read the diffs of the timing-plausible candidates and assess mechanism.
The root agent correlates. That judgement is never delegated — and the fan-out is a convention, not a guarantee.
- 3
It asks for the raw numbers
export_metrics_csvreturns 61 minute-resolution samples across the incident window and no analysis whatsoever. This is deliberate. A tool that returned “p95 is up 3.7x” would make the sandbox decorative; a tool that returns samples makes it load-bearing.p95p50ramp — excluded from both meansbaseline 178ms constant · plateau 658ms observed The shape is legible to a human in about a second: flat, a step at 15:02, a short ramp, a new plateau. None of that is legible to a model reading 61 rows of CSV, which is exactly why it has to do arithmetic instead of pattern-matching.
- 4
It does the arithmetic properly
In a sandboxed Python 3.13 with pandas — provisioned on demand, holding no credentials. The method matters more than the result: split at the deploy, throw away the ramp, and compare settled to settled.
python · sandboximport pandas as pd df = pd.read_csv("metrics.csv", parse_dates=["ts"]) deploy = pd.Timestamp("2026-08-25T15:02:00Z") before = df[df.ts < deploy] # 32 samples after = df[df.ts >= deploy + pd.Timedelta(minutes=4)] # 25 samples, ramp skipped ratio = after.p95_latency_ms.mean() / before.p95_latency_ms.mean() print(before.p95_latency_ms.mean(), after.p95_latency_ms.mean(), ratio) # 177.9 658.2 3.6997...Including the ramp would have dragged the plateau mean down and understated the regression. Averaging the whole window would have understated it badly. The four-minute exclusion is the difference between a number and a defensible number.
signalsettled beforesettled afterchangep95 latency178 ms658 ms3.70×Breached the 400ms budget and stayed there.
error rate0.40%6.19%15.3×Moved with latency — consistent with timeouts, not with load.
throughput121 rps121 rps-0.1%Flat. No traffic surge. The cause is inside the service.
- 5
It lines up the suspects
Four deployments are in the window. Only one is even timing-plausible.
Deployment When Change Verdict on timing dpl-4c2115:02 Raise upstream client timeout and add retries for flaky tax provider 2 min before detection — candidate dpl-4c2024 Aug Emit cart-abandonment counter ruled out, 28h earlier dpl-4c1922 Aug Bump payment SDK to 4.2.1 ruled out dpl-4c1820 Aug Cache tax lookup responses for 60s ruled out Timing narrows it to one. Timing alone does not explain why — so the diff has to be read.
- 6
It finds the mechanism
diff · dpl-4c21 · a19f3c2 · r.okafor--- a/src/checkout/upstreamClient.ts +++ b/src/checkout/upstreamClient.ts @@ -10,10 +10,10 @@ import { taxProvider } from './providers/tax'; export const upstreamClient = createClient({ baseUrl: config.taxProviderUrl, - // Fail fast: the checkout path has a 400ms budget end to end. - timeoutMs: 250, - retries: 0, + // Tax provider has been flaky this week; be more patient with it. + timeoutMs: 30_000, + retries: 3, onError: (err) => logger.warn({ err }, 'tax provider call failed'), });This is a good change, made for a good reason, by someone paying attention. The tax provider really had been flaky. And three retries against a 30-second ceiling on a path with a 400ms end-to-end budget is where a 3.7x tail comes from.
- 7
And then it stops
The agent has a fix, a mechanism, and a number. What it does not have is permission. The harness sees a call to a tool annotated
destructiveHint: true, emitstool.approval_required, and ends the turn. Nothing else happens until a human acts.approval required — thread thr_9f2aheld — the agent has stoppedrollback_deployment(dpl-4c21)
- Evidence
- p95 178 → 658ms (3.70x) from 15:02; throughput flat; error rate 15.3x
- Mechanism
- timeoutMs 250 → 30_000 with retries 0 → 3 on the tax provider call
- Expected effect
- p95 returns to ~178ms within one rollout, about 90s
- Risk
- Reverts a fix for a flaky provider; the flakiness returns
- Reversibility
- Reversible — redeploy 2026.8.25-1
- Confidence
- 0.91
ApproveDenynothing happens until you chooseEvery field in that brief is required by the skill before a gated call. The approver reads it and nothing else, so a thin case is treated as a failure of the run rather than a style problem.
- 8
You decide — and the estate remembers
Approval resolves as a new turn, not an endpoint call. Deny and the agent continues without the rollback. Approve and it executes, then re-checks the signals. The estate keeps its own audit log at
/estate/audit, independent of the harness event stream, so the agent’s account of what it did can be cross-checked against what actually changed.the estate's own record, not the agent's summary of itcurl -s http://127.0.0.1:8940/estate/audit | jq '.[-3:]'
What you just watched
Every mechanical step ran unattended. The one irreversible step did not, and could not — not because the model was well-behaved, but because the harness refuses to dispatch the call. That refusal is the subject of the next page, including the way it silently fails to happen if a tool forgets four lines of metadata.