Get started

Guided tour

One incident, start to finish. Every number on this page is one the agent derives itself — nothing in the estate names a cause, and nothing in the tool responses hands over a ratio.

It is 15:04 UTC. A synthetic probe has been failing its checkout budget for two minutes, and it is not recovering. That is all anyone knows.

INC-2048 · SEV-2 · investigating

Checkout p95 latency regression

Synthetic checkout probe breached its 400ms p95 budget and has stayed above it. Card authorisation success is unaffected; customers see slow checkouts and some 5xx.

service checkout-api · detected 15:04:00Z by synthetic-probe/checkout-p95

Step 0 — start the stack

Two processes and a harness. If you would rather read than run, skip ahead: everything below this point is reproducible, but nothing below this point requires you to have run it.

terminal 1 — ops MCP server on 127.0.0.1:8940
npm run dev:mcp
terminal 2 — the harness, on :8790
npx @truefoundry/trueforge@latest
terminal 3 — this UI, plus the operator console on :3000
npm run dev:web

What happens next, in order

  1. 1

    It orients

    Four read-only calls, no prompts: the incident, the service health, the deployment history, and the golden signals. Read-only tools are annotated readOnlyHint: true, so the harness never pauses on them. Investigation should never need a click.

    sentinel-agent — run · INC-2048
    15:07:02get_incidentINC-2048done
    15:07:04get_service_healthcheckout-apidone
    15:07:06performance-investigatorcharacterise onset + magnituderunning
    15:07:06deployment-investigatorenumerate changes in windowrunning
    15:07:07code-investigatorread timing-plausible diffsrunning
    15:07:19export_metrics_csv61 samples, rawdone
    15:07:31execpandas, split at deploydone
    15:07:44rollback_deploymentdpl-4c21awaiting you
  2. 2

    It splits into three

    Three subagents run concurrently with isolated contexts, and return conclusions rather than transcripts:

    • performance-investigator — characterise the symptom. When did it start, and how big is it?
    • deployment-investigator — enumerate every change in a generous window, and rule candidates in or out on timing alone.
    • code-investigator — read the diffs of the timing-plausible candidates and assess mechanism.

    The root agent correlates. That judgement is never delegated — and the fan-out is a convention, not a guarantee.

  3. 3

    It asks for the raw numbers

    export_metrics_csv returns 61 minute-resolution samples across the incident window and no analysis whatsoever. This is deliberate. A tool that returned “p95 is up 3.7x” would make the sandbox decorative; a tool that returns samples makes it load-bearing.

    checkout-api p95 latency, 14:30–15:30 UTC0200400600400ms budget15:02 · dpl-4c21baseline 177.9ms · n=32plateau 658.2ms · n=253.70×p95 regression14:3014:4515:0015:1515:30
    p95p50ramp — excluded from both meansbaseline 178ms constant · plateau 658ms observed

    The shape is legible to a human in about a second: flat, a step at 15:02, a short ramp, a new plateau. None of that is legible to a model reading 61 rows of CSV, which is exactly why it has to do arithmetic instead of pattern-matching.

  4. 4

    It does the arithmetic properly

    In a sandboxed Python 3.13 with pandas — provisioned on demand, holding no credentials. The method matters more than the result: split at the deploy, throw away the ramp, and compare settled to settled.

    python · sandbox
    import pandas as pd
    
    df     = pd.read_csv("metrics.csv", parse_dates=["ts"])
    deploy = pd.Timestamp("2026-08-25T15:02:00Z")
    
    before = df[df.ts <  deploy]                              # 32 samples
    after  = df[df.ts >= deploy + pd.Timedelta(minutes=4)]    # 25 samples, ramp skipped
    
    ratio = after.p95_latency_ms.mean() / before.p95_latency_ms.mean()
    print(before.p95_latency_ms.mean(), after.p95_latency_ms.mean(), ratio)
    # 177.9  658.2  3.6997...

    Including the ramp would have dragged the plateau mean down and understated the regression. Averaging the whole window would have understated it badly. The four-minute exclusion is the difference between a number and a defensible number.

    signal
    settled before
    settled after
    change
    p95 latency
    178 ms
    658 ms
    3.70×

    Breached the 400ms budget and stayed there.

    error rate
    0.40%
    6.19%
    15.3×

    Moved with latency — consistent with timeouts, not with load.

    throughput
    121 rps
    121 rps
    -0.1%

    Flat. No traffic surge. The cause is inside the service.

  5. 5

    It lines up the suspects

    Four deployments are in the window. Only one is even timing-plausible.

    DeploymentWhenChangeVerdict on timing
    dpl-4c2115:02Raise upstream client timeout and add retries for flaky tax provider2 min before detection — candidate
    dpl-4c2024 AugEmit cart-abandonment counterruled out, 28h earlier
    dpl-4c1922 AugBump payment SDK to 4.2.1ruled out
    dpl-4c1820 AugCache tax lookup responses for 60sruled out

    Timing narrows it to one. Timing alone does not explain why — so the diff has to be read.

  6. 6

    It finds the mechanism

    diff · dpl-4c21 · a19f3c2 · r.okafor
    --- a/src/checkout/upstreamClient.ts
    +++ b/src/checkout/upstreamClient.ts
    @@ -10,10 +10,10 @@ import { taxProvider } from './providers/tax';
    
     export const upstreamClient = createClient({
       baseUrl: config.taxProviderUrl,
    -  // Fail fast: the checkout path has a 400ms budget end to end.
    -  timeoutMs: 250,
    -  retries: 0,
    +  // Tax provider has been flaky this week; be more patient with it.
    +  timeoutMs: 30_000,
    +  retries: 3,
       onError: (err) => logger.warn({ err }, 'tax provider call failed'),
     });

    This is a good change, made for a good reason, by someone paying attention. The tax provider really had been flaky. And three retries against a 30-second ceiling on a path with a 400ms end-to-end budget is where a 3.7x tail comes from.

  7. 7

    And then it stops

    The agent has a fix, a mechanism, and a number. What it does not have is permission. The harness sees a call to a tool annotated destructiveHint: true, emits tool.approval_required, and ends the turn. Nothing else happens until a human acts.

    approval required — thread thr_9f2a
    held — the agent has stopped

    rollback_deployment(dpl-4c21)

    Evidence
    p95 178 → 658ms (3.70x) from 15:02; throughput flat; error rate 15.3x
    Mechanism
    timeoutMs 250 → 30_000 with retries 0 → 3 on the tax provider call
    Expected effect
    p95 returns to ~178ms within one rollout, about 90s
    Risk
    Reverts a fix for a flaky provider; the flakiness returns
    Reversibility
    Reversible — redeploy 2026.8.25-1
    Confidence
    0.91
    ApproveDenynothing happens until you choose

    Every field in that brief is required by the skill before a gated call. The approver reads it and nothing else, so a thin case is treated as a failure of the run rather than a style problem.

  8. 8

    You decide — and the estate remembers

    Approval resolves as a new turn, not an endpoint call. Deny and the agent continues without the rollback. Approve and it executes, then re-checks the signals. The estate keeps its own audit log at /estate/audit, independent of the harness event stream, so the agent’s account of what it did can be cross-checked against what actually changed.

    the estate's own record, not the agent's summary of it
    curl -s http://127.0.0.1:8940/estate/audit | jq '.[-3:]'

What you just watched

0
clicks required to reach a root cause
1
click required to change production
61
raw samples the agent had to reduce itself
0.91
stated confidence, with the evidence attached

Every mechanical step ran unattended. The one irreversible step did not, and could not — not because the model was well-behaved, but because the harness refuses to dispatch the call. That refusal is the subject of the next page, including the way it silently fails to happen if a tool forgets four lines of metadata.