How it works

Sandbox execution

The regression's magnitude is never handed to the agent. It has to be computed — which is the difference between an agent that reports a number and one that can be checked.

export_metrics_csv returns 61 minute-resolution samples and no analysis. If it returned “p95 is up 3.7x”, the sandbox would be decorative and the agent would be quoting a constant. Returning samples makes the arithmetic load-bearing.

The analysis, in full

python · running in the sandbox
import pandas as pd

df     = pd.read_csv("metrics.csv", parse_dates=["ts"])
deploy = pd.Timestamp("2026-08-25T15:02:00Z")

# 1. split at the change point
before = df[df.ts <  deploy]
# 2. skip the ramp — connection pools filling, retries piling up
after  = df[df.ts >= deploy + pd.Timedelta(minutes=4)]

# 3. compare settled to settled
p95_ratio = after.p95_latency_ms.mean() / before.p95_latency_ms.mean()
err_ratio = after.error_rate.mean()      / before.error_rate.mean()
rps_delta = after.rps.mean()             / before.rps.mean() - 1

print(f"p95 {p95_ratio:.2f}x  errors {err_ratio:.1f}x  throughput {rps_delta:+.1%}")
# p95 3.70x  errors 15.3x  throughput -0.1%
3.70×
p95 latency, settled baseline vs settled plateau
15.3×
error rate — moved with latency
-0.1%
throughput — did not move at all
57
samples used; 4 discarded as ramp

Three decisions in that script are doing real work, and each of them would be easy to get wrong:

  • Split at the change point, not the window midpoint. The deploy timestamp comes from list_recent_deployments, not from eyeballing the series.
  • Throw away the ramp. The four minutes after the deploy are a transient — pools filling, retries stacking. Including them drags the plateau mean down and understates the regression.
  • Check the signals that should not have moved. Flat throughput is what rules out a traffic surge. A conclusion that only looks at the signal that broke is a conclusion that cannot be falsified.
checkout-api p95 latency, 14:30–15:30 UTC0200400600400ms budget15:02 · dpl-4c21baseline 177.9ms · n=32plateau 658.2ms · n=253.70×p95 regression14:3014:4515:0015:1515:30
p95p50ramp — excluded from both meansbaseline 178ms constant · plateau 658ms observed

Where the code runs, and what it can reach

Generated code is untrusted by definition — it was written seconds ago by a model, in response to data. So the sandbox is provisioned on demand and holds nothing worth stealing.

Inside the sandbox
RuntimePython 3.13. No Node runtime.
Preinstalledpandas, requests, pydantic, openpyxl
Time limit60 seconds per exec call
CredentialsNone. Tool calls are bridged back to the harness, where the real keys live.
ProviderDaytona if configured, otherwise TrueForge’s LocalSandboxProvider

You almost certainly do not need Daytona

This project’s own README used to claim a Daytona API key was a hard requirement. That was wrong. TrueForge ships a LocalSandboxProvider on Linux and macOS that needs three host binaries and no external account at all.

bwrap, socat, rg — the only prerequisites
sudo apt install bubblewrap socat ripgrep
harness log, once they are present
info Local sandbox fallback is available {"platform":"linux","shell":"/usr/bin/bash","python":"/usr/bin/python3.10"}

npm run doctor then reports sandbox provider — none configured — local fallback confirmed active in harness log. A Daytona key is only required if that provider fails to start, and on Windows, where TrueForge cannot run directly — hence scripts/wsl-up.sh, which brings the whole stack up under WSL for exactly that reason.