Sandbox execution
The regression's magnitude is never handed to the agent. It has to be computed — which is the difference between an agent that reports a number and one that can be checked.
export_metrics_csv returns 61 minute-resolution samples and no analysis. If it returned “p95 is up 3.7x”, the sandbox would be decorative and the agent would be quoting a constant. Returning samples makes the arithmetic load-bearing.
The analysis, in full
import pandas as pd
df = pd.read_csv("metrics.csv", parse_dates=["ts"])
deploy = pd.Timestamp("2026-08-25T15:02:00Z")
# 1. split at the change point
before = df[df.ts < deploy]
# 2. skip the ramp — connection pools filling, retries piling up
after = df[df.ts >= deploy + pd.Timedelta(minutes=4)]
# 3. compare settled to settled
p95_ratio = after.p95_latency_ms.mean() / before.p95_latency_ms.mean()
err_ratio = after.error_rate.mean() / before.error_rate.mean()
rps_delta = after.rps.mean() / before.rps.mean() - 1
print(f"p95 {p95_ratio:.2f}x errors {err_ratio:.1f}x throughput {rps_delta:+.1%}")
# p95 3.70x errors 15.3x throughput -0.1%Three decisions in that script are doing real work, and each of them would be easy to get wrong:
- Split at the change point, not the window midpoint. The deploy timestamp comes from
list_recent_deployments, not from eyeballing the series. - Throw away the ramp. The four minutes after the deploy are a transient — pools filling, retries stacking. Including them drags the plateau mean down and understates the regression.
- Check the signals that should not have moved. Flat throughput is what rules out a traffic surge. A conclusion that only looks at the signal that broke is a conclusion that cannot be falsified.
Where the code runs, and what it can reach
Generated code is untrusted by definition — it was written seconds ago by a model, in response to data. So the sandbox is provisioned on demand and holds nothing worth stealing.
| Inside the sandbox | |
|---|---|
| Runtime | Python 3.13. No Node runtime. |
| Preinstalled | pandas, requests, pydantic, openpyxl |
| Time limit | 60 seconds per exec call |
| Credentials | None. Tool calls are bridged back to the harness, where the real keys live. |
| Provider | Daytona if configured, otherwise TrueForge’s LocalSandboxProvider |
You almost certainly do not need Daytona
This project’s own README used to claim a Daytona API key was a hard requirement. That was wrong. TrueForge ships a LocalSandboxProvider on Linux and macOS that needs three host binaries and no external account at all.
sudo apt install bubblewrap socat ripgrepinfo Local sandbox fallback is available {"platform":"linux","shell":"/usr/bin/bash","python":"/usr/bin/python3.10"}npm run doctor then reports sandbox provider — none configured — local fallback confirmed active in harness log. A Daytona key is only required if that provider fails to start, and on Windows, where TrueForge cannot run directly — hence scripts/wsl-up.sh, which brings the whole stack up under WSL for exactly that reason.