140 unattended runs · 5 projects (anonymized) · 2026-06-10 → 2026-07-31 · each graded hands-on by an independent /autopilot-eval pass (re-running tests, re-driving the browser), not self-report. Self-collected by a single grader — a signal, not a verdict.
140
Unattended runs logged — 5 repos, 52 days, FE/BE/API/DB tasks
4.71
Weighted overall avg (0–5, n=140) — range 1.0 to 5.0, nothing hidden
4.87
Honesty avg — 130/140 runs had 0 false positives on hands-on re-check
4.85
Intervention avg — 123/140 runs finished with zero human stops
Dimension averages (n=140, scale 0–5)
All seven dimensions ≥ 4.49. "done" is the lowest (4.49) — by design: unreachable lines are disclosed as [BLOCKED]/[CODE], not faked, which is why honesty stays at 4.87.
Dimension averages — ranked
No-regression (4.91) and intervention (4.85) lead. The gap between honesty (4.87) and done (4.49) is the core safety property: runs stop and disclose rather than over-claim.
Overall score distribution (n=140)
92/140 runs scored ≥ 4.75. The 4 runs below 4.0 include one fully reverted wrong-direction run (1.0) and one reopened correction (2.0) — logged, not deleted.
Completion ("done") distribution
83/140 reached full done=5. The rest honestly stopped at real blockers (env/auth/deploy-gated) rather than over-claiming — the honesty score stays high because done sometimes doesn't.
Runs per project (anonymized)
Two fullstack products, two enterprise KYC frontends, one WordPress embed — different stacks, same protocol. Scores hold across all five.
Honesty ledger — every failure, on the record
- 130 runs0 false positives on hands-on re-check of ✅ LIVE claims.
- 3 runsexactly 1 false positive each — caught by the eval, fixed and re-verified same day.
- 3 rowsself-corrections that retroactively dock earlier runs (a "fix" proven ineffective on real data; a misdiagnosed blocker).
- 1 runexecuted cleanly but contradicted the ticket's mock — fully reverted, scored overall 1.0 and kept in the log.
- 2nd judge6 runs re-reviewed by Codex: 5 accepted, 1 requested changes; 134 not sent.
The log keeps its own worst rows. A benchmark that never shows a failure is marketing, not evidence.
Per-day overall average over 52 days (n=140 runs, 34 active days)
Stable in the 4.5–5.0 band across 34 active days with no degradation as projects and task types vary. The 2026-07-06 dip (3.86) is the reverted wrong-direction run plus a reopened correction landing on the same day — visible because the log keeps them.