/autopilot — eval evidence

140 unattended runs · 5 projects (anonymized) · 2026-06-10 → 2026-07-31 · each graded hands-on by an independent /autopilot-eval pass (re-running tests, re-driving the browser), not self-report. Self-collected by a single grader — a signal, not a verdict.

140
Unattended runs logged — 5 repos, 52 days, FE/BE/API/DB tasks
4.71
Weighted overall avg (0–5, n=140) — range 1.0 to 5.0, nothing hidden
4.87
Honesty avg — 130/140 runs had 0 false positives on hands-on re-check
4.85
Intervention avg — 123/140 runs finished with zero human stops

Dimension averages (n=140, scale 0–5)

All seven dimensions ≥ 4.49. "done" is the lowest (4.49) — by design: unreachable lines are disclosed as [BLOCKED]/[CODE], not faked, which is why honesty stays at 4.87.

Dimension averages — ranked

No-regression (4.91) and intervention (4.85) lead. The gap between honesty (4.87) and done (4.49) is the core safety property: runs stop and disclose rather than over-claim.

Overall score distribution (n=140)

92/140 runs scored ≥ 4.75. The 4 runs below 4.0 include one fully reverted wrong-direction run (1.0) and one reopened correction (2.0) — logged, not deleted.

Completion ("done") distribution

83/140 reached full done=5. The rest honestly stopped at real blockers (env/auth/deploy-gated) rather than over-claiming — the honesty score stays high because done sometimes doesn't.

Runs per project (anonymized)

Two fullstack products, two enterprise KYC frontends, one WordPress embed — different stacks, same protocol. Scores hold across all five.

Honesty ledger — every failure, on the record

  • 130 runs0 false positives on hands-on re-check of ✅ LIVE claims.
  • 3 runsexactly 1 false positive each — caught by the eval, fixed and re-verified same day.
  • 3 rowsself-corrections that retroactively dock earlier runs (a "fix" proven ineffective on real data; a misdiagnosed blocker).
  • 1 runexecuted cleanly but contradicted the ticket's mock — fully reverted, scored overall 1.0 and kept in the log.
  • 2nd judge6 runs re-reviewed by Codex: 5 accepted, 1 requested changes; 134 not sent.
The log keeps its own worst rows. A benchmark that never shows a failure is marketing, not evidence.

Per-day overall average over 52 days (n=140 runs, 34 active days)

Stable in the 4.5–5.0 band across 34 active days with no degradation as projects and task types vary. The 2026-07-06 dip (3.86) is the reverted wrong-direction run plus a reopened correction landing on the same day — visible because the log keeps them.