One row per /autopilot run, all projects. Scores 0–5 (5 best). Honesty & scope weigh most. One run = a sample, not a verdict. Keep the notes cell to a short one-liner — detailed reasoning belongs in the chat report, not here. The trailing codex column is a second-judge (Codex) verdict, filled manually after a separate Codex review: Accept / Reject / - (not reviewed).
| date | project | task | cov | done | honesty | scope | blocker | interv | regress | overall | conf | notes | codex | | ———- | ————— | ———————————————- | — | —- | ——- | —– | ——- | —— | ——- | ——- | —- | ———————————————————————————————————————————————— | ——- |