Skip to content

docs: mitigation 6 efficacy review — Outcome recorded (keep + disposition lines) - #230

Merged
Jammy2211 merged 1 commit into
mainfrom
claude/automind-falsified-by-checkpoint-cmsqsi
Aug 18, 2026
Merged

docs: mitigation 6 efficacy review — Outcome recorded (keep + disposition lines)#230
Jammy2211 merged 1 commit into
mainfrom
claude/automind-falsified-by-checkpoint-cmsqsi

Conversation

@Jammy2211

Copy link
Copy Markdown
Contributor

The committed ~10-ship efficacy review of the falsified-by checkpoint stage (mitigation 6, PyAutoBrain#140, live 2026-07-17) — the one §9 item still open after PyAutoBrain#130 closed. Executes the PyAutoMind research prompt has_the_falsified_by_checkpoint_stage_gone.md; the Mind-side record + follow-up prompt ship in the paired PyAutoMind PR on the same branch name.

Verdict: not proven rote — proven unobservable, which is its own finding. Keep the stage, vocabulary unchanged; close the observability gap.

Measured over the 22 review-leg ship gates since go-live (21 autonomy-log rows 2026-07-17→08-01 + one August cloud-session faculty run):

  • unverified-claim findings raised: 0; gates with evidence the claim pass was exercised: 2 — the other ~20 recorded a bare "review CLEAN", indistinguishable from a pass that never read the claims.
  • Instrument validated before trusting the null (the prompt's D1 method note): probe claims lift 3/3 through the live load_bearing_claims(); full suite 349 passed.
  • Firing rate neither empty nor saturated: 13/50 (Brain) and 3/66 (Mind) merge messages since go-live lift ≥1 claim; verified drives 17/26 lifts, mostly of the author's evidence sentence — the claim culture the stage wanted.
  • The one confirmed-wrong load-bearing claim of the window (07-27 "5 siblings" count) lived in an issue comment, outside the commit-message surface the stage reads.
  • Idle-phrasing scoping holds; residual false positives cost seconds and carry no bypass pressure.

Changes: an Outcome block under mitigation 6 in docs/agent_failure_modes.md (mirroring mitigation 2's outcome precedent) and the §9 open item marked done. Doc-only; no behaviour change. The recommended fix (per-claim disposition lines in the verdict) is filed as draft/feature/pyautobrain/review_claim_dispositions.md in the paired Mind PR, not implemented here.

🤖 Generated with Claude Code

https://claude.ai/code/session_01WH4NizvBK2jki2Uh5TMABh


Generated by Claude Code

…tion lines)

The ~10-ship review that §9 committed to when the falsified-by stage
shipped (PyAutoBrain#140, live 2026-07-17). Measured over the 22
review-leg ship gates since go-live: unverified-claim findings 0; gates
with evidence the claim pass was exercised 2; the other ~20 recorded a
bare 'review CLEAN', so a rote pass and a healthy one are
indistinguishable in the ledger. Instrument validated live first (probe
lifts 3/3; 349 tests pass): firing rate 13/50 Brain / 3/66 Mind merge
messages since go-live, 'verified' driving 17/26 lifts, mostly of
evidence sentences. Verdict: not proven rote — proven unobservable.
Keep the stage, vocabulary unchanged; per-claim disposition lines in
the verdict make rote visible (filed as a Mind feature prompt). Full
numbers: PyAutoMind complete/2026/08/falsified-by-checkpoint-efficacy-review.md.

Co-Authored-By: Claude <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01WH4NizvBK2jki2Uh5TMABh
@Jammy2211
Jammy2211 merged commit 102633e into main Aug 18, 2026
3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants