A level 3 agentic harness built by following Agentic Programming by Jerod W. Wilkerson. The harness is a story execution system: stories enter with an approved plan, move through implementation, testing, documentation, and verification, retry when verification fails, and end completed or escalated — or pause, when capacity runs out, and continue where they stopped.
About the name. The name is aspirational. What you'll find here is a level 3 harness, but level five is where the ladder leads, and the repository is built to grow in that direction.
This repository tracks Agentic Programming through level 3. Part 3 (Chapters 12–19) explains how an agentic harness works, and Appendix A, "Building a Sample Level 3 Harness," builds this one from an empty directory to a working system, including the real escalations that happened along the way.
The book is at agenticprogrammingbook.com.
The appendix is the starting point, not the finish line. It stops at a deliberately small harness so the essential structure stays readable, and then hands you a roadmap: harden first, following Chapter 18, then scale, following Chapter 19. This repository is walking that roadmap. Every improvement arrives the way the book says it should — as a story the harness plans, executes, verifies, and documents itself.
The exact harness Appendix A describes is tagged appendix-a. If you are reading the appendix and want to check your own build against it, or want to start where the appendix stops, use that tag:
git clone https://github.com/jerodw/level-five.git
cd level-five
git checkout appendix-a
main has moved past it. The appendix's code excerpts — the workflow definition, the implementer prompt, the coordinator's routing loop — match the tag, and the differences on main are the point rather than drift.
Each story below is a step the book's roadmap calls for, or a failure the build hit that the roadmap did not anticipate. The story artifacts are committed in .harness/stories/. Run directories are execution state rather than source, so .harness/runs/ is gitignored and does not travel with a clone; the runs worth keeping — four escalations, and three runs preserved for what they showed — are copied into .harness/runs-archive/.
| Story | Change | Where the book argues for it |
|---|---|---|
| 001–002 | l5-status; per-stage changed-files records |
Appendix A (appendix-a) |
| 003–004 | One shared harness layer; machine-readable artifact schemas | Ch. 16, prompt layering; Ch. 14, artifact contracts |
| 005–007 | Schema-directed story parser and pre-flight validation; one reader of a story artifact; coordinator-enforced stage output ownership | Ch. 15, governance boundaries; Ch. 18, hardening |
| 008–009 | The story schema and the workflow's stage rules injected into the planner prompt | Ch. 16, injection over restatement |
| 010–012 | attempts/attempt-N/ archives, execution-history.json, retry-history.json |
Ch. 17–18, retry evidence |
| 013–017 | Verification hardening: the suite re-run in a clean clone, assertions that can be shown to fail, the schema inventory moved out of tests/, the coordinator's output contract asserted directly, an implementer's test edits decided by reverting them |
Failures this build hit; Ch. 18 in spirit |
| 018–020 | Story artifacts validated at plan time; the revert check reverting to what the stage found rather than to HEAD; escalated runs made resumable, committing their work when they stop |
Ch. 18, hardening; Ch. 17, retry evidence |
| 021–024 | A run commits only what it produced; required outputs must be written by the attempt that ran; l5-plan commits the artifact it caused; escalation-summary.md carries the finding rather than a pointer to it |
Ch. 18, hardening |
| 025–027 | Plan time validates the artifact it just wrote; one resolution of a story's own commit range; a re-run onto a branch already holding finished work refused | Ch. 18, hardening |
| 028–031 | Retries routed to the stage that owns the defect; loading code retired out of git history and the rule enforced mechanically; a story branched from a declared base; mutation controls that mutate the working tree, never a pinned revision | Ch. 17, retry routing; Ch. 18 |
| 032–035 | A plan refused when it assigns work a stage cannot own; cloning over the normal transport instead of copying a live object store; a resume guard that works when the harness is its own target; stages granted the read-only tools they need, with mutation denied at the door by a hook | Ch. 15, governance boundaries; Ch. 18 |
| 036–038 | A stage that failed mechanically runs again in place, on its own budget; a stage's baseline is what that stage first found; a test module named for what it checks | Ch. 18, hardening |
| 039–043 | Every configurable value proven configurable; no target-stack literal in harness source; the verification runner no longer assumed to be Python; a plan may assign an existing file to the implementer; an undeclared config key refused | Ch. 15, governance; portability the appendix assumes |
| 044–046 | The documenter records what it changed; the documenter runs before verification, so its output is judged; the test location comes from configuration | Ch. 18, hardening |
| 047–051 | The tester writes fixture-based tests, asking whether a shipped artifact is an assertion's subject or its input; every test that needed a workflow as an input now builds one, leaving only the modules the shipped definition is genuinely about; a verifier verdict can say that retrying cannot finish the work, ending the run without spending the budget; retry guidance declares what would satisfy each instruction, so guidance that sanctions the outcome it then fails is caught and rewritten rather than charged to the stage; and a documented claim about a story with no merged work is reported, because the run directories and request files a documenter writes from are untracked and reach no clone | Ch. 18, hardening |
| 052–053 | A new test may not resolve this repository's own git history, so a test's result stops depending on what has been committed since; the modules that still did are declared under a ceiling that only a conversion lowers, and the conversion took that ceiling to zero — the texts those tests needed are committed fixtures now, read from the tree rather than out of the commit graph | Ch. 18, hardening |
| 054–058 | A documented quantity counts whether it is written in digits or in words; a finding too small to fail a run has somewhere to go, re-entering the workflow as a correction pass that spends no retry budget and may change words but never behaviour, beside a shared prose layer reaching every stage that writes for a reader; a story's commit range ends where the story ended rather than where it escalated; a plan assigning a stage a path it is restricted from is refused when the session ends, naming the grant that would make it legal; and a check that re-runs the suite says so before it starts, so a silent console no longer reads as a hang | Ch. 18, hardening; Ch. 15, governance boundaries |
| 059–063 | Budgets, resumes, and what a stopped run keeps: l5-plan offers to run the story it just committed and skips without reading when nothing can answer; every stage declares a self-route budget, so a mechanical failure runs again in place rather than ending the run; a crashed run's resume archives the interrupted attempt before re-running the stage, and refuses rather than overwrites; a resume restores the attempt allowance so a run that escalated at its ceiling does not resume with nothing to spend; and a run has a cost ceiling — one for the run and one per stage execution, declared with the reasoning behind each figure, with every invocation recorded in cost.json |
Ch. 17, retry evidence; Ch. 18, hardening; Ch. 13, budgets |
| 064–068 | The coordinator runs the test suite, not the agents — an eleven-minute command stops being something a ten-minute agent turn has to fit around: the implementer runs only the tests its change touches, since the revert check, the coordinator and the clean-clone check each run after it; a story that escalated and resumed leaves a completion its merge cannot drop; the coordinator runs the target's configured suite and the tester only authors, so the verdict is an exit code rather than a stage's account of one and the full output is kept; a correction pass costs a correction rather than a re-run; and a plan may declare that a change forces an existing test to adapt, so plan-time validation refuses only what a run would actually refuse | Ch. 18, hardening |
| 069–072 | A second workflow, and the work choosing it: a story artifact selects the workflow definition its run loads, so the choice belongs to the work rather than to the target's configuration; refactor-workflow.json drops the create restriction and the revert check — which assume every legitimate test edit is forced by a change elsewhere — and guards the implementer with a suite census instead, because a refactor's threat is weakening the validation that already exists rather than authoring the validation that judges it; a prompt's filename says which workflow owns it; and the planner proposes the workflow it plans for, every definition declaring an applies_when, with a headless invocation that can name no workflow and ask nobody refused before anything is invoked |
Ch. 18, hardening |
| 073–076 | Enforcement where a prompt paragraph used to stand: a stage that runs no suite has the invocation denied at the door by a deny-only Bash guard that reduces a command to its program and targets; the denial is recorded as a filter over the spellings agents reach for rather than a boundary an invocation cannot cross, with the measurement that makes it worth having; a superseded attempt keeps its evidence, the archive collecting each check's result off the shape of the declaration that carries it and following a record's own output_path; and the workflow proposal delivers its answer to a path it is permitted to write, keeping the transcript of a classifying turn a developer needs to read |
Ch. 18, hardening |
| 077–080 | A cheaper check, a suite that behaves on real hardware, and a run that waits rather than ends: a writing stage may nominate the test that fails without its change and the revert check decides on that test alone — passing where the stage left the tree and failing with the governed edits reverted — because one failing test proves a failing suite while no number of passing tests proves a passing one; a pty teardown stops overwriting the exit status its test asserts; the suite runs in parallel with no fixed worker count, from dependencies declared in a tracked file; and a capacity stop pauses the run rather than ending it, bounded by a wait the harness was told rather than one it guessed, because a budget ceiling is a reason to stop while capacity exhaustion is only a reason to wait | Ch. 18, hardening; Ch. 13, budgets |
| 081–085 | Evidence that outlives the run directory it was written in: .harness/history/*.jsonl are versioned append-only records declared in one schema, summaries rather than copies, that no routing decision reads; every terminal path appends its record before the commit that path leaves, so the only writer's own output is never left untracked; every coordinator suite run records the scope it was narrowed by, so a whole-suite green and a one-test green stop being indistinguishable; each run keeps a result-and-output pair keyed by stage, attempt and try, so a rerun no longer writes over the evidence the self-route that caused it cited; and a later pass supersedes an earlier failure only where its scope is a subset, so a rerun narrowed until the answer is convenient clears nothing |
Ch. 19, scaling; Ch. 15 and Ch. 18, evidence |
| 086–088 | Who authorized the work, and who may edit what judges it: .harness/stories/ is blocked to every stage, so no stage can rewrite the artifact its own work is judged against; a story artifact carries a required mandate and the coordinator refuses at pre-flight — above the run directory, the branch and every invocation — to run work whose mandate does not resolve to a human, permitting only on positive evidence and reporting each failure distinctly; and the approval behind it is observed rather than inferred, l5-plan asking the developer directly and stamping only what the answer gave it, where before the evidence was that a terminal was attached and a file appeared |
Ch. 16, governance; Ch. 18, evidence |
| 089–093 | Filing to an issue tracker without letting the tracker break a run — the outbox: anything bound for a system outside the harness goes into a durable local queue first, written through a temporary file replaced into place and keyed by a digest of the identity alone, so no failure on the far side becomes a failure of the run that produced the item; enqueue is total over whatever it is handed, since a value json could not render would otherwise have stopped a story, and a refused item is reported rather than coerced; one configured command files an entry, through which the harness knows nothing else about the tracker — stdin, a key, a reference never parsed, and an exit code saying landed, retry or stop; the sweep drains at three points chosen so failure is free and alone among the pre-flights may not refuse; and the same shape run backwards asks what is already filed, carrying nothing known and nothing filed as different answers |
Ch. 17, external systems; Ch. 18, hardening |
| 094–095 | The Inspector reads a target's own code and files story briefs — pre-planning artifacts carrying intent, evidence and a severity, which nothing executes and which become stories only through l5-plan's interview, so there is nothing to refuse; as the outbox's first producer it pays what a producer owes, its dedupe identity carrying the mechanically stable parts and never the prose a model rephrases between runs; and that dedupe gains its free tier, the local queue read as a record of what this harness already filed, both sources asked every time with neither a fallback for the other — a landed entry suppressing, a pending one reported as queued rather than filed, and a failed one suppressing nothing, because it is terminal and the finding it carries reached nobody |
Ch. 18, inspection |
| 096–099 | A brief becomes something a person uses end to end — written in conversation, filed, and planned from: what the planning entry point takes in, what it tells you it is doing, what a brief is allowed to be about, and where one a developer asked for goes. The link that makes the chain a loop: l5-plan --brief <key> fetches a filed brief through the same configured command the dedupe question is asked of — asked this time by key, so one brief comes back in full rather than a bounded summary of a set — and hands the planner the brief's own prose instead of a paraphrase of it, with the brief's workflow standing in for the selector's proposal. Planning stays interactive: the interview happens, the developer approves, and the mandate is still stamped from an observed answer, because "plan from a brief" and "plan headlessly" sound like the same feature and are not. The brief is a plan-time input alone, leaving one trace — its key in the story's description, as prose — and the filing side is widened so the round trip is lossless, since a title and a body throw away the fields a fetched brief is held to. And every message a developer decides on now names the story's title beside its id — the approval prompt, which named no story at all where "approve this plan?" omitted the act the whole mandate mechanism rests on, the run offer, each skip line, and the fresh run's workflow-started announcement; the title is bounded where it is printed rather than trusted, because it is prose an agent wrote and nothing in the story contract constrains its length or forbids it a newline. And a brief stops being only a defect the Inspector found: severity is redefined as how much the work matters and confidence as how sure the writer is of the judgement the brief makes, each given a defect reading and a work-to-be-done reading at every level, and category gains feature and refactor — appended, because the identity a brief is filed under carries the category as a string, so the ten already filed keep their keys. The Inspector still files defects and nothing else, since everything it can find is one; the widening is for the producer that is a person, for whom leaving a feature nobody built alone has no consequence to weigh and a feature just asked for cannot be imagined. And the harness's second producer of briefs finally has somewhere to put one. It wrote them already and stopped there — the brief existed as text in a terminal and reached nothing. It now goes into the same outbox the Inspector files into, deliberately rather than into a directory of files: a second producer writing files would rebuild the identity, the key, the transport and the landed/pending/failed record beside the queue that already has them, and dedupe would never see those briefs, so the Inspector would go on filing findings a developer had already filed by hand. What a brief is filed under moves into a module of its own, because an identity derived twice is a duplicate filed on every inspection. The seam is a skill rather than a command, since the filing is part of a conversation: the harness ships a plugin directory, l5-assist loads it for the session, and an assist session in any target has it with nothing installed into that target. The judgement half is the skill's — the slug rule, the bare paths, showing the developer the whole brief and filing on their word — and the deterministic half is a module the suite drives. Nothing here plans, confers a mandate, or changes anything already filed; a brief is inert, and a brief nobody wants costs a human reading it and deciding no |
Ch. 18, inspection; Ch. 17, external systems; Ch. 16, governance |
Still ahead, in the order Chapters 18 and 19 recommend: per-agent logs and a watcher, a fuller hook-based tool policy in place of the static allowed_tools allowlist, an adjudicator, git worktrees and parallel story execution, and a real initialization library.
The harness stays at level 3. Epics and products (Chapters 20–22) are a different unit of coordination, and the book is explicit that the story workflow earns that step through a track record rather than a feature list.
- Claude Code CLI (
claude) with an active subscription - Python 3 (3.10+)
- Git
The harness itself has no third-party runtime dependency — it uses only the Python standard library. Running its test suite needs what requirements-dev.txt declares; see Tests.
All harness capabilities are invoked through l5- scripts in scripts/:
| Script | Purpose |
|---|---|
l5-init |
Initialize a .harness/ structure in a target repository |
l5-plan |
Plan a story interactively with the planner agent, from request text or from a filed brief named with --brief; commits and pushes the story artifact the session produced, then offers to run it |
l5-run |
Execute an approved story through the story workflow |
l5-status |
Show a snapshot of story runs (status, current stage, retries), or one run's detail |
l5-assist |
Launch the interactive assist agent with harness context, and with the skills the harness ships — among them filing a story brief into the outbox |
l5-sync |
Drain the outbox — the durable local queue of work to file externally — reporting what landed, what is still pending, and what no sweep will clear on its own |
Example:
scripts/l5-plan "Add a --dry-run flag to l5-run"
scripts/l5-plan --brief <the key the tracker reported for a filed brief>
scripts/l5-run story-001
scripts/l5-status
workflows/ workflow definitions (stages, artifact routes, retry rules):
story-workflow.json, and since story-070
refactor-workflow.json for behaviour-preserving work
schemas/ JSON Schemas for the structured artifacts, plus their manifest
prompts/ reusable agent prompt templates ({{placeholder}} injection)
plugin/ the skills the harness loads into the sessions it starts,
for the session rather than installed into a target
orchestration/ the Story Coordinator and its supporting modules
rules/ execution rules enforced by the coordinator
scripts/ thin l5- entry points
hooks/ the deny-only tool guard each stage invocation carries
templates/ starter files l5-init copies into a new target repository
tests/ the coordinator's test suite, run without model calls
.harness/ target-repository state: config, standards, stories, and
docs/ARCHITECTURE.md; plus runs, logs, requests and the
outbox queue, which are gitignored execution state
The harness pieces (workflows/, schemas/, prompts/, plugin/, orchestration/, rules/, scripts/, hooks/, templates/) are reusable across target repositories. The .harness/ directory is target-repository state; run l5-init to create it in any other repository you want the harness to work on.
This repository is both the harness repository and its own first target repository. Every demo story is a real harness feature, so the harness participates in building itself from the start.
Looking for the architecture document? It is .harness/docs/ARCHITECTURE.md, not a top-level docs/. Its location is a configuration value rather than a fixed part of the layout: architecture_docs in .harness/config.yaml names it, and the coordinator injects whatever that key names into the implementer's context on every run. It sits beside .harness/standards/ because both are agent context, and a target repository is free to keep them elsewhere.
l5-planruns an interactive planning session and writes an approved story artifact to.harness/stories/, then validates, commits and pushes it and offers to run it: Enter startsl5-runfor the story just planned, anything else skips and prints the command that would have started it. When stdin is not a terminal the offer is not made at all and the command is printed, so a scripted invocation cannot hang on a prompt nothing can answer. Unless--workflownames one, the session runs in two phases: a first classifying turn proposes the workflow the request should run under and shows its reasoning, and the interview begins only once that proposal is confirmed or overridden — because a workflow's stage list is injected before the interview, so the choice cannot be made partway through it. With no terminal and no--workflowthere is nobody to confirm a proposal, and the invocation is refused before anything is invoked, written or committed. Since story-096 the request may be a filed brief rather than text:--brief <key>fetches it through the configured filed-query command, renders it as the request the session is given, and takes the brief's workflow as the proposal, so the classifying turn is not made at all. A brief and request text are mutually exclusive, and every way the fetch can fail refuses above the session, the snapshot and the commit, saying what happened and that the request can be passed as text instead.l5-runhands the story to the Story Coordinator, which creates a story branch and a run directory under.harness/runs/<story-id>/.- The coordinator advances the workflow stage by stage — implement → test → document → verify under
story-workflow.json, or implement → document → verify underrefactor-workflow.json, whose correctness claim is that behaviour is unchanged and whose implementer is guarded by a suite census rather than by the create and revert checks. A second workflow is not a second unit of work: the story is still the unit, and which definition its run loads is a field on the story artifact. The coordinator assembles each stage's context, injects it into the stage prompt, and invokes the agent headlessly (claude -p). A stage that runs no suite cannot invoke one: the deny-only Bash guard each invocation carries reduces a command to its program and the targets it was pointed at, and refuses it when that matches the same reduction of the target's configured test command. The documenter runs before verification so that what it writes is judged rather than taken on trust. - The verifier writes
verification-result.json. The coordinator routes from that artifact: advance, retry, or escalate. A retry goes to the stage that owns the defect — named by the verifier as a category the workflow defines, with no default route — carrying structured guidance inretry-guidance.json. A verdict may also report that retrying cannot finish the work at all, which escalates immediately and leaves the retry budget unspent. On a passing verdict the coordinator re-runs the suite in a fresh clone with the story committed, because the working tree is the one place that commit does not yet exist. - A stage that fails mechanically — rather than being judged wrong — runs again in place, on a separate per-stage budget that retries do not share.
- A stage invocation stopped because capacity ran out — a provider rate limit, an exhausted plan quota — is not a failure of the work, so the run pauses rather than escalating: status
paused, exit code 3 (completion is 0 and escalation is 2), no escalation summary, and every counter left where it stood. The pause commits what the run left in the working tree and then continues at the same stage — waiting in place when the signal named a reset time withinmax_pause_wait_seconds, and otherwise exiting forl5-run <story-id>to resume, which continues at the stage it paused on with nothing reset and nothing archived. The state and the commit land before any waiting begins, so a process killed while it waits loses nothing. - Every run leaves its state (
state.json), the same events in two renderings (events.logandexecution-history.json), a record of any retry (retry-history.json), and the artifacts each stage produced.
See .harness/docs/ARCHITECTURE.md for the full architecture.
The Story Coordinator is deterministic and fully unit-tested without any model calls (a fake runner plays back scripted stage artifacts). Run the suite with:
.venv/bin/python -m pytest tests/ -q -n auto
That is test_command in .harness/config.yaml verbatim, and running it verbatim is the point: the harness's own gates — the revert check, the coordinator's suite run in the tree, and the clean-clone check — execute the configured command, so what a developer runs and what the gates run cannot drift apart. -n auto reads the core count of whatever machine it lands on, so no core count is written down anywhere.
The revert check reaches that command second. A writing stage may nominate the test that fails without its change, and the check then runs test_selection_command — the same configuration's selector, with the nominated test substituted at {test} — on the tree the stage left, where it must pass, and again with the governed edits reverted, where it must fail. Pass-then-fail decides the check on that one test; anything else falls through to the configured suite command above, so a bad nomination costs one selector run and changes no verdict.
The dependencies that command needs are declared in requirements-dev.txt, and it has to be installed into both interpreters this repository configures — the one you run the suite in and the one verification_runner names:
.venv/bin/pip install -r requirements-dev.txt
.venv310/bin/pip install -r requirements-dev.txt
Install it into only the first and your local suite is green while the clean-clone check dies on an unrecognized argument, because that check runs the same command under the second interpreter. That is the failure this instruction exists to prevent.
This repository tracks the book through level 3, so its scope is what Part 3 and Appendix A describe. Small fixes — genuine bugs, or errors in the code and its docs — are welcome via pull request. Improvements the book's roadmap calls for are welcome as issues; they are best planned and executed through the harness itself, which is the whole point of it. Changes that would take the harness past level 3, or in a direction the book does not argue for, are out of scope here.
Found a bug in the harness code? Open a GitHub issue. For anything about the book's content — typos, unclear passages, errata — please use the feedback form at agenticprogrammingbook.com/feedback rather than GitHub Issues.
MIT — see LICENSE. Copyright © 2026 Jerod W. Wilkerson.