An AI system that writes code the way a careful engineering team does: it plans, waits for a human to approve the plan, writes the code, tests it, has it checked by a quality gate, and fixes its own mistakes until everything passes.
flowchart LR
A[Plan] --> B{Human approves?}
B -- yes --> C[Write code]
B -- no --> X[Stop]
C --> D[Run tests]
D -- fail --> R{Real bug?}
R -- yes --> F[AI fixes it]
R -- flaky test --> E
D -- pass --> E[Quality gate]
E -- fail --> F
F --> D
E -- pass --> G[AI code review]
G --> H[Open pull request]
It was asked to add a feature to a codebase that had two problems deliberately planted in it.
- A human reviewed and approved the plan before any code was written.
- The tests caught a broken parser, and the system fixed it on its first attempt.
- The SonarQube quality gate caught 8 issues, including hardcoded credentials. The system fixed those too, which brought coverage from 87% to 97%.
- A second AI reviewer then found a real bug that both the tests and the quality gate had missed.
- About 6 minutes of work from approval to finish, with no human edits.
- AI agents that are safe to leave running: human approval before any code is written, retry limits, and recovery after a crash
- Several AI attempts racing in parallel, with the best one kept
- Built for real codebases, not just toy ones: it maps the repository, breaks large changes into approved steps, and reruns only the tests a step affects
- Guardrails a team would ask for: run limits, test runs in a locked-down container, a check for flaky tests and broken environments before any repair, and tracing that plugs into standard monitoring tools
- A benchmark of 15 tasks with hidden acceptance tests, to measure the system against a plain AI agent rather than describe it
- Claude Code, SonarQube, the Model Context Protocol (MCP), Docker, Git, Python, OpenTelemetry
- 313 automated tests at 93% coverage
No login and no Docker. A scripted AI stands in for Claude, and everything else is real: the git repository, the tests that fail, the fixes, and the human approval, which is you.
git clone https://github.com/AdityaPandey172/Agentic-Graph
cd Agentic-Graph
python -m venv .venv
.venv\Scripts\python.exe -m pip install -r requirements-dev.txt
.\aeg.cmd demo --openIt narrates each step in plain English, pauses for you to approve the plan,
and finishes by opening a one-page report of the run that anyone can read.
.\aeg.cmd report <run-id> produces the same page for any run, live ones
included.
The graph lives outside the model. The tempting design is one big agent with tools that decides what to do next. That version has no recovery points, no approval gate, and no graph. It is a chat loop with a prompt.
Here, a dull deterministic dispatcher owns control flow and the model is confined to the inside of individual nodes. There are no model calls, no clocks and no randomness in the routing. Three consequences fall out of that one constraint:
- A node is an executable. It reads a run directory, does one thing, writes one JSON state file, exits. A node that hangs or crashes cannot take the dispatcher with it, and its tool permissions are enforced from outside.
- The run directory is the state store. No database. A paused run is a
directory sitting on disk, and
catandlsare the admin tools. - Crash recovery and the human pause are the same primitive. Both are "stop, leave the directory, resume later", so the resume path is exercised constantly rather than only after a real crash.
Only six of the seventeen nodes call a model at all. The rest (the repository map, the test runner, the diagnosis, the quality gate, the join, the writer, the shipper) are ordinary code, because their answers have to mean the same thing every time they are asked.
| Python | 3.10+ (developed on 3.11) |
| Git | any recent version |
| An agentic CLI | Claude Code, logged in |
| Docker | for SonarQube, and for the MCP backed reviewer |
gh |
only for the ship node |
| Java | bundled with sonar-scanner, no separate install needed |
python -m venv .venv
.venv\Scripts\python.exe -m pip install -r requirements-dev.txtLog in to the agentic CLI:
claude auth login
claude auth status # expect "loggedIn": trueStart SonarQube:
docker compose -f infra/docker-compose.yml up -dWait for curl http://localhost:9000/api/system/status to report UP, then in
the UI create a Global Analysis Token and save it:
Set-Content -NoNewline -Encoding ascii infra\.sonar-token "PASTE_TOKEN_HERE"Configure the quality gate. This step is not optional. SonarQube's stock
gate only evaluates new code, so on a fresh project it passes no matter what
the agent writes, and a gate that cannot fail is not a gate. Create a gate with
conditions on Overall Code (violations > 0, coverage < 90) and make it
the default. Avoid any new_* condition: those measure against analysis
history, so the same code can get different verdicts on consecutive runs.
Install sonar-scanner into tools/, from
https://binaries.sonarsource.com/Distribution/sonar-scanner-cli/. The path is
configured in graph.yaml.
The fixture. fixtures/json-parser/ is a strict RFC 8259 JSON parser with
an 81 case test suite, the task the graph plans against, codes, tests, reviews
and repairs. It is its own git repository and is not tracked here.
Always go through .\aeg.cmd rather than calling python -m aeg.cli directly:
it starts the venv interpreter with a scrubbed environment, so a run means
the same thing regardless of what the surrounding shell has set.
.\aeg.cmd run --goal "Add a dumps() serializer, with tests" --workspace ..\workspaceThe run plans, then stops and waits for you:
.\aeg.cmd plan <run-id> # read it
.\aeg.cmd approve <run-id> --resume # edit it first if you like
.\aeg.cmd reject <run-id> -m "wrong approach"Then:
.\aeg.cmd list # every run, read straight off the manifests
.\aeg.cmd show <run-id> # per node table: outcome, duration, commit
.\aeg.cmd timeline <run-id> # the path the run actually took
.\aeg.cmd resume <run-id> # continue after a crash, or after approval
.\aeg.cmd report <run-id> # a one-page HTML report anyone can read
.\aeg.cmd demo # the whole loop with a scripted agent
.\aeg.cmd trace <run-id> # the run as an OpenTelemetry traceUseful flags: --from <node> enters the graph part way, --agent stub runs the
whole thing with a scripted agent, --graph <file> picks a pipeline,
--max-minutes caps a run's working time, --json on timeline for machine
readable output.
Routing is data, so choosing a pipeline is picking a file rather than editing code:
| File | Shape |
|---|---|
graph.yaml |
plan, approve, code, test, review, write |
graph-fanout.yaml |
plan, approve, fanout, join, write |
graph-candidate.yaml |
the sub graph one fanout candidate runs |
graph-ship.yaml |
the full pipeline, ending in advise and a pull request |
graph-local.yaml |
the same pipeline without the pull request, for local runs and evaluation |
graph-large.yaml |
context, a decomposed plan, subtasks, then the full checks once |
graph-subtask.yaml |
the sub graph one step of a decomposed plan runs |
A node reports an outcome from a closed vocabulary, and the graph decides where that leads. Nodes never name their successor, so a confused node cannot steer the run somewhere arbitrary. Retry ceilings live in the graph too, because they are a property of the pipeline rather than of the node that keeps failing:
limits:
fix:
max_attempts: 3
on_exceeded: give_upModel and CLI are per node. graph-ship.yaml plans and reviews on a lighter
model than it codes with, and agent: on a node names a different CLI entirely.
Nothing else changes when you swap either, because the only contract a node has
with the rest of the system is the state file it writes.
plan:
impl: aeg.nodes.plan
agent: claude_code
model: claude-sonnet-5| Node | Kind | What it does |
|---|---|---|
context |
deterministic | Maps the repository and ranks files against the goal, with reasons. |
plan |
agentic | Writes a plan. Read only tools: a planner that can edit code will edit code. |
decompose |
agentic | A plan as a list of small steps, under the name plan. |
await_approval |
deterministic | Reports waiting, approved or rejected. Does not block, the run parks. |
code |
agentic | Implements the approved plan, then commits. |
test |
deterministic | Runs pytest, optionally only the affected tests, optionally in a sandbox. |
diagnose |
deterministic | Reruns failures: flaky moves on, a broken environment parks, a bug goes to fix. |
review |
deterministic | sonar-scanner plus the web API. Gate status and findings. |
review_sonar_mcp |
agentic | Same verdict via the SonarQube MCP server, same payload shape. |
advise |
agentic | A second opinion, kept outside the deterministic gate. |
fix |
agentic | Repairs whatever the last gate rejected, with attempt history. |
fanout |
orchestration | N candidates in isolated worktrees, in parallel. |
join |
orchestration | Discards failures, picks a winner, fast forwards the trunk. |
subtasks |
orchestration | Runs approved steps in order, each as its own run, surviving crashes between them. |
write |
deterministic | Assembles out/. |
ship |
deterministic | Pushes the branch and opens a pull request. |
give_up |
deterministic | Ends a run and explains why. |
runs/<run-id>/
run.json goal, workspace, graph, agent
goal.md
approval.json the human decision, hashed against the plan
manifest.jsonl append only, fsynced: drives resume, list and timeline
main/
01-plan.json one file per node execution
plan.md editable before approval
02-await_approval.json
...
budget.json which limit stopped the run, if one did
candidates/<id>/ a fanout candidate is itself a run
subtasks/<nn-id>/ so is each step of a decomposed plan
workspaces/<id>/ its git worktree
out/ summary.md, plan.md, changes.diff, report.html, trace.json
Every file is plain JSON and readable on its own. 04-code.json tells you what
ran, under which tool scope and model, against which commit.
Node results are written atomically (temp file plus os.replace) and manifest
lines are fsynced, so a kill leaves at worst a torn final line, which is
discarded on read because it describes work whose completion was never
confirmed.
The subtler half is the workspace. State is transactional, a working tree is not. So every node that touches code commits, and the manifest records the SHA. Resume resets the tree to the last good checkpoint before re entering the interrupted node, comparing HEAD to the checkpoint rather than just checking for dirtiness, because a node that committed and then died leaves a clean tree one commit ahead with no manifest line to say so.
A filelock on the run directory stops a second dispatcher driving a run that is
already in flight.
ship is the exception, and it is worth reading
aeg/nodes/ship.py for it. You cannot un-open a pull
request, so a crash between "created" and "recorded" would make a resume open a
second one for the same work. The answer is not a retry counter but a natural
idempotency key: the branch name comes from the run id, and the node asks the
forge whether a pull request already exists for that branch before creating
anything. Re-running converges on the same pull request instead of multiplying
them.
review is deterministic on purpose. The same input produces identical findings
every time, which is what makes a quality gate meaningful and what lets the fix
loop tell progress from noise.
advise is an AI reviewer, and it does not have that property. Merging its
opinions into the same node would quietly destroy the repeatability, so it is a
separate node positioned after the gate, with its own setting:
advise:
config:
advisory_only: true # findings are recorded and attached to the pull
# request, but the run always proceedsSet advisory_only: false and serious findings route to fix like any other
failed gate, sharing the same repair budget. Set command: to run an external
reviewer such as Sonar's Gitar instead of the agentic CLI, emitting the same
findings JSON on stdout.
graph-large.yaml is the pipeline for changes too big for one pass, on
repositories too big to read whole. Four additions make that work, and each
one is deterministic:
- A repository map before planning. The
contextnode lists every file (throughgit ls-files, so ignored files stay out), extracts symbols and imports, and ranks files against the goal, following one hop through the import graph so a module's tests and callers come along. Every ranked file carries the reasons it was ranked. The map goes into the plan, code and fix prompts, and it is on disk for anyone asking why the agent looked where it did. - Decomposition. The planner returns a short approach and a JSON list of
steps. A person approves or edits the steps, and each one runs as its own
sub-run (
graph-subtask.yaml: code, tests, fix), landing as its own commits. A step that gives up stops the sequence rather than letting later steps build on it. - Only the affected tests inside each step. The selector picks tests that import a changed module or package, and runs the full suite whenever it cannot be sure: a changed conftest or pyproject, a deleted file, a module nothing imports. It records its reason either way. The full suite then runs once, on the trunk, over the whole change.
- A quality gate over the change. With
scope: changed, SonarQube analyses only the files the run changed, so a legacy codebase's existing debt does not fail every change, and the scan takes seconds rather than minutes.
The hard part is a crash between steps. Steps share one workspace, and resuming the trunk rewinds it to a checkpoint from before the first step. Finished steps are therefore pinned with git refs, and the subtasks node moves the workspace forward to the last finished step before continuing. A test kills the trunk inside that node and checks that the finished step is neither lost nor redone.
Budgets. A retry ceiling bounds one loop; a budget bounds the whole run, across every loop, step and candidate:
budget:
max_node_seconds: 1800
max_node_executions: 60
on_exceeded: give_upor per run, with aeg run --max-minutes 30. The dispatcher checks between
nodes, counts sub-runs too, and ignores time spent waiting for a person. When a
limit is reached it records budget.json and routes to give_up, which says
which limit was reached. Fanout candidates and decomposed steps inherit the
remaining limits.
A sandbox for the tests. Running the tests means running code a model just wrote. With a sandbox configured, the test node runs it in a throwaway container with no network, capped memory, CPU and processes, and only the workspace mounted. A container that outlives its time budget is killed by name, because killing the Docker client does not stop it.
test:
config:
sandbox:
image: aeg-sandbox:py311 # docker build -t aeg-sandbox:py311 infra/sandboxDiagnosis before repair. The fix loop assumes every failure is a bug. The
diagnose node reruns the failing tests once first: a flaky test that passes
moves on without anyone "fixing" correct code, and a failure caused by the
machine (a missing dependency, a service that is down, a full disk) parks the
run with blocked instead of burning three repair attempts. Fix the machine,
aeg resume, and the tests run again.
Skills. A node can be given a team's written standards:
code:
skills: [python-conventions, testing-standards]A repository's own .aeg/skills/ beats the defaults in skills/, and skills
already written for Claude Code in .claude/skills/ are picked up as they are.
Each node records which skills it used and a fingerprint of each, so a change
can be traced to the exact version of the standard it was written against. A
missing skill stops the node rather than quietly writing code without it.
Tracing. aeg trace <run-id> exports a run as OpenTelemetry, and
--endpoint sends it to any OTLP collector (Jaeger, Grafana Tempo, Honeycomb,
Datadog). The run is the root span, each node a child, candidates and steps nest
under the node that started them, and model calls carry the standard
gen_ai.* attributes. It is built from the manifest, so there is nothing extra
to instrument and nothing new to install.
One good run is an anecdote, so aeg eval measures the graph against a fixed
suite of 15 tasks: six bug fixes, seven features and two security fixes, all
on the same small project. Each task runs under two conditions:
- graph: the full pipeline in
graph-local.yaml. - baseline: the same model given the same goal once, with every tool, including running the tests itself.
Neither condition grades itself. Verification copies the final code, restores
the project's original tests so weakening one does not help, and adds hidden
acceptance tests the agent never saw. aeg eval validate proves every task
fails on its starting code and passes with its reference solution, so an agent
that does nothing cannot score. Results carry 95% intervals, and a dry run with
the scripted agent is labelled as measuring nothing.
.\aeg.cmd eval validate # every task is fair
.\aeg.cmd eval run --agent stub --review stub # dry run of the harness
.\aeg.cmd eval run --repeats 3 --quality # the real measurementSee evals/README.md for how tasks are built.
.venv\Scripts\python.exe -m pytest313 tests at 93% coverage. Very little is mocked: the tests spawn real node processes, run real pytest inside scratch git repositories, make real commits and worktrees, let processes hang so their time budgets are tested, kill the dispatcher mid node to prove resume works, and send traces to a real HTTP server.
Nothing external is required. No login, no network, no Docker.
The agentic CLI is replaced by a scripted stub, the reviewer by a stub, and
gh by a fake backed by a local bare remote, Docker by a fake that records its
arguments. The one test that needs a live
SonarQube skips itself when there is none.
Every node runs as its own process, so coverage has to follow child processes:
set COVERAGE_PROCESS_START=.coveragerc
.venv\Scripts\python.exe -m pytest --cov=aeg --cov-report=term-missingEach test file mirrors the acceptance criteria for its step, for example
tests/test_step5_review.py.
All ten steps are complete, plus the work that followed them, with 313 tests at 93% coverage.
| Step | What it added | Tests | Verified |
|---|---|---|---|
| 0 | Environment, and a JSON parser fixture with an 81 case suite | - | live: the gate goes red on a planted issue and green on clean code |
| 1 | Nodes as executables, the dispatcher, graph.yaml, the state contract |
7 | live |
| 2 | The test node, the fix loop, failure signatures, give_up |
7 | live: a planted regression failed a test, and fix repaired it in one attempt |
| 3 | Crash resume from the manifest, workspace checkpointing, list |
10 | live: resumed twice mid run, plus a real process tree kill |
| 4 | The human approval gate, approve / reject, plan hashing |
12 | live: a real plan was reviewed and approved before any code was written |
| 5 | The SonarQube quality gate, findings routed to fix |
15 | live: caught a hardcoded credential the tests were happy with |
| 6 | The same review through the SonarQube MCP server | 13 | live: reached the same verdict and the same findings as the deterministic reviewer |
| 7 | fanout and join, candidates racing in isolated worktrees |
12 | real SonarQube per candidate |
| 8 | The run timeline, including parallel candidates | 13 | rendered from real runs |
| 9 | ship, advise, per node CLI and model, usage tracking |
22 | live: planning and review on Sonnet 5, code and fixes on Opus 5 |
| - | How the agentic CLI is invoked | 7 | regression cover for a bug found in a live run |
| - | The command line interface | 30 | every command, including unknown run ids and bad arguments |
| - | Time budgets and process cleanup | 10 | regression cover for a 5 second budget that took 61 seconds |
| - | The evaluation harness and its 15 tasks | 50 | every task proven fair: fails at the start, passes when solved |
| - | aeg demo, aeg report, plain-English narration |
35 | the demo drives the real graph; reports escape model output |
| - | Repository map, decomposition, affected tests, scoped gate | 34 | a crash inside the steps loses no finished step |
| - | Budgets, sandbox, traces, diagnosis, skills | 36 | a real OTLP receiver, real flaky and broken-environment tests |
"Live" means a real model and real infrastructure.