Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Agentic Engineering Graph

An AI system that writes code the way a careful engineering team does: it plans, waits for a human to approve the plan, writes the code, tests it, has it checked by a quality gate, and fixes its own mistakes until everything passes.

flowchart LR
    A[Plan] --> B{Human approves?}
    B -- yes --> C[Write code]
    B -- no --> X[Stop]
    C --> D[Run tests]
    D -- fail --> R{Real bug?}
    R -- yes --> F[AI fixes it]
    R -- flaky test --> E
    D -- pass --> E[Quality gate]
    E -- fail --> F
    F --> D
    E -- pass --> G[AI code review]
    G --> H[Open pull request]
Loading

What it did on a real run

It was asked to add a feature to a codebase that had two problems deliberately planted in it.

  • A human reviewed and approved the plan before any code was written.
  • The tests caught a broken parser, and the system fixed it on its first attempt.
  • The SonarQube quality gate caught 8 issues, including hardcoded credentials. The system fixed those too, which brought coverage from 87% to 97%.
  • A second AI reviewer then found a real bug that both the tests and the quality gate had missed.
  • About 6 minutes of work from approval to finish, with no human edits.

What this demonstrates

  • AI agents that are safe to leave running: human approval before any code is written, retry limits, and recovery after a crash
  • Several AI attempts racing in parallel, with the best one kept
  • Built for real codebases, not just toy ones: it maps the repository, breaks large changes into approved steps, and reruns only the tests a step affects
  • Guardrails a team would ask for: run limits, test runs in a locked-down container, a check for flaky tests and broken environments before any repair, and tracing that plugs into standard monitoring tools
  • A benchmark of 15 tasks with hidden acceptance tests, to measure the system against a plain AI agent rather than describe it
  • Claude Code, SonarQube, the Model Context Protocol (MCP), Docker, Git, Python, OpenTelemetry
  • 313 automated tests at 93% coverage

Try it yourself in a minute

No login and no Docker. A scripted AI stands in for Claude, and everything else is real: the git repository, the tests that fail, the fixes, and the human approval, which is you.

git clone https://github.com/AdityaPandey172/Agentic-Graph
cd Agentic-Graph
python -m venv .venv
.venv\Scripts\python.exe -m pip install -r requirements-dev.txt
.\aeg.cmd demo --open

It narrates each step in plain English, pauses for you to approve the plan, and finishes by opening a one-page report of the run that anyone can read. .\aeg.cmd report <run-id> produces the same page for any run, live ones included.


How it works (for engineers)

The idea

The graph lives outside the model. The tempting design is one big agent with tools that decides what to do next. That version has no recovery points, no approval gate, and no graph. It is a chat loop with a prompt.

Here, a dull deterministic dispatcher owns control flow and the model is confined to the inside of individual nodes. There are no model calls, no clocks and no randomness in the routing. Three consequences fall out of that one constraint:

  • A node is an executable. It reads a run directory, does one thing, writes one JSON state file, exits. A node that hangs or crashes cannot take the dispatcher with it, and its tool permissions are enforced from outside.
  • The run directory is the state store. No database. A paused run is a directory sitting on disk, and cat and ls are the admin tools.
  • Crash recovery and the human pause are the same primitive. Both are "stop, leave the directory, resume later", so the resume path is exercised constantly rather than only after a real crash.

Only six of the seventeen nodes call a model at all. The rest (the repository map, the test runner, the diagnosis, the quality gate, the join, the writer, the shipper) are ordinary code, because their answers have to mean the same thing every time they are asked.

Requirements

Python 3.10+ (developed on 3.11)
Git any recent version
An agentic CLI Claude Code, logged in
Docker for SonarQube, and for the MCP backed reviewer
gh only for the ship node
Java bundled with sonar-scanner, no separate install needed

Setup

python -m venv .venv
.venv\Scripts\python.exe -m pip install -r requirements-dev.txt

Log in to the agentic CLI:

claude auth login
claude auth status        # expect "loggedIn": true

Start SonarQube:

docker compose -f infra/docker-compose.yml up -d

Wait for curl http://localhost:9000/api/system/status to report UP, then in the UI create a Global Analysis Token and save it:

Set-Content -NoNewline -Encoding ascii infra\.sonar-token "PASTE_TOKEN_HERE"

Configure the quality gate. This step is not optional. SonarQube's stock gate only evaluates new code, so on a fresh project it passes no matter what the agent writes, and a gate that cannot fail is not a gate. Create a gate with conditions on Overall Code (violations > 0, coverage < 90) and make it the default. Avoid any new_* condition: those measure against analysis history, so the same code can get different verdicts on consecutive runs.

Install sonar-scanner into tools/, from https://binaries.sonarsource.com/Distribution/sonar-scanner-cli/. The path is configured in graph.yaml.

The fixture. fixtures/json-parser/ is a strict RFC 8259 JSON parser with an 81 case test suite, the task the graph plans against, codes, tests, reviews and repairs. It is its own git repository and is not tracked here.

Usage

Always go through .\aeg.cmd rather than calling python -m aeg.cli directly: it starts the venv interpreter with a scrubbed environment, so a run means the same thing regardless of what the surrounding shell has set.

.\aeg.cmd run --goal "Add a dumps() serializer, with tests" --workspace ..\workspace

The run plans, then stops and waits for you:

.\aeg.cmd plan <run-id>                     # read it
.\aeg.cmd approve <run-id> --resume         # edit it first if you like
.\aeg.cmd reject <run-id> -m "wrong approach"

Then:

.\aeg.cmd list                  # every run, read straight off the manifests
.\aeg.cmd show <run-id>         # per node table: outcome, duration, commit
.\aeg.cmd timeline <run-id>     # the path the run actually took
.\aeg.cmd resume <run-id>       # continue after a crash, or after approval
.\aeg.cmd report <run-id>       # a one-page HTML report anyone can read
.\aeg.cmd demo                  # the whole loop with a scripted agent
.\aeg.cmd trace <run-id>        # the run as an OpenTelemetry trace

Useful flags: --from <node> enters the graph part way, --agent stub runs the whole thing with a scripted agent, --graph <file> picks a pipeline, --max-minutes caps a run's working time, --json on timeline for machine readable output.

The graphs

Routing is data, so choosing a pipeline is picking a file rather than editing code:

File Shape
graph.yaml plan, approve, code, test, review, write
graph-fanout.yaml plan, approve, fanout, join, write
graph-candidate.yaml the sub graph one fanout candidate runs
graph-ship.yaml the full pipeline, ending in advise and a pull request
graph-local.yaml the same pipeline without the pull request, for local runs and evaluation
graph-large.yaml context, a decomposed plan, subtasks, then the full checks once
graph-subtask.yaml the sub graph one step of a decomposed plan runs

A node reports an outcome from a closed vocabulary, and the graph decides where that leads. Nodes never name their successor, so a confused node cannot steer the run somewhere arbitrary. Retry ceilings live in the graph too, because they are a property of the pipeline rather than of the node that keeps failing:

limits:
  fix:
    max_attempts: 3
    on_exceeded: give_up

Model and CLI are per node. graph-ship.yaml plans and reviews on a lighter model than it codes with, and agent: on a node names a different CLI entirely. Nothing else changes when you swap either, because the only contract a node has with the rest of the system is the state file it writes.

plan:
  impl: aeg.nodes.plan
  agent: claude_code
  model: claude-sonnet-5

The nodes

Node Kind What it does
context deterministic Maps the repository and ranks files against the goal, with reasons.
plan agentic Writes a plan. Read only tools: a planner that can edit code will edit code.
decompose agentic A plan as a list of small steps, under the name plan.
await_approval deterministic Reports waiting, approved or rejected. Does not block, the run parks.
code agentic Implements the approved plan, then commits.
test deterministic Runs pytest, optionally only the affected tests, optionally in a sandbox.
diagnose deterministic Reruns failures: flaky moves on, a broken environment parks, a bug goes to fix.
review deterministic sonar-scanner plus the web API. Gate status and findings.
review_sonar_mcp agentic Same verdict via the SonarQube MCP server, same payload shape.
advise agentic A second opinion, kept outside the deterministic gate.
fix agentic Repairs whatever the last gate rejected, with attempt history.
fanout orchestration N candidates in isolated worktrees, in parallel.
join orchestration Discards failures, picks a winner, fast forwards the trunk.
subtasks orchestration Runs approved steps in order, each as its own run, surviving crashes between them.
write deterministic Assembles out/.
ship deterministic Pushes the branch and opens a pull request.
give_up deterministic Ends a run and explains why.

The run directory

runs/<run-id>/
  run.json            goal, workspace, graph, agent
  goal.md
  approval.json       the human decision, hashed against the plan
  manifest.jsonl      append only, fsynced: drives resume, list and timeline
  main/
    01-plan.json      one file per node execution
    plan.md           editable before approval
    02-await_approval.json
    ...
  budget.json         which limit stopped the run, if one did
  candidates/<id>/    a fanout candidate is itself a run
  subtasks/<nn-id>/   so is each step of a decomposed plan
  workspaces/<id>/    its git worktree
  out/                summary.md, plan.md, changes.diff, report.html, trace.json

Every file is plain JSON and readable on its own. 04-code.json tells you what ran, under which tool scope and model, against which commit.

How a run survives dying

Node results are written atomically (temp file plus os.replace) and manifest lines are fsynced, so a kill leaves at worst a torn final line, which is discarded on read because it describes work whose completion was never confirmed.

The subtler half is the workspace. State is transactional, a working tree is not. So every node that touches code commits, and the manifest records the SHA. Resume resets the tree to the last good checkpoint before re entering the interrupted node, comparing HEAD to the checkpoint rather than just checking for dirtiness, because a node that committed and then died leaves a clean tree one commit ahead with no manifest line to say so.

A filelock on the run directory stops a second dispatcher driving a run that is already in flight.

The one node that cannot be rewound

ship is the exception, and it is worth reading aeg/nodes/ship.py for it. You cannot un-open a pull request, so a crash between "created" and "recorded" would make a resume open a second one for the same work. The answer is not a retry counter but a natural idempotency key: the branch name comes from the run id, and the node asks the forge whether a pull request already exists for that branch before creating anything. Re-running converges on the same pull request instead of multiplying them.

Two reviewers, deliberately separate

review is deterministic on purpose. The same input produces identical findings every time, which is what makes a quality gate meaningful and what lets the fix loop tell progress from noise.

advise is an AI reviewer, and it does not have that property. Merging its opinions into the same node would quietly destroy the repeatability, so it is a separate node positioned after the gate, with its own setting:

advise:
  config:
    advisory_only: true    # findings are recorded and attached to the pull
                           # request, but the run always proceeds

Set advisory_only: false and serious findings route to fix like any other failed gate, sharing the same repair budget. Set command: to run an external reviewer such as Sonar's Gitar instead of the agentic CLI, emitting the same findings JSON on stdout.

Large codebases

graph-large.yaml is the pipeline for changes too big for one pass, on repositories too big to read whole. Four additions make that work, and each one is deterministic:

  • A repository map before planning. The context node lists every file (through git ls-files, so ignored files stay out), extracts symbols and imports, and ranks files against the goal, following one hop through the import graph so a module's tests and callers come along. Every ranked file carries the reasons it was ranked. The map goes into the plan, code and fix prompts, and it is on disk for anyone asking why the agent looked where it did.
  • Decomposition. The planner returns a short approach and a JSON list of steps. A person approves or edits the steps, and each one runs as its own sub-run (graph-subtask.yaml: code, tests, fix), landing as its own commits. A step that gives up stops the sequence rather than letting later steps build on it.
  • Only the affected tests inside each step. The selector picks tests that import a changed module or package, and runs the full suite whenever it cannot be sure: a changed conftest or pyproject, a deleted file, a module nothing imports. It records its reason either way. The full suite then runs once, on the trunk, over the whole change.
  • A quality gate over the change. With scope: changed, SonarQube analyses only the files the run changed, so a legacy codebase's existing debt does not fail every change, and the scan takes seconds rather than minutes.

The hard part is a crash between steps. Steps share one workspace, and resuming the trunk rewinds it to a checkpoint from before the first step. Finished steps are therefore pinned with git refs, and the subtasks node moves the workspace forward to the last finished step before continuing. A test kills the trunk inside that node and checks that the finished step is neither lost nor redone.

Guardrails for unattended runs

Budgets. A retry ceiling bounds one loop; a budget bounds the whole run, across every loop, step and candidate:

budget:
  max_node_seconds: 1800
  max_node_executions: 60
  on_exceeded: give_up

or per run, with aeg run --max-minutes 30. The dispatcher checks between nodes, counts sub-runs too, and ignores time spent waiting for a person. When a limit is reached it records budget.json and routes to give_up, which says which limit was reached. Fanout candidates and decomposed steps inherit the remaining limits.

A sandbox for the tests. Running the tests means running code a model just wrote. With a sandbox configured, the test node runs it in a throwaway container with no network, capped memory, CPU and processes, and only the workspace mounted. A container that outlives its time budget is killed by name, because killing the Docker client does not stop it.

test:
  config:
    sandbox:
      image: aeg-sandbox:py311   # docker build -t aeg-sandbox:py311 infra/sandbox

Diagnosis before repair. The fix loop assumes every failure is a bug. The diagnose node reruns the failing tests once first: a flaky test that passes moves on without anyone "fixing" correct code, and a failure caused by the machine (a missing dependency, a service that is down, a full disk) parks the run with blocked instead of burning three repair attempts. Fix the machine, aeg resume, and the tests run again.

Skills. A node can be given a team's written standards:

code:
  skills: [python-conventions, testing-standards]

A repository's own .aeg/skills/ beats the defaults in skills/, and skills already written for Claude Code in .claude/skills/ are picked up as they are. Each node records which skills it used and a fingerprint of each, so a change can be traced to the exact version of the standard it was written against. A missing skill stops the node rather than quietly writing code without it.

Tracing. aeg trace <run-id> exports a run as OpenTelemetry, and --endpoint sends it to any OTLP collector (Jaeger, Grafana Tempo, Honeycomb, Datadog). The run is the root span, each node a child, candidates and steps nest under the node that started them, and model calls carry the standard gen_ai.* attributes. It is built from the manifest, so there is nothing extra to instrument and nothing new to install.

Measuring it

One good run is an anecdote, so aeg eval measures the graph against a fixed suite of 15 tasks: six bug fixes, seven features and two security fixes, all on the same small project. Each task runs under two conditions:

  • graph: the full pipeline in graph-local.yaml.
  • baseline: the same model given the same goal once, with every tool, including running the tests itself.

Neither condition grades itself. Verification copies the final code, restores the project's original tests so weakening one does not help, and adds hidden acceptance tests the agent never saw. aeg eval validate proves every task fails on its starting code and passes with its reference solution, so an agent that does nothing cannot score. Results carry 95% intervals, and a dry run with the scripted agent is labelled as measuring nothing.

.\aeg.cmd eval validate                          # every task is fair
.\aeg.cmd eval run --agent stub --review stub    # dry run of the harness
.\aeg.cmd eval run --repeats 3 --quality         # the real measurement

See evals/README.md for how tasks are built.

Testing

.venv\Scripts\python.exe -m pytest

313 tests at 93% coverage. Very little is mocked: the tests spawn real node processes, run real pytest inside scratch git repositories, make real commits and worktrees, let processes hang so their time budgets are tested, kill the dispatcher mid node to prove resume works, and send traces to a real HTTP server.

Nothing external is required. No login, no network, no Docker. The agentic CLI is replaced by a scripted stub, the reviewer by a stub, and gh by a fake backed by a local bare remote, Docker by a fake that records its arguments. The one test that needs a live SonarQube skips itself when there is none.

Every node runs as its own process, so coverage has to follow child processes:

set COVERAGE_PROCESS_START=.coveragerc
.venv\Scripts\python.exe -m pytest --cov=aeg --cov-report=term-missing

Each test file mirrors the acceptance criteria for its step, for example tests/test_step5_review.py.

Status

All ten steps are complete, plus the work that followed them, with 313 tests at 93% coverage.

Step What it added Tests Verified
0 Environment, and a JSON parser fixture with an 81 case suite - live: the gate goes red on a planted issue and green on clean code
1 Nodes as executables, the dispatcher, graph.yaml, the state contract 7 live
2 The test node, the fix loop, failure signatures, give_up 7 live: a planted regression failed a test, and fix repaired it in one attempt
3 Crash resume from the manifest, workspace checkpointing, list 10 live: resumed twice mid run, plus a real process tree kill
4 The human approval gate, approve / reject, plan hashing 12 live: a real plan was reviewed and approved before any code was written
5 The SonarQube quality gate, findings routed to fix 15 live: caught a hardcoded credential the tests were happy with
6 The same review through the SonarQube MCP server 13 live: reached the same verdict and the same findings as the deterministic reviewer
7 fanout and join, candidates racing in isolated worktrees 12 real SonarQube per candidate
8 The run timeline, including parallel candidates 13 rendered from real runs
9 ship, advise, per node CLI and model, usage tracking 22 live: planning and review on Sonnet 5, code and fixes on Opus 5
- How the agentic CLI is invoked 7 regression cover for a bug found in a live run
- The command line interface 30 every command, including unknown run ids and bad arguments
- Time budgets and process cleanup 10 regression cover for a 5 second budget that took 61 seconds
- The evaluation harness and its 15 tasks 50 every task proven fair: fails at the start, passes when solved
- aeg demo, aeg report, plain-English narration 35 the demo drives the real graph; reports escape model output
- Repository map, decomposition, affected tests, scoped gate 34 a crash inside the steps loses no finished step
- Budgets, sandbox, traces, diagnosis, skills 36 a real OTLP receiver, real flaky and broken-environment tests

"Live" means a real model and real infrastructure.

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages