Skip to content

Latest commit

 

History

History
558 lines (467 loc) · 46 KB

File metadata and controls

558 lines (467 loc) · 46 KB

Otto — Unified AI-Development Framework

Status: Build complete (steps 1–7, 2026-06-08) · autonomous capabilities validated 13/13 + Workflow backend built & evidenced (2026-07-15, VALIDATION.md) · deferred layers: ROADMAP.md · v1.1 Name: Otto (repo otto-dev). Chosen 2026-06-08. Scope of this document: the overall design. The build follows the roadmap in §11; this doc is the decision-of-record.


1. Thesis

Three excellent, overlapping, standalone projects exist in this workspace:

Project What it really is Best at Weak at
spec-kit (GitHub, MIT) A methodology engine — Python CLI, durable artifact tree, governance constitution, a real YAML workflow engine. spec→plan→tasks, governance, deterministic multi-step automation review, ship, debug, maintain
agent-skills (Addy Osmani, MIT) A phase-complete process library — 23 SDLC skills, 3 review personas, slash commands, hooks, reference checklists. the entire back half: review, security, perf, a11y, CI/CD, ship, deprecation; anti-rationalization discipline no durable artifact/state engine; no governance file
skills (Matt Pocock, MIT) A philosophy-driven practice library — 18 skills, feedback-loop & alignment obsessed, domain-language aware. alignment ("grilling"), CONTEXT.md/ADR domain capture, diagnose, prototype, architecture deepening, handoff breadth; no spec/plan tooling; no ship/ops

They operate at different altitudes, so the right move is to layer them, not merge them. A fourth merged library would just re-create the overlap problem.

spec-kit       = SKELETON  — durable state, governance, the autonomous engine
agent-skills   = MUSCLE    — the practices, especially the back half of the lifecycle
Pocock skills  = MUSCLE    — alignment, domain language, diagnosis, architecture
Otto (new)     = NERVOUS SYSTEM — a registry + router + workflow graph that picks the
                 right muscle at the right moment, referencing all three, modifying none.

What Otto is NOT: not a fork, not a copy, not a 4th skill library, not a replacement for any of the three. It is the thin orchestration + routing + project-memory layer above them.


2. Locked design decisions

# Decision Choice Consequence
D1 Overlap resolution Best-of-breed + layer One primary unit per capability by demonstrated strength; layer a second only where genuinely complementary; document alternates with swap criteria.
D2 Tooling dependency Hybrid Adopt spec-kit's artifact tree + constitution as canonical conventions; use its CLI/workflow engine when installed, never hard-require it; orchestration layer is host-native instructions.
D3 Workflow shape Spine + recipe presets One lifecycle graph, enterable at any stage, with reusable specialist modules; named recipes are preset routes through the graph.
D4 Host target Claude Code-first, portable-ready Optimize for Claude Code (skills, slash commands, hooks, subagents/Workflow). Keep units host-portable since all three sources already are.
D5 Output location <LOCAL-PATH>\ New sibling repo. The three sources stay read-only and untouched.
D6 Modes Manual and hands-off, co-equal Manual/adhoc fits the architect→coder fenced-prompt split; hands-off chains stages with gates. Both are first-class.
D7 Coverage model Shape × Scope (2 axes) ~11 shapes (8 base + assess + extract + inception) × {standard, workspace}. Portfolio layer + domain profiles deferred (approved 2026-06-08); see §12.
D8 Collective wisdom Auto-append + distill Agent-authored learnings/ store (tiered, structured, evidence-dated): frictionless auto-capture, periodic /otto-distill curation (dedupe → promote → prune), verify-on-read. See §12 item 11.
D9 Definition of success Measured, not assumed Otto "working" = every non-trivial task leaves a gate-trail; learnings/ accrues real lessons; graceful degradation is observed; zero skipped-verification rework. Provisional until tuned against the first real run. See §12 item 14.
D10 Per-repo footprint Manifest + declared git disposition Per-repo setup ends with .otto-manifest.md (every path created/modified + disposition per row); each repo declares team mode (governance tracked — default) or private mode (footprint excluded via .git/info/exclude; zero trace in history). Consolidating into a single .otto/ dir was rejected — most paths are host/source-pinned. See §12 item 16.










3. Layered architecture

flowchart TB
    L4["<b>Layer 4 · PROJECT MEMORY</b><br/>constitution · CONTEXT.md · docs/adr · CLAUDE.md<br/><i>governs all stages</i>"]
    L3["<b>Layer 3 · WORKFLOWS</b><br/>the SPINE (enter at any stage) + specialist MODULES<br/>+ named RECIPES (preset routes) · manual + hands-off"]
    L2["<b>Layer 2 · ROUTER</b><br/>manual face (/flow dispatcher) + agentic face (hook)"]
    L1["<b>Layer 1 · CAPABILITY REGISTRY</b><br/>every unit across all 3, tagged by capability,<br/>stage · source · invocation · mode · prefer-when"]
    L0["<b>Layer 0 · SOURCES</b><br/>spec-kit · agent-skills · Pocock skills (read-only)"]
    L4 --> L3 --> L2 --> L1 --> L0
Loading





















4. Layer 0 — Sources (read-only)

The three repos are referenced by path and pinned commit. Never edited, copied, or vendored. Otto does not execute their code; it routes to units that must be installed in the host agent. Full pinned refs + install commands live in sources/sources.md.

Source Path Install into host (Claude Code) Units exposed
spec-kit <LOCAL-PATH> specify init per target repo (scaffolds .specify/, /speckit.* commands) /speckit.constitution|specify|clarify|plan|analyze|tasks|implement|checklist|taskstoissues; CLI; workflows/*.yml
agent-skills <LOCAL-PATH> plugin install (or copy skills/ into .claude/skills/) 23 skills; /spec /plan /build /test /review /code-simplify /ship; personas; reference checklists
Pocock skills <LOCAL-PATH> npx skills@latest add mattpocock/skills (select promoted) /grill-me /grill-with-docs /to-prd /to-issues /tdd /prototype /diagnose /improve-codebase-architecture /zoom-out /triage /handoff /git-guardrails-claude-code …

Read-only contract: sources/sources.md records each repo's path, pinned ref, license (all MIT), and the install command. A drift check (optional) warns if a pinned unit's invocation name changed upstream.

Realism note (consequence of D2): because the units live in the three host installs, Otto's value is selection + sequencing + memory + the install recipe — not re-implementation. If a source isn't installed, the router degrades gracefully to whatever is present and says so.


5. Layer 1 — Capability registry (overlap resolution)

The heart of D1. Every capability maps to one primary + optional layer/alternate with an explicit prefer-when. Several apparent overlaps dissolve once split into adjacent-but-distinct capabilities (e.g. elicit-intent ≠ clarify-spec; simplify-local ≠ improve-architecture; git-practice ≠ git-safety).

Legend — source: SK spec-kit · AS agent-skills · PK Pocock. Canonical machine-readable version: registry/capabilities.yaml.

Two orthogonal pins ride alongside each entry, each with its own policy companion and the same inherit-by-default rule (no field ⇒ inherit the session): compute: — model + effort (registry/compute.md, seeded 2026-07-18) — and tools: — least-privilege profile (registry/tool-scope.md, seeded 2026-07-21). Both are binding in hands-off runs on the Workflow backend and advisory elsewhere. The tables below list capability routing only; the pins live in the YAML.

Define / Align

Capability Primary Layer / Alternate Prefer-when
govern (principles) SK /speckit.constitution → constitution.md — always, once per project
ideate / refine AS idea-refine — rough idea needs diverge→converge
elicit-intent AS interview-me (one-Q-at-a-time → ~95% confidence) PK /grill-me (relentless stress-test) AS for requirement extraction; PK to break a confident-but-shaky plan
capture-domain-language PK /grill-with-docs → writes CONTEXT.md + ADRs — whenever domain vocabulary is fuzzy (unique to PK)
harvest-backlog OTTO inline — index the idea bundle's deferred scope into root BACKLOG.md (link, don't copy) — DEFINE input is an idea bundle (e.g. idea-to-build) carrying deferred scope
classify-bundle OTTO inline — read a dropped-in idea bundle, resolve its units by content, write .otto-bundle.yaml (tools: inspect — read-only over untrusted docs) — a bundle is present and its adapter is absent or stale
apply-bundle OTTO inline — execute the extraction map's seed/link rows into the memory stack; delegates link rows to harvest-backlog — a reviewed .otto-bundle.yaml has pending rows

The bundle pair is deliberately split along the read/write line — the shape extract-spec cannot express and F-12 defers. Both are consumed by the inception feeder; four extraction modes (seed · link · input · evidence) each carry their own bundle-vs-repo precedence rule. Design record: assessments/2026-07-25-idea-bundle-intake.md.

Spec

Capability Primary Layer / Alternate Prefer-when
author-spec SK /speckit.specify → specs/NNN/spec.md (+branch, [NEEDS CLARIFICATION]) layer AS spec-driven-development template depth (objective/commands/structure/code-style/testing/boundaries) into the SK spec; alt PK /to-prd SK for artifact-tracked SDD; PK /to-prd when target is a lightweight PRD/issue, no tree
clarify-spec SK /speckit.clarify (resolves markers in the written spec) — a spec exists with open [NEEDS CLARIFICATION]
validate-requirements SK /speckit.checklist → specs/NNN/checklists/*.md ("unit tests for English" — audits requirement clarity/completeness) — requirement quality is load-bearing: 0→1, external contracts, contested requirements. Catches confidently-written-but-ambiguous requirements that clarify-spec cannot (it only closes markers the author already flagged)

Plan

Capability Primary Layer / Alternate Prefer-when
plan-architecture SK /speckit.plan → plan.md, research.md, data-model.md, contracts/ (constitution gates) — always for non-trivial features
decompose-tasks SK /speckit.tasks → tasks.md ([P] parallel, story-sliced, TDD-ordered) layer AS planning-and-task-breakdown (vertical-slicing, dep graphs, acceptance criteria); alt PK /to-issues SK tasks.md in-repo; PK /to-issues when work lives in GitHub/GitLab issues
cross-check SK /speckit.analyze (spec↔plan↔tasks consistency) — before committing to implementation (unique to SK)

Build

Capability Primary Layer / Alternate Prefer-when
implement AS incremental-implementation (thin-slice discipline) driven by SK tasks.md//speckit.implement when present the two compose: SK supplies the ordered loop, AS supplies slice-sizing + per-slice verification tasks.md exists → SK loop + AS discipline; else AS standalone
tdd AS test-driven-development (+ Prove-It bug pattern + test-engineer persona) layer PK /tdd tracer-bullet framing (anti-horizontal-slice) AS as default; PK when user wants its tracer-bullet style
prototype PK /prototype (throwaway to answer one question) — spike/feasibility before committing (unique to PK)
source-ground AS source-driven-development (+ context7 MCP) — framework/library correctness matters (unique)
doubt-check AS doubt-driven-development (fresh-context adversarial review) complements PK grilling high-stakes/irreversible/unfamiliar decisions
context-engineer AS context-engineering (rules files) alt PK /handoff (compaction/transition) session setup / quality degradation
frontend-ui AS frontend-ui-engineering alt env frontend-design skill building user-facing UI
api-design AS api-and-interface-design — designing any public interface/contract

Verify

Capability Primary Layer / Alternate Prefer-when
debug PK /diagnose (feedback-loop → reproduce → hypothesize → instrument → fix → cleanup) layer AS debugging-and-error-recovery "guard with regression test" + Prove-It toss-up; PK primary for instrumentation rigor (D4 confirmed); AS to pair the guard test with TDD
browser-verify AS browser-testing-with-devtools (Chrome DevTools MCP: DOM/console/network/perf) — anything running in a browser (matches your DOM-verification standard; unique)

Review

Capability Primary Layer / Alternate Prefer-when
code-review AS code-review-and-quality + code-reviewer persona (5-axis) alt env /code-review, /review before merge
security-review AS security-and-hardening + security-auditor persona + checklist alt env /security-review untrusted input, auth, storage, integrations (unique)
perf-review AS performance-optimization + checklist — perf budgets / Core Web Vitals (unique)
a11y-review AS accessibility checklist — user-facing UI (unique)
simplify-local AS code-simplification (+ env /simplify) — working code that's harder to read than needed
improve-architecture PK /improve-codebase-architecture (deepening + ADRs + HTML report) — structural/architectural improvement (distinct from local simplify; unique)

Ship

Capability Primary Layer / Alternate Prefer-when
git-practice AS git-workflow-and-versioning (trunk-based, atomic commits) complements git-safety advisory; you do git manually
git-safety PK /git-guardrails-claude-code (PreToolUse hook blocks dangerous git) complements git-practice install once; high value given manual-git preference
ci-cd AS ci-cd-and-automation — pipelines / quality gates (unique)
ship-launch AS shipping-and-launch + /ship parallel persona fan-out → go/no-go — production launch (flagship orchestration; unique)
deprecate / migrate AS deprecation-and-migration alt PK migrate-to-shoehorn (narrow TS) sunsetting/migrating systems
document / adr AS documentation-and-adrs layer PK inline ADR writes from /grill-with-docs decisions + API changes


Maintain / Cross-cutting

Capability Primary Layer / Alternate Prefer-when
triage PK /triage (state-machine labels, issue-tracker aware) — incoming bugs/enhancements (unique)
handoff / compact PK /handoff — agent transition / context running out (unique)
understand / onboard PK /zoom-out + architecture map — unfamiliar code (unique)
orchestrate SK workflow.yml engine (when backbone present) alt Claude Code Workflow / conductor skill hands-off multi-step runs
personas AS code-reviewer · security-auditor · test-engineer — any review/ship fan-out

spec-kit v0.9.5 extension alternates (opt-in): the bug extension adds /speckit-bug-{assess,fix,test} as an artifact-tracked alternate for the bugfix path (on debug); the git extension adds read-only validate/remote helpers under git-practice (mutating git ops stay human-gated — honors D6); agent-context auto-refreshes the CLAUDE.md SPECKIT block (under context-engineer). Enable with specify extension add <id>.

Registry format: machine-readable registry/capabilities.yaml (one record per capability with primary, alternates[], stage, source, invocation, modes, prefer_when, requires_install) plus a generated human-readable view. The router and the drift-check both read it.


6. Layer 2 — Router

One selection brain, two faces.

Manual face — /flow dispatcher (adhoc). Given a one-line intent + a glance at repo state (does specs/ exist? a failing test? an unstaged diff? a constitution.md?), it answers: which recipe → which stage → which capability → the exact invocation, then offers to run it. You can also bypass it and invoke any capability directly. This is the mode that fits your architect→coder split: the architect's fenced prompt names a capability; the router resolves it to the right source unit and the current artifact context.

Agentic face — session-start hook (auto-select). Mirrors agent-skills' using-agent-skills pattern: injects the registry + routing policy at session start so the agent auto-selects the right unit when you describe a task, and — in hands-off mode — advances the spine through gates autonomously, stopping only when a gate fails or a decision genuinely needs you.

Routing policy (router/routing-policy.md) encodes: stage detection from artifacts, the prefer-when swap criteria from §5, graceful degradation when a source isn't installed, and the gate definitions that must pass before advancing.


7. Layer 3 — Workflows

The spine (enter at any stage)

flowchart LR
    M["<b>Layer 4 project memory</b> governs every stage<br/>constitution · CONTEXT.md · docs/adr · CLAUDE.md"]
    subgraph spine["Enter at ANY stage · each stage has a GATE"]
        direction LR
        S0["<b>0 FRAME</b><br/>govern · ideate"] --> S1["<b>1 DEFINE</b><br/>elicit · domain-language"] --> S2["<b>2 SPEC</b><br/>author · clarify"] --> S3["<b>3 PLAN</b><br/>plan · tasks · analyze"] --> S4["<b>4 BUILD</b><br/>implement · tdd · prototype<br/>source · doubt"] --> S5["<b>5 VERIFY</b><br/>debug · browser-verify"] --> S6["<b>6 REVIEW</b><br/>code · security · perf/a11y<br/>simplify · improve-arch"] --> S7["<b>7 SHIP</b><br/>git-* · ci-cd · launch · deprecate"] --> S8["<b>8 MAINTAIN</b><br/>triage · handoff · understand"]
    end
    M -. governs .-> spine
Loading

Full graph with per-stage gates and entry signals: workflows/spine.md.

  • Gates = per-stage exit criteria (verification-first, from agent-skills) plus /speckit.analyze style consistency checks. No advance without passing — this is where "seems right" is forbidden and evidence is required.
  • Specialist modules are callable from any stage, not just their home: debug during BUILD, review-fan-out before SHIP, improve-architecture during MAINTAIN, handoff anywhere context runs low.
















Recipes (preset routes through the spine)

Recipe Route Primary sources
greenfield (0→1) FRAME→DEFINE→SPEC→PLAN→BUILD→VERIFY→REVIEW→SHIP SK front, AS back
feature (existing repo) light DEFINE→mini-SPEC→PLAN/tasks→BUILD(tdd)→REVIEW→SHIP (skip FRAME) SK + AS
bugfix / incident enter VERIFY→diagnose→prove-it test→BUILD(fix)→REVIEW→SHIP PK + AS
hardening enter REVIEW→security+perf+a11y fan-out→BUILD(fixes)→re-REVIEW AS personas
refactor / architecture enter REVIEW→improve-architecture→ADR→PLAN(migration)→BUILD→VERIFY PK + SK
release enter SHIP→/ship fan-out→go/no-go→ci-cd→launch→monitor/rollback AS
onboarding / understand enter MAINTAIN→zoom-out→arch map→capture domain language→CONTEXT.md PK
migration / deprecation enter SHIP/MAINTAIN→deprecation-and-migration AS
assess (feeder) investigate→survey→findings.md→hand off to a recipe ENV+AS+PK
extract (feeder) reverse: understand/extract-spec→spec/docs→link ENV+PK+AS
inception (feeder) idea bundle → classify-bundle→apply-bundle→seeded memory stack + one first slice → greenfield OTTO

Two modes per recipe (D6)

  • Manual variant — the router suggests the next step; you (or your architect's fenced prompt) invoke each capability and inspect the gate. Pause/resume is free because state lives in artifacts.
  • Hands-off variant — stages chained with automatic gate checks. Execution backend, per the conductor's selection table: (a) Claude Code Workflow/subagents — workspace scope, heterogeneous steps, or parallel fan-out (built & evidenced 2026-07-15, §12 item 17; protocol: workflows/autonomous/workflow-backend.md); (b) spec-kit workflow.yml engine — linear, spec-kit-dominant flows; (c) the Otto conductor skill — linear single-repo runs (cheaper) and portable hosts; halts at any failed gate or human-decision point. See workflows/autonomous/conductor.md.

Coverage model: Shape × Scope (D7)

A task = Shape (what kind of work) × Scope (blast radius). This keeps coverage broad without recipe sprawl — the answer to "8 recipes felt too few": 8 was too few shapes, and scope was a missing axis.

  • Shape — the 8 base recipes + three feeders: assess (investigate → findings → scope; hands off to another recipe), extract (reverse: code/app → spec/docs) and inception (idea bundle → seeded repo + one first slice, 0→1 only). ~11 shapes.
  • Scope — an overlay on any shape: standard (one repo) or workspace (coordinated across N repos/platforms — a workspace spec + inter-repo contracts + per-repo sub-specs + dependency sequencing + a cross-repo integration gate). See workflows/scope.md.

~10 shapes × 2 scopes = broad coverage from a small file set. Deferred dials (approved 2026-06-08, §12 items 6–7): a Portfolio layer for the standing stream of work, and domain profiles for per-surface gate/deploy semantics (hardware OTA, ML/CV eval).


8. Layer 4 — Project-memory stack (the bridge)

The strongest synthesis: the three projects' "memory" files are complementary, not overlapping. Otto unifies them into one stack and defines which stage reads/writes each. No single source has all four.

File Owner convention Holds Written by Read by
.specify/memory/constitution.md spec-kit principles / governance FRAME (/speckit.constitution) every gate, esp. PLAN
CONTEXT.md (repo root) Pocock domain language / glossary DEFINE (/grill-with-docs) SPEC, BUILD, REVIEW
docs/adr/*.md shared decisions (append-only) any stage (documentation-and-adrs, grill-with-docs) PLAN, REVIEW, refactor
CLAUDE.md / AGENTS.md agent-skills agent operating rules, stack, commands, boundaries DEFINE/BUILD (context-engineering) every session (host auto-loads)
stack.yaml (Otto) Otto-native pinned runtimes/frameworks/packages + authoritative docs DEFINE/FRAME; /otto-sync source-ground, BUILD
learnings/ (Otto) Otto-native agent-authored lessons (auto-append, tiered, evidence-dated) any stage, automatically; curated by /otto-distill session start (relevant slice), as hints-to-verify

The first four unify the three sources' memory conventions; stack.yaml (Q4) and learnings/ (collective wisdom) are Otto-native additions. Otto ships starter templates for any repo missing one, and the router treats these as the durable context every stage consults.





9. Artifacts & state

Adopt spec-kit's specs/NNN-feature/ tree as the canonical state store (D2). Non-spec-kit outputs slot in by convention so the whole lifecycle is in one place and resumable:

specs/NNN-feature/
  spec.md  plan.md  tasks.md  research.md  data-model.md  contracts/   ← spec-kit
  reviews/{code,security,perf,a11y}.md                                 ← agent-skills review outputs
  architecture.html                                                    ← Pocock improve-architecture report
  handoff.md                                                           ← Pocock handoff
  decision-log.md                                                      ← gate trail (D9) + running decisions (your PLAN.md habit)

State-in-tree + the git branch = mid-lifecycle entry (the router reads existing artifacts to know which stage you're in) and pause/resume (stop after any gate, resume later).


10. Framework repo layout

otto-dev/
  README.md                  what it is · quickstart · the layering thesis
  DESIGN.md                  this document (architecture + decision log)
  USAGE.md                   how-to: modes · capabilities · recipes · worked scenarios
  EXTENDING.md               add capabilities · recipes · rules · sources
  VALIDATION.md              static-audit results + dry-run acceptance checklist
  ROADMAP.md                 deferred layers reference (Portfolio · domain profiles) + minor open items
  BACKLOG.md                 the deferred-work index — every row carries an act-on trigger
  CLAUDE.md                  agent operating rules for THIS repo (auto-loaded each session)
  registry/
    capabilities.yaml        Layer 1 machine-readable map          ← step 1
    compute.md               `compute:` policy — tiers T0–T4 · roles · remediation ladder
    tool-scope.md            `tools:` policy — least-privilege profiles (judge-never-fixer, untrusted ingest)
    capabilities.md          generated human view — DEFERRED (§12 item 2; the §5 table is the human view)
  router/
    flow.md                  manual dispatcher skill
    routing-policy.md        selection · swap criteria · gates · stop conditions · degradation
    session-start.(sh|ps1)   agentic-face hook (injects registry + policy)
  workflows/
    spine.md                 lifecycle graph · gates · gate classes · entry points · modules   ← step 1
    scope.md                 scope axis: standard ↔ workspace (multi-repo overlay)   ← step 3
    recipes/                 8 base + assess + extract + inception (11 preset routes)
    autonomous/              conductor.md (hands-off driver, run budget + stop conditions) + workflow-backend.md (segmented Workflow backend) + sample-workflow.yml
  memory/
    memory-stack.md          how the 6 memory members interlock + gate wiring
    distill.md               /otto-distill — curate the learnings store
    templates/               constitution · CONTEXT · ADR · CLAUDE.md · BACKLOG · decision-log · stack.yaml · learnings.md · otto-manifest · otto-bundle.yaml
  learnings/
    global.md                live cross-project wisdom store (member 6, global tier)
  assessments/               dated read-only assess records (loop-engineering · compute routing · repo init · external tools)
  specs/                     Otto's own feature runs (spec/plan/tasks/decision-log per NNN)
  dashboard/
    index.html               self-contained visual map of the framework
  sources/
    sources.md               the 3 repos · pinned refs · licenses · install commands   ← step 1
    sync.md                  /otto-sync — re-pin + scan + diff vs registry (incl. --drift)
  install/
    claude-code.md           exhaustive Claude Code guide — global vs per-repo
    portable.md              Cursor/Copilot/Gemini notes (D4 portable-ready)
    otto-session-start.(sh|ps1)  SessionStart hook (agentic face)
    otto-learnings-nudge.(sh|ps1) failure-time learnings reminder
    otto-init.(sh|ps1)           single-command per-repo setup
    register-skills.(sh|ps1)     copy Otto's skills into ~/.claude/skills/ + stamp compute pins
    verify-otto.sh               scripted Part-1 self-consistency audit (see VALIDATION.md)
    otto-gate.sh                 audit ONE run's gate trail (the D9 evidence record)

11. Build roadmap

  1. Spine + registry — capabilities.yaml, spine.md, sources.md. The skeleton everything hangs on. — ✅ done
  2. Router — flow.md manual dispatcher + routing-policy.md (gates, swap criteria, degradation). — ✅ done
  3. Recipes + scope — 11 preset routes (8 base + assess, extract, inception) + the workspace scope overlay (scope.md); manual variant first. — ✅ done
  4. Memory stack — memory-stack.md + templates; wire the 4 per-repo files + stack.yaml (pinned libs/docs → feeds source-ground) + the learnings/ collective-wisdom store (auto-append + read-slice) + workspace-level memory, into the gates. — ✅ done
  5. Hands-off — autonomous driver (conductor.md + sample-workflow.yml): primary backend = Claude Code Workflow for big/heterogeneous/multi-repo, spec-kit engine for linear, conductor as portable fallback. Plus USAGE.md (how-to + worked scenarios). — ✅ done
  6. Install + maintenance — Claude Code wire-up + portable notes; EXTENDING.md (add capabilities/recipes/rules/sources); /otto-sync (re-pin + scan sources → diff vs registry → propose updates; register new sources) + drift-check; /otto-distill (curate learnings/: dedupe → promote durable lessons to rules → prune stale). — ✅ done
  7. Validation — dry-run each recipe end-to-end on a throwaway repo; confirm graceful degradation when a source is absent. — ✅ done (static audit passed 2026-06-08; full dry-run executed 2026-07-15: 13/13 items PASS on — evidence + findings in VALIDATION.md)







12. Decision log & remaining open items

  1. Name — ✅ RESOLVED: Otto (repo otto-dev).
  2. Registry format — ✅ RESOLVED: registry/capabilities.yaml is canonical (machine-first); the .md view is deferred — the §5 table is the human view.
  3. Hands-off backend priority — ✅ DIRECTION SET (2026-06-08): Claude Code Workflow primary for big/heterogeneous/multi-repo; spec-kit engine for linear spec-kit-dominant flows; Otto conductor as portable fallback. (Backend table in workflows/autonomous/conductor.md; the Workflow backend itself built + evidenced 2026-07-15 — item 17.)
  4. Debug primary — ✅ RESOLVED: PK /diagnose primary, AS debugging-and-error-recovery layered.
  5. Build sequencing — ✅ Steps 1–7 complete (skeleton, router, recipes + scope, memory stack, autonomous driver + USAGE, install + maintenance, validation).
  6. Portfolio layer — ⏳ DEFERRED (approved 2026-06-08). A standing-work layer above the router (future Layer 5): intake → triage → prioritize → track a continuous backlog, wired to the connected ClickUp/Todoist + PK /triage. Addresses "debt outpacing paydown." Recorded so the analysis isn't lost. Full reference (role, evidence for the gap, build shape, sequencing): ROADMAP.md §1 (2026-07-15).
  7. Domain profiles — ⏳ DEFERRED (approved 2026-06-08). Per-surface gate/deploy semantics the code-centric sources don't supply — hardware/firmware OTA (device cohorts, bricking risk) and ML/CV (accuracy/precision/latency eval, datasets, drift). Ship pluggable gate profiles + extension points; seed real profiles later. Full reference: ROADMAP.md §2 (2026-07-15) — build evidence-first, when a real hardware/ML project is in front of Otto.

Adopted 2026-06-08 — build later (from the Q1–Q4 review)

  1. stack.yaml — tech-stack & doc manifest — ✅ ADOPTED; builds in step 4. Project-owned, pinned, human-readable; declares runtimes/frameworks/packages + version + authoritative docs; feeds source-ground. Fetch stays pluggable (context7 / WebFetch / offline) so it's tool-agnostic.
  2. EXTENDING.md — extension guide — ✅ BUILT (step 6). How to add a capability, author a recipe, add a rule, and register a source — deferring unit-authoring to the sources' own tools.
  3. /otto-sync — self-revision command — ✅ BUILT (step 6, sources/sync.md, incl. --drift). Re-pin source refs, scan unit inventories, diff vs the registry, report new/renamed/removed, propose registry + sources.md edits for approval; sub-flow to register a new source.
  4. Collective-wisdom store (learnings/) — ✅ ADOPTED 2026-06-08 (auto-append + distill, D8). Tiered (global in otto-dev + per-repo + stack-tagged); structured entries (trigger / lesson / evidence / date / confidence / scope). Auto-appended on non-obvious resolved issues (commit stays human-gated); read as a filtered, verify-first slice at session start. Store + wiring → step 4 ✅; /otto-distill curation → step 6 ✅ (memory/distill.md). Guards: bounded growth, version-scoped staleness, no wrong-lesson propagation.
  5. evaluate-alternatives capability — ✅ BUILT (in the registry). Ecosystem search (npm / pub.dev / crates: maintenance, popularity, size, license, API-fit, migration-cost) + context7 docs + stack.yaml pins. Added to the registry; invoked through the existing assess recipe → hands off swaps to migration. No new shape.

Maintenance log

  1. spec-kit pinned v0.1.10 → v0.9.5, extensions wired — 2026-06-08. Drift-checked read-only against the v0.9.5 tree: no breaking change to the 9 core commands, Claude skills-mode (/speckit-*), or the scaffold. Wired the v0.9.5 extensions as registry alternates: bug trio (alternate bugfix path, artifacts in .specify/bugs/<slug>/), git helpers (mutating ops human-gated — honors D6; auto_commit hook kept disabled), agent-context (plumbing under context-engineer). selftest/template/scaffold ignored as non-capabilities. Watch: v0.10.0 removes --ai/--no-git; v0.12.0 moves inline CLAUDE.md management to the agent-context extension.

Adopted 2026-07-02 — define "working" (post-assessment)

  1. Success criteria (D9) — ✅ ADOPTED 2026-07-02. Otto had decisions (D1–D8) but no definition of working, so its value was unfalsifiable and the standing temptation was to keep polishing the framework itself. "Otto is working" is now judged by:
  • Gate trail — every non-trivial task run through Otto leaves per-gate evidence in specs/NNN/decision-log.md (which gate, what evidence cleared it). No evidence → not done.
  • Live learnings — learnings/ accrues real, verified lessons from actual runs (not just install scars); target: the store is non-trivially larger after a month of use and /otto-distill has promoted ≥1 lesson into a rule.
  • Degradation observed — with a source uninstalled, the router substitutes the alternate and states it (or reports unavailable + the install cmd) — never a silent skip.
  • No skipped-verification rework — no defect reaches REVIEW that a named earlier gate (tests green, analyze clean, browser DOM) should have caught; if one does, the gate is strengthened, not just the defect fixed.
  • Real adoption — Otto is exercised on real tasks (measurable as specs/NNN/ trees + decision-logs created), not just self-referential edits. Provisional until validated against the first real end-to-end run (VALIDATION Part 2); tune the thresholds from what that run shows rather than guessing now. (Assessment finding #8.)
  1. Pre-first-run fixes (2nd assessment) — ✅ APPLIED 2026-07-02. Seven defects fixed before any real run; everything else deferred to run evidence:
  • modes semantics defined (registry header): hands-off = the conductor may run it autonomously; a manual-only step in a hands-off route is a pause point (or skipped if its gate already passes). Mechanical capabilities (debug, browser-verify, perf/a11y-review, …) got hands-off added; interactive-by-nature ones (interview, grill, ideate, prototype) stay manual-only. Recipes' hands-off sections aligned.
  • Analyze gate predicate unified: "zero CRITICAL /speckit.analyze findings" everywhere; the spec-kit engine renders it as a human gate (no exit code to auto-evaluate), the conductor/Workflow backends read the report and auto-advance.
  • Gate trail made concrete: memory/templates/decision-log.md (the artifact D9 measures) + install/otto-gate.sh (audits a run's trail: evidence present, result vocab, stalled stages); wired into routing-policy §D, spine, /flow, conductor, recipes.
  • Conductor bounded + backend status honest: max 2 remediation attempts per gate, then pause; the Claude Code Workflow backend is explicitly designed-not-built (needs a per-run script + per-subagent run-state injection) — hands-off defaults to the conductor skill until it exists.
  • Triviality floor (routing-policy A0): one repo, ~one file, reversible, no security surface, provable in one step → no recipe, do it directly with tests. The verification bar stays.
  • verify-otto.sh check 6: every slash invocation in the routing surface (router/, workflows/, USAGE) must resolve against the registry — catches stale-invocation drift mechanically.
  • Learnings loop hardened: canonical naming (<repo>/learnings.md file · global otto-dev/learnings/global.md; learnings/ names the store) + a once-per-session Stop-hook capture nudge (install/otto-learnings-nudge.{ps1,sh}, wiring in install/claude-code.md A3b).

Adopted 2026-07-15 — footprint legibility (D10)

  1. Footprint manifest + git-disposition modes — ✅ ADOPTED 2026-07-15. Problem: per-repo setup touches ~10 paths across the tree (.claude/, .specify/, specs/, docs/agents|adr/, root memory files, CLAUDE.md blocks) with no single record of what was touched — so excluding, uninstalling, and auditing the footprint were all done from memory, and keeping Otto out of a repo's git history had no supported path. Two changes:
  • .otto-manifest.md (template memory/templates/otto-manifest.md; install B8) — written as the last per-repo step: one row per created/modified path with owner, install step, disposition per mode, and regenerability. Exclude lines are generated from it, uninstall walks it bottom-up, audit reads it, /otto-sync diffs it against disk.
  • Git disposition is declared per repo (install B6): team mode (default; governance + artifacts tracked, regenerable/personal tooling ignored — the prior B6 behavior) or private mode (whole footprint excluded via .git/info/exclude — per-clone, never committed, so no committed line reveals Otto exists; optional core.excludesFile for per-machine coverage; Otto never proposes committing manifest rows). Known hard edge, documented: managed blocks written into an already-tracked CLAUDE.md can't be hidden by ignore rules — private mode skips block-writing and uses ~/.claude/CLAUDE.md / CLAUDE.local.md instead.
  • Rejected alternative: consolidating the footprint into a single .otto/ root dir with a one-line .gitignore entry. Rejected because (a) most of the footprint is pinned by the host or the sources Otto deliberately doesn't fork (D2/D5): .claude/ + CLAUDE.md (Claude Code), .specify/ + specs/ (spec-kit CLI), docs/agents/ (Pocock) — only stack.yaml and learnings.md are Otto-movable, a minority of the sprawl; (b) blanket-ignoring contradicts the memory stack — in team mode the constitution, specs, and ADRs are meant to be tracked; (c) for the privacy goal, a committed .gitignore line itself advertises Otto — .git/info/exclude is strictly more private and works regardless of layout. Folding the two movable files into .otto/ stays deferred as cosmetic (~50 path references across 17 framework files + back-compat shims in both session hooks for existing repos, for no functional gain); revisit only if the Otto-movable set grows.

Adopted 2026-07-15 — Workflow backend built from run evidence (closes the §7 backend gap)

  1. Claude Code Workflow backend — ✅ BUILT 2026-07-15 (workflows/autonomous/workflow-backend.md), from the evidence of the first real hands-off run ( dry-run, VALIDATION.md Part 2): the REVIEW personas fan-out already ran as parallel fresh-context subagents with self-contained prompts and structured verdicts — the backend generalizes exactly that pattern to whole stages. Architecture (the two blockers from the old "designed-not-built" note, resolved):
  • Per-run script = gate-bounded segments. A workflow script cannot ask the human anything mid-run, so every human-gated decision point (FRAME sign-off, spec acceptance, go/no-go, irreversible ops, git) is a segment boundary: the conductor authors one script per maximal human-gate-free run of route steps, executes it, returns to the main loop to pause at the gate, then authors the next segment. The conductor stays the driver; Workflow executes segments.
  • Run-state injection = a standard prompt header, not the SessionStart hook: every agent() prompt begins with an [OTTO RUN STATE] block (recipe · scope · segment · repo + feature-dir paths · stage · capability · gate predicate · memory-file paths · standing rules). Subagents need no session hook; the state travels in-band.
  • Gates in script code: step agents return schema-validated structured verdicts with evidence; the deterministic script applies the gate predicate, runs the ≤2-attempt remediation loop, and on persistent failure returns {halted: stage, evidence} for the conductor to pause on. Agents write their own artifacts + decision-log rows (they have tools); the conductor audits with otto-gate.sh at each segment end.
  • Selection unchanged (conductor.md table): workspace scope / heterogeneous steps / parallel fan-out → this backend; linear single-repo flows stay on the cheaper conductor-skill backend.
  • Rejected: (a) one monolithic workflow with simulated pauses — a script that "waits" for a human either blocks a background run or fakes the gate; segment chaining keeps D6's pause semantics literal. (b) SessionStart-hook injection into subagents — hooks carry only the static routing brain; per-run state must be per-prompt.

Adopted 2026-07-15 — loop-engineering hardening (from assessments/2026-07-15-loop-engineering.md)

  1. Context compaction, independent verification, failure-time learnings, gate classes — ✅ APPLIED 2026-07-15. Four of the assessment's seven gaps closed as protocol edits; the rest deferred to run evidence (consistent with item 15's build-from-evidence rule):
  • Conductor-skill compaction (G2): the one-window backend now compacts as a protocol step — /handoff at every human pause point and after heavy stages; never enter a new stage in a degraded window (conductor.md "Big-task handling"). Closes the context-rot gap on the cheap path.
  • Independent verification MUST (G3): SHOULD→MUST, all backends — mechanical-gate evidence is re-checked by a second agent or by re-running the named command; judgment gates get a judge separate from the actor (workflow-backend.md gate authoring; conductor.md step 3; routing-policy §D6). otto-gate.sh still audits trail structure only — evidence re-execution stays with the conductor; extending the script is deferred until a run shows the need.
  • Failure-time learnings (G6): on any failed gate, grep the learnings store for TRIGGER matches on the failure signature before remediation attempt 1 (memory-stack.md "Read (failure-time)"; conductor.md step 3 fail branch; [OTTO RUN STATE] rules). The TRIGGER field finally fires at the moment it was designed for.
  • Gate classes (G7): every spine gate tagged mechanical | judgment | human with per-class enforcement (spine.md legend + per-stage lines; routing-policy §D6) — judgment predicates ("~95% confidence") no longer flow through the same self-assessed verdict path as "tests green".
  • Run budget + stop conditions (G1, G4, G5) — closed 2026-07-21. The 2026-07-15 deferral ("revisit on run evidence") was re-examined by explicit call after an independent external audit reached G1 and G5 cold (assessments/2026-07-21). A hands-off run now carries a declared ceiling — token budget where the backend can read its own spend (Workflow only), structural bounds everywhere (stage-steps · total remediation attempts · stage revisits · compactions) — with three stop conditions: budget exhausted, stage entered a 4th time, and no progress (two consecutive attempts returning the same failure signature). G4 landed as a prerequisite, not scope creep: comparing signatures requires carrying attempt 1's attempted summary forward, which is exactly what G4 asked for. Protocol: conductor.md · enforcement: workflow-backend.md authoring step 4 · routing-policy §D6-7. Defaults are engineered priors; a halt is recorded in the decision-log narrative, never as a gate-trail row (closed vocabulary).
  • Least-privilege tool scope — added 2026-07-21. New tools: field beside compute: (policy: registry/tool-scope.md), mechanizing two rules Otto had only as prose: judge-never-fixer (review capabilities run verify — no Write/Edit) and untrusted ingest (survey/understand run inspect — no shell). Binding in hands-off, advisory elsewhere; register-skills does not stamp it yet, which is stated rather than implied. Surfaced by the external audit, not by the loop rubric.