Status: Build complete (steps 1–7, 2026-06-08) · autonomous capabilities validated 13/13 + Workflow backend built & evidenced (2026-07-15, VALIDATION.md) · deferred layers: ROADMAP.md · v1.1 Name: Otto (repo
otto-dev). Chosen 2026-06-08. Scope of this document: the overall design. The build follows the roadmap in §11; this doc is the decision-of-record.
Three excellent, overlapping, standalone projects exist in this workspace:
| Project | What it really is | Best at | Weak at |
|---|---|---|---|
| spec-kit (GitHub, MIT) | A methodology engine — Python CLI, durable artifact tree, governance constitution, a real YAML workflow engine. | spec→plan→tasks, governance, deterministic multi-step automation | review, ship, debug, maintain |
| agent-skills (Addy Osmani, MIT) | A phase-complete process library — 23 SDLC skills, 3 review personas, slash commands, hooks, reference checklists. | the entire back half: review, security, perf, a11y, CI/CD, ship, deprecation; anti-rationalization discipline | no durable artifact/state engine; no governance file |
| skills (Matt Pocock, MIT) | A philosophy-driven practice library — 18 skills, feedback-loop & alignment obsessed, domain-language aware. | alignment ("grilling"), CONTEXT.md/ADR domain capture, diagnose, prototype, architecture deepening, handoff |
breadth; no spec/plan tooling; no ship/ops |
They operate at different altitudes, so the right move is to layer them, not merge them. A fourth merged library would just re-create the overlap problem.
spec-kit = SKELETON — durable state, governance, the autonomous engine
agent-skills = MUSCLE — the practices, especially the back half of the lifecycle
Pocock skills = MUSCLE — alignment, domain language, diagnosis, architecture
Otto (new) = NERVOUS SYSTEM — a registry + router + workflow graph that picks the
right muscle at the right moment, referencing all three, modifying none.
What Otto is NOT: not a fork, not a copy, not a 4th skill library, not a replacement for any of the three. It is the thin orchestration + routing + project-memory layer above them.
| # | Decision | Choice | Consequence |
|---|---|---|---|
| D1 | Overlap resolution | Best-of-breed + layer | One primary unit per capability by demonstrated strength; layer a second only where genuinely complementary; document alternates with swap criteria. |
| D2 | Tooling dependency | Hybrid | Adopt spec-kit's artifact tree + constitution as canonical conventions; use its CLI/workflow engine when installed, never hard-require it; orchestration layer is host-native instructions. |
| D3 | Workflow shape | Spine + recipe presets | One lifecycle graph, enterable at any stage, with reusable specialist modules; named recipes are preset routes through the graph. |
| D4 | Host target | Claude Code-first, portable-ready | Optimize for Claude Code (skills, slash commands, hooks, subagents/Workflow). Keep units host-portable since all three sources already are. |
| D5 | Output location | <LOCAL-PATH>\ |
New sibling repo. The three sources stay read-only and untouched. |
| D6 | Modes | Manual and hands-off, co-equal | Manual/adhoc fits the architect→coder fenced-prompt split; hands-off chains stages with gates. Both are first-class. |
| D7 | Coverage model | Shape × Scope (2 axes) | ~11 shapes (8 base + assess + extract + inception) × {standard, workspace}. Portfolio layer + domain profiles deferred (approved 2026-06-08); see §12. |
| D8 | Collective wisdom | Auto-append + distill | Agent-authored learnings/ store (tiered, structured, evidence-dated): frictionless auto-capture, periodic /otto-distill curation (dedupe → promote → prune), verify-on-read. See §12 item 11. |
| D9 | Definition of success | Measured, not assumed | Otto "working" = every non-trivial task leaves a gate-trail; learnings/ accrues real lessons; graceful degradation is observed; zero skipped-verification rework. Provisional until tuned against the first real run. See §12 item 14. |
| D10 | Per-repo footprint | Manifest + declared git disposition | Per-repo setup ends with .otto-manifest.md (every path created/modified + disposition per row); each repo declares team mode (governance tracked — default) or private mode (footprint excluded via .git/info/exclude; zero trace in history). Consolidating into a single .otto/ dir was rejected — most paths are host/source-pinned. See §12 item 16. |
flowchart TB
L4["<b>Layer 4 · PROJECT MEMORY</b><br/>constitution · CONTEXT.md · docs/adr · CLAUDE.md<br/><i>governs all stages</i>"]
L3["<b>Layer 3 · WORKFLOWS</b><br/>the SPINE (enter at any stage) + specialist MODULES<br/>+ named RECIPES (preset routes) · manual + hands-off"]
L2["<b>Layer 2 · ROUTER</b><br/>manual face (/flow dispatcher) + agentic face (hook)"]
L1["<b>Layer 1 · CAPABILITY REGISTRY</b><br/>every unit across all 3, tagged by capability,<br/>stage · source · invocation · mode · prefer-when"]
L0["<b>Layer 0 · SOURCES</b><br/>spec-kit · agent-skills · Pocock skills (read-only)"]
L4 --> L3 --> L2 --> L1 --> L0
The three repos are referenced by path and pinned commit. Never edited, copied, or vendored. Otto does not execute their code; it routes to units that must be installed in the host agent. Full pinned refs + install commands live in sources/sources.md.
| Source | Path | Install into host (Claude Code) | Units exposed |
|---|---|---|---|
| spec-kit | <LOCAL-PATH> |
specify init per target repo (scaffolds .specify/, /speckit.* commands) |
/speckit.constitution|specify|clarify|plan|analyze|tasks|implement|checklist|taskstoissues; CLI; workflows/*.yml |
| agent-skills | <LOCAL-PATH> |
plugin install (or copy skills/ into .claude/skills/) |
23 skills; /spec /plan /build /test /review /code-simplify /ship; personas; reference checklists |
| Pocock skills | <LOCAL-PATH> |
npx skills@latest add mattpocock/skills (select promoted) |
/grill-me /grill-with-docs /to-prd /to-issues /tdd /prototype /diagnose /improve-codebase-architecture /zoom-out /triage /handoff /git-guardrails-claude-code … |
Read-only contract: sources/sources.md records each repo's path, pinned ref, license (all MIT), and the install command. A drift check (optional) warns if a pinned unit's invocation name changed upstream.
Realism note (consequence of D2): because the units live in the three host installs, Otto's value is selection + sequencing + memory + the install recipe — not re-implementation. If a source isn't installed, the router degrades gracefully to whatever is present and says so.
The heart of D1. Every capability maps to one primary + optional layer/alternate with an explicit prefer-when. Several apparent overlaps dissolve once split into adjacent-but-distinct capabilities (e.g. elicit-intent ≠ clarify-spec; simplify-local ≠ improve-architecture; git-practice ≠ git-safety).
Legend — source: SK spec-kit · AS agent-skills · PK Pocock. Canonical machine-readable version: registry/capabilities.yaml.
Two orthogonal pins ride alongside each entry, each with its own policy companion and the same
inherit-by-default rule (no field ⇒ inherit the session): compute: — model + effort
(registry/compute.md, seeded 2026-07-18) — and tools: — least-privilege
profile (registry/tool-scope.md, seeded 2026-07-21). Both are binding in
hands-off runs on the Workflow backend and advisory elsewhere. The tables below list capability
routing only; the pins live in the YAML.
| Capability | Primary | Layer / Alternate | Prefer-when |
|---|---|---|---|
| govern (principles) | SK /speckit.constitution → constitution.md |
— | always, once per project |
| ideate / refine | AS idea-refine |
— | rough idea needs diverge→converge |
| elicit-intent | AS interview-me (one-Q-at-a-time → ~95% confidence) |
PK /grill-me (relentless stress-test) |
AS for requirement extraction; PK to break a confident-but-shaky plan |
| capture-domain-language | PK /grill-with-docs → writes CONTEXT.md + ADRs |
— | whenever domain vocabulary is fuzzy (unique to PK) |
| harvest-backlog | OTTO inline — index the idea bundle's deferred scope into root BACKLOG.md (link, don't copy) |
— | DEFINE input is an idea bundle (e.g. idea-to-build) carrying deferred scope |
| classify-bundle | OTTO inline — read a dropped-in idea bundle, resolve its units by content, write .otto-bundle.yaml (tools: inspect — read-only over untrusted docs) |
— | a bundle is present and its adapter is absent or stale |
| apply-bundle | OTTO inline — execute the extraction map's seed/link rows into the memory stack; delegates link rows to harvest-backlog |
— | a reviewed .otto-bundle.yaml has pending rows |
The bundle pair is deliberately split along the read/write line — the shape extract-spec cannot
express and F-12 defers. Both are consumed by the inception feeder; four extraction modes
(seed · link · input · evidence) each carry their own bundle-vs-repo precedence rule. Design record:
assessments/2026-07-25-idea-bundle-intake.md.
| Capability | Primary | Layer / Alternate | Prefer-when |
|---|---|---|---|
| author-spec | SK /speckit.specify → specs/NNN/spec.md (+branch, [NEEDS CLARIFICATION]) |
layer AS spec-driven-development template depth (objective/commands/structure/code-style/testing/boundaries) into the SK spec; alt PK /to-prd |
SK for artifact-tracked SDD; PK /to-prd when target is a lightweight PRD/issue, no tree |
| clarify-spec | SK /speckit.clarify (resolves markers in the written spec) |
— | a spec exists with open [NEEDS CLARIFICATION] |
| validate-requirements | SK /speckit.checklist → specs/NNN/checklists/*.md ("unit tests for English" — audits requirement clarity/completeness) |
— | requirement quality is load-bearing: 0→1, external contracts, contested requirements. Catches confidently-written-but-ambiguous requirements that clarify-spec cannot (it only closes markers the author already flagged) |
| Capability | Primary | Layer / Alternate | Prefer-when |
|---|---|---|---|
| plan-architecture | SK /speckit.plan → plan.md, research.md, data-model.md, contracts/ (constitution gates) |
— | always for non-trivial features |
| decompose-tasks | SK /speckit.tasks → tasks.md ([P] parallel, story-sliced, TDD-ordered) |
layer AS planning-and-task-breakdown (vertical-slicing, dep graphs, acceptance criteria); alt PK /to-issues |
SK tasks.md in-repo; PK /to-issues when work lives in GitHub/GitLab issues |
| cross-check | SK /speckit.analyze (spec↔plan↔tasks consistency) |
— | before committing to implementation (unique to SK) |
| Capability | Primary | Layer / Alternate | Prefer-when |
|---|---|---|---|
| implement | AS incremental-implementation (thin-slice discipline) driven by SK tasks.md//speckit.implement when present |
the two compose: SK supplies the ordered loop, AS supplies slice-sizing + per-slice verification | tasks.md exists → SK loop + AS discipline; else AS standalone |
| tdd | AS test-driven-development (+ Prove-It bug pattern + test-engineer persona) |
layer PK /tdd tracer-bullet framing (anti-horizontal-slice) |
AS as default; PK when user wants its tracer-bullet style |
| prototype | PK /prototype (throwaway to answer one question) |
— | spike/feasibility before committing (unique to PK) |
| source-ground | AS source-driven-development (+ context7 MCP) |
— | framework/library correctness matters (unique) |
| doubt-check | AS doubt-driven-development (fresh-context adversarial review) |
complements PK grilling | high-stakes/irreversible/unfamiliar decisions |
| context-engineer | AS context-engineering (rules files) |
alt PK /handoff (compaction/transition) |
session setup / quality degradation |
| frontend-ui | AS frontend-ui-engineering |
alt env frontend-design skill |
building user-facing UI |
| api-design | AS api-and-interface-design |
— | designing any public interface/contract |
| Capability | Primary | Layer / Alternate | Prefer-when |
|---|---|---|---|
| debug | PK /diagnose (feedback-loop → reproduce → hypothesize → instrument → fix → cleanup) |
layer AS debugging-and-error-recovery "guard with regression test" + Prove-It |
toss-up; PK primary for instrumentation rigor (D4 confirmed); AS to pair the guard test with TDD |
| browser-verify | AS browser-testing-with-devtools (Chrome DevTools MCP: DOM/console/network/perf) |
— | anything running in a browser (matches your DOM-verification standard; unique) |
| Capability | Primary | Layer / Alternate | Prefer-when |
|---|---|---|---|
| code-review | AS code-review-and-quality + code-reviewer persona (5-axis) |
alt env /code-review, /review |
before merge |
| security-review | AS security-and-hardening + security-auditor persona + checklist |
alt env /security-review |
untrusted input, auth, storage, integrations (unique) |
| perf-review | AS performance-optimization + checklist |
— | perf budgets / Core Web Vitals (unique) |
| a11y-review | AS accessibility checklist | — | user-facing UI (unique) |
| simplify-local | AS code-simplification (+ env /simplify) |
— | working code that's harder to read than needed |
| improve-architecture | PK /improve-codebase-architecture (deepening + ADRs + HTML report) |
— | structural/architectural improvement (distinct from local simplify; unique) |
| Capability | Primary | Layer / Alternate | Prefer-when |
|---|---|---|---|
| git-practice | AS git-workflow-and-versioning (trunk-based, atomic commits) |
complements git-safety | advisory; you do git manually |
| git-safety | PK /git-guardrails-claude-code (PreToolUse hook blocks dangerous git) |
complements git-practice | install once; high value given manual-git preference |
| ci-cd | AS ci-cd-and-automation |
— | pipelines / quality gates (unique) |
| ship-launch | AS shipping-and-launch + /ship parallel persona fan-out → go/no-go |
— | production launch (flagship orchestration; unique) |
| deprecate / migrate | AS deprecation-and-migration |
alt PK migrate-to-shoehorn (narrow TS) |
sunsetting/migrating systems |
| document / adr | AS documentation-and-adrs |
layer PK inline ADR writes from /grill-with-docs |
decisions + API changes |
| Capability | Primary | Layer / Alternate | Prefer-when |
|---|---|---|---|
| triage | PK /triage (state-machine labels, issue-tracker aware) |
— | incoming bugs/enhancements (unique) |
| handoff / compact | PK /handoff |
— | agent transition / context running out (unique) |
| understand / onboard | PK /zoom-out + architecture map |
— | unfamiliar code (unique) |
| orchestrate | SK workflow.yml engine (when backbone present) |
alt Claude Code Workflow / conductor skill | hands-off multi-step runs |
| personas | AS code-reviewer · security-auditor · test-engineer | — | any review/ship fan-out |
spec-kit v0.9.5 extension alternates (opt-in): the bug extension adds /speckit-bug-{assess,fix,test} as an artifact-tracked alternate for the bugfix path (on debug); the git extension adds read-only validate/remote helpers under git-practice (mutating git ops stay human-gated — honors D6); agent-context auto-refreshes the CLAUDE.md SPECKIT block (under context-engineer). Enable with specify extension add <id>.
Registry format: machine-readable registry/capabilities.yaml (one record per capability with primary, alternates[], stage, source, invocation, modes, prefer_when, requires_install) plus a generated human-readable view. The router and the drift-check both read it.
One selection brain, two faces.
Manual face — /flow dispatcher (adhoc). Given a one-line intent + a glance at repo state (does specs/ exist? a failing test? an unstaged diff? a constitution.md?), it answers: which recipe → which stage → which capability → the exact invocation, then offers to run it. You can also bypass it and invoke any capability directly. This is the mode that fits your architect→coder split: the architect's fenced prompt names a capability; the router resolves it to the right source unit and the current artifact context.
Agentic face — session-start hook (auto-select). Mirrors agent-skills' using-agent-skills pattern: injects the registry + routing policy at session start so the agent auto-selects the right unit when you describe a task, and — in hands-off mode — advances the spine through gates autonomously, stopping only when a gate fails or a decision genuinely needs you.
Routing policy (router/routing-policy.md) encodes: stage detection from artifacts, the prefer-when swap criteria from §5, graceful degradation when a source isn't installed, and the gate definitions that must pass before advancing.
flowchart LR
M["<b>Layer 4 project memory</b> governs every stage<br/>constitution · CONTEXT.md · docs/adr · CLAUDE.md"]
subgraph spine["Enter at ANY stage · each stage has a GATE"]
direction LR
S0["<b>0 FRAME</b><br/>govern · ideate"] --> S1["<b>1 DEFINE</b><br/>elicit · domain-language"] --> S2["<b>2 SPEC</b><br/>author · clarify"] --> S3["<b>3 PLAN</b><br/>plan · tasks · analyze"] --> S4["<b>4 BUILD</b><br/>implement · tdd · prototype<br/>source · doubt"] --> S5["<b>5 VERIFY</b><br/>debug · browser-verify"] --> S6["<b>6 REVIEW</b><br/>code · security · perf/a11y<br/>simplify · improve-arch"] --> S7["<b>7 SHIP</b><br/>git-* · ci-cd · launch · deprecate"] --> S8["<b>8 MAINTAIN</b><br/>triage · handoff · understand"]
end
M -. governs .-> spine
Full graph with per-stage gates and entry signals: workflows/spine.md.
- Gates = per-stage exit criteria (verification-first, from agent-skills) plus
/speckit.analyzestyle consistency checks. No advance without passing — this is where "seems right" is forbidden and evidence is required. - Specialist modules are callable from any stage, not just their home:
debugduring BUILD,review-fan-outbefore SHIP,improve-architectureduring MAINTAIN,handoffanywhere context runs low.
| Recipe | Route | Primary sources |
|---|---|---|
| greenfield (0→1) | FRAME→DEFINE→SPEC→PLAN→BUILD→VERIFY→REVIEW→SHIP | SK front, AS back |
| feature (existing repo) | light DEFINE→mini-SPEC→PLAN/tasks→BUILD(tdd)→REVIEW→SHIP (skip FRAME) | SK + AS |
| bugfix / incident | enter VERIFY→diagnose→prove-it test→BUILD(fix)→REVIEW→SHIP |
PK + AS |
| hardening | enter REVIEW→security+perf+a11y fan-out→BUILD(fixes)→re-REVIEW | AS personas |
| refactor / architecture | enter REVIEW→improve-architecture→ADR→PLAN(migration)→BUILD→VERIFY |
PK + SK |
| release | enter SHIP→/ship fan-out→go/no-go→ci-cd→launch→monitor/rollback |
AS |
| onboarding / understand | enter MAINTAIN→zoom-out→arch map→capture domain language→CONTEXT.md |
PK |
| migration / deprecation | enter SHIP/MAINTAIN→deprecation-and-migration |
AS |
| assess (feeder) | investigate→survey→findings.md→hand off to a recipe |
ENV+AS+PK |
| extract (feeder) | reverse: understand/extract-spec→spec/docs→link |
ENV+PK+AS |
| inception (feeder) | idea bundle → classify-bundle→apply-bundle→seeded memory stack + one first slice → greenfield |
OTTO |
- Manual variant — the router suggests the next step; you (or your architect's fenced prompt) invoke each capability and inspect the gate. Pause/resume is free because state lives in artifacts.
- Hands-off variant — stages chained with automatic gate checks. Execution backend, per the conductor's selection table: (a) Claude Code Workflow/subagents — workspace scope, heterogeneous steps, or parallel fan-out (built & evidenced 2026-07-15, §12 item 17; protocol:
workflows/autonomous/workflow-backend.md); (b) spec-kitworkflow.ymlengine — linear, spec-kit-dominant flows; (c) the Otto conductor skill — linear single-repo runs (cheaper) and portable hosts; halts at any failed gate or human-decision point. Seeworkflows/autonomous/conductor.md.
A task = Shape (what kind of work) × Scope (blast radius). This keeps coverage broad without recipe sprawl — the answer to "8 recipes felt too few": 8 was too few shapes, and scope was a missing axis.
- Shape — the 8 base recipes + three feeders:
assess(investigate → findings → scope; hands off to another recipe),extract(reverse: code/app → spec/docs) andinception(idea bundle → seeded repo + one first slice, 0→1 only). ~11 shapes. - Scope — an overlay on any shape: standard (one repo) or workspace (coordinated across N repos/platforms — a workspace spec + inter-repo contracts + per-repo sub-specs + dependency sequencing + a cross-repo integration gate). See
workflows/scope.md.
~10 shapes × 2 scopes = broad coverage from a small file set. Deferred dials (approved 2026-06-08, §12 items 6–7): a Portfolio layer for the standing stream of work, and domain profiles for per-surface gate/deploy semantics (hardware OTA, ML/CV eval).
The strongest synthesis: the three projects' "memory" files are complementary, not overlapping. Otto unifies them into one stack and defines which stage reads/writes each. No single source has all four.
| File | Owner convention | Holds | Written by | Read by |
|---|---|---|---|---|
.specify/memory/constitution.md |
spec-kit | principles / governance | FRAME (/speckit.constitution) |
every gate, esp. PLAN |
CONTEXT.md (repo root) |
Pocock | domain language / glossary | DEFINE (/grill-with-docs) |
SPEC, BUILD, REVIEW |
docs/adr/*.md |
shared | decisions (append-only) | any stage (documentation-and-adrs, grill-with-docs) |
PLAN, REVIEW, refactor |
CLAUDE.md / AGENTS.md |
agent-skills | agent operating rules, stack, commands, boundaries | DEFINE/BUILD (context-engineering) |
every session (host auto-loads) |
stack.yaml (Otto) |
Otto-native | pinned runtimes/frameworks/packages + authoritative docs | DEFINE/FRAME; /otto-sync |
source-ground, BUILD |
learnings/ (Otto) |
Otto-native | agent-authored lessons (auto-append, tiered, evidence-dated) | any stage, automatically; curated by /otto-distill |
session start (relevant slice), as hints-to-verify |
The first four unify the three sources' memory conventions; stack.yaml (Q4) and learnings/ (collective wisdom) are Otto-native additions. Otto ships starter templates for any repo missing one, and the router treats these as the durable context every stage consults.
Adopt spec-kit's specs/NNN-feature/ tree as the canonical state store (D2). Non-spec-kit outputs slot in by convention so the whole lifecycle is in one place and resumable:
specs/NNN-feature/
spec.md plan.md tasks.md research.md data-model.md contracts/ ← spec-kit
reviews/{code,security,perf,a11y}.md ← agent-skills review outputs
architecture.html ← Pocock improve-architecture report
handoff.md ← Pocock handoff
decision-log.md ← gate trail (D9) + running decisions (your PLAN.md habit)
State-in-tree + the git branch = mid-lifecycle entry (the router reads existing artifacts to know which stage you're in) and pause/resume (stop after any gate, resume later).
otto-dev/
README.md what it is · quickstart · the layering thesis
DESIGN.md this document (architecture + decision log)
USAGE.md how-to: modes · capabilities · recipes · worked scenarios
EXTENDING.md add capabilities · recipes · rules · sources
VALIDATION.md static-audit results + dry-run acceptance checklist
ROADMAP.md deferred layers reference (Portfolio · domain profiles) + minor open items
BACKLOG.md the deferred-work index — every row carries an act-on trigger
CLAUDE.md agent operating rules for THIS repo (auto-loaded each session)
registry/
capabilities.yaml Layer 1 machine-readable map ← step 1
compute.md `compute:` policy — tiers T0–T4 · roles · remediation ladder
tool-scope.md `tools:` policy — least-privilege profiles (judge-never-fixer, untrusted ingest)
capabilities.md generated human view — DEFERRED (§12 item 2; the §5 table is the human view)
router/
flow.md manual dispatcher skill
routing-policy.md selection · swap criteria · gates · stop conditions · degradation
session-start.(sh|ps1) agentic-face hook (injects registry + policy)
workflows/
spine.md lifecycle graph · gates · gate classes · entry points · modules ← step 1
scope.md scope axis: standard ↔ workspace (multi-repo overlay) ← step 3
recipes/ 8 base + assess + extract + inception (11 preset routes)
autonomous/ conductor.md (hands-off driver, run budget + stop conditions) + workflow-backend.md (segmented Workflow backend) + sample-workflow.yml
memory/
memory-stack.md how the 6 memory members interlock + gate wiring
distill.md /otto-distill — curate the learnings store
templates/ constitution · CONTEXT · ADR · CLAUDE.md · BACKLOG · decision-log · stack.yaml · learnings.md · otto-manifest · otto-bundle.yaml
learnings/
global.md live cross-project wisdom store (member 6, global tier)
assessments/ dated read-only assess records (loop-engineering · compute routing · repo init · external tools)
specs/ Otto's own feature runs (spec/plan/tasks/decision-log per NNN)
dashboard/
index.html self-contained visual map of the framework
sources/
sources.md the 3 repos · pinned refs · licenses · install commands ← step 1
sync.md /otto-sync — re-pin + scan + diff vs registry (incl. --drift)
install/
claude-code.md exhaustive Claude Code guide — global vs per-repo
portable.md Cursor/Copilot/Gemini notes (D4 portable-ready)
otto-session-start.(sh|ps1) SessionStart hook (agentic face)
otto-learnings-nudge.(sh|ps1) failure-time learnings reminder
otto-init.(sh|ps1) single-command per-repo setup
register-skills.(sh|ps1) copy Otto's skills into ~/.claude/skills/ + stamp compute pins
verify-otto.sh scripted Part-1 self-consistency audit (see VALIDATION.md)
otto-gate.sh audit ONE run's gate trail (the D9 evidence record)
- Spine + registry —
capabilities.yaml,spine.md,sources.md. The skeleton everything hangs on. — ✅ done - Router —
flow.mdmanual dispatcher +routing-policy.md(gates, swap criteria, degradation). — ✅ done - Recipes + scope — 11 preset routes (8 base +
assess,extract,inception) + the workspace scope overlay (scope.md); manual variant first. — ✅ done - Memory stack —
memory-stack.md+ templates; wire the 4 per-repo files +stack.yaml(pinned libs/docs → feedssource-ground) + thelearnings/collective-wisdom store (auto-append + read-slice) + workspace-level memory, into the gates. — ✅ done - Hands-off — autonomous driver (
conductor.md+sample-workflow.yml): primary backend = Claude Code Workflow for big/heterogeneous/multi-repo, spec-kit engine for linear, conductor as portable fallback. PlusUSAGE.md(how-to + worked scenarios). — ✅ done - Install + maintenance — Claude Code wire-up + portable notes;
EXTENDING.md(add capabilities/recipes/rules/sources);/otto-sync(re-pin + scan sources → diff vs registry → propose updates; register new sources) +drift-check;/otto-distill(curatelearnings/: dedupe → promote durable lessons to rules → prune stale). — ✅ done - Validation — dry-run each recipe end-to-end on a throwaway repo; confirm graceful degradation when a source is absent. — ✅ done (static audit passed 2026-06-08; full dry-run executed 2026-07-15: 13/13 items PASS on — evidence + findings in
VALIDATION.md)
- Name — ✅ RESOLVED: Otto (repo
otto-dev). - Registry format — ✅ RESOLVED:
registry/capabilities.yamlis canonical (machine-first); the.mdview is deferred — the §5 table is the human view. - Hands-off backend priority — ✅ DIRECTION SET (2026-06-08): Claude Code Workflow primary for big/heterogeneous/multi-repo; spec-kit engine for linear spec-kit-dominant flows; Otto conductor as portable fallback. (Backend table in
workflows/autonomous/conductor.md; the Workflow backend itself built + evidenced 2026-07-15 — item 17.) - Debug primary — ✅ RESOLVED: PK
/diagnoseprimary, ASdebugging-and-error-recoverylayered. - Build sequencing — ✅ Steps 1–7 complete (skeleton, router, recipes + scope, memory stack, autonomous driver + USAGE, install + maintenance, validation).
- Portfolio layer — ⏳ DEFERRED (approved 2026-06-08). A standing-work layer above the router (future Layer 5): intake → triage → prioritize → track a continuous backlog, wired to the connected ClickUp/Todoist + PK
/triage. Addresses "debt outpacing paydown." Recorded so the analysis isn't lost. Full reference (role, evidence for the gap, build shape, sequencing): ROADMAP.md §1 (2026-07-15). - Domain profiles — ⏳ DEFERRED (approved 2026-06-08). Per-surface gate/deploy semantics the code-centric sources don't supply — hardware/firmware OTA (device cohorts, bricking risk) and ML/CV (accuracy/precision/latency eval, datasets, drift). Ship pluggable gate profiles + extension points; seed real profiles later. Full reference: ROADMAP.md §2 (2026-07-15) — build evidence-first, when a real hardware/ML project is in front of Otto.
stack.yaml— tech-stack & doc manifest — ✅ ADOPTED; builds in step 4. Project-owned, pinned, human-readable; declares runtimes/frameworks/packages + version + authoritative docs; feedssource-ground. Fetch stays pluggable (context7 / WebFetch / offline) so it's tool-agnostic.EXTENDING.md— extension guide — ✅ BUILT (step 6). How to add a capability, author a recipe, add a rule, and register a source — deferring unit-authoring to the sources' own tools./otto-sync— self-revision command — ✅ BUILT (step 6,sources/sync.md, incl.--drift). Re-pin source refs, scan unit inventories, diff vs the registry, report new/renamed/removed, propose registry +sources.mdedits for approval; sub-flow to register a new source.- Collective-wisdom store (
learnings/) — ✅ ADOPTED 2026-06-08 (auto-append + distill, D8). Tiered (global in otto-dev + per-repo + stack-tagged); structured entries (trigger / lesson / evidence / date / confidence / scope). Auto-appended on non-obvious resolved issues (commit stays human-gated); read as a filtered, verify-first slice at session start. Store + wiring → step 4 ✅;/otto-distillcuration → step 6 ✅ (memory/distill.md). Guards: bounded growth, version-scoped staleness, no wrong-lesson propagation. evaluate-alternativescapability — ✅ BUILT (in the registry). Ecosystem search (npm / pub.dev / crates: maintenance, popularity, size, license, API-fit, migration-cost) + context7 docs +stack.yamlpins. Added to the registry; invoked through the existingassessrecipe → hands off swaps tomigration. No new shape.
- spec-kit pinned
v0.1.10→v0.9.5, extensions wired — 2026-06-08. Drift-checked read-only against the v0.9.5 tree: no breaking change to the 9 core commands, Claude skills-mode (/speckit-*), or the scaffold. Wired the v0.9.5 extensions as registry alternates:bugtrio (alternate bugfix path, artifacts in.specify/bugs/<slug>/),githelpers (mutating ops human-gated — honors D6; auto_commit hook kept disabled),agent-context(plumbing undercontext-engineer).selftest/template/scaffoldignored as non-capabilities. Watch: v0.10.0 removes--ai/--no-git; v0.12.0 moves inlineCLAUDE.mdmanagement to theagent-contextextension.
- Success criteria (D9) — ✅ ADOPTED 2026-07-02. Otto had decisions (D1–D8) but no definition of working, so its value was unfalsifiable and the standing temptation was to keep polishing the framework itself. "Otto is working" is now judged by:
- Gate trail — every non-trivial task run through Otto leaves per-gate evidence in
specs/NNN/decision-log.md(which gate, what evidence cleared it). No evidence → not done. - Live learnings —
learnings/accrues real, verified lessons from actual runs (not just install scars); target: the store is non-trivially larger after a month of use and/otto-distillhas promoted ≥1 lesson into a rule. - Degradation observed — with a source uninstalled, the router substitutes the alternate and states it (or reports unavailable + the install cmd) — never a silent skip.
- No skipped-verification rework — no defect reaches REVIEW that a named earlier gate (tests green, analyze clean, browser DOM) should have caught; if one does, the gate is strengthened, not just the defect fixed.
- Real adoption — Otto is exercised on real tasks (measurable as
specs/NNN/trees + decision-logs created), not just self-referential edits. Provisional until validated against the first real end-to-end run (VALIDATION Part 2); tune the thresholds from what that run shows rather than guessing now. (Assessment finding #8.)
- Pre-first-run fixes (2nd assessment) — ✅ APPLIED 2026-07-02. Seven defects fixed before any real run; everything else deferred to run evidence:
modessemantics defined (registry header):hands-off= the conductor may run it autonomously; a manual-only step in a hands-off route is a pause point (or skipped if its gate already passes). Mechanical capabilities (debug, browser-verify, perf/a11y-review, …) gothands-offadded; interactive-by-nature ones (interview, grill, ideate, prototype) stay manual-only. Recipes' hands-off sections aligned.- Analyze gate predicate unified: "zero CRITICAL
/speckit.analyzefindings" everywhere; the spec-kit engine renders it as a human gate (no exit code to auto-evaluate), the conductor/Workflow backends read the report and auto-advance. - Gate trail made concrete:
memory/templates/decision-log.md(the artifact D9 measures) +install/otto-gate.sh(audits a run's trail: evidence present, result vocab, stalled stages); wired into routing-policy §D, spine, /flow, conductor, recipes. - Conductor bounded + backend status honest: max 2 remediation attempts per gate, then pause; the Claude Code Workflow backend is explicitly designed-not-built (needs a per-run script + per-subagent run-state injection) — hands-off defaults to the conductor skill until it exists.
- Triviality floor (routing-policy A0): one repo, ~one file, reversible, no security surface, provable in one step → no recipe, do it directly with tests. The verification bar stays.
- verify-otto.sh check 6: every slash invocation in the routing surface (router/, workflows/, USAGE) must resolve against the registry — catches stale-invocation drift mechanically.
- Learnings loop hardened: canonical naming (
<repo>/learnings.mdfile · globalotto-dev/learnings/global.md;learnings/names the store) + a once-per-session Stop-hook capture nudge (install/otto-learnings-nudge.{ps1,sh}, wiring in install/claude-code.md A3b).
- Footprint manifest + git-disposition modes — ✅ ADOPTED 2026-07-15. Problem: per-repo setup
touches ~10 paths across the tree (
.claude/,.specify/,specs/,docs/agents|adr/, root memory files,CLAUDE.mdblocks) with no single record of what was touched — so excluding, uninstalling, and auditing the footprint were all done from memory, and keeping Otto out of a repo's git history had no supported path. Two changes:
.otto-manifest.md(templatememory/templates/otto-manifest.md; install B8) — written as the last per-repo step: one row per created/modified path with owner, install step, disposition per mode, and regenerability. Exclude lines are generated from it, uninstall walks it bottom-up, audit reads it,/otto-syncdiffs it against disk.- Git disposition is declared per repo (install B6): team mode (default; governance +
artifacts tracked, regenerable/personal tooling ignored — the prior B6 behavior) or private
mode (whole footprint excluded via
.git/info/exclude— per-clone, never committed, so no committed line reveals Otto exists; optionalcore.excludesFilefor per-machine coverage; Otto never proposes committing manifest rows). Known hard edge, documented: managed blocks written into an already-trackedCLAUDE.mdcan't be hidden by ignore rules — private mode skips block-writing and uses~/.claude/CLAUDE.md/CLAUDE.local.mdinstead. - Rejected alternative: consolidating the footprint into a single
.otto/root dir with a one-line.gitignoreentry. Rejected because (a) most of the footprint is pinned by the host or the sources Otto deliberately doesn't fork (D2/D5):.claude/+CLAUDE.md(Claude Code),.specify/+specs/(spec-kit CLI),docs/agents/(Pocock) — onlystack.yamlandlearnings.mdare Otto-movable, a minority of the sprawl; (b) blanket-ignoring contradicts the memory stack — in team mode the constitution, specs, and ADRs are meant to be tracked; (c) for the privacy goal, a committed.gitignoreline itself advertises Otto —.git/info/excludeis strictly more private and works regardless of layout. Folding the two movable files into.otto/stays deferred as cosmetic (~50 path references across 17 framework files + back-compat shims in both session hooks for existing repos, for no functional gain); revisit only if the Otto-movable set grows.
- Claude Code Workflow backend — ✅ BUILT 2026-07-15 (
workflows/autonomous/workflow-backend.md), from the evidence of the first real hands-off run ( dry-run, VALIDATION.md Part 2): the REVIEW personas fan-out already ran as parallel fresh-context subagents with self-contained prompts and structured verdicts — the backend generalizes exactly that pattern to whole stages. Architecture (the two blockers from the old "designed-not-built" note, resolved):
- Per-run script = gate-bounded segments. A workflow script cannot ask the human anything mid-run, so every human-gated decision point (FRAME sign-off, spec acceptance, go/no-go, irreversible ops, git) is a segment boundary: the conductor authors one script per maximal human-gate-free run of route steps, executes it, returns to the main loop to pause at the gate, then authors the next segment. The conductor stays the driver; Workflow executes segments.
- Run-state injection = a standard prompt header, not the SessionStart hook: every
agent()prompt begins with an[OTTO RUN STATE]block (recipe · scope · segment · repo + feature-dir paths · stage · capability · gate predicate · memory-file paths · standing rules). Subagents need no session hook; the state travels in-band. - Gates in script code: step agents return schema-validated structured verdicts with evidence;
the deterministic script applies the gate predicate, runs the ≤2-attempt remediation loop, and
on persistent failure returns
{halted: stage, evidence}for the conductor to pause on. Agents write their own artifacts + decision-log rows (they have tools); the conductor audits withotto-gate.shat each segment end. - Selection unchanged (conductor.md table): workspace scope / heterogeneous steps / parallel fan-out → this backend; linear single-repo flows stay on the cheaper conductor-skill backend.
- Rejected: (a) one monolithic workflow with simulated pauses — a script that "waits" for a human either blocks a background run or fakes the gate; segment chaining keeps D6's pause semantics literal. (b) SessionStart-hook injection into subagents — hooks carry only the static routing brain; per-run state must be per-prompt.
- Context compaction, independent verification, failure-time learnings, gate classes — ✅ APPLIED 2026-07-15. Four of the assessment's seven gaps closed as protocol edits; the rest deferred to run evidence (consistent with item 15's build-from-evidence rule):
- Conductor-skill compaction (G2): the one-window backend now compacts as a protocol step —
/handoffat every human pause point and after heavy stages; never enter a new stage in a degraded window (conductor.md "Big-task handling"). Closes the context-rot gap on the cheap path. - Independent verification MUST (G3): SHOULD→MUST, all backends — mechanical-gate evidence is
re-checked by a second agent or by re-running the named command; judgment gates get a judge
separate from the actor (workflow-backend.md gate authoring; conductor.md step 3;
routing-policy §D6).
otto-gate.shstill audits trail structure only — evidence re-execution stays with the conductor; extending the script is deferred until a run shows the need. - Failure-time learnings (G6): on any failed gate, grep the learnings store for TRIGGER
matches on the failure signature before remediation attempt 1 (memory-stack.md
"Read (failure-time)"; conductor.md step 3 fail branch;
[OTTO RUN STATE]rules). The TRIGGER field finally fires at the moment it was designed for. - Gate classes (G7): every spine gate tagged mechanical | judgment | human with per-class enforcement (spine.md legend + per-stage lines; routing-policy §D6) — judgment predicates ("~95% confidence") no longer flow through the same self-assessed verdict path as "tests green".
- Run budget + stop conditions (G1, G4, G5) — closed 2026-07-21. The 2026-07-15 deferral
("revisit on run evidence") was re-examined by explicit call after an independent external
audit reached G1 and G5 cold (assessments/2026-07-21). A hands-off run now carries a declared
ceiling — token budget where the backend can read its own spend (Workflow only), structural
bounds everywhere (stage-steps · total remediation attempts · stage revisits · compactions) —
with three stop conditions: budget exhausted, stage entered a 4th time, and no progress
(two consecutive attempts returning the same failure signature). G4 landed as a prerequisite,
not scope creep: comparing signatures requires carrying attempt 1's
attemptedsummary forward, which is exactly what G4 asked for. Protocol: conductor.md · enforcement: workflow-backend.md authoring step 4 · routing-policy §D6-7. Defaults are engineered priors; a halt is recorded in the decision-log narrative, never as a gate-trail row (closed vocabulary). - Least-privilege tool scope — added 2026-07-21. New
tools:field besidecompute:(policy: registry/tool-scope.md), mechanizing two rules Otto had only as prose: judge-never-fixer (review capabilities runverify— no Write/Edit) and untrusted ingest (survey/understandruninspect— no shell). Binding in hands-off, advisory elsewhere;register-skillsdoes not stamp it yet, which is stated rather than implied. Surfaced by the external audit, not by the loop rubric.