Skip to content

Latest commit

 

History

History
93 lines (77 loc) · 10.4 KB

File metadata and controls

93 lines (77 loc) · 10.4 KB

Otto — Validation (roadmap step 7)

Two parts: (1) a static consistency audit Otto runs on itself (done — results below), and (2) a dry-run acceptance checklist to run once the sources are installed in a host.

Part 1 — Static consistency audit · ✅ PASS (2026-06-08)

Check Method Result
Internal links resolve script: extract each Markdown link target, resolve it relative to its file, test existence 0 broken across all repo files
Capability registry count - id: in registry/capabilities.yaml 41 capabilities (39 at first validation; +harvest-backlog, +validate-requirements)
Capability index ↔ registry diff USAGE.md §6 index against the registry ids exact match (41 = 41)
Recipes present every recipe linked from USAGE.md resolves 11 / 11 (8 base + assess + extract + inception)
Stale counts / terms grep for four/38/4 memory 1 found + fixed — USAGE.md §10 "four durable files" → "six memory members"
Doc invocations ↔ registry (added 2026-07-02) every slash invocation in router/ + workflows/ + USAGE resolves against the registry (brace-expanded, dot/hyphen-normalized) 0 unresolved (verified to fail on a poisoned copy)
Registered skills ↔ sources (added 2026-07-15, finding R2) each ~/.claude/skills/ copy diffed against its OTTO_HOME source (unit list parsed from register-skills.ps1; footer stripped); SKIPs cleanly on repo-only hosts proven red→green — flagged all 14 stale copies on first run, clean after re-registration; has since caught live drift twice
Stamped pins ↔ registry (added 2026-07-18, spec 002) each pinned skill copy's model:/effort: frontmatter compared against its capability's registry compute: block, plus a BOM/frontmatter byte check proven red→green — tampering /handoff effort high→low emitted PIN DRIFT + AUDIT: FAIL, restored by re-registration
Dashboard ↔ registry (added 2026-07-21) dashboard/index.html hand-mirrors the registry; compare capability cards, stage caps[]/route c[] arrays, and the PINS map against capabilities.yaml proven red→green — 5 injected drift classes each caught; caught one live inconsistency on its first run (ship-launch pin not leading with its model)

Re-run the whole Part-1 audit anytime — it's scripted now, 9 checks (links + capability count + index↔registry match + recipe presence + stale terms + invocation↔registry cross-check + registered-skill drift + stamped-pin drift + dashboard↔registry drift), exiting non-zero on any hard failure:

bash install/verify-otto.sh          # pair with `/otto-sync --drift` for the upstream-drift half

Wire it into a pre-commit hook or a scheduled tick so consistency (and drift) are caught mechanically, not by memory. A run's gate trail has its own auditor: bash install/otto-gate.sh specs/NNN.

Part 2 — Dry-run acceptance checklist (run with sources installed)

Prereq: install the sources you'll exercise (sources/sources.md) in a throwaway repo. Walk one recipe per shape, in both modes where relevant, and confirm each gate:

  • manual routing — PASS 2026-07-15 (). Pre-FRAME probe correctly routed to FRAME (finding R1: B-table must check constitution authoredness, not file existence — init always seeds a placeholder); post-FRAME re-probe → feature/SPEC as expected.
  • triviality floor — PASS 2026-07-15. A0 caught the trivial probe; no recipe card.
  • gate trail — PASS 2026-07-15. 9 evidenced rows incl. a FAIL→PASS remediation pair; otto-gate.sh → PASS (also on the bugfix run's 4-row trail).
  • bugfix (manual) — PASS 2026-07-15 (planted defect, labeled SIMULATED in specs/002-*/decision-log.md). Failing test reproduced BEFORE the fix; 1-char fix → green; commit proposed only.
  • feature (hands-off) — PASS 2026-07-15 via /otto-conductor greenfield: (same gates + FRAME). 4 human pauses (constitution, DEFINE manual-only, spec acceptance, go/no-go + proposed commit); never auto-committed. REVIEW personas fan-out caught 3 probe-verified blockers (ReDoS, day-bounds snap escape, whitespace title) that a 41-test green suite missed — remediated in 1 of the 2 bounded attempts.
  • assess — PASS 2026-07-15. specs/003-*/findings.md: coverage stated, 4 evidenced findings, handoff → feature recipe.
  • extract — PASS 2026-07-15. specs/004-*/extracted-spec.md from the untested prototype: 4 assumptions marked ⚠, divergence table linking prototype ↔ shipped core.
  • workspace — PASS 2026-07-15. Effort <EFFORT> across (core producer) + -api (API consumer): B′ selected workspace scope; coordination dir holds manifest, workspace spec with the POST / contract + fixture-based agreement mechanism, dependency graph, and 3 workspace gates (contract agreed / cross-repo integration / coordinated rollout); per-repo sub-specs linked both ways. Per-repo BUILD out of dry-run scope by design.
  • memory — PASS 2026-07-15. Seeded without overwriting; stack.yaml pin caught real drift (node 22 vs pinned 20 LTS) — updated at SHIP with human approval.
  • learnings — PASS 2026-07-15 (one caveat). Auto-append: observed live twice (7 global + 3 project lessons from real resolutions). Slice-read: observed live in a fresh session via the A4 CLAUDE.md route; SessionStart hook injection confirmed NOT surfacing in the VS Code extension (the documented A3 caveat — A4 is the sanctioned baseline there). Nudge at-most-once: proven end-to-end at script level against the real 551 KB session transcript (block emitted, marker created, repeat no-ops); live delivery in the extension requires a FULL VS Code restart after hook install (same class as the env-inheritance learning) — residual: observe one nudge in the first substantial post-restart session.
  • pause points — PASS 2026-07-15. DEFINE gate failed on missing CONTEXT.md → conductor paused at manual-only /grill-with-docs rather than faking it.
  • maintenance — PASS 2026-07-15. Sync: sources at pins; proposals emitted (re-register skills for R2 drift; add --script ps to B1; verify-otto check 7) and NOT applied. Distill: curation proposed, NOT applied.
  • degradation — PASS 2026-07-15. PK diagnose hidden → card substituted the AS layer partner with explicit Missing line + install command; SK alternate correctly skipped (prefer_when unmet).

Full evidence: <LOCAL-PATH>\validation-log.md (dry run of 2026-07-15; the throwaway repo's log is not included in the public release). Findings — all applied 2026-07-15 (same-day follow-up): R1 authoredness check now in routing-policy B-table + spine FRAME entry signal; R2 verify-otto.sh check 7 added (registered-skill drift; proven red→green — flagged all 14 stale copies, clean after re-registration) + skills re-registered; R3 backfill convention documented in conductor.md step 5; B1 --script ps|sh now in install B1 + sources.md (regression-tested scaffolding the second throwaway); distill curation applied to learnings/global.md (encoding lessons merged; --script lesson promoted to rule). A3b Stop hook installed in user settings. Remaining residual (one observation, not a blocker): after a full VS Code restart, the first substantial session should show exactly one learnings nudge — everything up to harness delivery is verified.

Workflow backend evidenced 2026-07-15 (same throwaways): the <EFFORT> workspace effort ran as the backend's first evidence run — 7 agents, 0 errors; parallel PLANs, sequenced producer→consumer BUILDs (RED-first honored by the build agent unprompted-in-detail), review fan-out with execution probes, live cross-repo integration gate (11/11 fixtures over real HTTP, sha256-matched both sides, WS-4 grep clean). All three gate trails (specs/005, specs/001, workspace log) audit PASS. One protocol defect found + fixed in-run (W1: run-state header must mandate the decision-log template structure — heading, PASS|FAIL|WAIVED vocabulary, no angle-bracket tokens in evidence). Workspace gate (c) rollout/rollback deliberately left PENDING (human call — no deployment surface on throwaways).

Part 3 — External check (2026-07-21)

Every result above is Otto grading Otto with Otto's own gates. A single independent scorer was run for contrast: npx @cobusgreyling/loop-audit (read-only, no knowledge of Otto's conventions).

16/100 → 25/100 after the session's edits; level L0; exit 2. The score is not meaningful — its rubric keys on filenames Otto does not use (STATE.md, LOOP.md, gate.yaml, patterns/registry.yaml), and each was verified as a false negative against the filesystem (specs/NNN/decision-log.md, spine.md+conductor.md, registry/capabilities.yaml). The delta was the value, and it confirmed four things Otto's own audits do not check:

Signal Outcome
cost.*, governance.stallDetection Reached G1 and G5 cold, matching the 2026-07-15 self-assessment. Both closed 2026-07-21
governance.toolScope New — no least-privilege scope existed; closed via registry/tool-scope.md
agentsMd, github New — no root CLAUDE.md (added) and no .github/ (open, BACKLOG F-05)
loopActivity The only check Otto passed, on git-commit evidence — dynamic proof rather than files on disk

Failing every static-config check while passing the sole dynamic-proof check is the inverse of the cargo-cult failure mode, and the strongest single argument for adopting a loopActivity-style check into verify-otto.sh (BACKLOG F-06). Full analysis: assessments/2026-07-21-external-tool-integration.md (not included in the public release).

Not yet addressed: the circularity of Parts 1–2 is noted here but not written up as a limitation of those results — that is BACKLOG F-07, deliberately left open rather than quietly folded in.

Status

Otto build is complete (steps 1–7) and the Part-2 dry-run is executed: 13/13 PASS (2026-07-15), including the Workflow backend's first evidence run. All seven loop-hardening findings (G1–G7) are closed as of 2026-07-21. Deferred by choice (DESIGN §12 items 6–7, full reference ROADMAP.md): the Portfolio layer and domain profiles. Live worklist with act-on triggers: BACKLOG.md.