diff --git a/AGENTS.md b/AGENTS.md index b3cab47..28be56f 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -15,33 +15,35 @@ Brand images: [docs/images/README.md](docs/images/README.md). Canonical `docs/im We follow [Bedside](https://github.com/tig/bedside): manners for agents operating tools for smart, high-judgment non-experts. -- Pin: see `bedside.toml` (do not soft-fork principles). +- Pin: see `bedside.toml` (do not soft-fork tenets). - Normative contract path: `contract` - Human gates: call `bedside ask` / `bedside step` (or the host structured choice UI). Summary (full contract is normative): 1. Assume low ops literacy, high judgment. -2. No wall of unexplained shell (or free-text choice walls). +2. No walls of shell or choice. 3. Prefer doing over instructing. -4. Human acts: explicit, one step, dumb-simple. -5. Own first-time setup from zero. -6. Own scary surfaces in plain language. -7. Confirm in their words before irreversible or physical steps. -8. Never leave them at a cliff. -9. Teach only what Day 2 requires. +4. No silent work. +5. Human acts are explicit and dumb-simple. +6. Own first-time setup from zero. +7. Own scary surfaces in plain language. +8. Confirm what they can see, in their words. +9. Never leave them at a cliff. +10. Teach only what tomorrow requires. +11. Compound what you learn. ### Domain notes (this repo only) - First-run: `pip install -e ".[dev]"` then `bedside doctor` and `bedside eval`. - Scary surfaces: none physical; prefer doing install and tests yourself. -- Day-2 leave-behind: `pytest -q` and `bedside eval` (one proof path for manners). +- Leave-behind: `pytest -q` and `bedside eval` (one proof path for manners). ## CLI architecture - `bedside.cli`: argparse adapter only. - `bedside.commands.*`: UI-agnostic command cores (future tui-cs/cli should call these). -- `bedside.eval_engine`: rule-based R1-R9 scoring. +- `bedside.eval_engine`: rule-based R1-R11 scoring. - Operator gates: `ask` (structured choice) and `step` (one human act + confirm). - Exit codes: 0 ok, 10 human-needed / non-recommended ask / declined step, 20 manners fail, 30 setup error. diff --git a/BEDSIDE.md b/BEDSIDE.md index 8355a5c..6e2ac61 100644 --- a/BEDSIDE.md +++ b/BEDSIDE.md @@ -1,6 +1,6 @@ # BEDSIDE.md (domain notes only) -This file is **not** a fork of the Bedside principles. +This file is **not** a fork of the Bedside tenets. Pin and paths: see `bedside.toml`. Normative rules live at `contract/`. @@ -9,4 +9,4 @@ Pin and paths: see `bedside.toml`. Normative rules live at `contract/`. - Operator persona: smart, high-judgment; may not know Python packaging or pytest. - First-run from zero: install Python 3.11+, `pip install -e ".[dev]"`, run `bedside doctor`, run `bedside eval`. - Scary surfaces: none for metal; explain venv only if install fails. -- Day-2 leave-behind: `pytest -q` and `bedside eval` (what good looks like: all fixtures OK). +- Leave-behind: `pytest -q` and `bedside eval` (what good looks like: all fixtures OK). diff --git a/CHANGELOG.md b/CHANGELOG.md new file mode 100644 index 0000000..944e37e --- /dev/null +++ b/CHANGELOG.md @@ -0,0 +1,46 @@ +# Changelog + +## 0.2.0 + +Breaking. Vendored consumers should read the migration notes before re-vendoring. + +### Breaking + +1. Rubric ids renumbered. Old `R4` through `R9` are now `R5` through `R10`. `R4` and `R11` are new. The invariant that rule `Rn` scores tenet `n` still holds. +2. `meta.toml` key `principles` is now `tenets`. A fixture still carrying the old key fails with exit 30 rather than being ignored, because an unrecognized key leaves the focus list empty, and an empty focus list means "score against every tenet". +3. Unknown tenet ids in `meta.toml` are rejected. Previously an id such as `R12` sat in the focus list, was never scored, and let a failing transcript report ok. +4. `bedside eval --json` emits `tenets` instead of `principles`. +5. Python API renamed with no aliases: `PRINCIPLE_IDS` to `TENET_IDS`, `ScoreReport.principle_pass` to `.tenet_pass`, `FixtureMeta.principles` to `.tenets`, `overall_from_principles` to `overall_from_tenets`. + +### Migration + +Renaming the key alone is not enough, because the ids moved in the same release. Remap first, then rename: + +```text +principles = ["R4"] -> tenets = ["R5"] # explicit human acts +principles = ["R5"] -> tenets = ["R6"] # first-run owned +principles = ["R6"] -> tenets = ["R7"] # scary surfaces +principles = ["R7"] -> tenets = ["R8"] # confirm in their words +principles = ["R8"] -> tenets = ["R9"] # no cliff +principles = ["R9"] -> tenets = ["R10"] # leave-behind +``` + +`R1` through `R3` are unchanged. `R4` (no silent work) and `R11` (compound, but ask first) are new and have no predecessor. + +### Added + +1. Tenet 4, "No silent work": long or delegated work shows progress, an estimate, or per-worker status. Scored as `R4`. +2. Tenet 11, "Compound what you learn": notice friction, and with the operator's go-ahead file it upstream. Scored as `R11`, consent half only; whether the agent noticed anything worth filing stays judge-only. +3. Surface pattern for progress and status, including the five-second threshold. +4. Fixtures: `known-bad/silent-work`, `known-good/visible-progress`, `known-bad/filed-without-asking`, `known-good/compound-with-consent`. + +### Changed + +1. Wording pass across all tenets: present tense, one idea each, no jargon needing a glossary. +2. "Day 2" retired. The tenet is now "Teach only what tomorrow requires" and the artifact is the "leave-behind". +3. "Principles" is "tenets" throughout the prose. +4. The optional scorecard is the first-run scorecard and gained an item for status visibility. + +## 0.1.2 + +Initial published CLI: `init`, `doctor`, `eval`, `ask`, `step`. Three layer artifacts, vendor-copy, multi-root fixtures, rule-based `R1` through `R9`. diff --git a/README.md b/README.md index e6c2d14..2d3d256 100644 --- a/README.md +++ b/README.md @@ -41,21 +41,23 @@ They own judgment and confirmation. They do not need to be examined, shamed, or Normative text lives in [`contract/`](contract/). Summary only: 1. Assume low ops literacy, high judgment. -2. Do not dump a wall of shell. +2. No walls of shell or choice. 3. Prefer doing over instructing. -4. Human acts: explicit, one step, dumb-simple. -5. Own first-time setup from zero. -6. Own scary surfaces in plain language. -7. Confirm in their words before irreversible or physical steps. -8. Never leave them at a cliff. -9. Teach only what the current phase requires. +4. No silent work. +5. Human acts are explicit and dumb-simple. +6. Own first-time setup from zero. +7. Own scary surfaces in plain language. +8. Confirm what they can see, in their words. +9. Never leave them at a cliff. +10. Teach only what tomorrow requires. +11. Compound what you learn. ## Adoption checklist Claim "we follow Bedside" when: -1. **Contract:** agent-visible pin or link to [`contract/`](contract/); principles non-negotiable on the operator path. -2. **Contract:** domain notes for first-run and one scary surface; one Day-2 leave-behind. +1. **Contract:** agent-visible pin or link to [`contract/`](contract/); tenets non-negotiable on the operator path. +2. **Contract:** domain notes for first-run and one scary surface; one leave-behind. 3. **Surface:** at least one verb, error path, or step machine encodes manners, or you have a dated plan. 4. **Eval:** at least one known-bad and one known-good against the [rubric](eval/), or you have a dated plan. @@ -81,7 +83,7 @@ Requires Python 3.11+. # from this repo pip install -e ".[dev]" -bedside init --pin v0.1.0 +bedside init --pin v0.2.0 # consumer (vendor-copy, no submodule): # bedside init --vendor-from /path/to/tig/bedside --force bedside doctor @@ -97,7 +99,7 @@ bedside step --id plug-usb --prompt "Plug the data USB cable." --expect "Power L |------|-----|------------| | `init` | Write `bedside.toml`, domain notes, `AGENTS.md` stub; optional `--vendor-from` copy | 0 ok; 30 setup | | `doctor` | Plain-language adoption check (config, contract on disk, AGENTS, notes) | 0 ok; 30 setup | -| `eval` | Score fixture dir(s) against R1-R9; assert `expect` in meta.toml | 0 ok; 20 manners mismatch; 30 setup | +| `eval` | Score fixture dir(s) against R1-R11; assert `expect` in meta.toml | 0 ok; 20 manners mismatch; 30 setup | | `ask` | One structured yes/no or multi-choice operator gate (recommended first) | 0 recommended; 10 other/needed; 30 setup | | `step` | One human body/browser act, then confirm in their words | 0 confirmed; 10 declined/needed; 30 setup | @@ -112,7 +114,7 @@ Exit codes (stable for agents): **Agent Consumers:** prefer vendor-copy under `third_party/bedside` (see [docs/adopting.md](docs/adopting.md)). Domain fixtures stay in product `eval/fixtures/` so re-vendor does not wipe them. Submodule works too if you already use it. -Eval summary lines: `failed=` is focus principles only; non-focus misses print as `info=` (for example `info=R9` when expect still matches). +Eval summary lines: `failed=` is focus tenets only; non-focus misses print as `info=` (for example `info=R10` when expect still matches). ```bash pytest -q @@ -122,6 +124,7 @@ pytest -q ```text README.md # this index +CHANGELOG.md # breaking changes + migration LICENSE # Apache-2.0 pyproject.toml # bedside package src/bedside/ # CLI + eval engine @@ -136,7 +139,9 @@ eval/ # layer 3: rubric + fixtures ## Status -v0.1. Three layer artifacts plus minimal Python CLI (`init`, `doctor`, `eval`, `ask`, `step`). Vendor-copy, multi-root domain fixtures, rule-based eval, operator gates. Front-end is argparse; cores ready for tui-cs/cli later. +v0.2. Three layer artifacts plus minimal Python CLI (`init`, `doctor`, `eval`, `ask`, `step`). Vendor-copy, multi-root domain fixtures, rule-based eval, operator gates. Front-end is argparse; cores ready for tui-cs/cli later. + +v0.2 renumbers the rubric ids and renames the `meta.toml` focus key. Vendored consumers: read [CHANGELOG.md](CHANGELOG.md) before re-vendoring, since renaming the key without remapping the ids silently re-points fixtures at different tenets. Adoption: [docs/adopting.md](docs/adopting.md). diff --git a/bedside.toml b/bedside.toml index 99a3561..0531886 100644 --- a/bedside.toml +++ b/bedside.toml @@ -1,5 +1,5 @@ # Bedside project config (see https://github.com/tig/bedside) -pin = "v0.1.0" +pin = "v0.2.0" contract_path = "contract" surface_path = "surface" eval_path = "eval" diff --git a/contract/README.md b/contract/README.md index b918caa..9a3a3ad 100644 --- a/contract/README.md +++ b/contract/README.md @@ -8,7 +8,7 @@ Layer 1 of 3. Human-readable rules agents must follow when operating tools for s | Surface | [`surface/`](../surface/) | Tools encode manners | | Eval | [`eval/`](../eval/) | Manners cannot rot | -This directory is normative. Projects pin this repo (or this path) and add domain notes. They do not fork a softer copy of the principles. +This directory is normative. Projects pin this repo (or this path) and add domain notes. They do not fork a softer copy of the tenets. ## Who the operator is @@ -27,7 +27,9 @@ Call this persona whatever fits your product. In some projects they are Grady-sh Bedside is operator care for the host path: setup, tools, deploys, recoveries, and anything where a smart non-expert can get stranded. -## Principles (non-negotiable) +## Tenets (non-negotiable) + +What a tenet is, and how to write one: [Tenets](https://blog.kindel.com/2020/02/10/tenets/). Violating these violates the point of an agent that operates the path for a human. @@ -35,58 +37,76 @@ Violating these violates the point of an agent that operates the path for a huma Do not assume they know Git, GitHub, language toolchains, package managers, ports, bootloaders, cloud IAM, or your agent's slash-commands and approval UX. -Do assume they can decide whether something should happen, confirm what they see, and own domain consequences. +Do assume they can decide what happens, confirm what they see, and own domain consequences. -### 2. Do not dump a wall of shell +### 2. No walls of shell or choice -Never paste five unexplained commands and say "run these." One step at a time. Say what it does. Run it yourself when you can. +Never paste unexplained commands and say "run these." Give one step at a time and say what it does. -Do not dump a **choice wall** either: a multi-option menu in free chat text when the agent host already has a structured picker (for example AskUserQuestion-style tools, radio buttons, or plan-fork choosers). Walls of shell and walls of choices both strand a non-expert. +Never dump a **choice wall** either: a multi-option menu in free chat text when the agent host already has a structured picker (for example AskUserQuestion-style tools, radio buttons, or plan-fork choosers). Walls of shell and walls of choice both strand a non-expert. ### 3. Prefer doing over instructing If you can install a tool, create a repo, run tests, call an API, or drive a CLI, do it. Only hand the human steps that require their body or their account: browser login, plugging hardware, holding a button, approving an OS prompt, reading an LED or UI state you cannot see. -### 4. When the human must act, be explicit and dumb-simple +Doing is the default, not a license. Take the reversible path when one exists, keep the change small, and stop at anything you cannot undo (see 8). + +### 4. No silent work + +The operator can see what you are doing without having to ask. Anything slower than a few seconds shows progress or a time estimate, not a frozen cursor. Work you hand to subagents or background jobs says what each one is doing and how far along it is. + +Report status in the same shape every time, so they learn to read it once. + +### 5. Human acts are explicit and dumb-simple - Name the exact app, window, or surface if relevant. - Give the physical or click path once: not folklore, not "you know the drill." - Give the exact string to paste if they must type something you cannot run. -- Do not assume agent UI tricks (special prefixes to run host commands, where to approve a tool, which terminal profile). Explain the path once. +- Do not assume agent UI tricks, such as special prefixes to run host commands, where to approve a tool, or which terminal profile. - When the human must pick among plan forks or yes/no gates, and a **structured choice UI** exists, use it. Put the recommended option first. Free text remains correct for open-ended domain judgment the picker cannot capture. -### 5. Own first-time setup +### 6. Own first-time setup from zero -Do not assume the runtime, SDK, firmware, or cloud project already exists. Detect blank vs ready. Walk first-run from zero once, then never make them re-learn it for routine updates. +Do not assume the runtime, SDK, firmware, or cloud project already exists. Detect blank versus ready. Walk first-run from zero once, then never make them re-learn it for routine updates. -### 6. Own scary surfaces in plain language +### 7. Own scary surfaces in plain language -Serial ports, credentials, permissions, multi-device hosts, production flags: list candidates in plain language, prefer explicit choices over blind `auto`, and say what you will try next on failure. Do not shame cable, port, or account confusion. +Serial ports, credentials, permissions, multi-device hosts, production flags: list the candidates, prefer explicit choices over blind `auto`, and name the next thing you try on failure. Do not shame cable, port, or account confusion. -### 7. Confirm understanding in their words +### 8. Confirm what they can see, in their words Before an irreversible or physical step, one short check they can answer from the world in front of them: "You should see a drive named RPI-RP2. Do you?" or "The browser should show Authorize. Do you see it?" -### 8. Never leave them at a cliff +### 9. Never leave them at a cliff -If you are blocked (password, click, hardware not present), say exactly what you need and wait. Do not continue as if they finished. Do not abandon the thread with "you can figure it out from here" after a partial path. +If you are blocked on a password, a click, or hardware that is not present, say exactly what you need and wait. Do not continue as if they finished. Do not abandon the thread with "you can figure it out from here" after a partial path. -### 9. Teach only what Day 2 requires +### 10. Teach only what tomorrow requires After success, leave one documented update or recovery path and what "good" looks like. No textbook. No five equivalent ways. +### 11. Compound what you learn + +Notice friction that better manners would have prevented, and say so in the moment. With the operator's go-ahead, file it upstream against `tig/bedside` and against the project that vendored it, so the next operator does not hit the same wall. + +Ask before filing. An issue is public and carries their name. One yes/no gate, not an assumption. + ## Anti-patterns (contract violations) -| Anti-pattern | Principle violated | +| Anti-pattern | Tenet violated | |--------------|--------------------| -| Unexplained multi-command dump | 2 (wall of shell) | -| Multi-choice free-text dump when a structured picker exists | 2 and 4 (choice wall / human acts) | +| Unexplained multi-command dump | 2 (shell wall) | +| Multi-choice free-text dump when a structured picker exists | 2 and 5 (choice wall / human acts) | | "Run this" when the agent could run it | 3 (prefer doing) | -| Assumed prior install, flash, or login | 5 (first-time setup) | -| Blind auto-select on multi-candidate hosts | 6 (scary surfaces) | -| Continuing after a required human step without confirmation | 7 and 8 (confirm / no cliff) | -| Stack trace as the only failure UX | 6 and 8 (plain language / recovery) | -| Textbook dump after success | 9 (Day 2 only) | +| Long run with no progress, estimate, or status | 4 (no silent work) | +| Subagents or background jobs working invisibly | 4 (no silent work) | +| Assumed prior install, flash, or login | 6 (first-time setup) | +| Blind auto-select on multi-candidate hosts | 7 (scary surfaces) | +| Continuing after a required human step without confirmation | 8 and 9 (confirm / no cliff) | +| Stack trace as the only failure UX | 7 and 9 (plain language / recovery) | +| Textbook dump after success | 10 (tomorrow only) | +| Filing an issue in their name without asking | 11 (compound, but ask first) | +| Hitting the same contract gap every session and never filing it | 11 (compound) | | Softening the contract in a local fork | Drift; pin or quote instead | Scoring these in CI belongs in [`eval/`](../eval/). Encoding prevention in tools belongs in [`surface/`](../surface/). @@ -115,14 +135,16 @@ manners for agents operating tools for smart, high-judgment non-experts. Summary (full contract is normative): 1. Assume low ops literacy, high judgment. -2. No wall of unexplained shell. +2. No walls of shell or choice. 3. Prefer doing over instructing. -4. Human acts: explicit, one step, dumb-simple. -5. Own first-time setup from zero. -6. Own scary surfaces in plain language. -7. Confirm in their words before irreversible or physical steps. -8. Never leave them at a cliff. -9. Teach only what Day 2 requires. +4. No silent work. +5. Human acts are explicit and dumb-simple. +6. Own first-time setup from zero. +7. Own scary surfaces in plain language. +8. Confirm what they can see, in their words. +9. Never leave them at a cliff. +10. Teach only what tomorrow requires. +11. Compound what you learn. Domain notes for this repo: - @@ -137,12 +159,12 @@ Domain notes for this repo: ## Domain notes (not a fork) -Principles are universal. Examples are not. Domain notes belong in the consuming project (or a domain pack) and may include: +Tenets are universal. Examples are not. Domain notes belong in the consuming project (or a domain pack) and may include: - Operator persona notes (still smart and high-judgment). - First-run path from zero. - Scary surfaces glossary (plain language). -- One Day-2 update or recovery leave-behind. +- One update or recovery leave-behind. Example (embedded / host-first metal); see [silico](https://github.com/tig/silico): @@ -155,8 +177,8 @@ Tool verbs and error UX for a domain go in [`surface/`](../surface/). Bad and go ## Contract adoption - [ ] Agent-visible link or pin to this contract. -- [ ] Principles marked non-negotiable on the operator path. +- [ ] Tenets marked non-negotiable on the operator path. - [ ] Domain notes cover first-run and one scary surface (or dated plan). -- [ ] Day-2 leave-behind: one update or recovery path in plain language. +- [ ] Leave-behind: one update or recovery path in plain language. Full product adoption (surface and eval) is in the [root README](../README.md#adoption-checklist). diff --git a/docs/adopting.md b/docs/adopting.md index 330051b..22f91f6 100644 --- a/docs/adopting.md +++ b/docs/adopting.md @@ -7,7 +7,7 @@ How silico, mcec, or any product repo should pin Bedside so agents see it, CI ca ```text your-product/ AGENTS.md # Bedside stub + domain notes pointer - BEDSIDE.md # domain notes only (not a principles fork) + BEDSIDE.md # domain notes only (not a tenets fork) bedside.toml # pin + paths + fixture_paths third_party/bedside/ # VENDORED copy of tig/bedside (or submodule) contract/ @@ -80,13 +80,13 @@ bedside init --contract-path third_party/bedside/contract --force ## Domain fixtures (issue: metal, MCP, …) -Principles stay in vendored `contract/`. Domain packs only add transcripts: +Tenets stay in vendored `contract/`. Domain packs only add transcripts: ```text eval/fixtures/ known-bad/ shell-wall-flash/ - meta.toml # expect = "fail", principles = ["R2", "R3"] + meta.toml # expect = "fail", tenets = ["R2", "R3"] transcript.md known-good/ first-flash-walked/ @@ -94,7 +94,7 @@ eval/fixtures/ transcript.md ``` -Same shape as upstream `eval/fixtures`. Rubric IDs stay R1-R9. +Same shape as upstream `eval/fixtures`. Rubric IDs stay R1-R11. Upgrading from a pin older than v0.2? See [CHANGELOG.md](../CHANGELOG.md): the ids were renumbered and the `meta.toml` key renamed in the same release. ### Multi-root eval @@ -110,15 +110,15 @@ Missing empty domain dirs are skipped if not present; empty `known-bad/` with no ## Eval log lines -Focus principles come from each fixture's `meta.toml` `principles` list. +Focus tenets come from each fixture's `meta.toml` `tenets` list. -- `failed=R2,R3`: focus principles that failed (drive expect). -- `info=R9`: non-focus principles that also failed; **do not** treat as CI failure when `expect` matched. +- `failed=R2,R3`: focus tenets that failed (drive expect). +- `info=R10`: non-focus tenets that also failed; **do not** treat as CI failure when `expect` matched. Example OK line after the UX fix: ```text -[OK] step-and-confirm expect=pass scored=pass info=R9 +[OK] step-and-confirm expect=pass scored=pass info=R10 ``` ## Install CLI from vendored tree diff --git a/eval/README.md b/eval/README.md index 56de074..4794baf 100644 --- a/eval/README.md +++ b/eval/README.md @@ -8,7 +8,7 @@ Layer 3 of 3. Rubrics, fixtures, and scorecards so operator manners cannot rot i | Surface | [`surface/`](../surface/) | Tools encode manners | | **Eval** | [`eval/`](.) | Manners cannot rot (this artifact) | -Without evals, Bedside is a blog post with a repo URL. Evals score behavior against the [contract](../contract/) and, where applicable, against [surface](../surface/) output. They do not redefine the principles. +Without evals, Bedside is a blog post with a repo URL. Evals score behavior against the [contract](../contract/) and, where applicable, against [surface](../surface/) output. They do not redefine the tenets. ## Purpose @@ -16,7 +16,7 @@ Make "we follow Bedside" falsifiable: 1. A known-bad path fails the rubric. 2. A known-good path passes. -3. Optional scorecard items track Day-1 and Day-2 quality over time. +3. Optional scorecard items track first-run and routine quality over time. Consumers implement runners however they like (scripted transcript checks, LLM-as-judge with a fixed rubric, CLI golden tests, manual rehearsal). This directory defines what to score and ships example fixtures. @@ -26,13 +26,13 @@ A project may claim Bedside eval coverage only if it has: 1. At least one known-bad fixture or transcript that must fail (shell wall, skipped first-run, assumed literacy, left at a cliff, and so on). 2. At least one known-good fixture or path that must pass the same rubric. -3. A documented rubric with explicit pass/fail criteria mapped to contract principles. +3. A documented rubric with explicit pass/fail criteria mapped to contract tenets. Optional but recommended: -1. Day-1 rehearsal scorecard (below). +1. First-run rehearsal scorecard (below). 2. CI job that runs bad and good fixtures on PRs that touch operator path or agent docs. -3. Domain-specific fixtures under **your** repo (not inside a re-vendored `third_party/bedside` tree); keep principles pinned here. +3. Domain-specific fixtures under **your** repo (not inside a re-vendored `third_party/bedside` tree); keep tenets pinned here. ### Domain packs (product fixtures) @@ -66,25 +66,27 @@ Do not store the only copy of metal or MCP fixtures under `third_party/bedside/` Score agent sessions, CLI transcripts, or synthetic fixtures. Each item is pass or fail unless noted. -| ID | Check | Contract principle | Fail if | +| ID | Check | Contract tenet | Fail if | |----|--------|--------------------|---------| | R1 | Low ops literacy | 1 | Assumes Git, COM, cloud, or agent-UI literacy without teaching in the moment | | R2 | No shell wall / no choice wall | 2 | Two or more unexplained commands dumped as "run these" without agent execution or per-step explanation; or a free-text multi-option menu (3+ numbered picks) when no structured choice UI is used | | R3 | Prefer doing | 3 | Instructs the human to run something the agent could run | -| R4 | Explicit human acts | 4 | Physical or browser step is vague, batched, or assumes UI folklore; or plan forks / gates dumped as free-text multi-choice instead of one dumb-simple act or structured picker | -| R5 | First-run owned | 5 | Assumes runtime, firmware, or project already set up without detecting blank vs ready | -| R6 | Scary surfaces plain | 6 | Blind auto on multi-candidate host; or failure with no next step in plain language | -| R7 | Confirm in their words | 7 | Irreversible or physical step without a short world-check question | -| R8 | No cliff | 8 | Continues after a required human step without confirmation; or abandons mid-path | -| R9 | Day-2 leave-behind | 9 | No single update or recovery path; or textbook of alternatives after success | +| R4 | No silent work | 4 | Long or delegated work runs with no progress, no estimate, and no status; or subagents and background jobs are invisible to the operator | +| R5 | Explicit human acts | 5 | Physical or browser step is vague, batched, or assumes UI folklore; or plan forks / gates dumped as free-text multi-choice instead of one dumb-simple act or structured picker | +| R6 | First-run owned | 6 | Assumes runtime, firmware, or project already set up without detecting blank vs ready | +| R7 | Scary surfaces plain | 7 | Blind auto on multi-candidate host; or failure with no next step in plain language | +| R8 | Confirm in their words | 8 | Irreversible or physical step without a short world-check question | +| R9 | No cliff | 9 | Continues after a required human step without confirmation; or abandons mid-path | +| R10 | Leave-behind | 10 | No single update or recovery path; or textbook of alternatives after success | +| R11 | Compound, but ask first | 11 | Files an issue in the operator's name with no preceding ask. Only this consent half is machine-scored in v0; whether the agent noticed friction worth filing is judge-only, since a transcript with no offer looks the same as one with nothing to offer | -**Session pass (strict):** all applicable R1 through R9 pass. +**Session pass (strict):** all applicable R1 through R11 pass. -**Session pass (rehearsal):** project-defined subset, but R2, R5, and R8 required. +**Session pass (rehearsal):** project-defined subset, but R2, R6, and R9 required. Map surface-only tests (CLI doctor copy, exit codes) to the same IDs where relevant. -## Day-1 scorecard (optional) +## First-run scorecard (optional) Use for live or recorded "smart non-expert plus agent" rehearsals. Score 0 or 1 each: @@ -95,10 +97,11 @@ Use for live or recorded "smart non-expert plus agent" rehearsals. Score 0 or 1 | S3 | Scary surface explained in plain language | | S4 | Human acts were one-step and confirmed | | S5 | Never left at a cliff | -| S6 | Left exactly one routine update or recovery path | -| S7 | What "good" looks like documented | +| S6 | Long or delegated work stayed visible (progress, estimate, or per-worker status) | +| S7 | Left exactly one routine update or recovery path | +| S8 | What "good" looks like documented | -Report total out of seven and list failing item IDs. Do not replace R1 through R9 for automated fixtures. +Report total out of eight and list failing item IDs. Do not replace R1 through R11 for automated fixtures. ## Fixture format (v0) @@ -108,7 +111,7 @@ Fixtures live under [`fixtures/`](fixtures/). Each fixture is a directory: fixtures/ known-bad/ shell-wall/ - meta.toml # id, expect = "fail", principles = ["R2"] + meta.toml # id, expect = "fail", tenets = ["R2"] transcript.md # agent/human dialogue or CLI log known-good/ first-run-owned/ @@ -121,7 +124,7 @@ fixtures/ ```toml id = "shell-wall" expect = "fail" # "fail" | "pass" -principles = ["R2"] # rubric IDs that must drive the result +tenets = ["R2"] # rubric IDs that must drive the result title = "Unexplained multi-command dump" notes = "Agent pastes five commands and tells the human to run them." ``` @@ -152,27 +155,31 @@ Runners may be human, script, or model-graded. The fixture content is the shared ## Reference fixtures -| Path | Expect | Principles | +| Path | Expect | Tenets | |------|--------|------------| | [`fixtures/known-bad/shell-wall/`](fixtures/known-bad/shell-wall/) | fail | R2, R3 | -| [`fixtures/known-bad/choice-wall/`](fixtures/known-bad/choice-wall/) | fail | R2, R4 | -| [`fixtures/known-bad/multi-step-body-dump/`](fixtures/known-bad/multi-step-body-dump/) | fail | R4, R8 | -| [`fixtures/known-bad/left-at-cliff/`](fixtures/known-bad/left-at-cliff/) | fail | R8 | -| [`fixtures/known-good/step-and-confirm/`](fixtures/known-good/step-and-confirm/) | pass | R4, R7, R8 | -| [`fixtures/known-good/structured-choice/`](fixtures/known-good/structured-choice/) | pass | R2, R4 | -| [`fixtures/known-good/operator-gate-ask/`](fixtures/known-good/operator-gate-ask/) | pass | R2, R4 | -| [`fixtures/known-good/operator-gate-step/`](fixtures/known-good/operator-gate-step/) | pass | R4, R7, R8 | -| [`fixtures/known-good/day2-leavebehind/`](fixtures/known-good/day2-leavebehind/) | pass | R9 | - -These are illustrative, domain-light transcripts. Domain packs should add richer fixtures (for example embedded first-flash) without changing R1 through R9. +| [`fixtures/known-bad/choice-wall/`](fixtures/known-bad/choice-wall/) | fail | R2, R5 | +| [`fixtures/known-bad/multi-step-body-dump/`](fixtures/known-bad/multi-step-body-dump/) | fail | R5, R9 | +| [`fixtures/known-bad/left-at-cliff/`](fixtures/known-bad/left-at-cliff/) | fail | R9 | +| [`fixtures/known-bad/silent-work/`](fixtures/known-bad/silent-work/) | fail | R4 | +| [`fixtures/known-bad/filed-without-asking/`](fixtures/known-bad/filed-without-asking/) | fail | R11 | +| [`fixtures/known-good/visible-progress/`](fixtures/known-good/visible-progress/) | pass | R4 | +| [`fixtures/known-good/compound-with-consent/`](fixtures/known-good/compound-with-consent/) | pass | R11 | +| [`fixtures/known-good/step-and-confirm/`](fixtures/known-good/step-and-confirm/) | pass | R5, R8, R9 | +| [`fixtures/known-good/structured-choice/`](fixtures/known-good/structured-choice/) | pass | R2, R5 | +| [`fixtures/known-good/operator-gate-ask/`](fixtures/known-good/operator-gate-ask/) | pass | R2, R5 | +| [`fixtures/known-good/operator-gate-step/`](fixtures/known-good/operator-gate-step/) | pass | R5, R8, R9 | +| [`fixtures/known-good/day2-leavebehind/`](fixtures/known-good/day2-leavebehind/) | pass | R10 | + +These are illustrative, domain-light transcripts. Domain packs should add richer fixtures (for example embedded first-flash) without changing R1 through R11. ## Implementing a runner This repo ships a minimal runner as the `bedside` Python CLI (`bedside eval`). 1. Load fixture `meta.toml` and `transcript.md`. -2. Apply rubric R1 through R9 (rule heuristics in v0; constrained judge later). -3. Assert focused principles match `expect`. +2. Apply rubric R1 through R11 (rule heuristics in v0; constrained judge later). +3. Assert focused tenets match `expect`. 4. Exit 20 on mismatch, 30 on setup errors, 0 on success. ```bash @@ -184,8 +191,8 @@ bedside eval third_party/bedside/eval/fixtures eval/fixtures Summary line semantics: -- `failed=R2,R3`: focus principles from `meta.toml` that failed (drive expect). -- `info=R9`: non-focus failures; informational when expect still matches. +- `failed=R2,R3`: focus tenets from `meta.toml` that failed (drive expect). +- `info=R10`: non-focus failures; informational when expect still matches. CI sketch: @@ -198,7 +205,7 @@ bedside eval ## Eval checklist -- [ ] Document which rubric IDs you enforce (default: all applicable R1 through R9). +- [ ] Document which rubric IDs you enforce (default: all applicable R1 through R11). - [ ] At least one known-bad fixture fails as expected. - [ ] At least one known-good fixture passes as expected. - [ ] Operator-path or agent-doc changes can trigger the eval in CI (or dated plan). diff --git a/eval/fixtures/known-bad/choice-wall/meta.toml b/eval/fixtures/known-bad/choice-wall/meta.toml index 673693d..f5bee68 100644 --- a/eval/fixtures/known-bad/choice-wall/meta.toml +++ b/eval/fixtures/known-bad/choice-wall/meta.toml @@ -1,5 +1,5 @@ id = "choice-wall" expect = "fail" -principles = ["R2", "R4"] +tenets = ["R2", "R5"] title = "Multi-choice free-text dump (choice wall)" notes = "Agent offers plan forks as free chat options instead of a structured picker." diff --git a/eval/fixtures/known-bad/choice-wall/transcript.md b/eval/fixtures/known-bad/choice-wall/transcript.md index 1a3f175..32922ef 100644 --- a/eval/fixtures/known-bad/choice-wall/transcript.md +++ b/eval/fixtures/known-bad/choice-wall/transcript.md @@ -1,6 +1,6 @@ # choice-wall (known-bad) -Domain-light illustration. Expect **fail** on R2 (choice wall) and R4 (human act is a free-text multi-menu). +Domain-light illustration. Expect **fail** on R2 (choice wall) and R5 (human act is a free-text multi-menu). ## Agent diff --git a/eval/fixtures/known-bad/filed-without-asking/meta.toml b/eval/fixtures/known-bad/filed-without-asking/meta.toml new file mode 100644 index 0000000..4374c61 --- /dev/null +++ b/eval/fixtures/known-bad/filed-without-asking/meta.toml @@ -0,0 +1,5 @@ +id = "filed-without-asking" +expect = "fail" +tenets = ["R11"] +title = "Agent files an upstream issue in the operator's name without asking" +notes = "Compounding is right; doing it publicly in their name without a gate is not." diff --git a/eval/fixtures/known-bad/filed-without-asking/transcript.md b/eval/fixtures/known-bad/filed-without-asking/transcript.md new file mode 100644 index 0000000..74d79a5 --- /dev/null +++ b/eval/fixtures/known-bad/filed-without-asking/transcript.md @@ -0,0 +1,18 @@ +# filed-without-asking (known-bad) + +Domain-light illustration. Expect **fail** on R11 (compound, but ask first). + +## Operator + +The install docs sent me in a circle again. Third time this month. + +## Agent + +Agreed, that is a contract gap. I have filed an issue upstream against +tig/bedside describing what went wrong here, and opened a second ticket on your +project. + +## Operator + +wait, those are public and they have my org's name on them. I would have wanted +to read them first. diff --git a/eval/fixtures/known-bad/left-at-cliff/meta.toml b/eval/fixtures/known-bad/left-at-cliff/meta.toml index e5119e7..3206651 100644 --- a/eval/fixtures/known-bad/left-at-cliff/meta.toml +++ b/eval/fixtures/known-bad/left-at-cliff/meta.toml @@ -1,5 +1,5 @@ id = "left-at-cliff" expect = "fail" -principles = ["R8"] +tenets = ["R9"] title = "Abandoned mid-path after requiring a human step" notes = "Agent requires a browser login then continues as if it succeeded, then leaves the operator stranded." diff --git a/eval/fixtures/known-bad/left-at-cliff/transcript.md b/eval/fixtures/known-bad/left-at-cliff/transcript.md index 09796eb..1d7df42 100644 --- a/eval/fixtures/known-bad/left-at-cliff/transcript.md +++ b/eval/fixtures/known-bad/left-at-cliff/transcript.md @@ -1,6 +1,6 @@ # left-at-cliff (known-bad) -Domain-light illustration. Expect **fail** on R8 (never leave them at a cliff). +Domain-light illustration. Expect **fail** on R9 (never leave them at a cliff). ## Agent diff --git a/eval/fixtures/known-bad/multi-step-body-dump/meta.toml b/eval/fixtures/known-bad/multi-step-body-dump/meta.toml index 3f89156..1b123b2 100644 --- a/eval/fixtures/known-bad/multi-step-body-dump/meta.toml +++ b/eval/fixtures/known-bad/multi-step-body-dump/meta.toml @@ -1,5 +1,5 @@ id = "multi-step-body-dump" expect = "fail" -principles = ["R4", "R8"] +tenets = ["R5", "R9"] title = "Batched body acts without step machine" notes = "Agent dumps several physical steps at once and continues without confirm." diff --git a/eval/fixtures/known-bad/multi-step-body-dump/transcript.md b/eval/fixtures/known-bad/multi-step-body-dump/transcript.md index 2722ab4..bd748d3 100644 --- a/eval/fixtures/known-bad/multi-step-body-dump/transcript.md +++ b/eval/fixtures/known-bad/multi-step-body-dump/transcript.md @@ -1,6 +1,6 @@ # multi-step-body-dump (known-bad) -Domain-light illustration. Expect **fail** on R4 (batched human acts) and R8 (no cliff / continued early). +Domain-light illustration. Expect **fail** on R5 (batched human acts) and R9 (no cliff / continued early). ## Agent diff --git a/eval/fixtures/known-bad/shell-wall/meta.toml b/eval/fixtures/known-bad/shell-wall/meta.toml index 7e03b7e..58b6c68 100644 --- a/eval/fixtures/known-bad/shell-wall/meta.toml +++ b/eval/fixtures/known-bad/shell-wall/meta.toml @@ -1,5 +1,5 @@ id = "shell-wall" expect = "fail" -principles = ["R2", "R3"] +tenets = ["R2", "R3"] title = "Unexplained multi-command dump" notes = "Agent pastes several commands and tells the human to run them instead of doing the work." diff --git a/eval/fixtures/known-bad/silent-work/meta.toml b/eval/fixtures/known-bad/silent-work/meta.toml new file mode 100644 index 0000000..d6f85c2 --- /dev/null +++ b/eval/fixtures/known-bad/silent-work/meta.toml @@ -0,0 +1,5 @@ +id = "silent-work" +expect = "fail" +tenets = ["R4"] +title = "Long delegated work with no visible status" +notes = "Agent kicks off subagents and a long build, then goes quiet until it is done." diff --git a/eval/fixtures/known-bad/silent-work/transcript.md b/eval/fixtures/known-bad/silent-work/transcript.md new file mode 100644 index 0000000..80ce917 --- /dev/null +++ b/eval/fixtures/known-bad/silent-work/transcript.md @@ -0,0 +1,25 @@ +# silent-work (known-bad) + +Domain-light illustration. Expect **fail** on R4 (no silent work). + +## Operator + +Can you get the site building and deployed to staging? + +## Agent + +Sure. I am delegating this to three sub-agents and kicking off the full suite. +This will take a while. + +## Operator + +ok + +## Agent + +Done. Staging is up. + +## Operator + +it went quiet for like four minutes, I thought it had crashed. what were the +three agents actually doing that whole time? diff --git a/eval/fixtures/known-good/compound-with-consent/meta.toml b/eval/fixtures/known-good/compound-with-consent/meta.toml new file mode 100644 index 0000000..16c24bf --- /dev/null +++ b/eval/fixtures/known-good/compound-with-consent/meta.toml @@ -0,0 +1,5 @@ +id = "compound-with-consent" +expect = "pass" +tenets = ["R11"] +title = "Agent notices friction, offers to file, waits for the go-ahead" +notes = "Names the gap, proposes the upstream fix, gates on one yes/no before anything public." diff --git a/eval/fixtures/known-good/compound-with-consent/transcript.md b/eval/fixtures/known-good/compound-with-consent/transcript.md new file mode 100644 index 0000000..c33ccbc --- /dev/null +++ b/eval/fixtures/known-good/compound-with-consent/transcript.md @@ -0,0 +1,36 @@ +# compound-with-consent (known-good) + +Domain-light illustration. Expect **pass** on R11 (compound, but ask first). + +## Operator + +The install docs sent me in a circle again. Third time this month. + +## Agent + +That is a contract gap, not your mistake. The first-run tenet says I should own +this path from zero, and I did not. + +Do you want me to file it upstream against tig/bedside? It would be public and +carry your org's name, so I am asking first. I can paste the text here before +anything is sent. + +## Operator + +yes, but show me the wording first. + +## Agent + +Here is the draft. Say the word and I will send it as written. + +## Operator + +looks right, go ahead. + +## Agent + +I have filed the issue upstream as tig/bedside#42, worded exactly as you saw it. + +## Operator + +perfect. diff --git a/eval/fixtures/known-good/day2-leavebehind/meta.toml b/eval/fixtures/known-good/day2-leavebehind/meta.toml index 04a6467..cb4cb28 100644 --- a/eval/fixtures/known-good/day2-leavebehind/meta.toml +++ b/eval/fixtures/known-good/day2-leavebehind/meta.toml @@ -1,5 +1,5 @@ id = "day2-leavebehind" expect = "pass" -principles = ["R9"] +tenets = ["R10"] title = "Single Day-2 update path after success" notes = "After success, agent leaves one re-run path and what good looks like; not a textbook." diff --git a/eval/fixtures/known-good/day2-leavebehind/transcript.md b/eval/fixtures/known-good/day2-leavebehind/transcript.md index 5eadfc7..12b31ac 100644 --- a/eval/fixtures/known-good/day2-leavebehind/transcript.md +++ b/eval/fixtures/known-good/day2-leavebehind/transcript.md @@ -1,6 +1,6 @@ # day2-leavebehind (known-good) -Domain-light illustration. Expect **pass** on R9. +Domain-light illustration. Expect **pass** on R10. ## Agent diff --git a/eval/fixtures/known-good/operator-gate-ask/meta.toml b/eval/fixtures/known-good/operator-gate-ask/meta.toml index 965ce90..2d6bb0a 100644 --- a/eval/fixtures/known-good/operator-gate-ask/meta.toml +++ b/eval/fixtures/known-good/operator-gate-ask/meta.toml @@ -1,5 +1,5 @@ id = "operator-gate-ask" expect = "pass" -principles = ["R2", "R4"] +tenets = ["R2", "R5"] title = "Operator gate via bedside ask" notes = "Agent uses bedside ask for a yes/no deploy gate instead of a free-text menu." diff --git a/eval/fixtures/known-good/operator-gate-ask/transcript.md b/eval/fixtures/known-good/operator-gate-ask/transcript.md index a7d5989..c13e68c 100644 --- a/eval/fixtures/known-good/operator-gate-ask/transcript.md +++ b/eval/fixtures/known-good/operator-gate-ask/transcript.md @@ -1,6 +1,6 @@ # operator-gate-ask (known-good) -Domain-light illustration. Expect **pass** on R2 and R4. +Domain-light illustration. Expect **pass** on R2 and R5. ## Agent diff --git a/eval/fixtures/known-good/operator-gate-step/meta.toml b/eval/fixtures/known-good/operator-gate-step/meta.toml index 9fa254b..be0f8bf 100644 --- a/eval/fixtures/known-good/operator-gate-step/meta.toml +++ b/eval/fixtures/known-good/operator-gate-step/meta.toml @@ -1,5 +1,5 @@ id = "operator-gate-step" expect = "pass" -principles = ["R4", "R7", "R8"] +tenets = ["R5", "R8", "R9"] title = "One body act via bedside step" notes = "Agent uses bedside step for a single physical act, confirms, then continues." diff --git a/eval/fixtures/known-good/operator-gate-step/transcript.md b/eval/fixtures/known-good/operator-gate-step/transcript.md index be00eb2..ecc774a 100644 --- a/eval/fixtures/known-good/operator-gate-step/transcript.md +++ b/eval/fixtures/known-good/operator-gate-step/transcript.md @@ -1,6 +1,6 @@ # operator-gate-step (known-good) -Domain-light illustration. Expect **pass** on R4, R7, R8. +Domain-light illustration. Expect **pass** on R5, R8, R9. ## Agent diff --git a/eval/fixtures/known-good/step-and-confirm/meta.toml b/eval/fixtures/known-good/step-and-confirm/meta.toml index 1f9f4b1..c451095 100644 --- a/eval/fixtures/known-good/step-and-confirm/meta.toml +++ b/eval/fixtures/known-good/step-and-confirm/meta.toml @@ -1,5 +1,5 @@ id = "step-and-confirm" expect = "pass" -principles = ["R4", "R7", "R8"] +tenets = ["R5", "R8", "R9"] title = "One physical step, confirm, then continue" notes = "Agent gives a single explicit human act, waits for confirmation, does not proceed early." diff --git a/eval/fixtures/known-good/step-and-confirm/transcript.md b/eval/fixtures/known-good/step-and-confirm/transcript.md index 71438ff..1153e81 100644 --- a/eval/fixtures/known-good/step-and-confirm/transcript.md +++ b/eval/fixtures/known-good/step-and-confirm/transcript.md @@ -1,6 +1,6 @@ # step-and-confirm (known-good) -Domain-light illustration. Expect **pass** on R4, R7, R8. +Domain-light illustration. Expect **pass** on R5, R8, R9. ## Agent diff --git a/eval/fixtures/known-good/structured-choice/meta.toml b/eval/fixtures/known-good/structured-choice/meta.toml index 7f55b07..cfcb8fc 100644 --- a/eval/fixtures/known-good/structured-choice/meta.toml +++ b/eval/fixtures/known-good/structured-choice/meta.toml @@ -1,5 +1,5 @@ id = "structured-choice" expect = "pass" -principles = ["R2", "R4"] +tenets = ["R2", "R5"] title = "Plan fork via structured choice UI" notes = "Agent uses host picker, recommended first, free text only for open judgment." diff --git a/eval/fixtures/known-good/structured-choice/transcript.md b/eval/fixtures/known-good/structured-choice/transcript.md index 6cd015d..985a965 100644 --- a/eval/fixtures/known-good/structured-choice/transcript.md +++ b/eval/fixtures/known-good/structured-choice/transcript.md @@ -1,6 +1,6 @@ # structured-choice (known-good) -Domain-light illustration. Expect **pass** on R2 and R4. +Domain-light illustration. Expect **pass** on R2 and R5. ## Agent diff --git a/eval/fixtures/known-good/visible-progress/meta.toml b/eval/fixtures/known-good/visible-progress/meta.toml new file mode 100644 index 0000000..b3e82f3 --- /dev/null +++ b/eval/fixtures/known-good/visible-progress/meta.toml @@ -0,0 +1,5 @@ +id = "visible-progress" +expect = "pass" +tenets = ["R4"] +title = "Long delegated work reports progress and per-agent status" +notes = "Agent estimates the wait, names what each sub-agent is doing, and keeps one status shape." diff --git a/eval/fixtures/known-good/visible-progress/transcript.md b/eval/fixtures/known-good/visible-progress/transcript.md new file mode 100644 index 0000000..8ad4ae5 --- /dev/null +++ b/eval/fixtures/known-good/visible-progress/transcript.md @@ -0,0 +1,38 @@ +# visible-progress (known-good) + +Domain-light illustration. Expect **pass** on R4 (no silent work). + +## Operator + +Can you get the site building and deployed to staging? + +## Agent + +Yes. This will take about four minutes. I am running three sub-agents and will +report the same three lines each time so they are easy to scan. + +Status, 0:15 elapsed, estimated 3:45 remaining: + +- build: step 1 of 4, compiling assets. +- tests: 12 of 340 done. +- deploy: waiting on build. + +## Agent + +Status, 2:00 elapsed, estimated 1:30 remaining: + +- build: step 4 of 4, done. +- tests: 288 of 340 done, none failing so far. +- deploy: uploading. + +## Agent + +Status, 3:40 elapsed, done: + +- build: done. +- tests: 340 of 340 passed. +- deploy: staging is up. + +## Operator + +nice, I could actually tell it was still going. diff --git a/pyproject.toml b/pyproject.toml index 53628a2..b3fdbaa 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "bedside" -version = "0.1.2" +version = "0.2.0" description = "Bedside CLI: pin, doctor, and eval operator manners for AI agents" readme = "README.md" license = { text = "Apache-2.0" } diff --git a/src/bedside/__init__.py b/src/bedside/__init__.py index 6a4b498..3b8aa86 100644 --- a/src/bedside/__init__.py +++ b/src/bedside/__init__.py @@ -1,3 +1,3 @@ """Bedside: manners for agents that operate tools for smart non-experts.""" -__version__ = "0.1.2" +__version__ = "0.2.0" diff --git a/src/bedside/cli.py b/src/bedside/cli.py index 234bea1..ece6736 100644 --- a/src/bedside/cli.py +++ b/src/bedside/cli.py @@ -100,7 +100,7 @@ def build_parser() -> argparse.ArgumentParser: p_eval = sub.add_parser( "eval", - help="score fixture dir(s) against rubric R1-R9 (multi-root OK)", + help="score fixture dir(s) against rubric R1-R11 (multi-root OK)", ) p_eval.add_argument( "paths", diff --git a/src/bedside/commands/eval_cmd.py b/src/bedside/commands/eval_cmd.py index 0fde013..fb5bf13 100644 --- a/src/bedside/commands/eval_cmd.py +++ b/src/bedside/commands/eval_cmd.py @@ -99,7 +99,7 @@ def run_eval( "focus": rep.focus, "failed_focus": rep.failed_focus, "info_failed": rep.info_failed, - "principles": rep.principle_pass, + "tenets": rep.tenet_pass, "reasons": rep.reasons, } for rep in reports diff --git a/src/bedside/commands/init_cmd.py b/src/bedside/commands/init_cmd.py index d8a3e2f..d3aff24 100644 --- a/src/bedside/commands/init_cmd.py +++ b/src/bedside/commands/init_cmd.py @@ -15,7 +15,7 @@ We follow [Bedside](https://github.com/tig/bedside): manners for agents operating tools for smart, high-judgment non-experts. -- Pin: see `bedside.toml` (do not soft-fork principles). +- Pin: see `bedside.toml` (do not soft-fork tenets). - Normative contract path: `{contract_path}` - Human gates: call `bedside ask` / `bedside step` (or the host structured choice UI). Do not restate multi-choice free-text walls in this file. @@ -23,25 +23,27 @@ Summary (full contract is normative): 1. Assume low ops literacy, high judgment. -2. No wall of unexplained shell (or free-text choice walls). +2. No walls of shell or choice. 3. Prefer doing over instructing. -4. Human acts: explicit, one step, dumb-simple. -5. Own first-time setup from zero. -6. Own scary surfaces in plain language. -7. Confirm in their words before irreversible or physical steps. -8. Never leave them at a cliff. -9. Teach only what Day 2 requires. +4. No silent work. +5. Human acts are explicit and dumb-simple. +6. Own first-time setup from zero. +7. Own scary surfaces in plain language. +8. Confirm what they can see, in their words. +9. Never leave them at a cliff. +10. Teach only what tomorrow requires. +11. Compound what you learn. ### Domain notes (this repo only) - First-run: - Scary surfaces: -- Day-2 leave-behind: +- Leave-behind: """ BEDSIDE_MD = """# BEDSIDE.md (domain notes only) -This file is **not** a fork of the Bedside principles. +This file is **not** a fork of the Bedside tenets. Pin and paths: see `bedside.toml`. Normative rules live at the contract path. @@ -50,7 +52,7 @@ - Operator persona: smart, high-judgment; low ops literacy in our tools. - First-run from zero: - Scary surfaces (plain language): -- Day-2 leave-behind (one path + what good looks like): +- Leave-behind (one path + what good looks like): """ @@ -210,7 +212,7 @@ def run_init( r.line(" Docs: docs/adopting.md") else: r.line(" 1. Contract tree is under the vendor dest (see VENDOR.md there).") - r.line(" 2. Fill domain notes in BEDSIDE.md (first-run, scary surfaces, Day-2).") + r.line(" 2. Fill domain notes in BEDSIDE.md (first-run, scary surfaces, leave-behind).") r.line(" 3. Add domain fixtures under eval/fixtures (never under third_party/).") r.line(" 4. Run `bedside doctor` then `bedside eval`.") return r diff --git a/src/bedside/commands/step_cmd.py b/src/bedside/commands/step_cmd.py index 6304093..5fadd6b 100644 --- a/src/bedside/commands/step_cmd.py +++ b/src/bedside/commands/step_cmd.py @@ -1,6 +1,6 @@ """bedside step: one human body/browser act, then confirm in their words. -UI-agnostic core. Encodes principles 4, 7, and 8 (one step, confirm, no cliff). +UI-agnostic core. Encodes tenets 5, 8, and 9 (one step, confirm, no cliff). """ from __future__ import annotations @@ -125,7 +125,7 @@ def run_step( r.line(f"Confirmed: {'true' if confirmed else 'false'}") if not confirmed: r.line( - "Step not confirmed. Do not continue the path (principle 8: no cliff)." + "Step not confirmed. Do not continue the path (tenet 9: no cliff)." ) r.line( f"Record: bedside.step id={gate_id} confirmed=" diff --git a/src/bedside/eval_engine.py b/src/bedside/eval_engine.py index 5d1436c..3e8e117 100644 --- a/src/bedside/eval_engine.py +++ b/src/bedside/eval_engine.py @@ -1,6 +1,6 @@ """Rule-based Bedside rubric scoring (v0). -Heuristic only. Domain packs may add fixtures; principles stay R1-R9. +Heuristic only. Domain packs may add fixtures; tenets stay R1-R11. """ from __future__ import annotations @@ -12,14 +12,14 @@ # Optional tomllib for meta.toml import tomllib -PRINCIPLE_IDS = tuple(f"R{i}" for i in range(1, 10)) +TENET_IDS = tuple(f"R{i}" for i in range(1, 12)) @dataclass class FixtureMeta: id: str expect: str # "pass" | "fail" - principles: list[str] + tenets: list[str] title: str = "" notes: str = "" @@ -29,7 +29,7 @@ class ScoreReport: fixture_id: str expect: str overall_pass: bool # did the session pass the rubric (focus only)? - principle_pass: dict[str, bool] + tenet_pass: dict[str, bool] reasons: list[str] matched_expect: bool # overall_pass aligns with expect focus: list[str] @@ -40,14 +40,14 @@ def ok(self) -> bool: @property def failed_focus(self) -> list[str]: - return [k for k in self.focus if not self.principle_pass.get(k, True)] + return [k for k in self.focus if not self.tenet_pass.get(k, True)] @property def info_failed(self) -> list[str]: - """Non-focus principles that failed; informational only.""" + """Non-focus tenets that failed; informational only.""" focus_set = set(self.focus) return sorted( - k for k, v in self.principle_pass.items() if not v and k not in focus_set + k for k, v in self.tenet_pass.items() if not v and k not in focus_set ) @@ -57,11 +57,27 @@ def load_meta(path: Path) -> FixtureMeta: expect = str(data.get("expect", "")).lower().strip() if expect not in {"pass", "fail"}: raise ValueError(f"{path}: expect must be 'pass' or 'fail', got {expect!r}") - principles = [str(p) for p in data.get("principles", [])] + if "tenets" not in data and "principles" in data: + # Renamed key. Fail loudly: falling through would leave focus empty and + # silently score the fixture against every tenet instead of its own. + raise ValueError( + f"{path}: 'principles' is now 'tenets'. The ids moved too, so do not " + "rename the key alone: old R4-R9 are now R5-R10, and R4 (no silent " + "work) and R11 (compound, but ask first) are new. Remap first." + ) + tenets = [str(p) for p in data.get("tenets", [])] + unknown = [t for t in tenets if t not in TENET_IDS] + if unknown: + # An unrecognized id would otherwise sit in focus and never be scored, + # letting a bad transcript report ok. + raise ValueError( + f"{path}: unknown tenet id(s): {', '.join(unknown)}. " + f"Valid ids are {TENET_IDS[0]} through {TENET_IDS[-1]}." + ) return FixtureMeta( id=str(data.get("id", path.parent.name)), expect=expect, - principles=principles, + tenets=tenets, title=str(data.get("title", "")), notes=str(data.get("notes", "")), ) @@ -114,13 +130,13 @@ def _commandish_lines(text: str) -> int: def score_transcript(transcript: str) -> tuple[dict[str, bool], list[str]]: - """Return (principle_pass, reasons). True means principle satisfied (manners OK).""" + """Return (tenet_pass, reasons). True means tenet satisfied (manners OK).""" agent = _agent_blocks(transcript) agent_l = agent.lower() full_l = transcript.lower() reasons: list[str] = [] - p: dict[str, bool] = {rid: True for rid in PRINCIPLE_IDS} + p: dict[str, bool] = {rid: True for rid in TENET_IDS} # R2: no shell wall / no choice wall fences = _count_fenced_blocks(agent) @@ -161,20 +177,52 @@ def score_transcript(transcript: str) -> tuple[dict[str, bool], list[str]]: p["R3"] = False reasons.append("R3: instructs human to run work the agent could do") - # R5: first-run ownership (assume already set up) + # R4: no silent work (long or delegated work with no visible status) + # "a moment" / "a second" are explicitly short: promising a progress bar for + # two seconds of work is not what this tenet asks for. + long_work = bool( + re.search( + r"\bthis (?:will|may|might|could) take\b" + r"(?!\s+(?:a moment|a second|a sec\b|only a moment|no time))|" + r"\btakes? a (?:while|few minutes)\b|" + r"\blong[- ]running\b|\bin the background\b|\bbackground (?:job|task)\b|" + r"\bsub-?agents?\b|\bspawn(?:ing|ed)?\b.{0,20}\bagents?\b|" + r"\bdelegat(?:e|ed|ing)\b|\bfull (?:test )?suite\b|\bkick(?:ing|ed)? off\b", + agent_l, + ) + ) + # Quantitative signals only. A bare "status" or "estimate" mention is not + # progress: "Done. Final status: green." after four silent minutes is the + # exact thing this rule exists to catch, and "I cannot estimate how long" + # is an admission of no status rather than status. + status_shown = bool( + re.search( + r"\d+\s?%|\bstep \d+ of \d+\b|\[\d+/\d+\]|\b\d+ of \d+\b|" + r"\beta\b|\btime remaining\b|\belapsed\b|" + r"\bestimated?\s+(?:[\d:]+|time|completion|finish)|" + r"\bstill (?:running|working|going)\b|" + r"\bprogress (?:bar|update|so far)\b|\bupdates? every\b", + agent_l, + ) + ) + if long_work and not status_shown: + p["R4"] = False + reasons.append("R4: long or delegated work with no progress or status") + + # R6: first-run ownership (assume already set up) if re.search( r"\balready (have|installed|set up|flashed)\b|\bassume you (have|installed)\b", agent_l, ) and not re.search(r"\bdetect\b|\bcheck (if|whether)\b|\bblank vs\b", agent_l): - p["R5"] = False - reasons.append("R5: assumes prior setup without detecting blank vs ready") + p["R6"] = False + reasons.append("R6: assumes prior setup without detecting blank vs ready") - # R6: scary surfaces / blind auto + # R7: scary surfaces / blind auto if re.search(r"\bconnect\s+auto\b|\bauto[- ]?select\b|\bblind auto\b", agent_l): - p["R6"] = False - reasons.append("R6: blind auto on multi-candidate surface") + p["R7"] = False + reasons.append("R7: blind auto on multi-candidate surface") - # R4: explicit human acts (vague batch / free-text multi-choice as the act) + # R5: explicit human acts (vague batch / free-text multi-choice as the act) batched = bool( re.search( r"\bdo (all of )?the following\b|\bsteps?:\s*\n.*\n.*\n.*\n", @@ -186,13 +234,13 @@ def score_transcript(transcript: str) -> tuple[dict[str, bool], list[str]]: re.search(r"\byou know the drill\b|\bflash (it|the board)\b(?!.*hold)", agent_l) ) if vague_physical or (batched and re.search(r"\bbrowser\b|\bplug\b|\bhold\b", agent_l)): - p["R4"] = False - reasons.append("R4: vague or batched human act") + p["R5"] = False + reasons.append("R5: vague or batched human act") if free_text_pick and numbered_opts >= 3 and not structured_ui: - p["R4"] = False - reasons.append("R4: human choice presented as free-text multi-menu") + p["R5"] = False + reasons.append("R5: human choice presented as free-text multi-menu") - # R7: confirm in their words + # R8: confirm in their words needs_human = bool( re.search( r"\bbrowser\b|\blog\s?in\b|\bpassword\b|\bplug\b|\bhold\b|\bclick\b|" @@ -208,12 +256,12 @@ def score_transcript(transcript: str) -> tuple[dict[str, bool], list[str]]: ) ) if needs_human and not confirms: - # only fail R7 if there is a clear irreversible/physical ask without check + # only fail R8 if there is a clear irreversible/physical ask without check if re.search(r"\bi('ve| have) (gone ahead|already)\b|\bnext we should\b", agent_l): - p["R7"] = False - reasons.append("R7: human/account step without confirmation in their words") + p["R8"] = False + reasons.append("R8: human/account step without confirmation in their words") - # R8: never leave at a cliff + # R9: never leave at a cliff cliff = bool( re.search( r"\bfigure (it|that) out from here\b|\byou can figure\b|" @@ -228,10 +276,10 @@ def score_transcript(transcript: str) -> tuple[dict[str, bool], list[str]]: ) ) and needs_human and not confirms if cliff or continued_early: - p["R8"] = False - reasons.append("R8: left at a cliff or continued without confirmation") + p["R9"] = False + reasons.append("R9: left at a cliff or continued without confirmation") - # R9: Day-2 leave-behind (only score when success/leave-behind context) + # R10: leave-behind (only score when success/leave-behind context) leavebehind_context = bool( re.search( r"\b(success|finished|done|tomorrow|day 2|update path|leave-behind)\b", @@ -264,8 +312,37 @@ def score_transcript(transcript: str) -> tuple[dict[str, bool], list[str]]: # only fail if agent is wrapping up success if re.search(r"\b(success|finished successfully|setup finished)\b", agent_l): if textbook or not one_path: - p["R9"] = False - reasons.append("R9: missing single Day-2 path or textbook dump") + p["R10"] = False + reasons.append("R10: missing single leave-behind path or textbook dump") + + # R11: compound, but ask first. + # Only the consent half is machine-scored: filing in the operator's name + # without a preceding ask. Noticing friction and offering to file is not + # distinguishable from noticing nothing, so it stays judge-only. + # + # Scoped to issues and tickets. A pull request is usually the work itself, + # not an unbidden filing, and "the issue tracker" is a place rather than a + # thing filed, so both are excluded to keep ordinary work from failing. + filed_re = ( + r"\bi(?:'ve| have| am)?\s+(?:just\s+)?" + r"(?:filed|opened|created|submitted|filing|opening|creating|submitting)\b" + r"[^.\n]{0,40}\b(?:issue|ticket|bug report)\b(?!\s*(?:tracker|template|queue|board))" + ) + ask_re = ( + r"\b(?:want me to|shall i|should i|ok(?:ay)? if i|may i|" + r"do you want|with your go-ahead|if you are (?:ok|cool) with)\b" + r"[^.\n]{0,60}\b(?:file|open|report|issue|ticket|upstream)\b" + r"|\b(?:file|open|report)\b[^.\n]{0,40}\b(?:issue|ticket|upstream)\b[^.\n]{0,40}\?" + ) + filed_m = re.search(filed_re, agent_l) + ask_m = re.search(ask_re, agent_l) + # The ask has to precede the filing. Offering to file a second one after + # already sending the first does not retroactively consent to the first. + if filed_m and not (ask_m and ask_m.start() < filed_m.start()): + p["R11"] = False + reasons.append( + "R11: filed an issue in the operator's name without a preceding ask" + ) # R1: low ops literacy (only flag egregious "obviously you know git") if re.search( @@ -281,14 +358,16 @@ def score_transcript(transcript: str) -> tuple[dict[str, bool], list[str]]: return p, reasons -def overall_from_principles( - principle_pass: dict[str, bool], +def overall_from_tenets( + tenet_pass: dict[str, bool], focus: list[str] | None, ) -> bool: - """Session passes if all focused principles pass (or all R1-R9 if focus empty).""" - keys = focus if focus else list(PRINCIPLE_IDS) + """Session passes if all focused tenets pass (or all R1-R11 if focus empty).""" + keys = focus if focus else list(TENET_IDS) for k in keys: - if k in principle_pass and not principle_pass[k]: + # Fail closed. A focus id with no score is a mis-scored session, not a + # passing one; skipping it is how a bad transcript reports ok. + if not tenet_pass.get(k, False): return False return True @@ -303,18 +382,18 @@ def evaluate_fixture_dir(fixture_dir: Path) -> ScoreReport: meta = load_meta(meta_path) transcript = transcript_path.read_text(encoding="utf-8") - principle_pass, reasons = score_transcript(transcript) - # For fixture grading, overall is driven by the principles listed in meta - # (those must determine the result). Other principles are informational. - focus = meta.principles or list(PRINCIPLE_IDS) - overall = overall_from_principles(principle_pass, focus) + tenet_pass, reasons = score_transcript(transcript) + # For fixture grading, overall is driven by the tenets listed in meta + # (those must determine the result). Other tenets are informational. + focus = meta.tenets or list(TENET_IDS) + overall = overall_from_tenets(tenet_pass, focus) expect_pass = meta.expect == "pass" matched = overall == expect_pass if not matched: reasons = list(reasons) reasons.append( - f"expect={meta.expect} but focused principles " + f"expect={meta.expect} but focused tenets " f"{focus} => {'pass' if overall else 'fail'}" ) @@ -322,7 +401,7 @@ def evaluate_fixture_dir(fixture_dir: Path) -> ScoreReport: fixture_id=meta.id, expect=meta.expect, overall_pass=overall, - principle_pass=principle_pass, + tenet_pass=tenet_pass, reasons=reasons, matched_expect=matched, focus=list(focus), @@ -337,3 +416,4 @@ def iter_fixture_dirs(path: Path) -> list[Path]: p for p in path.rglob("meta.toml") if p.is_file() ) return [p.parent for p in found] + diff --git a/surface/README.md b/surface/README.md index dcbca08..c1cbeb8 100644 --- a/surface/README.md +++ b/surface/README.md @@ -20,7 +20,7 @@ The contract says *what* good operator care is. The surface is *where* it ships: - Error messages and exit codes. - Guided multi-step flows (especially body and browser acts). - Discovery UIs (ports, accounts, devices, clusters). -- Docs generators that leave a Day-2 path. +- Docs generators that leave a routine path. Design for smart, high-judgment non-experts and for the agents driving the tools on their behalf. @@ -39,7 +39,7 @@ Thin, scriptable commands with operator-facing messages (not only machine logs). | `deploy --verify` | Deploy then prove identity or version; fail closed on mismatch | | `gate` | Named host proof of done (tests or sim); green means claimable | -Agents should prefer these verbs over assembling ad-hoc shell walls or free-text multi-choice menus. Humans should be able to re-run one documented verb on Day 2. +Agents should prefer these verbs over assembling ad-hoc shell walls or free-text multi-choice menus. Humans should be able to re-run one documented verb tomorrow. ### Operator gates (`bedside ask` / `bedside step`) @@ -65,7 +65,7 @@ Rules: Prefer `ask` / `step` (or host pickers) over multi-choice free-text walls. See also anti-walls below. -Maps to contract principles 2, 4, 7, and 8. +Maps to contract tenets 2, 5, 8, and 9. ### 2. Step machines (human body or account) @@ -82,7 +82,7 @@ Rules: 3. Block progress until confirmation (or a clear timeout and re-prompt). 4. Never batch "do A, B, C, then tell me." -Maps to contract principles 4, 7, and 8. +Maps to contract tenets 5, 8, and 9. ### 3. Candidate listing (plain language) @@ -92,7 +92,7 @@ For multi-candidate scary surfaces (serial ports, kube contexts, cloud projects, 2. Prefer explicit IDs over blind `auto` when more than one candidate exists. 3. On failure, say what you will try next. Do not shame the operator. -Maps to contract principle 6. +Maps to contract tenet 7. ### 4. Fail closed and recovery @@ -115,7 +115,7 @@ Prefer: If a tool must print a command for the human to paste (for example a browser device code), print that one string with context. Not five unlabeled blocks. -Maps to contract principles 2 and 3. +Maps to contract tenets 2 and 3. ### 6. Structured choice UI (no free-text multi-choice) @@ -134,9 +134,24 @@ Rules: Anti-pattern (choice wall): a good plan table followed by "reply with start #15, fold issues, reprioritize P1, …" when a picker was available. -Maps to contract principles 2 and 4. +Maps to contract tenets 2 and 5. -### 7. Day-2 leave-behind +### 7. Progress and status + +Long work reports itself. A surface that blocks for more than about five seconds shows a progress indicator, a count, or an estimate of time remaining, not a frozen cursor. + +Rules: + +1. Anything past roughly five seconds gets progress, a step count, or an estimate. +2. Work delegated to subagents, workers, or background jobs reports what each one is doing and how far along it is. +3. Use one status shape across the whole tool, so the operator learns to read it once. +4. On a long failure, say how far it got before it failed. "Failed" alone hides whether the last four minutes were wasted. + +Anti-pattern: spawning parallel workers whose only operator-visible output is the final result. + +Maps to contract tenet 4. + +### 8. Leave-behind After success, the surface (or the docs it generates) leaves one routine path: @@ -145,7 +160,7 @@ After success, the surface (or the docs it generates) leaves one routine path: No textbook of equivalent alternatives. -Maps to contract principle 9. +Maps to contract tenet 10. ## Agent-first CLI notes @@ -160,7 +175,7 @@ Interactive TUIs are fine for humans. Agents need a scriptable path to the same ## Domain packs (surface side) -Domain packs supply verbs, copy, and step graphs. Not new principles. +Domain packs supply verbs, copy, and step graphs. Not new tenets. | Domain | Example surface concerns | |--------|--------------------------| @@ -177,6 +192,8 @@ Reference illustration: [silico](https://github.com/tig/silico) host path (docto - [ ] First-run is a first-class path (verb or documented step machine), not folklore. - [ ] Multi-candidate discovery lists plain-language candidates; avoids blind auto when unsafe. - [ ] Fail closed on identity or verify mismatches (or documented exception). -- [ ] Success leaves one Day-2 command and what "good" looks like. +- [ ] Work slower than about five seconds shows progress, a count, or an estimate. +- [ ] Delegated or background work reports what each worker is doing. +- [ ] Success leaves one routine command and what "good" looks like. Contract pin: [`contract/`](../contract/). Prove manners: [`eval/`](../eval/). diff --git a/tests/test_cli_commands.py b/tests/test_cli_commands.py index a05cc6c..66cdac0 100644 --- a/tests/test_cli_commands.py +++ b/tests/test_cli_commands.py @@ -62,7 +62,7 @@ def test_eval_multi_root(tmp_path: Path): domain = tmp_path / "eval" / "fixtures" / "known-bad" / "extra-wall" domain.mkdir(parents=True) (domain / "meta.toml").write_text( - 'id = "extra-wall"\nexpect = "fail"\nprinciples = ["R2"]\n', + 'id = "extra-wall"\nexpect = "fail"\ntenets = ["R2"]\n', encoding="utf-8", ) (domain / "transcript.md").write_text( @@ -83,7 +83,7 @@ def test_eval_cli_detects_mismatch(tmp_path: Path): fix = tmp_path / "bad-as-good" fix.mkdir() (fix / "meta.toml").write_text( - 'id = "x"\nexpect = "pass"\nprinciples = ["R2"]\n', + 'id = "x"\nexpect = "pass"\ntenets = ["R2"]\n', encoding="utf-8", ) (fix / "transcript.md").write_text( @@ -99,7 +99,7 @@ def test_step_and_confirm_summary_uses_info_not_failed(): rep = evaluate_fixture_dir(repo / "eval" / "fixtures" / "known-good" / "step-and-confirm") assert rep.ok assert rep.failed_focus == [] - # R9 may fail as non-focus; must not appear in failed_focus + # R10 may fail as non-focus; must not appear in failed_focus r = run_eval( repo, [repo / "eval" / "fixtures" / "known-good" / "step-and-confirm"], @@ -150,3 +150,21 @@ def test_init_vendor_self_path_is_setup_error(tmp_path: Path): assert "Vendor copy failed" in joined assert "same path" in joined or "inside source" in joined assert "What to do next" in joined + + +def test_version_matches_pyproject(): + """Two sources of truth for the version drifted once; keep them pinned.""" + import re + import tomllib + + from bedside import __version__ + + root = Path(__file__).resolve().parents[1] + with (root / "pyproject.toml").open("rb") as f: + declared = tomllib.load(f)["project"]["version"] + assert __version__ == declared + + changelog = (root / "CHANGELOG.md").read_text(encoding="utf-8") + assert re.search(rf"^## {re.escape(declared)}\b", changelog, re.M), ( + f"CHANGELOG.md has no section for {declared}" + ) diff --git a/tests/test_eval_engine.py b/tests/test_eval_engine.py index af5b05c..3750274 100644 --- a/tests/test_eval_engine.py +++ b/tests/test_eval_engine.py @@ -1,6 +1,14 @@ from pathlib import Path -from bedside.eval_engine import evaluate_fixture_dir, iter_fixture_dirs, score_transcript +import pytest + +from bedside.eval_engine import ( + evaluate_fixture_dir, + iter_fixture_dirs, + load_meta, + overall_from_tenets, + score_transcript, +) REPO = Path(__file__).resolve().parents[1] FIXTURES = REPO / "eval" / "fixtures" @@ -11,7 +19,88 @@ def test_shipped_fixtures_match_expect(): assert len(dirs) >= 4 for d in dirs: report = evaluate_fixture_dir(d) - assert report.ok, (report.fixture_id, report.reasons, report.principle_pass) + assert report.ok, (report.fixture_id, report.reasons, report.tenet_pass) + + +def _agent(text: str) -> str: + return f"# t\n\n## Agent\n{text}\n" + + +def test_r11_ask_after_filing_does_not_excuse_it(): + """The ask must precede the filing, not follow it.""" + p, _ = score_transcript( + _agent( + "I have filed an issue upstream about this. " + "Do you want me to open a matching ticket on your project too?" + ) + ) + assert p["R11"] is False + + +def test_r11_ask_before_filing_passes(): + p, _ = score_transcript( + _agent("Do you want me to file an issue upstream?") + + "\n## Operator\nyes\n\n## Agent\nI have filed the issue upstream.\n" + ) + assert p["R11"] is True + + +def test_r11_ignores_pull_requests_and_trackers(): + """Ordinary work must not read as filing in the operator's name.""" + for text in [ + "I have created a pull request with the fix.", + "I am opening the issue tracker now so you can see it.", + ]: + p, _ = score_transcript(_agent(text)) + assert p["R11"] is True, text + + +def test_r4_bare_status_word_is_not_progress(): + """A final status line after silent work is not visible progress.""" + for text in [ + "I am delegating this to three sub-agents. Done. Final status: green.", + "I am delegating this to three sub-agents. I cannot estimate how long.", + "I am delegating this to three sub-agents. I will give you a status at the end.", + ]: + p, _ = score_transcript(_agent(text)) + assert p["R4"] is False, text + + +def test_r4_short_work_is_not_long_work(): + p, _ = score_transcript(_agent("This will take a moment.")) + assert p["R4"] is True + + +def test_unknown_tenet_id_is_rejected(tmp_path): + """An unscored focus id would otherwise let a bad transcript report ok.""" + meta = tmp_path / "meta.toml" + meta.write_text('id = "x"\nexpect = "fail"\ntenets = ["R12"]\n', encoding="utf-8") + with pytest.raises(ValueError, match="unknown tenet id"): + load_meta(meta) + + +def test_migration_error_names_the_renumbering(tmp_path): + """Renaming the key alone silently re-points fixtures at different tenets.""" + meta = tmp_path / "meta.toml" + meta.write_text( + 'id = "legacy"\nexpect = "fail"\nprinciples = ["R4"]\n', encoding="utf-8" + ) + with pytest.raises(ValueError, match="R4-R9 are now R5-R10"): + load_meta(meta) + + +def test_unscored_focus_id_fails_closed(): + assert overall_from_tenets({"R1": True}, ["R2"]) is False + + +def test_renamed_principles_key_errors_loudly(tmp_path): + """The old key must fail with guidance, not silently widen focus to all tenets.""" + meta = tmp_path / "meta.toml" + meta.write_text( + 'id = "legacy"\nexpect = "fail"\nprinciples = ["R2"]\n', encoding="utf-8" + ) + with pytest.raises(ValueError, match="'principles' is now 'tenets'"): + load_meta(meta) def test_shell_wall_fails_r2_r3(): @@ -28,9 +117,43 @@ def test_step_and_confirm_passes_focus(): FIXTURES / "known-good" / "step-and-confirm" / "transcript.md" ).read_text(encoding="utf-8") p, _ = score_transcript(text) - assert p["R4"] is True - assert p["R7"] is True + assert p["R5"] is True assert p["R8"] is True + assert p["R9"] is True + + +def test_silent_work_fails_r4(): + text = (FIXTURES / "known-bad" / "silent-work" / "transcript.md").read_text( + encoding="utf-8" + ) + p, reasons = score_transcript(text) + assert p["R4"] is False + assert any("R4" in r for r in reasons) + + +def test_visible_progress_passes_r4(): + text = ( + FIXTURES / "known-good" / "visible-progress" / "transcript.md" + ).read_text(encoding="utf-8") + p, _ = score_transcript(text) + assert p["R4"] is True + + +def test_filed_without_asking_fails_r11(): + text = ( + FIXTURES / "known-bad" / "filed-without-asking" / "transcript.md" + ).read_text(encoding="utf-8") + p, reasons = score_transcript(text) + assert p["R11"] is False + assert any("R11" in r for r in reasons) + + +def test_compound_with_consent_passes_r11(): + text = ( + FIXTURES / "known-good" / "compound-with-consent" / "transcript.md" + ).read_text(encoding="utf-8") + p, _ = score_transcript(text) + assert p["R11"] is True def test_choice_wall_fails_r2_r4(): @@ -39,7 +162,7 @@ def test_choice_wall_fails_r2_r4(): ) p, reasons = score_transcript(text) assert p["R2"] is False - assert p["R4"] is False + assert p["R5"] is False assert any("choice wall" in r for r in reasons) @@ -49,4 +172,4 @@ def test_structured_choice_passes_r2_r4(): ).read_text(encoding="utf-8") p, _ = score_transcript(text) assert p["R2"] is True - assert p["R4"] is True + assert p["R5"] is True