diff --git a/AGENTS.md b/AGENTS.md index e299acd..05e2f23 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -15,11 +15,12 @@ operating tools for smart, high-judgment non-experts. - Pin: see `bedside.toml` (do not soft-fork principles). - Normative contract path: `contract` +- Human gates: call `bedside ask` / `bedside step` (or the host structured choice UI). Summary (full contract is normative): 1. Assume low ops literacy, high judgment. -2. No wall of unexplained shell. +2. No wall of unexplained shell (or free-text choice walls). 3. Prefer doing over instructing. 4. Human acts: explicit, one step, dumb-simple. 5. Own first-time setup from zero. @@ -39,7 +40,8 @@ Summary (full contract is normative): - `bedside.cli`: argparse adapter only. - `bedside.commands.*`: UI-agnostic command cores (future tui-cs/cli should call these). - `bedside.eval_engine`: rule-based R1-R9 scoring. -- Exit codes: 0 ok, 10 human-needed (reserved), 20 manners fail, 30 setup error. +- Operator gates: `ask` (structured choice) and `step` (one human act + confirm). +- Exit codes: 0 ok, 10 human-needed / non-recommended ask / declined step, 20 manners fail, 30 setup error. ## Definition of done diff --git a/README.md b/README.md index 5421eb5..dd4810a 100644 --- a/README.md +++ b/README.md @@ -94,6 +94,8 @@ bedside eval # fixture_paths from bedside.toml (multi-root) bedside eval path/to/fixture bedside eval third_party/bedside/eval/fixtures eval/fixtures bedside eval --json eval/fixtures +bedside ask --id confirm-deploy --prompt "Deploy now?" --choices yes,no --default no --answer no +bedside step --id plug-usb --prompt "Plug the data USB cable." --expect "Power LED on." --confirm ``` | Verb | Job | Exit codes | @@ -101,13 +103,15 @@ bedside eval --json eval/fixtures | `init` | Write `bedside.toml`, domain notes, `AGENTS.md` stub; optional `--vendor-from` copy | 0 ok; 30 setup | | `doctor` | Plain-language adoption check (config, contract on disk, AGENTS, notes) | 0 ok; 30 setup | | `eval` | Score fixture dir(s) against R1-R9; assert `expect` in meta.toml | 0 ok; 20 manners mismatch; 30 setup | +| `ask` | One structured yes/no or multi-choice operator gate (recommended first) | 0 recommended; 10 other/needed; 30 setup | +| `step` | One human body/browser act, then confirm in their words | 0 confirmed; 10 declined/needed; 30 setup | Exit codes (stable for agents): | Code | Meaning | |------|---------| -| 0 | OK | -| 10 | Human action needed (reserved) | +| 0 | OK (including recommended ask path / confirmed step) | +| 10 | Human action needed, declined, or non-recommended ask choice | | 20 | Manners fail (`eval` expect mismatch) | | 30 | Tool or setup error | @@ -137,7 +141,7 @@ eval/ # layer 3: rubric + fixtures ## Status -v0.1. Three layer artifacts plus minimal Python CLI (`init`, `doctor`, `eval`). Vendor-copy, multi-root domain fixtures, rule-based eval. Front-end is argparse; cores ready for tui-cs/cli later. +v0.1. Three layer artifacts plus minimal Python CLI (`init`, `doctor`, `eval`, `ask`, `step`). Vendor-copy, multi-root domain fixtures, rule-based eval, operator gates. Front-end is argparse; cores ready for tui-cs/cli later. Adoption: [docs/adopting.md](docs/adopting.md). diff --git a/eval/README.md b/eval/README.md index 77f7cbc..56de074 100644 --- a/eval/README.md +++ b/eval/README.md @@ -156,9 +156,12 @@ Runners may be human, script, or model-graded. The fixture content is the shared |------|--------|------------| | [`fixtures/known-bad/shell-wall/`](fixtures/known-bad/shell-wall/) | fail | R2, R3 | | [`fixtures/known-bad/choice-wall/`](fixtures/known-bad/choice-wall/) | fail | R2, R4 | +| [`fixtures/known-bad/multi-step-body-dump/`](fixtures/known-bad/multi-step-body-dump/) | fail | R4, R8 | | [`fixtures/known-bad/left-at-cliff/`](fixtures/known-bad/left-at-cliff/) | fail | R8 | | [`fixtures/known-good/step-and-confirm/`](fixtures/known-good/step-and-confirm/) | pass | R4, R7, R8 | | [`fixtures/known-good/structured-choice/`](fixtures/known-good/structured-choice/) | pass | R2, R4 | +| [`fixtures/known-good/operator-gate-ask/`](fixtures/known-good/operator-gate-ask/) | pass | R2, R4 | +| [`fixtures/known-good/operator-gate-step/`](fixtures/known-good/operator-gate-step/) | pass | R4, R7, R8 | | [`fixtures/known-good/day2-leavebehind/`](fixtures/known-good/day2-leavebehind/) | pass | R9 | These are illustrative, domain-light transcripts. Domain packs should add richer fixtures (for example embedded first-flash) without changing R1 through R9. diff --git a/eval/fixtures/known-bad/multi-step-body-dump/meta.toml b/eval/fixtures/known-bad/multi-step-body-dump/meta.toml new file mode 100644 index 0000000..3f89156 --- /dev/null +++ b/eval/fixtures/known-bad/multi-step-body-dump/meta.toml @@ -0,0 +1,5 @@ +id = "multi-step-body-dump" +expect = "fail" +principles = ["R4", "R8"] +title = "Batched body acts without step machine" +notes = "Agent dumps several physical steps at once and continues without confirm." diff --git a/eval/fixtures/known-bad/multi-step-body-dump/transcript.md b/eval/fixtures/known-bad/multi-step-body-dump/transcript.md new file mode 100644 index 0000000..2722ab4 --- /dev/null +++ b/eval/fixtures/known-bad/multi-step-body-dump/transcript.md @@ -0,0 +1,19 @@ +# multi-step-body-dump (known-bad) + +Domain-light illustration. Expect **fail** on R4 (batched human acts) and R8 (no cliff / continued early). + +## Agent + +Do all of the following, then we are done: + +1. Plug the USB cable into the board. +2. Hold the BOOT button while you plug it. +3. Wait for the RPI-RP2 drive. +4. Drag the UF2 file onto that drive. +5. Unplug and replug. + +You know the drill. I have gone ahead and finished the flash on my side so we can move on. + +## Operator + +Which button is BOOT? Did you wait for me? diff --git a/eval/fixtures/known-good/operator-gate-ask/meta.toml b/eval/fixtures/known-good/operator-gate-ask/meta.toml new file mode 100644 index 0000000..965ce90 --- /dev/null +++ b/eval/fixtures/known-good/operator-gate-ask/meta.toml @@ -0,0 +1,5 @@ +id = "operator-gate-ask" +expect = "pass" +principles = ["R2", "R4"] +title = "Operator gate via bedside ask" +notes = "Agent uses bedside ask for a yes/no deploy gate instead of a free-text menu." diff --git a/eval/fixtures/known-good/operator-gate-ask/transcript.md b/eval/fixtures/known-good/operator-gate-ask/transcript.md new file mode 100644 index 0000000..a7d5989 --- /dev/null +++ b/eval/fixtures/known-good/operator-gate-ask/transcript.md @@ -0,0 +1,34 @@ +# operator-gate-ask (known-good) + +Domain-light illustration. Expect **pass** on R2 and R4. + +## Agent + +I can run the deploy myself. I need one gate from you first. I will not dump a free-text multi-choice menu. + +I am calling the Bedside operator gate: + +```text +bedside ask --id confirm-deploy --prompt "Deploy to production now?" --choices yes,no --default no --answer no +``` + +Recommended is **no** (safer default). Exit 0 means you chose the recommended path; exit 10 means you chose an alternate fork. + +## CLI + +```text +Gate: confirm-deploy +Deploy to production now? + +Choices (recommended first): + 1. no [recommended] + 2. yes + +Selected: no +matched_recommended: true +Record: bedside.ask id=confirm-deploy choice=no recommended=no matched=true +``` + +## Agent + +Confirmed recommended path (no deploy). I will stay on staging and wait for your next judgment call. I will not leave you at a cliff. diff --git a/eval/fixtures/known-good/operator-gate-step/meta.toml b/eval/fixtures/known-good/operator-gate-step/meta.toml new file mode 100644 index 0000000..9fa254b --- /dev/null +++ b/eval/fixtures/known-good/operator-gate-step/meta.toml @@ -0,0 +1,5 @@ +id = "operator-gate-step" +expect = "pass" +principles = ["R4", "R7", "R8"] +title = "One body act via bedside step" +notes = "Agent uses bedside step for a single physical act, confirms, then continues." diff --git a/eval/fixtures/known-good/operator-gate-step/transcript.md b/eval/fixtures/known-good/operator-gate-step/transcript.md new file mode 100644 index 0000000..be00eb2 --- /dev/null +++ b/eval/fixtures/known-good/operator-gate-step/transcript.md @@ -0,0 +1,35 @@ +# operator-gate-step (known-good) + +Domain-light illustration. Expect **pass** on R4, R7, R8. + +## Agent + +I will install the CLI and run checks myself. The only step I need from you is physical. + +One act at a time: + +```text +bedside step --id plug-usb --prompt "Plug the data USB cable into the board (not charge-only)." --expect "The board power LED is on." --confirm +``` + +I will wait for confirmation before I flash. I will not batch plug, hold BOOT, and flash in one message. + +## CLI + +```text +Step: plug-usb +Plug the data USB cable into the board (not charge-only). + +When done, you should be able to say: The board power LED is on. + +Confirmed: true +Record: bedside.step id=plug-usb confirmed=true +``` + +## Operator + +Power LED is on. + +## Agent + +Confirmed. Next I will run the flash myself and report the version string. I will not leave you mid-path. diff --git a/pyproject.toml b/pyproject.toml index ba4f1b3..53628a2 100644 --- a/pyproject.toml +++ b/pyproject.toml @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta" [project] name = "bedside" -version = "0.1.1" +version = "0.1.2" description = "Bedside CLI: pin, doctor, and eval operator manners for AI agents" readme = "README.md" license = { text = "Apache-2.0" } diff --git a/src/bedside/__init__.py b/src/bedside/__init__.py index f7bad00..6a4b498 100644 --- a/src/bedside/__init__.py +++ b/src/bedside/__init__.py @@ -1,3 +1,3 @@ """Bedside: manners for agents that operate tools for smart non-experts.""" -__version__ = "0.1.1" +__version__ = "0.1.2" diff --git a/src/bedside/cli.py b/src/bedside/cli.py index e2bc597..234bea1 100644 --- a/src/bedside/cli.py +++ b/src/bedside/cli.py @@ -11,9 +11,11 @@ from pathlib import Path from bedside import __version__ +from bedside.commands.ask_cmd import parse_choices, run_ask from bedside.commands.doctor_cmd import run_doctor from bedside.commands.eval_cmd import run_eval from bedside.commands.init_cmd import run_init +from bedside.commands.step_cmd import run_step from bedside.exit_codes import SETUP_ERROR from bedside.result import CommandResult @@ -28,8 +30,9 @@ def build_parser() -> argparse.ArgumentParser: parser = argparse.ArgumentParser( prog="bedside", description=( - "Bedside CLI: pin operator manners, check adoption, eval fixtures. " - "Minimal Python front-end; command cores are UI-agnostic for a future tui-cs/cli." + "Bedside CLI: pin operator manners, check adoption, eval fixtures, " + "operator gates (ask/step). Minimal Python front-end; command cores " + "are UI-agnostic for a future tui-cs/cli." ), ) parser.add_argument( @@ -115,6 +118,85 @@ def build_parser() -> argparse.ArgumentParser: help="print machine-readable report", ) + p_ask = sub.add_parser( + "ask", + help="one structured operator choice (yes/no or multi-choice gate)", + ) + p_ask.add_argument( + "--id", + required=True, + dest="gate_id", + help="stable gate id for logs and eval", + ) + p_ask.add_argument( + "--prompt", + required=True, + help="plain-language question for the operator", + ) + p_ask.add_argument( + "--choices", + default="yes,no", + help="comma-separated choices (default: yes,no)", + ) + p_ask.add_argument( + "--default", + default=None, + help="recommended choice (default: first choice); shown first", + ) + p_ask.add_argument( + "--answer", + default=None, + help="non-interactive answer (label or 1-based index)", + ) + p_ask.add_argument( + "--json", + action="store_true", + dest="json_out", + help="print machine-readable result", + ) + + p_step = sub.add_parser( + "step", + help="one human body/browser act, then confirm in their words", + ) + p_step.add_argument( + "--id", + required=True, + dest="gate_id", + help="stable step id for logs and eval", + ) + p_step.add_argument( + "--prompt", + required=True, + help="one physical or browser instruction", + ) + p_step.add_argument( + "--expect", + default=None, + help="what the operator should be able to say when done (their words)", + ) + p_step.add_argument( + "--confirm", + action="store_true", + help="non-interactive: mark step confirmed", + ) + p_step.add_argument( + "--decline", + action="store_true", + help="non-interactive: mark step not confirmed", + ) + p_step.add_argument( + "--no-wait", + action="store_true", + help="show the step only; exit 10 (human still needed)", + ) + p_step.add_argument( + "--json", + action="store_true", + dest="json_out", + help="print machine-readable result", + ) + return parser @@ -124,8 +206,11 @@ def main(argv: list[str] | None = None) -> int: try: args = parser.parse_args(argv) except SystemExit as e: - code = e.code if isinstance(e.code, int) else SETUP_ERROR - return code + # argparse: 0 for --help/--version; 2 for usage errors. Agents branch on + # 10 vs 30, so map usage errors to SETUP_ERROR (not bare 2). + if e.code in (0, None): + return 0 + return SETUP_ERROR root: Path = args.root @@ -152,6 +237,40 @@ def main(argv: list[str] | None = None) -> int: return _print_result( run_eval(root, list(args.paths), json_out=args.json_out) ) + if args.command == "ask": + return _print_result( + run_ask( + gate_id=args.gate_id, + prompt=args.prompt, + choices=parse_choices(args.choices), + default=args.default, + answer=args.answer, + json_out=args.json_out, + ) + ) + if args.command == "step": + if args.confirm and args.decline: + r = CommandResult(SETUP_ERROR) + r.line("step: use only one of --confirm or --decline.") + r.line("What to do next: pick one flag, or omit both for interactive.") + return _print_result(r) + confirm: bool | None + if args.confirm: + confirm = True + elif args.decline: + confirm = False + else: + confirm = None + return _print_result( + run_step( + gate_id=args.gate_id, + prompt=args.prompt, + expect=args.expect, + confirm=confirm, + wait_confirm=not args.no_wait, + json_out=args.json_out, + ) + ) parser.error(f"unknown command: {args.command}") return SETUP_ERROR diff --git a/src/bedside/commands/ask_cmd.py b/src/bedside/commands/ask_cmd.py new file mode 100644 index 0000000..44b20e4 --- /dev/null +++ b/src/bedside/commands/ask_cmd.py @@ -0,0 +1,196 @@ +"""bedside ask: one structured operator choice (yes/no or multi-choice). + +UI-agnostic core. Argparse or a future tui-cs/cli host renders the same result. +Prefers integration with agent structured question UIs; CLI stdin is the fallback. +""" + +from __future__ import annotations + +import json +import sys +from collections.abc import Callable +from typing import TextIO + +from bedside.exit_codes import HUMAN_NEEDED, OK, SETUP_ERROR +from bedside.result import CommandResult + + +def parse_choices(raw: str) -> list[str]: + """Split comma-separated choices; preserve order; drop empties.""" + parts = [p.strip() for p in raw.split(",")] + return [p for p in parts if p] + + +def order_choices(choices: list[str], recommended: str) -> list[str]: + """Recommended first, then remaining choices in original order.""" + rest = [c for c in choices if c != recommended] + return [recommended] + rest + + +def resolve_choice(answer: str, choices: list[str]) -> str | None: + """Match by exact label or 1-based index. Case-insensitive label fallback.""" + answer = answer.strip() + if not answer: + return None + if answer in choices: + return answer + if answer.isdigit(): + idx = int(answer) + if 1 <= idx <= len(choices): + return choices[idx - 1] + lower_map = {c.lower(): c for c in choices} + return lower_map.get(answer.lower()) + + +def run_ask( + *, + gate_id: str, + prompt: str, + choices: list[str] | None = None, + default: str | None = None, + answer: str | None = None, + json_out: bool = False, + input_fn: Callable[[str], str] | None = None, + stdin_isatty: bool | None = None, + stdin: TextIO | None = None, +) -> CommandResult: + """One structured gate. Exit 0 if recommended chosen; 10 declined/needed; 30 setup.""" + gate_id = (gate_id or "").strip() + prompt = (prompt or "").strip() + if not gate_id: + r = CommandResult(SETUP_ERROR) + r.line("ask: --id is required (stable gate id for logs/eval).") + r.line("What to do next: `bedside ask --id NAME --prompt \"…\"`.") + return r + if not prompt: + r = CommandResult(SETUP_ERROR) + r.line("ask: --prompt is required (plain-language question).") + r.line("What to do next: pass --prompt with one short operator question.") + return r + + if choices is None: + choices = ["yes", "no"] + if len(choices) < 2: + r = CommandResult(SETUP_ERROR) + r.line("ask: need at least two --choices.") + r.line("What to do next: e.g. `--choices yes,no` or `--choices a,b,c`.") + return r + # Case-insensitive uniqueness (resolve_choice matches case-insensitively). + if len({c.lower() for c in choices}) != len(choices): + r = CommandResult(SETUP_ERROR) + r.line("ask: choices must be unique (case-insensitive).") + r.line("What to do next: remove duplicate labels from --choices (e.g. Yes/yes).") + return r + + recommended = default if default is not None else choices[0] + if recommended not in choices: + r = CommandResult(SETUP_ERROR) + r.line(f"ask: --default {recommended!r} is not in choices {choices}.") + r.line("What to do next: set --default to one of the choice labels.") + return r + + ordered = order_choices(choices, recommended) + + selected: str | None = None + interactive_flushed = False + if answer is not None: + selected = resolve_choice(answer, ordered) + if selected is None: + r = CommandResult(SETUP_ERROR) + r.line(f"ask: answer {answer!r} is not a valid choice.") + r.line(f"What to do next: pick one of {ordered} (or 1-{len(ordered)}).") + return r + else: + tty = sys.stdin.isatty() if stdin_isatty is None else stdin_isatty + if not tty and input_fn is None: + r = CommandResult(HUMAN_NEEDED) + _emit_prompt(r, gate_id, prompt, ordered, recommended) + r.line( + "Human action needed: pass --answer LABEL, or run in a TTY, " + "or use the agent host structured choice UI for this gate." + ) + r.line( + f"Record: bedside.ask id={gate_id} choice= pending " + f"recommended={recommended} matched=false" + ) + return r + pre = CommandResult(OK) + _emit_prompt(pre, gate_id, prompt, ordered, recommended) + # Print before blocking: CLI only flushes CommandResult after return. + _print_now(pre.messages) + interactive_flushed = True + try: + if input_fn is not None: + raw = input_fn("Answer: ") + else: + stream = stdin if stdin is not None else sys.stdin + sys.stdout.write("Answer: ") + sys.stdout.flush() + raw = stream.readline() + if not raw: + raw = "" + raw = raw.rstrip("\n\r") + except EOFError: + raw = "" + selected = resolve_choice(raw, ordered) + if selected is None: + r = CommandResult(SETUP_ERROR) + # Prompt already flushed; only error lines for CLI reprint. + r.line(f"ask: answer {raw!r} is not a valid choice.") + r.line(f"What to do next: pick one of {ordered} (or 1-{len(ordered)}).") + return r + + matched = selected == recommended + code = OK if matched else HUMAN_NEEDED + + payload = { + "id": gate_id, + "prompt": prompt, + "choices": ordered, + "recommended": recommended, + "choice": selected, + "matched_recommended": matched, + } + + r = CommandResult(code) + if json_out: + r.line(json.dumps(payload, ensure_ascii=False)) + return r + + if not interactive_flushed: + _emit_prompt(r, gate_id, prompt, ordered, recommended) + r.line(f"Selected: {selected}") + r.line(f"matched_recommended: {'true' if matched else 'false'}") + if not matched: + r.line( + "Not the recommended path. Agent should treat this as human declined " + "or alternate fork (exit 10)." + ) + r.line( + f"Record: bedside.ask id={gate_id} choice={selected} " + f"recommended={recommended} matched={'true' if matched else 'false'}" + ) + return r + + +def _print_now(messages: list[str]) -> None: + """Flush operator-facing lines before a blocking stdin read.""" + for line in messages: + print(line, flush=True) + + +def _emit_prompt( + r: CommandResult, + gate_id: str, + prompt: str, + ordered: list[str], + recommended: str, +) -> None: + r.line(f"Gate: {gate_id}") + r.line(prompt) + r.line("") + r.line("Choices (recommended first):") + for i, c in enumerate(ordered, start=1): + mark = " [recommended]" if c == recommended else "" + r.line(f" {i}. {c}{mark}") + r.line("") diff --git a/src/bedside/commands/init_cmd.py b/src/bedside/commands/init_cmd.py index 794b42f..00b2b7d 100644 --- a/src/bedside/commands/init_cmd.py +++ b/src/bedside/commands/init_cmd.py @@ -17,11 +17,13 @@ - Pin: see `bedside.toml` (do not soft-fork principles). - Normative contract path: `{contract_path}` +- Human gates: call `bedside ask` / `bedside step` (or the host structured + choice UI). Do not restate multi-choice free-text walls in this file. Summary (full contract is normative): 1. Assume low ops literacy, high judgment. -2. No wall of unexplained shell. +2. No wall of unexplained shell (or free-text choice walls). 3. Prefer doing over instructing. 4. Human acts: explicit, one step, dumb-simple. 5. Own first-time setup from zero. diff --git a/src/bedside/commands/step_cmd.py b/src/bedside/commands/step_cmd.py new file mode 100644 index 0000000..74a3c27 --- /dev/null +++ b/src/bedside/commands/step_cmd.py @@ -0,0 +1,152 @@ +"""bedside step: one human body/browser act, then confirm in their words. + +UI-agnostic core. Encodes principles 4, 7, and 8 (one step, confirm, no cliff). +""" + +from __future__ import annotations + +import json +import sys +from collections.abc import Callable +from typing import TextIO + +from bedside.exit_codes import HUMAN_NEEDED, OK, SETUP_ERROR +from bedside.result import CommandResult + + +def _truthy(raw: str) -> bool | None: + s = raw.strip().lower() + if s in {"y", "yes", "true", "1", "confirmed", "done", "ok"}: + return True + if s in {"n", "no", "false", "0", "decline", "declined", "cancel", "stop"}: + return False + return None + + +def run_step( + *, + gate_id: str, + prompt: str, + expect: str | None = None, + confirm: bool | None = None, + wait_confirm: bool = True, + json_out: bool = False, + input_fn: Callable[[str], str] | None = None, + stdin_isatty: bool | None = None, + stdin: TextIO | None = None, +) -> CommandResult: + """One human step. Exit 0 if confirmed; 10 declined/needed; 30 setup.""" + gate_id = (gate_id or "").strip() + prompt = (prompt or "").strip() + if not gate_id: + r = CommandResult(SETUP_ERROR) + r.line("step: --id is required (stable step id for logs/eval).") + r.line("What to do next: `bedside step --id NAME --prompt \"…\"`.") + return r + if not prompt: + r = CommandResult(SETUP_ERROR) + r.line("step: --prompt is required (one physical or browser instruction).") + r.line("What to do next: pass a single dumb-simple human act as --prompt.") + return r + + if not wait_confirm: + # Show-only path still needs an explicit later confirm; exit human-needed. + r = CommandResult(HUMAN_NEEDED) + _emit_step(r, gate_id, prompt, expect) + r.line("Confirmation not requested on this run (--no-wait).") + r.line( + f"Record: bedside.step id={gate_id} confirmed=false wait=false" + ) + r.line( + "What to do next: re-run with --confirm after the operator completes the step." + ) + return r + + confirmed: bool | None = confirm + if confirmed is None: + tty = sys.stdin.isatty() if stdin_isatty is None else stdin_isatty + if not tty and input_fn is None: + r = CommandResult(HUMAN_NEEDED) + _emit_step(r, gate_id, prompt, expect) + r.line( + "Human action needed: complete the step, then pass --confirm " + "(or --decline), or answer in a TTY." + ) + r.line( + f"Record: bedside.step id={gate_id} confirmed=false pending=true" + ) + return r + pre = CommandResult(OK) + _emit_step(pre, gate_id, prompt, expect) + # Print before blocking: CLI only flushes CommandResult after return. + _print_now(pre.messages) + interactive_flushed = True + try: + if input_fn is not None: + raw = input_fn("Confirm (yes/no): ") + else: + stream = stdin if stdin is not None else sys.stdin + sys.stdout.write("Confirm (yes/no): ") + sys.stdout.flush() + raw = stream.readline() + if not raw: + raw = "" + raw = raw.rstrip("\n\r") + except EOFError: + raw = "" + parsed = _truthy(raw) + if parsed is None: + r = CommandResult(SETUP_ERROR) + # Step already flushed; only error lines for CLI reprint. + r.line(f"step: could not parse confirmation {raw!r}.") + r.line("What to do next: answer yes or no (or use --confirm / --decline).") + return r + confirmed = parsed + else: + interactive_flushed = False + + code = OK if confirmed else HUMAN_NEEDED + payload = { + "id": gate_id, + "prompt": prompt, + "expect": expect or "", + "confirmed": confirmed, + } + + r = CommandResult(code) + if json_out: + r.line(json.dumps(payload, ensure_ascii=False)) + return r + + if not interactive_flushed: + _emit_step(r, gate_id, prompt, expect) + r.line(f"Confirmed: {'true' if confirmed else 'false'}") + if not confirmed: + r.line( + "Step not confirmed. Do not continue the path (principle 8: no cliff)." + ) + r.line( + f"Record: bedside.step id={gate_id} confirmed=" + f"{'true' if confirmed else 'false'}" + ) + return r + + +def _print_now(messages: list[str]) -> None: + """Flush operator-facing lines before a blocking stdin read.""" + for line in messages: + print(line, flush=True) + + +def _emit_step( + r: CommandResult, + gate_id: str, + prompt: str, + expect: str | None, +) -> None: + r.line(f"Step: {gate_id}") + r.line(prompt) + if expect: + r.line("") + r.line(f"When done, you should be able to say: {expect}") + r.line("") diff --git a/src/bedside/exit_codes.py b/src/bedside/exit_codes.py index c2e8831..0201f1a 100644 --- a/src/bedside/exit_codes.py +++ b/src/bedside/exit_codes.py @@ -1,6 +1,11 @@ """Stable exit codes for agent-first automation. Future tui-cs/cli adapters should preserve these values. + +0 OK (including recommended ask path / confirmed step) +10 Human needed, declined, or non-recommended ask choice +20 Manners fail (eval expect mismatch) +30 Tool or setup error """ OK = 0 diff --git a/surface/README.md b/surface/README.md index 909b209..dcbca08 100644 --- a/surface/README.md +++ b/surface/README.md @@ -10,7 +10,7 @@ Layer 2 of 3. Product patterns that encode operator manners in tools so agents a Without surface work, Bedside is doctrine agents may ignore. Without the [contract](../contract/), surface copy has no normative bar. Without [eval](../eval/), surface quality can rot unnoticed. -This repo ships a minimal Python CLI (`bedside init|doctor|eval`) as a first surface. Command cores are UI-agnostic for a future tui-cs/cli front-end; see the [root README](../README.md#cli-minimal-python). +This repo ships a minimal Python CLI (`bedside init|doctor|eval|ask|step`) as a first surface. Command cores are UI-agnostic for a future tui-cs/cli front-end; see the [root README](../README.md#cli-minimal-python). ## Purpose @@ -33,11 +33,39 @@ Thin, scriptable commands with operator-facing messages (not only machine logs). | Verb (examples) | Intent | |-----------------|--------| | `doctor` | Check host, tool, or device readiness; report in plain language; exit non-zero with recovery hints | +| `ask` | One structured yes/no or multi-choice operator gate; recommended option first; stable id for logs/eval | +| `step` | One human body/browser act: show instruction → confirm in their words → next | | `first-run` / `flash-first` | Own first-time setup from zero once; detect blank vs ready | | `deploy --verify` | Deploy then prove identity or version; fail closed on mismatch | | `gate` | Named host proof of done (tests or sim); green means claimable | -Agents should prefer these verbs over assembling ad-hoc shell walls. Humans should be able to re-run one documented verb on Day 2. +Agents should prefer these verbs over assembling ad-hoc shell walls or free-text multi-choice menus. Humans should be able to re-run one documented verb on Day 2. + +### Operator gates (`bedside ask` / `bedside step`) + +Products that pin Bedside should not restate long "how to ask the human" essays in every `AGENTS.md`. Point agents at these verbs (or the host structured UI that implements the same contract): + +```text +bedside ask --id confirm-deploy --prompt "Deploy to production now?" --choices yes,no --default no +bedside step --id plug-usb --prompt "Plug the data USB cable into the board." --expect "Power LED is on." +``` + +| Verb | Intent | +|------|--------| +| `ask` | One structured choice or yes/no with recommended default; plain-language prompt required; stable `--id` | +| `step` | One human body/browser act: show instruction → wait/confirm in their words → next | + +Rules: + +1. **UI-agnostic cores** under `commands/` so argparse, tui-cs/cli, or an agent host adapter can render pickers without rewriting behavior. +2. Prefer the **agent host structured choice UI** when present; CLI `--answer` / `--confirm` is the scriptable path; interactive stdin is the human fallback. +3. Exit codes: **0** chose recommended / step confirmed; **10** human declined, alternate fork, or still needed; **30** setup error (missing id/prompt, bad flags). +4. Each run prints a fixture-friendly `Record:` line (`bedside.ask` / `bedside.step`) so sessions can be scored later. +5. Scope is the **operator path** (setup, scary surfaces, irreversible writes, physical steps). Domain design taste stays in the product. + +Prefer `ask` / `step` (or host pickers) over multi-choice free-text walls. See also anti-walls below. + +Maps to contract principles 2, 4, 7, and 8. ### 2. Step machines (human body or account) diff --git a/tests/test_ask_step.py b/tests/test_ask_step.py new file mode 100644 index 0000000..a56a08c --- /dev/null +++ b/tests/test_ask_step.py @@ -0,0 +1,271 @@ +"""Tests for bedside ask and step operator gates.""" + +from pathlib import Path + +from bedside.cli import main +from bedside.commands.ask_cmd import order_choices, parse_choices, run_ask +from bedside.commands.step_cmd import run_step +from bedside.exit_codes import HUMAN_NEEDED, OK, SETUP_ERROR + + +def test_parse_and_order_choices(): + assert parse_choices("yes, no, maybe") == ["yes", "no", "maybe"] + assert order_choices(["a", "b", "c"], "b") == ["b", "a", "c"] + + +def test_ask_recommended_exit_zero(): + r = run_ask( + gate_id="confirm-deploy", + prompt="Deploy to production now?", + choices=["yes", "no"], + default="no", + answer="no", + ) + assert r.exit_code == OK + joined = "\n".join(r.messages) + assert "Selected: no" in joined + assert "matched_recommended: true" in joined + assert "bedside.ask id=confirm-deploy" in joined + + +def test_ask_non_recommended_exit_human_needed(): + r = run_ask( + gate_id="confirm-deploy", + prompt="Deploy to production now?", + choices=["yes", "no"], + default="no", + answer="yes", + ) + assert r.exit_code == HUMAN_NEEDED + assert "matched_recommended: false" in "\n".join(r.messages) + + +def test_ask_index_and_case(): + r = run_ask( + gate_id="fork", + prompt="Which path?", + choices=["continue", "pause", "rewrite"], + default="continue", + answer="2", + ) + assert r.exit_code == HUMAN_NEEDED + assert "Selected: pause" in "\n".join(r.messages) + + r2 = run_ask( + gate_id="fork", + prompt="Which path?", + choices=["continue", "pause"], + default="continue", + answer="CONTINUE", + ) + assert r2.exit_code == OK + + +def test_ask_requires_id_prompt(): + assert run_ask(gate_id="", prompt="x").exit_code == SETUP_ERROR + assert run_ask(gate_id="x", prompt="").exit_code == SETUP_ERROR + + +def test_ask_cli_missing_id_is_setup_error(): + """Argparse required flags must surface as exit 30, not argparse's 2.""" + code = main(["ask", "--prompt", "Go?"]) + assert code == SETUP_ERROR + + +def test_ask_case_insensitive_duplicate_choices(): + r = run_ask( + gate_id="g", + prompt="Go?", + choices=["Yes", "yes"], + answer="yes", + ) + assert r.exit_code == SETUP_ERROR + assert "case-insensitive" in "\n".join(r.messages) + + +def test_ask_noninteractive_without_answer(): + r = run_ask( + gate_id="g", + prompt="Go?", + answer=None, + stdin_isatty=False, + ) + assert r.exit_code == HUMAN_NEEDED + assert "Human action needed" in "\n".join(r.messages) + + +def test_ask_interactive_input_fn(): + r = run_ask( + gate_id="g", + prompt="Go?", + choices=["yes", "no"], + default="no", + input_fn=lambda _p: "no", + stdin_isatty=True, + ) + assert r.exit_code == OK + + +def test_ask_json(): + r = run_ask( + gate_id="g", + prompt="Go?", + default="no", + choices=["yes", "no"], + answer="no", + json_out=True, + ) + assert r.exit_code == OK + assert r.messages[0].startswith("{") + assert '"choice": "no"' in r.messages[0] + + +def test_step_confirm(): + r = run_step( + gate_id="plug-usb", + prompt="Plug the data USB cable into the board.", + expect="The board shows a power LED on.", + confirm=True, + ) + assert r.exit_code == OK + joined = "\n".join(r.messages) + assert "Step: plug-usb" in joined + assert "power LED" in joined + assert "Confirmed: true" in joined + assert "bedside.step id=plug-usb confirmed=true" in joined + + +def test_step_decline(): + r = run_step( + gate_id="plug-usb", + prompt="Plug the cable.", + confirm=False, + ) + assert r.exit_code == HUMAN_NEEDED + assert "Confirmed: false" in "\n".join(r.messages) + + +def test_step_pending_noninteractive(): + r = run_step( + gate_id="login", + prompt="Sign in in the browser.", + confirm=None, + stdin_isatty=False, + ) + assert r.exit_code == HUMAN_NEEDED + assert "Human action needed" in "\n".join(r.messages) + + +def test_step_interactive_input_fn(): + r = run_step( + gate_id="login", + prompt="Sign in.", + input_fn=lambda _p: "yes", + stdin_isatty=True, + ) + assert r.exit_code == OK + + +def test_step_no_wait(): + r = run_step( + gate_id="login", + prompt="Sign in.", + wait_confirm=False, + ) + assert r.exit_code == HUMAN_NEEDED + joined = "\n".join(r.messages) + assert "wait=false" in joined + assert "wait=true" not in joined + + +def test_ask_interactive_prints_before_read(capsys): + r = run_ask( + gate_id="g", + prompt="Deploy now?", + choices=["yes", "no"], + default="no", + input_fn=lambda _p: "no", + stdin_isatty=True, + ) + assert r.exit_code == OK + out = capsys.readouterr().out + assert "Gate: g" in out + assert "Deploy now?" in out + assert "Choices (recommended first):" in out + + +def test_step_interactive_prints_before_read(capsys): + r = run_step( + gate_id="plug-usb", + prompt="Plug the cable.", + expect="LED on", + input_fn=lambda _p: "yes", + stdin_isatty=True, + ) + assert r.exit_code == OK + out = capsys.readouterr().out + assert "Step: plug-usb" in out + assert "Plug the cable." in out + assert "LED on" in out + + +def test_ask_cli(): + code = main( + [ + "ask", + "--id", + "confirm-deploy", + "--prompt", + "Deploy?", + "--choices", + "yes,no", + "--default", + "no", + "--answer", + "no", + ] + ) + assert code == OK + + +def test_step_cli_confirm_decline_mutex(): + code = main( + [ + "step", + "--id", + "x", + "--prompt", + "Do it", + "--confirm", + "--decline", + ] + ) + assert code == SETUP_ERROR + + +def test_step_cli_confirm(): + code = main( + [ + "step", + "--id", + "plug-usb", + "--prompt", + "Plug USB", + "--expect", + "LED on", + "--confirm", + ] + ) + assert code == OK + + +def test_operator_gate_fixtures_match(): + """Shipped ask/step fixtures must exist and match expect.""" + repo = Path(__file__).resolve().parents[1] + from bedside.commands.eval_cmd import run_eval + + for name in ("operator-gate-ask", "operator-gate-step"): + path = repo / "eval" / "fixtures" / "known-good" / name + assert path.is_dir(), f"missing shipped fixture: {path}" + r = run_eval(repo, [path]) + assert r.exit_code == OK, (name, r.messages)