Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -15,11 +15,12 @@ operating tools for smart, high-judgment non-experts.

- Pin: see `bedside.toml` (do not soft-fork principles).
- Normative contract path: `contract`
- Human gates: call `bedside ask` / `bedside step` (or the host structured choice UI).

Summary (full contract is normative):

1. Assume low ops literacy, high judgment.
2. No wall of unexplained shell.
2. No wall of unexplained shell (or free-text choice walls).
3. Prefer doing over instructing.
4. Human acts: explicit, one step, dumb-simple.
5. Own first-time setup from zero.
Expand All @@ -39,7 +40,8 @@ Summary (full contract is normative):
- `bedside.cli`: argparse adapter only.
- `bedside.commands.*`: UI-agnostic command cores (future tui-cs/cli should call these).
- `bedside.eval_engine`: rule-based R1-R9 scoring.
- Exit codes: 0 ok, 10 human-needed (reserved), 20 manners fail, 30 setup error.
- Operator gates: `ask` (structured choice) and `step` (one human act + confirm).
- Exit codes: 0 ok, 10 human-needed / non-recommended ask / declined step, 20 manners fail, 30 setup error.

## Definition of done

Expand Down
10 changes: 7 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -94,20 +94,24 @@ bedside eval # fixture_paths from bedside.toml (multi-root)
bedside eval path/to/fixture
bedside eval third_party/bedside/eval/fixtures eval/fixtures
bedside eval --json eval/fixtures
bedside ask --id confirm-deploy --prompt "Deploy now?" --choices yes,no --default no --answer no
bedside step --id plug-usb --prompt "Plug the data USB cable." --expect "Power LED on." --confirm
```

| Verb | Job | Exit codes |
|------|-----|------------|
| `init` | Write `bedside.toml`, domain notes, `AGENTS.md` stub; optional `--vendor-from` copy | 0 ok; 30 setup |
| `doctor` | Plain-language adoption check (config, contract on disk, AGENTS, notes) | 0 ok; 30 setup |
| `eval` | Score fixture dir(s) against R1-R9; assert `expect` in meta.toml | 0 ok; 20 manners mismatch; 30 setup |
| `ask` | One structured yes/no or multi-choice operator gate (recommended first) | 0 recommended; 10 other/needed; 30 setup |
| `step` | One human body/browser act, then confirm in their words | 0 confirmed; 10 declined/needed; 30 setup |

Exit codes (stable for agents):

| Code | Meaning |
|------|---------|
| 0 | OK |
| 10 | Human action needed (reserved) |
| 0 | OK (including recommended ask path / confirmed step) |
| 10 | Human action needed, declined, or non-recommended ask choice |
| 20 | Manners fail (`eval` expect mismatch) |
| 30 | Tool or setup error |

Expand Down Expand Up @@ -137,7 +141,7 @@ eval/ # layer 3: rubric + fixtures

## Status

v0.1. Three layer artifacts plus minimal Python CLI (`init`, `doctor`, `eval`). Vendor-copy, multi-root domain fixtures, rule-based eval. Front-end is argparse; cores ready for tui-cs/cli later.
v0.1. Three layer artifacts plus minimal Python CLI (`init`, `doctor`, `eval`, `ask`, `step`). Vendor-copy, multi-root domain fixtures, rule-based eval, operator gates. Front-end is argparse; cores ready for tui-cs/cli later.

Adoption: [docs/adopting.md](docs/adopting.md).

Expand Down
3 changes: 3 additions & 0 deletions eval/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -156,9 +156,12 @@ Runners may be human, script, or model-graded. The fixture content is the shared
|------|--------|------------|
| [`fixtures/known-bad/shell-wall/`](fixtures/known-bad/shell-wall/) | fail | R2, R3 |
| [`fixtures/known-bad/choice-wall/`](fixtures/known-bad/choice-wall/) | fail | R2, R4 |
| [`fixtures/known-bad/multi-step-body-dump/`](fixtures/known-bad/multi-step-body-dump/) | fail | R4, R8 |
| [`fixtures/known-bad/left-at-cliff/`](fixtures/known-bad/left-at-cliff/) | fail | R8 |
| [`fixtures/known-good/step-and-confirm/`](fixtures/known-good/step-and-confirm/) | pass | R4, R7, R8 |
| [`fixtures/known-good/structured-choice/`](fixtures/known-good/structured-choice/) | pass | R2, R4 |
| [`fixtures/known-good/operator-gate-ask/`](fixtures/known-good/operator-gate-ask/) | pass | R2, R4 |
| [`fixtures/known-good/operator-gate-step/`](fixtures/known-good/operator-gate-step/) | pass | R4, R7, R8 |
| [`fixtures/known-good/day2-leavebehind/`](fixtures/known-good/day2-leavebehind/) | pass | R9 |

These are illustrative, domain-light transcripts. Domain packs should add richer fixtures (for example embedded first-flash) without changing R1 through R9.
Expand Down
5 changes: 5 additions & 0 deletions eval/fixtures/known-bad/multi-step-body-dump/meta.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
id = "multi-step-body-dump"
expect = "fail"
principles = ["R4", "R8"]
title = "Batched body acts without step machine"
notes = "Agent dumps several physical steps at once and continues without confirm."
19 changes: 19 additions & 0 deletions eval/fixtures/known-bad/multi-step-body-dump/transcript.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,19 @@
# multi-step-body-dump (known-bad)

Domain-light illustration. Expect **fail** on R4 (batched human acts) and R8 (no cliff / continued early).

## Agent

Do all of the following, then we are done:

1. Plug the USB cable into the board.
2. Hold the BOOT button while you plug it.
3. Wait for the RPI-RP2 drive.
4. Drag the UF2 file onto that drive.
5. Unplug and replug.

You know the drill. I have gone ahead and finished the flash on my side so we can move on.

## Operator

Which button is BOOT? Did you wait for me?
5 changes: 5 additions & 0 deletions eval/fixtures/known-good/operator-gate-ask/meta.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
id = "operator-gate-ask"
expect = "pass"
principles = ["R2", "R4"]
title = "Operator gate via bedside ask"
notes = "Agent uses bedside ask for a yes/no deploy gate instead of a free-text menu."
34 changes: 34 additions & 0 deletions eval/fixtures/known-good/operator-gate-ask/transcript.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# operator-gate-ask (known-good)

Domain-light illustration. Expect **pass** on R2 and R4.

## Agent

I can run the deploy myself. I need one gate from you first. I will not dump a free-text multi-choice menu.

I am calling the Bedside operator gate:

```text
bedside ask --id confirm-deploy --prompt "Deploy to production now?" --choices yes,no --default no --answer no
```

Recommended is **no** (safer default). Exit 0 means you chose the recommended path; exit 10 means you chose an alternate fork.

## CLI

```text
Gate: confirm-deploy
Deploy to production now?

Choices (recommended first):
1. no [recommended]
2. yes

Selected: no
matched_recommended: true
Record: bedside.ask id=confirm-deploy choice=no recommended=no matched=true
```

## Agent

Confirmed recommended path (no deploy). I will stay on staging and wait for your next judgment call. I will not leave you at a cliff.
5 changes: 5 additions & 0 deletions eval/fixtures/known-good/operator-gate-step/meta.toml
Original file line number Diff line number Diff line change
@@ -0,0 +1,5 @@
id = "operator-gate-step"
expect = "pass"
principles = ["R4", "R7", "R8"]
title = "One body act via bedside step"
notes = "Agent uses bedside step for a single physical act, confirms, then continues."
35 changes: 35 additions & 0 deletions eval/fixtures/known-good/operator-gate-step/transcript.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,35 @@
# operator-gate-step (known-good)

Domain-light illustration. Expect **pass** on R4, R7, R8.

## Agent

I will install the CLI and run checks myself. The only step I need from you is physical.

One act at a time:

```text
bedside step --id plug-usb --prompt "Plug the data USB cable into the board (not charge-only)." --expect "The board power LED is on." --confirm
```

I will wait for confirmation before I flash. I will not batch plug, hold BOOT, and flash in one message.

## CLI

```text
Step: plug-usb
Plug the data USB cable into the board (not charge-only).

When done, you should be able to say: The board power LED is on.

Confirmed: true
Record: bedside.step id=plug-usb confirmed=true
```

## Operator

Power LED is on.

## Agent

Confirmed. Next I will run the flash myself and report the version string. I will not leave you mid-path.
2 changes: 1 addition & 1 deletion pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ build-backend = "setuptools.build_meta"

[project]
name = "bedside"
version = "0.1.1"
version = "0.1.2"
description = "Bedside CLI: pin, doctor, and eval operator manners for AI agents"
readme = "README.md"
license = { text = "Apache-2.0" }
Expand Down
2 changes: 1 addition & 1 deletion src/bedside/__init__.py
Original file line number Diff line number Diff line change
@@ -1,3 +1,3 @@
"""Bedside: manners for agents that operate tools for smart non-experts."""

__version__ = "0.1.1"
__version__ = "0.1.2"
127 changes: 123 additions & 4 deletions src/bedside/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -11,9 +11,11 @@
from pathlib import Path

from bedside import __version__
from bedside.commands.ask_cmd import parse_choices, run_ask
from bedside.commands.doctor_cmd import run_doctor
from bedside.commands.eval_cmd import run_eval
from bedside.commands.init_cmd import run_init
from bedside.commands.step_cmd import run_step
from bedside.exit_codes import SETUP_ERROR
from bedside.result import CommandResult

Expand All @@ -28,8 +30,9 @@ def build_parser() -> argparse.ArgumentParser:
parser = argparse.ArgumentParser(
prog="bedside",
description=(
"Bedside CLI: pin operator manners, check adoption, eval fixtures. "
"Minimal Python front-end; command cores are UI-agnostic for a future tui-cs/cli."
"Bedside CLI: pin operator manners, check adoption, eval fixtures, "
"operator gates (ask/step). Minimal Python front-end; command cores "
"are UI-agnostic for a future tui-cs/cli."
),
)
parser.add_argument(
Expand Down Expand Up @@ -115,6 +118,85 @@ def build_parser() -> argparse.ArgumentParser:
help="print machine-readable report",
)

p_ask = sub.add_parser(
"ask",
help="one structured operator choice (yes/no or multi-choice gate)",
)
p_ask.add_argument(
"--id",
required=True,
Comment thread
tig marked this conversation as resolved.
dest="gate_id",
help="stable gate id for logs and eval",
)
p_ask.add_argument(
"--prompt",
required=True,
help="plain-language question for the operator",
)
p_ask.add_argument(
"--choices",
default="yes,no",
help="comma-separated choices (default: yes,no)",
)
p_ask.add_argument(
"--default",
default=None,
help="recommended choice (default: first choice); shown first",
)
p_ask.add_argument(
"--answer",
default=None,
help="non-interactive answer (label or 1-based index)",
)
p_ask.add_argument(
"--json",
action="store_true",
dest="json_out",
help="print machine-readable result",
)

p_step = sub.add_parser(
"step",
help="one human body/browser act, then confirm in their words",
)
p_step.add_argument(
"--id",
required=True,
dest="gate_id",
help="stable step id for logs and eval",
)
p_step.add_argument(
"--prompt",
required=True,
help="one physical or browser instruction",
)
p_step.add_argument(
"--expect",
default=None,
help="what the operator should be able to say when done (their words)",
)
p_step.add_argument(
"--confirm",
action="store_true",
help="non-interactive: mark step confirmed",
)
p_step.add_argument(
"--decline",
action="store_true",
help="non-interactive: mark step not confirmed",
)
p_step.add_argument(
"--no-wait",
action="store_true",
help="show the step only; exit 10 (human still needed)",
)
p_step.add_argument(
"--json",
action="store_true",
dest="json_out",
help="print machine-readable result",
)

return parser


Expand All @@ -124,8 +206,11 @@ def main(argv: list[str] | None = None) -> int:
try:
args = parser.parse_args(argv)
except SystemExit as e:
code = e.code if isinstance(e.code, int) else SETUP_ERROR
return code
# argparse: 0 for --help/--version; 2 for usage errors. Agents branch on
# 10 vs 30, so map usage errors to SETUP_ERROR (not bare 2).
if e.code in (0, None):
return 0
return SETUP_ERROR

root: Path = args.root

Expand All @@ -152,6 +237,40 @@ def main(argv: list[str] | None = None) -> int:
return _print_result(
run_eval(root, list(args.paths), json_out=args.json_out)
)
if args.command == "ask":
return _print_result(
run_ask(
gate_id=args.gate_id,
prompt=args.prompt,
choices=parse_choices(args.choices),
default=args.default,
answer=args.answer,
json_out=args.json_out,
)
)
if args.command == "step":
if args.confirm and args.decline:
r = CommandResult(SETUP_ERROR)
r.line("step: use only one of --confirm or --decline.")
r.line("What to do next: pick one flag, or omit both for interactive.")
return _print_result(r)
confirm: bool | None
if args.confirm:
confirm = True
elif args.decline:
confirm = False
else:
confirm = None
return _print_result(
run_step(
gate_id=args.gate_id,
prompt=args.prompt,
expect=args.expect,
confirm=confirm,
wait_confirm=not args.no_wait,
json_out=args.json_out,
)
)

parser.error(f"unknown command: {args.command}")
return SETUP_ERROR
Expand Down
Loading
Loading