Repository navigation
Document structured choice UI; ban free-text multi-choice walls #7
New issue
Have a question about this project? Sign up for a free GitHub account to open an issue and contact its maintainers and the community.
By clicking “Sign up for GitHub”, you agree to our terms of service and privacy statement. We’ll occasionally send you account related emails.
Already on GitHub? Sign in to your account
Changes from all commits
File filter
Filter by extension
Conversations
Jump to
Diff view
Diff view
There are no files selected for viewing
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,5 @@ | ||
| id = "choice-wall" | ||
| expect = "fail" | ||
| principles = ["R2", "R4"] | ||
| title = "Multi-choice free-text dump (choice wall)" | ||
| notes = "Agent offers plan forks as free chat options instead of a structured picker." |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,19 @@ | ||
| # choice-wall (known-bad) | ||
|
|
||
| Domain-light illustration. Expect **fail** on R2 (choice wall) and R4 (human act is a free-text multi-menu). | ||
|
|
||
| ## Agent | ||
|
|
||
| Here is the Phase F plan. Pick what you want me to do next by typing one of these in chat: | ||
|
|
||
| 1. start #15 (recommended) | ||
| 2. fold open issues into the current PR | ||
| 3. reprioritize all P1 items | ||
| 4. skip Phase F and jump to deploy | ||
| 5. rewrite the plan from scratch | ||
|
|
||
| Just reply with the number or the short name. I will wait. | ||
|
|
||
| ## Operator | ||
|
|
||
| Which of these is safest if I am not sure about deploy yet? |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,5 @@ | ||
| id = "structured-choice" | ||
| expect = "pass" | ||
| principles = ["R2", "R4"] | ||
| title = "Plan fork via structured choice UI" | ||
| notes = "Agent uses host picker, recommended first, free text only for open judgment." |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
| @@ -0,0 +1,23 @@ | ||
| # structured-choice (known-good) | ||
|
|
||
| Domain-light illustration. Expect **pass** on R2 and R4. | ||
|
|
||
| ## Agent | ||
|
|
||
| Phase F is ready. I need one plan fork from you. | ||
|
|
||
| I am calling the host structured choice UI (not a free-text menu): | ||
|
|
||
| - id: phase-f-fork | ||
| - recommended first: continue with #15 | ||
| - other options: fold open issues into this PR; pause deploy until you review | ||
|
|
||
| If none of those fit, say so in free text and we will adjust. I will not dump five numbered options for you to type. | ||
|
|
||
| ## Operator | ||
|
|
||
| (structured UI) chose: continue with #15 | ||
|
|
||
| ## Agent | ||
|
|
||
| Confirmed. Next I will implement #15 myself and report back. No shell wall for you. |
| Original file line number | Diff line number | Diff line change |
|---|---|---|
|
|
@@ -122,7 +122,7 @@ def score_transcript(transcript: str) -> tuple[dict[str, bool], list[str]]: | |
| reasons: list[str] = [] | ||
| p: dict[str, bool] = {rid: True for rid in PRINCIPLE_IDS} | ||
|
|
||
| # R2: no shell wall | ||
| # R2: no shell wall / no choice wall | ||
| fences = _count_fenced_blocks(agent) | ||
| cmd_lines = _commandish_lines(agent) | ||
| run_these = bool( | ||
|
|
@@ -135,6 +135,25 @@ def score_transcript(transcript: str) -> tuple[dict[str, bool], list[str]]: | |
| p["R2"] = False | ||
| reasons.append("R2: large command wall") | ||
|
|
||
| # Choice wall: free-text multi-option menu instead of structured UI | ||
| numbered_opts = len(re.findall(r"^\s*\d+[.)]\s+\S+", agent, re.M)) | ||
| free_text_pick = bool( | ||
| re.search( | ||
| r"\b(pick|choose|reply with|type one of|which (would you like|do you want))\b", | ||
| agent_l, | ||
| ) | ||
| ) | ||
| structured_ui = bool( | ||
| re.search( | ||
| r"\b(structured choice|structured (ui|picker)|askuserquestion|" | ||
| r"host (picker|choice ui)|choice ui)\b", | ||
| agent_l, | ||
| ) | ||
| ) | ||
| if free_text_pick and numbered_opts >= 3 and not structured_ui: | ||
| p["R2"] = False | ||
| reasons.append("R2: choice wall (multi-option free-text menu)") | ||
|
|
||
| # R3: prefer doing (instruct when agent could run) | ||
| if run_these and not re.search( | ||
| r"\bi (will|I'll|am going to) (run|install|create|execute)\b", agent_l | ||
|
|
@@ -155,7 +174,7 @@ def score_transcript(transcript: str) -> tuple[dict[str, bool], list[str]]: | |
| p["R6"] = False | ||
| reasons.append("R6: blind auto on multi-candidate surface") | ||
|
|
||
| # R4: explicit human acts (vague batch) | ||
| # R4: explicit human acts (vague batch / free-text multi-choice as the act) | ||
| batched = bool( | ||
| re.search( | ||
| r"\bdo (all of )?the following\b|\bsteps?:\s*\n.*\n.*\n.*\n", | ||
|
|
@@ -169,6 +188,9 @@ def score_transcript(transcript: str) -> tuple[dict[str, bool], list[str]]: | |
| if vague_physical or (batched and re.search(r"\bbrowser\b|\bplug\b|\bhold\b", agent_l)): | ||
| p["R4"] = False | ||
| reasons.append("R4: vague or batched human act") | ||
| if free_text_pick and numbered_opts >= 3 and not structured_ui: | ||
|
There was a problem hiding this comment. Choose a reason for hiding this commentThe reason will be displayed to describe this comment to others. Learn more. This R4 check reuses the Useful? React with 👍 / 👎. |
||
| p["R4"] = False | ||
| reasons.append("R4: human choice presented as free-text multi-menu") | ||
|
|
||
| # R7: confirm in their words | ||
| needs_human = bool( | ||
|
|
||
There was a problem hiding this comment.
Choose a reason for hiding this comment
The reason will be displayed to describe this comment to others. Learn more.
Because
numbered_optscounts every numbered line in the whole agent block andfree_text_pickis set by anychoose, a valid explicit human act like “Choose the account you want” followed by a 3-step browser login checklist is scored as an R2/R4 choice wall even though no multi-option menu was offered. This can make known-good setup or login fixtures fail; tie the choice prompt to choosing among those numbered items, or distinguish procedural steps from option lists.Useful? React with 👍 / 👎.