Skip to content

Commit b576ba8

Browse files
MShokryclaude
andcommitted
docs: changelog for v0.9.0
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_017tZMT6285UmrKjA27wmzCT
1 parent 739ef18 commit b576ba8

3 files changed

Lines changed: 91 additions & 25 deletions

File tree

‎CHANGELOG.md‎

Lines changed: 43 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -10,6 +10,49 @@ update workflow. Version bumps mean: **MAJOR** = state-file contract /
1010
role authority / script interface changed; **MINOR** = new template,
1111
script, flag, or role rule; **PATCH** = prose and docs.
1212

13+
## v0.9.0 — 2026-09-27
14+
15+
- `[safety]` Synced this toolkit's own opencode v2 migration into the
16+
scaffolded pipeline templates: the shell permission key renamed `bash:` →
17+
`shell:` in `builder.md.tmpl`/`reviewer.md.tmpl`/`tester.md.tmpl` (under
18+
the old key, none of a template's shell `deny`/`ask` rules matched
19+
anything under opencode v2, so `rm -rf /*`, `sudo *`, `curl *`, etc. were
20+
silently unenforced); `oc.sh.tmpl` now authenticates with
21+
`OPENCODE_PASSWORD` (v2's `serve` is always password-gated, v1 had none),
22+
polls via `opencode api` against v2's `/api/*` routes, aborts via the
23+
renamed `/api/session/<id>/interrupt` endpoint, and picks the *newest*
24+
message for watchdog progress instead of the oldest (v2 returns messages
25+
newest-first; the old indexing could silently track the wrong message and
26+
let a wedged run look like it was still progressing).
27+
- `[docs]` Corrected an earlier `oc.sh.tmpl`/agent-template comment claiming
28+
`--auto` lets a deny-listed shell command run anyway on opencode v2 —
29+
re-verified live via `GET /api/agent/<name>` and could not be reproduced;
30+
left the correction in place rather than deleting the original claim
31+
outright.
32+
- `[process]` `bin/init.sh` now checks `opencode --version` on every
33+
scaffold or `--update` run and warns (non-fatally) when it's missing, v1,
34+
or newer than the v2 these templates assume, explaining specifically what
35+
breaks on v1 (password auth, the permission-key rename, the `opencode
36+
api` subcommand, `--server` replacing the removed `--attach`/`--dir`).
37+
`toolkit-update.md.tmpl` now tells the lead to surface this warning to
38+
the user rather than let it scroll past.
39+
- `[process]` `bin/init.sh`'s `--builder-model`, `--reviewer-model`,
40+
`--reviewer-fallback-model`, and `--tester-model` are no longer required —
41+
`apply_defaults()` now fills them with `opencode-go/glm-5.3-flash` /
42+
`opencode-go/minimax-m2.7` / `opencode-go/deepseek-v4-flash` /
43+
`hcnsec/auto` when omitted, the same lineup two independently scaffolded
44+
projects (resto-agent, relationship) landed on. `--project-name` remains
45+
the only required flag; every flag can still be passed explicitly to
46+
override the defaults. Backward compatible — a call that already passed
47+
all four flags behaves identically.
48+
- `[safety]` `init.sh` now also adds `.agents/.oc-password` (the local
49+
opencode v2 server password `scripts/team.sh` generates) to a scaffolded
50+
project's `.gitignore`, alongside the existing `.oc-port` and Claude
51+
session-id entries.
52+
- `[docs]` README's quickstart and `docs/MODELS.md`'s recommended lineup
53+
updated to show this lineup as the documented default instead of the
54+
older Kimi/GLM-5.2/Sonnet-fallback/DeepSeek-Flash recommendation.
55+
1356
## v0.8.0 — 2026-09-13
1457

1558
- `[process]` Fresh scaffolds now support Codex as the project lead when Claude

‎README.md‎

Lines changed: 22 additions & 8 deletions
Original file line numberDiff line numberDiff line change
@@ -168,20 +168,34 @@ opencode models # see what's actually configured before picking models
168168
~/tools/agent-toolkit/bin/init.sh \
169169
--target . \
170170
--project-name "my-project" \
171-
--claude-model sonnet \
172-
--builder-model "hcnsec/auto" \
173-
--reviewer-model "hcnsec/glm-5.3" \
174-
--reviewer-fallback-model sonnet \
175-
--tester-model "hcnsec/auto" \
176171
--test-dir e2e
177172
```
178173

174+
`--project-name` is the only required flag. `--claude-model`,
175+
`--builder-model`, `--reviewer-model`, `--reviewer-fallback-model`, and
176+
`--tester-model` all default to the lineup two independent real projects
177+
converged on: Claude Sonnet lead/planner, `opencode-go/glm-5.3-flash`
178+
builder, `opencode-go/minimax-m2.7` reviewer,
179+
`opencode-go/deepseek-v4-flash` reviewer fallback, `hcnsec/auto` tester.
180+
Override any of them per project once `opencode models` shows your server's
181+
actual list differs — these are a starting point, not a guarantee those
182+
exact ids still exist for you:
183+
184+
```bash
185+
~/tools/agent-toolkit/bin/init.sh \
186+
--target . \
187+
--project-name "my-project" \
188+
--builder-model "<vendor/model>" \
189+
--reviewer-model "<vendor/model, different family than builder>" \
190+
--reviewer-fallback-model "<vendor/model, different family again>" \
191+
--tester-model "<vendor/model>"
192+
```
193+
179194
`--reviewer-model` and `--reviewer-fallback-model` should be **different
180195
model families** — the fallback is what the pipeline switches to when
181196
`builder` implements and would otherwise share a vendor with the default
182-
reviewer, which would defeat cross-vendor independence. The `hcnsec/auto`
183-
values above are flag *shape* only — run `opencode models`, pin real
184-
strings, and do not use `auto` for the reviewer.
197+
reviewer, which would defeat cross-vendor independence. Do not use `auto`
198+
for the reviewer.
185199

186200
Cost/quality picks (Kimi implementer, GLM reviewer, DeepSeek Flash
187201
tester, Claude Sonnet lead/planner/fallback), and why one OpenCode

‎docs/MODELS.md‎

Lines changed: 26 additions & 17 deletions
Original file line numberDiff line numberDiff line change
@@ -41,48 +41,57 @@ request volume per 5h window, expensive → cheap:
4141
| Kimi K3 | Very high on Go (~110 req / 5h) | Rare “actually hard” implementer, not default |
4242
| GLM-5.3 | High (~220 req / 5h, tighter cap than 5.2) | Best GLM for **one-shot** review |
4343
| GLM-5.1 / 5.2 | Mid–high (~880 req / 5h) | Review, or planner if Claude credits are tight |
44-
| Kimi K2.6 / K2.7 Code | Mid (~1k–1.3k req / 5h) | Default implementer — agentic loops |
44+
| GLM-5.3-flash | Mid, cheaper than full 5.3 | **Default implementer** — agentic loops |
45+
| Minimax M2.7 | Mid, no published Go cap seen yet | **Default reviewer** — different lab than GLM, one/two calls |
46+
| Kimi K2.6 / K2.7 Code | Mid (~1k–1.3k req / 5h) | Implementer alternative — agentic loops |
4547
| DeepSeek V4 Pro | Low | Implementer when the window is tight |
46-
| DeepSeek V4 Flash | Lowest (~30k req / 5h) | Tester only |
48+
| DeepSeek V4 Flash | Lowest (~30k req / 5h) | **Default reviewer fallback**; also fine as tester |
4749

4850
These counts are provider estimates, not a promise. GLM-5.3 is usually
4951
the same *job* as 5.2 at a worse credit rate — prefer 5.2 unless 5.3 is
5052
clearly better on your list.
5153

5254
## Recommended lineup
5355

54-
For this toolkit's own work (bash + sed, markdown templates, process
55-
rules, no app runtime):
56+
`bin/init.sh`'s own defaults (`apply_defaults()` — override any of them
57+
with the matching flag per project). This is the lineup two independent
58+
real projects landed on separately, then converged on deliberately:
5659

5760
| Role | Model | Why |
5861
| --- | --- | --- |
5962
| **Lead** | Claude Sonnet | Already the Claude Code session. Opus does not improve dispatch. |
6063
| **Planner** | Claude Sonnet | One shot. A cheaper model writes mushy ACs or starts designing the implementation. Drop to GLM-5.2 only if the Claude pool is the bottleneck. |
61-
| **Implementer** (`builder` / `senior-dev`) | **Kimi K2.7 Code** (or K2.6) | Cost/quality for surgical template + `init.sh` work. GLM-5.2 only if Kimi is missing. Not GLM-5.3, not K3, not Opus as the default loop. |
62-
| **Reviewer** | **GLM-5.3** or **5.2** | Different family than Kimi. One/two calls, so GLM's per-call cost is acceptable. Catches process bugs (leftover placeholders, `render()` skip-if-exists, permission YAML ≠ enforcement) better than Flash. |
63-
| **Reviewer fallback** | Claude Sonnet (or DeepSeek Pro) | A **third** family, and preferably a **second gateway** (see below). Used when builder would share a family with the default reviewer. |
64-
| **Tester** | DeepSeek V4 Flash | Smoke `init.sh` into `/tmp`, assert no `__[A-Z_]*__`, second run does not clobber. Flash is enough. |
64+
| **Implementer** (`builder` / `senior-dev`) | **`opencode-go/glm-5.3-flash`** | Cost/quality for the implementer loop's tool-call volume. Kimi K2.7 Code is a solid alternative if GLM-5.3-flash is missing or underperforming on your codebase. |
65+
| **Reviewer** | **`opencode-go/minimax-m2.7`** | Different family than the builder's GLM. One/two calls, so per-call cost is acceptable; a different lab's weights catch blind spots the implementer's own family shares. |
66+
| **Reviewer fallback** | **`opencode-go/deepseek-v4-flash`** | A **third** family. Used when the implementer (e.g. `senior-dev` or a differently-configured builder) would otherwise share a vendor with the default reviewer. |
67+
| **Tester** | **`hcnsec/auto`** | One shot + a smoke command; faithful reporting matters more than reasoning power here, so a router is fine for this role specifically (never for the reviewer — see below). |
6568

66-
`init.sh` shape (replace with strings from `opencode models`):
69+
`init.sh` shape (these are the defaults — passing no model flags at all
70+
gets you exactly this; replace with strings from `opencode models` if your
71+
server's list differs):
6772

6873
```bash
6974
--claude-model sonnet \
70-
--builder-model "<aggregator>/kimi-k2.7-code" \
71-
--reviewer-model "<aggregator>/glm-5.2" \
72-
--reviewer-fallback-model sonnet \
73-
--tester-model "<aggregator>/deepseek-v4-flash"
75+
--builder-model "opencode-go/glm-5.3-flash" \
76+
--reviewer-model "opencode-go/minimax-m2.7" \
77+
--reviewer-fallback-model "opencode-go/deepseek-v4-flash" \
78+
--tester-model "hcnsec/auto"
7479
```
7580

7681
If you implement with GLM, reviewer must **not** be GLM (5.2 vs 5.3
77-
does not count). Switch reviewer to Kimi or DeepSeek Pro.
82+
does not count) — that's exactly why the default reviewer is Minimax, a
83+
different lab, rather than another GLM snapshot. If your implementer
84+
changes family (e.g. Kimi), re-check that the reviewer still differs from
85+
it; switch to the fallback family if it does not.
7886

7987
Tight credits: DeepSeek Pro implementer, GLM-5.2 reviewer, Flash tester.
8088
High-stakes (`render()`, permission blocks, state-file contract): keep
81-
Kimi on implementer; optionally fire the Sonnet fallback as a second
82-
review pass — still cheaper than implementing on Opus.
89+
the default implementer; optionally fire the Sonnet reviewer fallback as
90+
a second review pass — still cheaper than implementing on Opus.
8391

8492
Do not: Claude on builder **and** reviewer; K3/Opus as default
85-
implementer; Flash as reviewer; `auto` anywhere independence matters.
93+
implementer; Flash as reviewer; `auto` anywhere independence matters
94+
(reviewer, never tester).
8695

8796
## Provider: one aggregator for workers, not one tool per lab
8897

0 commit comments

Comments
 (0)