paper/: the working preprint draft, ported under export discipline - #3
Conversation
The draft was written pre-rename and pre-export. It is REWRITTEN here rather
than copied: the old tool name is gone from prose, titles and the author line
("The ripwire authors" — no personal names by policy), and every measured
number was reconciled against docs/EVALS.md and the bench/ files it names.
Four numbers did not survive reconciliation, and the paper says so in place
rather than quietly swapping them:
- The headline held-out strict file@10 was 66.7%. That figure belongs to the
BASELINE ARM of the anchor-hop experiment at that experiment's binary, and
EVALS.md §8 refuses it in the role the draft used it for. §4.1 now leads
with the published 60.9% (bench/locbench/README.md) and 66.7% appears only
in §6.2, labelled as the arm it is.
- The mention anchor's single-file stratum was +4.9pp. The machine-generated
scoreboard gives 86.3% vs 82.4% = +3.9pp, which is also EVALS.md §4's
figure. The draft's number was arithmetic error.
- "LLM agents buy roughly +20pp over this floor" is removed: the subtraction
depended on the retired 66.7%, and it compared a symbol-max-pooled held-out
slice against a file-document full-set number across two papers — the exact
comparison §3.1 forbids.
- The private-corpus C++ absolutes are not restated at all (EVALS.md §8);
§4.3 carries EVALS.md §7's framing instead, which is the publishable form.
The two-tier gate is also demoted from "policy" to what GATE_DECISION.md
actually says it is — a proposal, opt-in in the comparator — while keeping the
true and narrower claim: it was pre-registered and applied to §6's two rejects.
The negative result is the centerpiece and keeps its detail, now with the
sequel EVALS.md §7 carries (the +0.41pp / LB +0.00pp anchor candidate, also
rejected). §9 names the second contribution being folded in — the audit loop,
the two-flavour CI build, sibling-completeness, and the gated-claims machinery
that scans this very file.
Every table names the in-repo artifact that pins it; all cited paths verified
present in this tree.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
LINEAGE.md's "what is actually new" closed on a paper "in preparation" with nothing to open; it now names the directory. Sentence shape unchanged, so readmedriftcheck.sh's count arms are untouched — re-run anyway, ALL PASS (E8 additionally re-verifies every repo-relative path LINEAGE cites, now 31). docs/README.md's "Outside this directory" list gains the paper/ row, with the status stated in the row itself so the index cannot imply a finished artifact. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A paper is the maximal fabricated-flag surface: long, written in one pass,
quoting flags in running prose rather than in command blocks, and re-read for
argument rather than for syntax — so a plausible-but-wrong flag survives longer
there than in a deck, which is the failure this gate was built for.
Follows the file's own V3 MED-1 convention rather than only bumping the total:
paper is a tagged glob FAMILY with its own paperCount and its own non-empty
assertion, so the family cannot vanish behind a total the other four still
satisfy. Total floor 38 -> 40.
RED FIRST, as the convention requires: with paper/ scratch-removed the new
assertion fires before the total floor gets a chance to —
deckcheck: 0 paper/*.md prose source(s) found — refusing to run
EXIT=2
and with it restored the gate scans 44 files, ALL PASS. It earned its keep on
the first green run: it caught a metasyntactic "--flag" in the paper's own §9
sentence about this very gate. Fixed in the prose, NOT allowlisted — a
placeholder is not a claim, and an allowlist entry for one would be the hiding
place this gate's rot check exists to prevent.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Warning Review limit reached
Next review available in: 36 minutes You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. Review details⚙️ Run configurationConfiguration used: Organization UI Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (3)
📝 WalkthroughSummary by CodeRabbit
WalkthroughThe change adds a ripwire preprint covering methodology, benchmarks, results, experiments, limitations, reproducibility, and artifacts. Repository documentation links to the draft. ChangesPreprint documentation and validation
Estimated code review effort: 3 (Moderate) | ~20 minutes 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 4
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@paper/PREPRINT.md`:
- Around line 24-26: Qualify the “sub-second” claim by specifying the applicable
latency statistic and cache condition, noting the 1,623.8 ms cold p95 rather
than presenting it as unconditional. Apply the same wording in paper/PREPRINT.md
lines 24-26 and paper/README.md lines 6-8 so the abstract and paper index
consistently describe the evaluated ranker.
- Around line 301-306: Remove the private-corpus numeric results and the
comparison against the public SFML figures from the paragraph beginning “The
public-versus-private divergence.” Retain only the statement that the private
corpus was removed, its results are unreproducible, and docs/EVALS.md §8 does
not publish private absolutes.
- Around line 314-327: Update the comparison description in PREPRINT.md to
reconcile the “four arms” wording with the five table rows by identifying the
aider repo-map no-persona row as an auxiliary sensitivity run, or by describing
the comparison as five arms. Explicitly state which measured rows underpin the
win/loss and latency claims.
In `@paper/README.md`:
- Around line 9-11: Update the machine-checking claim in the paper text to say
“selected publication claims” instead of “every published claim,” unless the
implementation can document and demonstrate complete binary-check coverage for
all listed benchmarks, metrics, gates, and command references.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: Organization UI
Review profile: CHILL
Plan: Pro Plus
Run ID: cb3f2804-2071-4be9-8894-f19d0d741aa6
📒 Files selected for processing (5)
docs/LINEAGE.mddocs/README.mdpaper/PREPRINT.mdpaper/README.mdtest/deckcheck.sh
| design space: a **deterministic, zero-runtime-dependency, sub-second, symbol-granularity** static | ||
| ranker (ripwire) that uses no model, no network, and no randomness, and whose output is | ||
| byte-identical run-to-run. We contribute: **(1)** a corrected LocBench evaluation methodology for |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win
Qualify the “sub-second” headline.
The paper reports a 1,623.8 ms cold p95 for an evaluated candidate. State whether “sub-second” means warm p95, median latency, or another specific path. Do not present it as an unconditional property of the evaluated ranker.
paper/PREPRINT.md#L24-L26: add the latency statistic and cache condition to the abstract.paper/README.md#L6-L8: use the same qualified wording in the paper index.
📍 Affects 2 files
paper/PREPRINT.md#L24-L26(this comment)paper/README.md#L6-L8
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@paper/PREPRINT.md` around lines 24 - 26, Qualify the “sub-second” claim by
specifying the applicable latency statistic and cache condition, noting the
1,623.8 ms cold p95 rather than presenting it as unconditional. Apply the same
wording in paper/PREPRINT.md lines 24-26 and paper/README.md lines 6-8 so the
abstract and paper index consistently describe the evaluated ranker.
| Four local-first, no-API-key arms, identical checkouts, identical gold, and the **same metric code | ||
| imported unmodified** from the LocBench harness; zero exclusions in any arm. Versions pinned in the | ||
| report: aider-chat 0.86.2 (repo-map, fed the issue through its own ident-extraction), | ||
| codebase-memory-mcp 0.9.0 (its shipping BM25 graph-search path), graphifyy 0.9.15 (local BFS | ||
| traversal, `PYTHONHASHSEED=0` — see below). The slice is deliberately hard: 40 of 60 instances have | ||
| multi-file gold. | ||
|
|
||
| | arm | file@1 | file@3 | file@5 | file@10 | any@10 | median wall | | ||
| |---|---|---|---|---|---|---| | ||
| | **ripwire `--for`** | 5.0% | **18.3%** | **26.7%** | **36.7%** | **75.0%** | **0.074 s** (warm, pre-built index) | | ||
| | codebase-memory-mcp | **6.7%** | 10.0% | 16.7% | 26.7% | 66.7% | 1.14 s | | ||
| | graphify (BFS order) | 0.0% | 3.3% | 5.0% | 21.7% | 41.7% | 5.8 s | | ||
| | aider repo-map (personalized) | 0.0% | 1.7% | 6.7% | 13.3% | 33.3% | 2.5 s | | ||
| | aider repo-map (no persona) | 0.0% | 1.7% | 3.3% | 10.0% | 26.7% | — | |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win
Clarify the head-to-head arm count.
The prose says the comparison has four arms, but the table contains five measured rows because aider repo-map appears with and without a persona. Label the no-persona row as an auxiliary sensitivity run, or describe the comparison as five arms. State which row supports the win/loss and latency claims.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@paper/PREPRINT.md` around lines 314 - 327, Update the comparison description
in PREPRINT.md to reconcile the “four arms” wording with the five table rows by
identifying the aider repo-map no-persona row as an auxiliary sensitivity run,
or by describing the comparison as five arms. Explicitly state which measured
rows underpin the win/loss and latency claims.
…d three provenance sentences F1: §3.2×2 + §7 claimed zstd/jq/ponyc live in bench/multiswe/dataset.lock; the lock's languages list is ['cpp'] and its own README says the C split needs a separate mine. Now: minable by the same harness, not in the lock. F2: the REGISTER was wrong — EVALS §8 said 'train slice' for a number its own pinning file books under heldout_acceptance; fixed there (n=243) and noted in §6.2 per the port's own header rule. F3: 'does not restate them' was false four lines under two rounded restatements; the sentence now says exactly what is and is not carried. F4: §1(c)'s universal hedged; C1 quotes the hedge it actually uses; the multiswe README H1 hedged at the pin's destination. Plus V9's two nits (§4.2 modifier, §4.5 wall/inst pointer). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…k caught it The paper/ scanning arm flagged --languages (a mining-harness flag, not a ripwire flag) four commits after it was added for exactly this class. The pipe in my gate chain had swallowed deckcheck's exit code; the follow-up commit is the penalty. Prose now names the capability, not the flag. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…d now says so Measured on this repository: 20 of the 40 default-preset panel rows were src/main.cpp symbols, every one of them reporting the identical `hrank=0 churn=47`. Churn is mined per PATH, so every symbol in a file carries that file's churn=/hrank= verbatim and a symbol in a churny file collects the historical family for free, with no property of its own. That is the family's real unit, not a bug — ensemble.h has always ranked it by churn alone, deliberately, so it cannot agree with structural by construction. What was wrong is the legend a reader meets first: it said "historical (git change frequency)" and stopped, in a report whose entire premise is evidence families that must be able to DISAGREE. A reader counting an inherited file fact as per-symbol corroboration is reading the fam= column wrong, and non-negotiable #3 makes that a defect in the output rather than a caveat for the docs. Both legends now name the unit where they introduce the family, and both state the consequence: file evidence INHERITED by the row, not the row's own history. Nothing else moves — no ranking, no threshold, no family set, no preset. The historical family's exclusion from `strict` and its churn-alone ranking are unchanged and still measured, not re-argued. Gate: test/qualitypanelcheck.sh arm (N), new, and in that order on purpose — (N1) MEASURES the property from the binary's own output (both hub.cpp rows must carry one identical historical evidence string, with the once-committed file firing historical on no row as the control, so (N1) pins a shared value rather than a constant one), and only then (N2) asserts the disclosure in both the quality-panel and --ensemble legends. A gate that pinned the sentence alone would keep passing if the family's unit ever changed underneath it. Both (N2) arms verified red against the pre-change binary. Verified: qualitypanelcheck.sh PASS standalone; full suite pargates.py -j 6 ALL PASS (362 gates, 3 standing skips); determinism double-run diff clean; xmllint clean. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…call above the shadowing local no longer vanishes (kParserVer 55)
Fourth A5 iteration, and the first refutation that was a RECALL loss rather
than over-suppression: iteration 2 scoped suppression to the declaring block
but started the span at the block's opening BRACE, so a real call written
ABOVE the local was silently eaten. Verifier attack4 —
void key() { }
void caller() { key(); int key = 0; (void)key; }
— went base 2 (one right, one false) → postS/postSfix/postSfix2 count=0. Zero
is the wrong answer twice over: it breaks the prereg's non-regression #5 and
CLAUDE.md non-negotiable #3, and 79df919's "within-block declaration-point
ordering, disclosed" disclosed the wrong thing — a disclosed floor may
under-suppress, never delete a real site.
THE DECLARATION POINT SHIPPED IS THE END BYTE OF THE COMPLETE DECLARATOR,
which is C++ [basic.scope.pdecl] exactly (immediately after the complete
declarator, before its initializer; for a structured binding, after its
identifier-list — the outermost declarator's end byte is both). Chosen over
the end of the declaration STATEMENT because it is the only boundary that
gets a multi-declarator statement right in both directions:
`int a = probe(), probe = 0, b = probe;` keeps the call in a's initializer
AND suppresses b's read. It is a byte the grammar hands us, not an estimate;
what remains is the pre-existing name-matching floor, now merely visible
above the point too (a kept pre-declaration site is kept, never resolved).
enclosingBlockSpan becomes enclosingShadowScope and reports plainBlock, so
narrowing applies to an ordinary block ONLY: control-statement header
declarations (iteration 3) and every whole-scope shape — definition/lambda/
catch parameters, captures and init-captures, range-for variables — keep
their whole-statement/whole-body spans, because their names ARE in scope from
the start and narrowing them would re-mint iterations 1-3's false positives.
Un-masked defect, fixed in the same diff: narrowing stops covering the
declaration LINE, which exposed that isDeclSiteName never saw a `declaration`
parent at all — `int key;` leaked its own declared name as a role="read", and
`int a, key;` leaked the second one even where the parent type IS listed,
because a `declaration` carries one `declarator` FIELD PER NAME and
ts_node_child_by_field_name returns the first. Iterations 1-3 could not
observe either: the block-start span suppressed the declaration line along
with everything else.
Verifier fixtures: attack4 0→1 (the genuine call, and only it);
attack2a-c/3a_for/3a_if/3a_extent/3a_nestedwrap/3b_catch/3c_goto/
3d_genericlambda and all eleven attack5_* hold their counts exactly.
test/shadowcheck.sh grows 31→50 arms and each new arm is proven red on a
binary that fails it: 4 on the 54 binary (pre-declaration call in --uses and
in --callers, multi-declarator pre-declarator call, structured-binding
declaration), 2 on a boundary mutant that uses the statement end (b's read,
`int probe = probe;`), 2 on the un-fixed narrowing build (bare and
second-declarator name leaks), 9 on a mutant that unifies the whole-scope
shapes onto the enclosing block (param, range-for, catch, init-capture,
sibling blocks) — none was green everywhere.
Bite on src (54 → 55): --uses=key 10→10 with a byte-identical site list,
line 44→42, name 294→290, run 6→5, empty 1551→1551, count 11→11. ZERO sites
reappeared — src contains no pre-declaration-call instance of the bug class —
and all seven that vanished are declaration-site names hand-checked at
source: `std::string line;` (recall.h:400,457), `std::string name, eq, s;`
(arch.h:376), `std::string_view file, name;` (graph.h:2139,2433),
`std::string run;` (search.h:691), `std::string_view fileFilter, name;`
(layout.h:2359) — the leak this commit closes, not shadow suppression.
Span values and the reference population change → kParserVer 54→55 + quality.h
mirror + qschemetrip re-pin. EVALS gate count unchanged (no new gate script).
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ness into SARIF Adversarial-verifier round on the just-merged idea-harvest wave. Five findings, four fixed as briefed and one refuted by measurement. F1 (MAJOR) — silent built-in subtree prunes. Three prune paths converge on ingest.cpp's disable_recursion_pending(): a --exclude match, the committed denylist (kCrawlSkipDirs), and the CMakeCache.txt build-output sentinel. Only the first incremented a counter, so `--skipped` over a tree with node_modules/ reported EVERY counter zero while whole subtrees had been dropped — the honesty contract's own failure mode one level up from the class §L1 closed. Built-in prunes now land in their own CrawlSkips::prunedDirs, emitted as pruned_dirs= on the <skipped> root and NOT folded into excluded_dirs=: "a rule you passed" and "a rule this build carries" are different answers to "why is my tree missing", and one counter cannot say both. The legend carries the same "contents UNKNOWN, not zero, and in no count here" language for the new class. Multi-root sums it. F2 (MAJOR) — SARIF dropped two facts the XML states, and both are the difference between "ran and found nothing" and "never ran here". tool.driver.rules was byte-identical whether --lint-select= kept a rule or not, and inert rules read as ran-clean. Deselected rules now carry defaultConfiguration.enabled=false (SARIF's own "catalogued, not enabled"; kept rules omit it and inherit the standard's default of true); every rule carries properties.applicable, true as well as false, because a SARIF document travels without a legend and an absent property would be read as true; runs[0].properties mirrors the XML root's selected="K of N" plus the raw select=/ignore=, and only under a real selection. F3 (MINOR) — REFUTED AS BRIEFED, closed on its other half. The brief asked for unindexed=/unindexed_exts= clauses in the DEFAULT MAP legend. Measured against a pristine HEAD binary: on src/ at --max-tokens=500 the map's fixed floor is 1173 B against a 1180 B allowance — 7 bytes of headroom — so any clause, conditional or not, reds test/tokenbudgetcheck.sh arm #3; the shortest honest wording is ~150 B. serialize.h's own note already recorded this constraint and it holds. What WAS genuinely undefined everywhere is unindexed_exts=, the TOP-6 cap bit, and that is now defined in the --skipped legend beside unindexed=, which is under no token budget. The map legend is left byte-identical, the refutation is recorded at buildUnindexedAttr with the measurement, and skipreasoncheck arm (9) now pins the negative so a clause cannot land there silently. F4 (MINOR) — .ripwire_quality_acks carried 8 duplicate (kind,key) rows from a three-way merge, violating its own "a file this binary wrote never contains a duplicate key" invariant. Deduped by quality.h::readAckRecords' D2 rule verbatim (max(ackNow) wins, strict > so a tie keeps the first line); --quality-delta output verified BYTE-IDENTICAL before and after. No key lost, no ratchet floor lowered. F5 (MINOR) — --recall's per-doc lines= is the selected section range even after the byte budget cuts that body mid-section. One clause on the header, charged only to a run that truncated something and appended after the last attribute so the name=value tail stays scannable. No behaviour change. Gates, RED against a pristine HEAD binary built from `git archive HEAD` and green after: test/skipreasoncheck.sh arms (8) pruned_dirs and (9) unindexed_exts; test/sarifcheck.sh arms 8 (selection crossover), 9 (applicability crossover) and 10 (mutation control — the unfiltered C-family run must show the negative of both). Every quality regression this round opened was FIXED, not acked: emitLintSarif's complexity 11->20 / verbosity 44->69 / params and api-surface 5->6 all back to baseline via writeSarifRuleDecl + writeSarifResult + writeSarifRunProperties and one SarifRunProperties struct instead of a parameter per disclosure; emitRunLintSarif 8->12 params back to 8 by taking the MainDispatch every other handler already takes; formatRecallHeader 55->67 LOC back by hoisting the rationale above the function. The 8 acked rows are short-horizon-churn=self only — this change's own edit window on the verbs a disclosure gap cannot be closed without touching. pargates 415 gates: 412 pass, 3 skip, 0 fail. --quality-delta exit 0, gating=0. --skipped / --lint --sarif / --recall deterministic run-to-run; XML xmllint-clean; SARIF parses as JSON. README --callers line numbers and docs/COMMANDS.md regenerated from the binary. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…ng (R9)
E6 / the wave-2 growth round reproduced "<bodies> ABSENT ENTIRELY" instead
of shown="0" on four of five --token-budget=3000 --for queries across four
corpora (routing note R9, exp-e6.md A5): the honesty contract ("a zero
means none found, never none exists", CONTRIBUTING.md #3) applies to
elements, not only counts, and this violated it twice over.
src/main.cpp buildForAutoBodies (the --for / MCP `for` T3 auto-bundle
path) had two branches that discarded the element instead of keeping it:
- autoBodyIds.empty() (no positive-score candidates at all) returned
without ever calling packBodies, so no <bodies> render existed to keep.
Now calls packBodies with the (empty) candidate list, which renders the
honest "<bodies shown="0" total="0" capped="0"></bodies>" shell on its
own -- no second hand-rolled tag.
- autoEmitted.kept.empty() (candidates existed, none fit) had ALREADY
rendered that same honest shell via the chargeSection() call just
above, then threw it away ("drop the empty section whole"). Now keeps
it; only the attr= reason on <ctx> is new information.
src/packtask.h packTaskBundleText's bodiesStr had the identical gap: the
`else` of `if (!bodyIds.empty() && bodiesBudget >= kPackTaskSectionFloor)`
left bodiesStr empty, so the whole element vanished. Fixed with a bare,
byte-bounded "<bodies shown="0" total="N" capped="0|1"></bodies>" tag --
deliberately NOT routed through packBodies here: an earlier version of
this fix did call packBodies (truncateOversizedFirst=false, reusing its
per-item "body omitted (over budget)" comments), but those comments are
UNBUDGETED and pushed a --token-budget=2000 run from 5428 B allowance to
5692 B (packtaskcheck.sh's own ceiling arm caught it). The bare tag is a
small, fixed cost, same shape restatePackTaskBodiesWrapper already
hand-formats a few lines below for the same reason.
src/packtask.h::packTaskListSection (backing <callers>/<far>/<notes>/
<tests>) has the identical defect and is DELIBERATELY left alone here --
a separate, larger blast radius; a follow-up task is spawned for it.
test/packtaskcheck.sh's pre-existing "budget EXCEEDED" failure at
--token-budget=2000 is ALSO pre-existing and unrelated (verified against a
clean pre-lane baseline binary); a separate follow-up task is spawned for
that too.
Gate: test/bodiesshowncheck.sh (red-first; registered in
test/regression.sh in this commit). Six arms across all four scenarios
(--for/--pack-task x candidates-exist-but-none-fit/no-candidates-at-all):
<bodies shown="0" total=N capped=C> present and arithmetic, the existing
bundle=/reason= attributes unchanged, the non-empty case still reports a
real shown=N (proving this isn't an always-shown=0 false pass),
well-formed + deterministic output. DECISIVE ARM PROVEN: run against a
pre-fix binary (git-stash-isolated build of this commit's changes) --
4 of 6 arms FAIL there with "no <bodies shown=\"0\" ...> tag found — the
element is ABSENT (the exact R9 bug)".
Existing gates updated to match the new, more honest shape (all verified
against the pre-fix baseline to confirm they previously encoded the OLD,
buggy behavior as correct):
- test/forautobodycheck.sh: the "no body fits whole" arm asserted
<bodies> was WHOLLY ABSENT; now asserts shown="0" total="N" capped="1"
is PRESENT.
- test/packtaskcheck.sh: the tiny-budget arm asserted a bare "omitted
(budget)" report line with no <bodies> element; now asserts <bodies
shown="0" total="N" capped="1"> is present, and the report line reads
"kept 0 of N" (packTaskBundleText's listStatus() derives the label
from bodiesStr.empty(), which is no longer true on this path -- "kept
0 of N" is a more honest label than the retired "omitted", not a
regression).
- test/estchargecheck.sh: the weak="1" arm assumed a flat-rate
(bytes/2.50) document; a weak query now also carries the empty
<bodies> shell (charged at the body rate, 3.80), so it switched from
selfconsistent() to mixedconsistent() (already proven correct for the
non-weak auto-bodies case just above it in the same file).
Verified: truncvocabcheck.sh (--lint payload cap), sarifcheck.sh,
lintscopecheck.sh, skippedcheck.sh, ripwirepubliccheck.sh all still ALL
PASS on this change.
…ment Two items from the stack-graphs recon lane (#1 RefRole::Type use-sites, #2 namespaceCompatible candidate filtering), result-free: baseline edges=13755 ambiguous=5466 re-derived on this lane's own tree, the six success criteria, the failure criteria that revert each item, and the pre-fix reference binary the gates are recorded red against. Item #3 of that report (depth-2 receiver-chain resolution) is out of scope for this round and is not registered here.
…ing it Result-free registration for lane J2 (item #3 of the stack-graphs recon). Baseline re-derived on this lane's own cd30104 build against a pristine detached checkout; success criteria #3a-#3e and the failure criteria that revert the lane are fixed before any feature code exists. Two recon claims re-derived and recorded as corrections: the composeEdges hoist is unnecessary (buildFieldNarrowTables already runs ahead of the resolve loop over the same isCompose stream), and Rule 1's bareCish path currently WRONG-narrows `this->field.method()` as a bare unqualified call — removing that can raise ambiguity, which is why #3b is a whole-tree measurement rather than an assumption.
…ng it The depth-2 lane's REJECT (lane/depth2-chains @ 21f75a9) found two real resolver bugs and an instrument that could not see them: ambiguous= nets a removed wrong pin against a recovered call. This round registers the instrument the lane's finding asked for — recovered-edge and corrected-target counts via a fixture gate (test/chainguardcheck.sh) and a two-binary whole-tree audit (test/edgediff.py) — plus exact whole-tree predictions taken from the lane's own measured capture-widening-only arm. Scope re-justified from c3ebbd8 rather than restored: the widening alone IS the fix (five recv==None guards stop misfiring); Rule 4 stays out per the lane's finding #3. Result-free by construction; the RESULT lands after the instrument runs. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…itized tier A shipped-but-not-authored file recites a subsystem's vocabulary in a three-line document, which is the shape BM25's length normalization rewards most — so a vendored admin script and a numbered migration took two of five top slots on a database-backend task, and the module an agent would actually edit never appeared. They now carry the SAME 0.35 factor decks and generated captures already had: down-weighted, never dropped, still findable by name. Extends the existing table rather than growing a second component list, and carries the existing factor rather than a second uncalibrated knob. Two rows were DELETED from the first draft of that table after checking whether they could ever fire: vendor/, node_modules/ and dist/ are already pruned whole by the crawl, and *.min.js is already denylisted by filename, so tier rows for them would be dead weight that reads like coverage. test/vendoredassetcheck.sh pins that reasoning as its last arm — if a *.min.js file ever becomes a symbol, the arm goes red and the rule becomes worth adding. Gate is red-first on the lane base (the migration is #1, the asset #3, the real implementation last) and green here, and it carries the two near misses a prefix match would break: migrations/__init__.py is unnumbered and staticfiles/ is not static/. Measured on the registered external slice: vendored slots in top-5 6/60 -> 0/60 (band [0,2]), gold-in-top-5 6/12 -> 7/12 (guard band [-1,+4]). All simultaneous floors byte-identical.
da61bac --test-gate=da61bac..HEAD (def-over-decl lane, finding: a ref range the verb cannot parse) reported changed="0" at exit 0 — a silent zero on unparseable input, breaching non-negotiable #3. New testgaterefusecheck.sh holds the verb to --quality-delta's refusal standard (qdrefpaircheck arms (C)): exit 1, the offending token NAMED, an adjacent probe offered — over A..B, A...B, the half-typed --test-gate=, a no-such-file token, a mixed list, and a comma-only list. testgatecheck arm (d), which pinned the defective silent zero, flips to the refusal; the honest changed="0" keeps coverage via a new clean-git-tree arm (d2). Verified red before the code: 7 of 8 arms fail on the da61bac binary (sha256 275ecfbb…), only the valid-FILES control is green. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…nt changed=0 --test-gate=da61bac..HEAD (a ref range the verb has no grammar for) reported changed="0" at exit 0 — read by a caller as 'your change touches nothing', a non-negotiable #3 breach found by the def-over-decl lane. Now held to --quality-delta's refusal standard (qdrefpaircheck arms (C)): exit 1, the offending token named verbatim, an adjacent probe offered — the range shapes (A..B, A...B) get the git diff --name-only expansion that turns the range into the FILES the verb takes; a no-such-file token gets the --skipped pointer; refusal is per-token, so a mixed list cannot hide a bad item behind a good one. The half-typed --test-gate= flips EmptyValue::Meaningful→Refuse (same ruling as --dmm=/--quality-delta=: an unset shell variable must not silently gate a different question). changedMaskFromList grows a checked twin (ChangedList: mask + first bad item + item count) and delegates to it — one loop, two surfaces, no drift; --situ and the MCP verb keep their lenient contract. The refusal body lives in situ.h::testGateRefusesFileList; the dispatch keeps only the branch, acked (+3 ccx) in .ripwire_quality_acks per the parseArgs flag-arm precedent. testgaterefusecheck.sh: 8/8 arms green (7 red at da61bac). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… old walk Makes gitignorecheck (RED at 4f6e601) green. ripwire is named for ripgrep, whose defining default is that ignored files are not searched; the crawl walked them. On this repository's own root that meant crawling every gitignored checkout under bench/external, which is where the 2 GiB-cap cache thrash and the 300 s gate timeouts of several past rounds came from. THE WALK, AND WHY IT SHELLS OUT. One `git -C <root> -c core.quotepath=false ls-files --others --ignored --exclude-standard --directory -z` per root, not a .gitignore matcher of our own. A native matcher has to be bug-compatible with git across nested .gitignore precedence, negation, **, the trailing-slash directory form, core.excludesFile, .git/info/exclude and macOS case folding — and the moment it diverges it deletes source from a corpus while claiming the repository asked for that. Measured: 0.020 s (ugrep), 0.030 s (rocksdb), 0.150 s (duckdb), 0.032 s (this root). `--directory` is what keeps that flat — a wholly ignored directory is ONE entry and ONE prune, never an enumeration. This is not a new dependency: gitmine.h, crossref.h, prcontext.h, quality.h and binstale.h already popen git, and G3 is about what the binary LINKS. Its absence degrades the same way — full walk, and --skipped says so. WHAT IS DISCLOSED, because a corpus that shrinks silently is the defect this replaces. ignored_files= (exact: files the rules covered) and ignored_dirs= (subtrees they pruned — the walk stopped there, contents UNKNOWN not zero, the excluded_dirs=/pruned_dirs= contract), both ABSENT when nothing was cut, so test/golden.xml is byte-identical and every argvdiff vector survives. The map legend clause is charged only to a map that carries the attributes (tokenbudgetcheck arm #3 leaves seven bytes of headroom on `src`). --skipped rows both classes and names the mode in ignore_mode=: git / off / unavailable / root-ignored. ORDERING IS THE CONTRACT, as it already was for the extension-vs-exclude pair. The ignore test runs LAST — after the extension, the --exclude match and the built-in denylist — so ignored= only ever counts a file that would OTHERWISE HAVE BEEN INDEXED (which is what lets the accounting invariant read indexed= + oversize= + excluded= + ignored=), ignoredDirs= counts only subtrees no existing rule had pruned, and unsupported_ext=/unindexed= mean exactly what they meant before. FOUR CASES KEEP THE FULL WALK. No git work tree, no git binary, --no-ignore — and the one the registration did not name, found by building it: a root that is ITSELF inside an ignored subtree (`ripwire build/` in a repo that ignores build/). git answers "./" — everything — and honouring that returns an EMPTY map for a directory the user pointed at deliberately. That answer is recognised and refused; ignore_mode="root-ignored" says why. A TRACKED file matching a .gitignore pattern also stays indexed, because git ignores nothing it tracks. CACHE: deliberately untouched. The blob is keyed per FILE, so a --no-ignore run writes a SUPERSET and a default run can only read back what it crawled (gitignorecheck arm 10 pins it). The residual is speed, not correctness, and it is in the ledger — same shape --exclude has always had, which is why neither is in the cache key. Files: src/ingest_crawl.h (the probe + the two extracted recorders), src/model.h (IgnoreMode + the two counters and row vectors), src/ingest.h + src/ingest.cpp + src/main.cpp (the flag reaches the crawl through one trailing defaulted parameter), src/cli.h (--no-ignore, help, flag-arm ledger 204 -> 205), src/serialize.h (buildIgnoredAttr + the conditional legend), src/verbs_report.h (--skipped header, rows, legend), src/workspace.h (per-root merge; the merged mode is the WEAKEST across roots, seeded from the first part). Also: test/deckcheck_allowlist.txt retires the --no-ignore exemption in the commit that ships the flag, and .ripwire_quality_acks records the seven gating quality-delta findings with the reason (one deliberate contract change, three self-churn rows, and what remained of collectSources after probeIgnoreSet/recordDirPrune were extracted: +3 ccx / +11 LOC, from +15/+43 inline). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…at 2 Battery #3 on the merged round-5 tree reddened two gates that are not about goto at all. Both pinned the literal 2 as scaffolding for what they actually test, and the flow-sensitive slice lane added a THIRD goto in a fixture — test/sliceflowsensfix/disclosed.cpp's cd03, the deliberately disclosed "goto falls through, untracked" case the lane's legend now names. - matchcapturecheck asserts that an EXPLICIT @capture is left alone (no auto_captured=) and that a trailing `;` comment does not defeat the scan. It now derives the population once, from the same --match it is validating, and fails loudly if that derivation yields nothing. - lintbudgetcheck asserts that a built-in rule's count stays a clean total beside a saturating same-named USER rule. It already computed GOTO_TRUTH for the arm above; the cap arm now uses it too. This is the fixture-addition gate footprint, the same shape as the flag-addition one: a gate that pins a whole-tree population makes every future fixture its business. Neither gate lost coverage — both still fail if the invariant they exist for breaks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… old walk Makes gitignorecheck (RED at 1bbd9ea) green. ripwire is named for ripgrep, whose defining default is that ignored files are not searched; the crawl walked them. On this repository's own root that meant crawling every gitignored checkout under bench/external, which is where the 2 GiB-cap cache thrash and the 300 s gate timeouts of several past rounds came from. THE WALK, AND WHY IT SHELLS OUT. One `git -C <root> -c core.quotepath=false ls-files --others --ignored --exclude-standard --directory -z` per root, not a .gitignore matcher of our own. A native matcher has to be bug-compatible with git across nested .gitignore precedence, negation, **, the trailing-slash directory form, core.excludesFile, .git/info/exclude and macOS case folding — and the moment it diverges it deletes source from a corpus while claiming the repository asked for that. Measured: 0.020 s (ugrep), 0.030 s (rocksdb), 0.150 s (duckdb), 0.032 s (this root). `--directory` is what keeps that flat — a wholly ignored directory is ONE entry and ONE prune, never an enumeration. This is not a new dependency: gitmine.h, crossref.h, prcontext.h, quality.h and binstale.h already popen git, and G3 is about what the binary LINKS. Its absence degrades the same way — full walk, and --skipped says so. WHAT IS DISCLOSED, because a corpus that shrinks silently is the defect this replaces. ignored_files= (exact: files the rules covered) and ignored_dirs= (subtrees they pruned — the walk stopped there, contents UNKNOWN not zero, the excluded_dirs=/pruned_dirs= contract), both ABSENT when nothing was cut, so test/golden.xml is byte-identical and every argvdiff vector survives. The map legend clause is charged only to a map that carries the attributes (tokenbudgetcheck arm #3 leaves seven bytes of headroom on `src`). --skipped rows both classes and names the mode in ignore_mode=: git / off / unavailable / root-ignored. ORDERING IS THE CONTRACT, as it already was for the extension-vs-exclude pair. The ignore test runs LAST — after the extension, the --exclude match and the built-in denylist — so ignored= only ever counts a file that would OTHERWISE HAVE BEEN INDEXED (which is what lets the accounting invariant read indexed= + oversize= + excluded= + ignored=), ignoredDirs= counts only subtrees no existing rule had pruned, and unsupported_ext=/unindexed= mean exactly what they meant before. FOUR CASES KEEP THE FULL WALK. No git work tree, no git binary, --no-ignore — and the one the registration did not name, found by building it: a root that is ITSELF inside an ignored subtree (`ripwire build/` in a repo that ignores build/). git answers "./" — everything — and honouring that returns an EMPTY map for a directory the user pointed at deliberately. That answer is recognised and refused; ignore_mode="root-ignored" says why. A TRACKED file matching a .gitignore pattern also stays indexed, because git ignores nothing it tracks. CACHE: deliberately untouched. The blob is keyed per FILE, so a --no-ignore run writes a SUPERSET and a default run can only read back what it crawled (gitignorecheck arm 10 pins it). The residual is speed, not correctness, and it is in the ledger — same shape --exclude has always had, which is why neither is in the cache key. Files: src/ingest_crawl.h (the probe + the two extracted recorders), src/model.h (IgnoreMode + the two counters and row vectors), src/ingest.h + src/ingest.cpp + src/main.cpp (the flag reaches the crawl through one trailing defaulted parameter), src/cli.h (--no-ignore, help, flag-arm ledger 204 -> 205), src/serialize.h (buildIgnoredAttr + the conditional legend), src/verbs_report.h (--skipped header, rows, legend), src/workspace.h (per-root merge; the merged mode is the WEAKEST across roots, seeded from the first part). Also: test/deckcheck_allowlist.txt retires the --no-ignore exemption in the commit that ships the flag, and .ripwire_quality_acks records the seven gating quality-delta findings with the reason (one deliberate contract change, three self-churn rows, and what remained of collectSources after probeIgnoreSet/recordDirPrune were extracted: +3 ccx / +11 LOC, from +15/+43 inline). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…at 2 Battery #3 on the merged round-5 tree reddened two gates that are not about goto at all. Both pinned the literal 2 as scaffolding for what they actually test, and the flow-sensitive slice lane added a THIRD goto in a fixture — test/sliceflowsensfix/disclosed.cpp's cd03, the deliberately disclosed "goto falls through, untracked" case the lane's legend now names. - matchcapturecheck asserts that an EXPLICIT @capture is left alone (no auto_captured=) and that a trailing `;` comment does not defeat the scan. It now derives the population once, from the same --match it is validating, and fails loudly if that derivation yields nothing. - lintbudgetcheck asserts that a built-in rule's count stays a clean total beside a saturating same-named USER rule. It already computed GOTO_TRUTH for the arm above; the cap arm now uses it too. This is the fixture-addition gate footprint, the same shape as the flag-addition one: a gate that pins a whole-tree population makes every future fixture its business. Neither gate lost coverage — both still fail if the invariant they exist for breaks. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
… — --situ joins its three siblings
H6 (capture-audit 2026-09-04, lens 6 F1). `--situ=src/nosuch.h` printed "0 changed file(s) — nothing to
analyze" at exit 0, and the MCP twin `situational_awareness{files:…}` returned all-empty arrays with a green
`_fresh: ok`, while --affected / --test-gate / --exercises refused the identical path with flag + problem +
example. An agent that misspells the file it just edited was told its edit has no blast radius and no tests
to run — the false zero non-negotiable #3 forbids, and the JSON form is the worse of the two because a
result reads as an ANSWER to every caller that only checks for an `error` key.
The refusal was not missing, it was wired to one arm: situ.h's testGateRefusesFileList already spoke the
whole triple over the SAME ChangedList and the SAME grammar. It is now parameterised by the surface's own
spelling (fileListRefusalText: `lead` + `selector`) so --situ, --test-gate and the MCP `files` field share
one text, and the MCP arm returns it in a -32602 instead of an empty report.
Second half: the path near-miss. The MCP `cochange` refusal was the only surface in the tool that suggested
the nearest indexed PATH ("did you mean './src/graph.h'?"); every CLI FILE selector had none, and --affected
ran the SYMBOL suggester over a path and answered `--affected=tow.c` with "did you mean 'TOOLS'?" (F18) with
`two.c` sitting in the file list. The candidate search moves to didyoumean.h::nearestIndexedFile — one
suggester, both surfaces — and --situ / --test-gate / --affected / --exercises / --cochange all append it.
--cochange's CLI refusal also gains the example its MCP twin had (F11).
Gate: test/fileselectorrefusecheck.sh (new, registered in test/regression.sh). RED on the pre-fix binary —
A/B fail on --situ (exit 0, no refusal), C fails on all five (no did-you-mean), E fails (MCP result with
empty arrays). GREEN after. testgaterefusecheck / affectedcheck / situdiffcheck / cochange* / didyoumeancheck
/ selectorrefusecheck / mcpcontractcheck all still pass.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…(sites_l=), not just where the caller is defined capture-audit 2026-09-04, finding P3 / M21 (lens 8 #3, lens 0). A `<c>` row's p= is the CALLER'S DEFINITION line — `<c n="benchScores" p="bench/bench_radix_ab.cpp:133" incompatible="1"/>` for an edit whose sites are 146 and 152. The lines an agent has to open were printed by exactly one other verb, --uses, so "the contract moved" -> "where do I fix it" cost a second call whose answer this document already held. editCheckCallSites() is ONE pass over ing.references plus one sort — never a per-row rescan (editCheckIncompatibleFlags' own rule; a widely-shared name has hundreds of caller rows) — and it is paid for only when a row will carry the attribute. Its filters are --uses' OWN (collectUseSites): role=="call", never a compose edge, a doc mention or a wikilink. That is what lets the gate assert the two verbs equal without either knowing about the other. It deliberately does NOT filter on argCountKnown: a spread call site cannot be PROVEN incompatible but is still a line to open, and --uses prints it — the legend says so. SPELLING: the finding named the attribute `sites=`. It is emitted as `sites_l=` instead, because `sites` is already this tool's word for a COUNT of reference sites (--context-ratio's <s>/<f> rows, --uses' amb_sites=/call_sites_of_name=) and a LINE LIST under that spelling is a new same-name-different-type collision — the §P8 class test/attrvocabcheck.sh exists to prevent, and whose stated rule is that the meaning with zero consumers moves. `_l` is the tool's line-number suffix (--hotspots' top_l=). Reasoning recorded here, not silently. Gate: test/editcheckcheck.sh arm (k) — the PROPERTY, not a literal. Every row flagged incompatible="1" must carry sites_l= equal to the ascending unique set of --uses role="call" lines whose in_id names that caller, over a fixture that exercises both the multi-site and the single-site shape (asserted non-vacuous), plus the counterpart rule that an unflagged row carries no sites_l=. RED on the pre-fix binary: caller twice: sites_l="" but --uses role=call lines are "5,6" caller once: sites_l="" but --uses role=call lines are "12" FAIL (k) incompatible rows do not carry the call sites --uses prints GREEN after. MCP edit_check rides the same assembler (mcpclidiffcheck green). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…oom, --tree, --external-surface; gate defaultceilingcheck
Capture-audit 2026-09-04, lane L7, plan item P4 (§5a decision 4 accepted: the first rows stay, the tail is
windowed and disclosed). Lens 8 measured the defaults on this tree: --pr-context=main~1 659,824 B (~165K tokens,
217 files), --zoom 433,867 B, --tree 187,209 B, --around=SYM 61,891 B, --external-surface 67,862 B led by sh
builtins. Usage data (§5a): --pr-context=REF 0 uses, --around 9 in 20 days.
--pr-context budgeted BY DEFAULT: kPrDefaultBudgetTokens = 8000 (budget_tokens= budget_default="1" on the
root; --token-budget/--max-tokens set it explicitly). The trim ladder runs first (detail, never
files); when even the structural floor of every changed file exceeds it, the FILES are windowed in
blast-radius order (files_shown= + the quintet, next="--pr-context[=REF] --offset=N"); --offset/
--limit page them (pr-context joins the paging set). PrBudget replaces the size_t budget param
(same arity). The shaping guard lets this one verb both page and shape by budget; --top-k on it
is its kShapingVerbs notice, not the paging refusal.
--around default depth 1 (root depth= discloses; --around-depth=2 restores): 63,471 → 6,297 B here
--zoom levels_shown="2" of levels= (a module AT the cut carries children=; --zoom-levels=N, 0 = all)
over the 40 largest top modules (shown=/capped=/total=/next_offset=, next="--zoom --offset=40"):
444,200 → 8,063 B here
--tree the 80 best files (kTreeRowCap; --limit raises): 189,209 → 11,818 B; the plan said 400 — 400 rows
are 54,160 B on this repo, which the same plan's ≤12,000 B gate forbids; 80 ≈ 11.5 KB
--external-surface 100 rows (kExternalSurfaceRowCap) and the sh BUILTINS dropped and counted
(builtins_excluded=; externalnames.h kShellBuiltinNames — POSIX + bash builtins, never grep/sed/git,
which ARE the surface); --include-builtins keeps them: 68,437 → 5,569 B
Every legend defines its new attributes where the reader meets them; every cut carries next=.
Two new flags (--zoom-levels=, --include-builtins; kTotalFlagArms 207 → 209), each refusing alone naming its verb.
Gate test/defaultceilingcheck.sh (regression.sh), RED first on the wave-2 binary (30 rows): each ≤ 12,000 B at
defaults on this repo (pr-context: est_tokens ≤ 8000 on a 120-file fixture, the file window firing and --offset=N
continuing it), every cut disclosed with capped="1" + a next= that runs, the restoring flags restoring, the first 30
rows of --tree/--external-surface byte-identical to the uncapped listing's, every --around row a depth-2 row.
Re-pins, each to the new contract: prbudgetcheck #1 (default budget disclosed) / #2 (large budget: no
budget_default=) / #3 (a file cut is disclosed, never silent); treecheck (default = min(80,total) rows, "drops
nothing" on --limit=1000000); usescheck 5b (the per-language split asserted with --include-builtins; the default's
drop asserted beside it); w3fixlegendcheck §7/§11 (the default external-surface/tree roots are disclosed windows);
seedboundscheck (depth 1); mcpcontractcheck TWIN (--pr-context: no twin); the --max-tokens/--token-budget guard
sentences (pr-context named, budgetpolicycheck's parser shape kept); README flag/gate counts (179 / 537).
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…tet only (no files_shown= twin); paging sweep enumerates it pageview.h rule 6: the noun-prefixed shown_/capped forms are for SECONDARY listings; the changed files are the pr-context root's primary listing, so a cut carries shown=/capped=/total=/has_more=/next_offset=/offset=/limit= and nothing else — P4 had emitted files_shown= beside pageDisclosure's shown= for one fact (two names, the M15 class). Legend, --help, defaultceilingcheck (5) and prbudgetcheck #3 read shown=. pagingsweepcheck (L) table gains --pr-context (it joined the --help HONORED-by set), so the M2 quintet rule now covers its window too. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ers a confident zero (F3) `--affected`'s file reading is a substring PATTERN, so one item can match several files: `geo.py` also matches `test/check_geo.py`. Every matched file's symbols became seeds, and transitiveCallers returns "reached, MINUS the seeds" — so the very tests that reach the change were subtracted out of their own answer. On lane-L8's sandbox (2026-09-04, found-not-fixed #1) `--affected=geo.py` reported seeds="6" tests="0" reached="0" while `--affected=./geo.py` reported reached="1" over the same bytes. A confidently wrong ZERO in the verb whose whole job is telling an agent which tests to run, with the reason it was wrong nowhere in the output — non-negotiable #3's class exactly. DECIDED: ANSWER, do not refuse (METHODOLOGY §9 #3 — a disclosed answer is terminal, a refusal costs the agent a round trip to learn a spelling). The two facts that put a test file in the answer are different and are now kept apart: - a test cannot REACH a change it is part of → its own symbols are not seeds of the caller walk; - the argument MATCHED it → it changed, so run it, and its row says so with seed_kind="test". Dropping those symbols from the seed SET never hides them from the traversal — they are still found as callers — so the answer is byte-identical for every argument that matches no test file. The partition runs over the DEDUPED seed set and therefore over both readings: a bare symbol name that resolves into a test file is the same fact (`--affected=check_geo.cpp` now answers tests="1" reached="0", "run the test you changed", instead of tests="0"). On this repo, `--affected=test/affectedcheck.sh` now names that gate with run="bash test/affectedcheck.sh" instead of reporting no tests at all. Honesty in attributes, not prose (§9 #4): root gains seed_test_files="N" (a zero means none matched), rows gain seed_kind="test"; both defined in the legend and in --help. Structure: the partition (partitionAffectedSeedsByTestPath) and the answer assembly (affectedAnswer/AffectedAnswer) live in testmap.h next to the seeding they constrain, which also removes the two complexity regressions the inline version carried (runAffected 21→29 → back under, resolveAffectedSeeds 21→25 → back under). Gate: test/affectedcheck.sh gains a colliding fixture pair (src/geo.cpp + test/check_geo.cpp) and 11 arms — the collision spelling must answer the same tests/reached as the unambiguous one, seeds= must prove the pattern really matched more (non-vacuity), seed_test_files= 1 vs 0, seed_kind= on the matched row and NOT on a merely-reached row, the test-only pattern, the legend, and the standing rule that tests="0" while seed_test_files>0 is a failure. RED pre-fix: 8 of 11 arms fail, headed by "--affected=geo.cpp disagrees with --affected=src/geo.cpp: tests=0/1 reached=0/2". GREEN: 33 PASS. Verified: affectedcheck testscopecheck rootrelemitcheck testrowruncheck receiptpostcheck exercisescheck legendcoveragecheck compactlegendcheck a9disclosurecheck floormarkcheck runhintcheck selectorchaincheck fileselectorrefusecheck shellgateindexcheck testmacrocheck edithandlehintcheck dispatchordercheck w3fixlegendcheck argvdiffcheck flagtablecheck — all rc=0, no gate needed re-pinning. Determinism ×2 byte-identical, xmllint clean, ASan/UBSan/LSan clean on seven --affected paths (collision, test-only, unambiguous, multi-item, symbol reading). --edit-check: resolveAffectedSeeds and runAffected both status="unchanged". --quality-delta regressions=0 gating=0 (7 rows acked, both reasons this lane's own). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…de batch too (M1) The capture-audit landed the compact dialect OPT-IN (§5a decision 3, P1) and registered the default flip as follow-up #3. Opt-in made the right posture a feature an agent has to already know about: measured on this repo, the ten-verb MCP edit loop was paying 30,839 B of legend prose per pass and 2,866 B is the same answer. So `legend` absent now means COMPACT, `legend:"full"` restores the historic legend byte-for-byte, and `legend:"compact"` stays accepted (a no-op spelling of the default — refusing a request for what you already do is the worst kind of refusal). THE BILL, this repo's src/, one call per verb, MCP, `legend:"full"` vs the new default: verb legend full legend dflt whole doc full whole doc dflt doc saving edit_check 5,863 357 11,703 6,228 46.8% impact 3,660 434 7,824 4,625 40.9% uses 3,807 387 12,903 9,508 26.3% slice 4,500 281 4,917 724 85.3% whereis 4,496 294 15,831 11,657 26.4% flags 854 159 17,047 16,378 3.9% doc_drift 5,291 199 18,271 13,209 27.7% lego 931 249 1,261 604 52.1% exemplar 687 243 1,024 609 40.5% path_between 750 263 1,048 586 44.1% TEN-VERB EDIT LOOP 30,839 2,866 -90.7% analyze 1,349 483 43,494 42,652 1.9% explore 1,738 373 12,502 11,167 10.7% from_trace 2,002 290 5,024 3,343 33.5% connect 1,400 345 1,810 783 56.7% owners 745 201 878 361 58.9% stray_content 2,267 173 17,648 15,588 11.7% batch 363 142 21,224 14,435 32.0% ALL 17 40,703 4,873 194,409 152,457 21.6% ~7.0K tokens off a ten-verb loop, per pass, per session. SCOPE: exactly the seventeen verbs that DECLARE `legend` (mcpVerbDeclaresLegend, read from kMcpVerbFields — the same table the unknown-field refusal dispatches from, so a verb joins the family with one edit). The other fourteen are untouched: their payloads are JSON or plain text with no legend, or — `for` — carry a header that is already its own short legend whose comments are DATA (A10). A default that silently rewrote those would be changing answers the argument was never offered for. Inside `batch` that distinction is load-bearing, not cosmetic: batchcheck (a) asserts every sub-answer is byte-identical to its STANDALONE verb, and compacting a sub-answer whose standalone twin is full breaks exactly that. TWO THINGS THE FLIP FORCED, both of them real fixes rather than accommodations: 1. THE POSTURE REACHES INSIDE CDATA (applyCompactToBatchSubs, called by BOTH batch front doors). Both surfaces apply the dialect to the finished document and both must leave CDATA alone — a CDATA section is DATA and may hold JSON or plain text. So a compacted batch shipped its own legend short and up to sixteen sub-answers' legends long, on the one verb where legend duplication costs most. The CLI had the same hole under `--batch --legend=compact` and it is closed in the same commit (the owner's CLI-outranks-MCP ruling): measured on the fixture, uses+slice, 8,840 B full -> 8,645 B envelope-only -> 2,282 B with the sub-answers. `slice` takes BOTH steps, in the standalone dispatch's order — its own emitter builds the compact form (runBatchSub's new compactLegend parameter, which decides where schema= sits on the root) and the layer then replaces the legend TEXT. Doing only the first step gave a batched slice 1,542 B against the standalone verb's 606 B: identical rows, different legend. batchcheck (h) caught it. 2. dropped_positive= MOVED FROM PROSE TO THE ROOT (packtask.h). A2 put --pack-task's dropped_positive in the ledger COMMENT while its --for twin put it on the <ctx> root, and the asymmetry was invisible until compact became the default: the dialect strips prose, so a COMPLETENESS fact vanished from the default answer — measured on src/ at budget_tokens=300, dropped_positive="2" under legend:"full" and NOTHING under the default (droppedpositivecheck #6 went red on exactly this). METHODOLOGY §9 #4: honesty lives in attributes, precisely so a legend posture can never decide whether a caller is told. Same spelling as the --for twin (one name on both roots, mcpattrparitycheck), spliced onto BOTH root spellings including the ladder's route-dropped rebuild and the partition slice, and defined in the compact legend's completeness vocabulary. The ledger's per-section counts stay prose, which compact is entitled to drop. WHERE AN AGENT READS IT: the `legend` field description, "a STRING legend posture: compact (the default) or full" — one string spliced into all seventeen inputSchema stanzas, which is the per-tool sentence an MCP client actually shows, and it names both values and the default in the SAME 54 bytes the old wording spent. tools/list is byte-identical at 40,841 B, measured before and after. Deliberately NOT a clause added to seventeen tool descriptions: mcpmanifestcheck's own registered rule is that the ceiling moves for a declared argument's obliged description and never for prose, and prose there would cost ~680 B against 159 B of headroom. --help's --legend paragraph states the split (MCP default compact, CLI default full). GATE, WRITTEN FIRST: compactlegendcheck arm (N), the family read from src/mcprefusal.h rather than listed, so a verb joining without a call in the gate FAILS instead of skipping. Four assertions per verb: every posture answers; the DEFAULT is byte-identical to legend:"compact"; legend:"full" carries strictly more legend bytes; the PAYLOAD is byte-identical between postures. RED pre-flip: all seventeen — "(N) analyze: the DEFAULT is not compact — default 1161 B of legend, compact 290 B" and sixteen more. GREEN post-flip: all seventeen, e.g. "(N) edit_check: default == compact (357 B legend), legend:\"full\" restores 5863 B, payload byte-identical". RE-PINS — twelve gates, each to the NEW contract, stated precisely, in this commit (rule 10b: the cost is recorded here, it is not a reason to keep the shape). Two different re-pins, and which one a gate gets is decided by what the gate ASSERTS, not by which was easier: · compact <-> compact, because the gate compares the two surfaces' DEFAULT paths and that is now compact on one side: mcpslicecheck (13) and sliceflowsenscheck (7) add --legend=compact to their CLI operand (sliceflowsenscheck keeps $FB full for arms (5)/(6) and takes a second, compact operand). · full <-> full, because the gate reads the FULL document by construction and compact <-> compact would make it compare nothing: mcpclidiffcheck (LENS 1 extracts the first `<ELEM attr="…">` tag, and a compact legend names its elements in bare shorthand `<iface n= p= defs=>` — zero attributes, and the gate's own "the probe is broken" guard would fire; LENS 3 quotes a CLI legend CLAUSE the compact dialect does not carry), mcpw2fixcheck, mcptranchecheck, mcpw3fixcheck, floormarkcheck (it compares disclosure TAILS), mcpattrparitycheck (whose slice block gains a `default (M1: compact)` row asked raw — the flip itself, pinned across both surfaces), mcpverbscheck (two shape assertions quote the full document's shape; the pack_task alias arm now also proves that `legend` survives alias resolution), editpreviewcheck (N). · batchcheck: three root assertions that pinned attribute ORDER (`<batch n="16" requested="20" …`) are re-pinned to the FACTS, asserted individually in any order — strictly stronger, since the old spelling would also have passed with two of those numbers swapped. · compactlegendcheck (U): --batch joins --for as a named exemption from the atomic-payload arm, for the opposite reason to --for's — its CDATA holds sub-answers whose legends legitimately moved. The property is not dropped: arm (N) re-asserts it on the batch after flattening the CDATA and normalising the two attributes leg.py already normalises on a single root (schema=, est_tokens=). VERIFIED: all 98 --mcp-touching gates green (regression.sh excepted — it is the meta-runner). --edit-check runBatchSub: status="contract-change" change="params" callers="2" incompatible="0". --quality-delta regressions="0" gating="0" acked="24", every ack my own row with its reason in the ledger. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ob_sha, ONE next=; the stderr line names it and nothing else
Terminality round A (2026-09-05), lane E, E2. On a runner without a Read-before-edit policy the agent stops on the
receipt (baseline: 12/12 windows terminal). So the receipt must hold what a Read would have shown, or the stop is blind:
region {start,end,context,text,capped:false} — the applied lines plus context=3 each side, the bytes as they are
ON DISK; over kReceiptRegionBudgetBytes (2048) it carries head + tail (whole lines) + elided_lines with
capped:true — never a silently truncated text
blob_sha git's blob id of the written bytes (== `git hash-object FILE`; infra/gitblob.h, a hand-rolled SHA-1 as git
IDENTITY, zero dependencies) — "what is on disk" without reading the file
next exactly ONE pasteable follow-up (METHODOLOGY §9 #3), read off the fold: a contract-change with broken
callers → --uses=FILE:SYM; else the first tests_to_run run= recipe; else --test-gate=FILE; under
--no-post-check → --edit-check=FILE:SYM (the one call that shows the state). region and blob_sha need no
index and ride under --no-post-check too.
Same keys on the MCP twins (one engine). The stderr line now repeats the receipt's next and names no second command
(the old line prescribed --edit-check AND --affected — two calls the receipt already answered). The `ripwire wrap`
blurb no longer says "then --edit-check=SYM" after an edit (it coached the redundant check the meter counts).
Gate: test/receiptpostcheck.sh arms 10-17 (NEW, RED pre-fix on all eight): region == disk bytes, blob_sha == git
hash-object, one next= that runs and follows the rule, the stderr names only it, MCP key-set parity, the budget
head/tail/elided_lines, --no-post-check keeps the free half, the insert family, the blurb. Re-pinned in this commit
to the one-next contract: edithandlehintcheck (the printed next is the receipt's, names the RESOLVED file, never a
handle, and RUNS — asserted verbatim under --no-post-check where next IS --edit-check=FILE:SYM), rootrelemitcheck
ARM7 (next read off `next:`, its spelling == the receipt's file, and --edit-check=<receipt file>:sym runs).
postCheckJson: +1 defaulted out-param (callers 2, incompatible 0); runEditVerb unchanged.
--quality-delta regressions=0 gating=0 (22 rows acked by exact name with reasons). 17-gate sibling sweep green.
Receipt deterministic x2 (stale_index stamp aside), valid JSON.
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
… of its body (V1 R2+N4) Gate first: test/prbudgetcheck.sh arm #9 recounts the delivered bytes at the tool's one markup rate against the number on the root, for BOTH --pr-context roots and every budget rung. RED on 044d4d0 with nine failing rows: (#9) empty-diff root (clean tree): est_tokens=278 but the document is 4894 B = 1958 tokens (17.60 B/tok) (#9) working-tree root (1 file): est_tokens=517 but the document is 5463 B = 2185 tokens (10.56 B/tok) (#9) default-budget root (8000): est_tokens=1123 but the document is 7670 B = 3068 tokens ( 6.82 B/tok) (#9) --max-tokens 100000/20000/8000/4000/2000: 1123 vs 3061/3061/3060/3060/3060 GREEN after, delta 0 on all eight roots: est_tokens == round( emitted bytes / 2.50 ) exactly. WHY the number was wrong. prEstTokens() modelled ceil( (BODY bytes + kPrHeaderOverheadBytes 560) / kMinBytesPerToken 2.36 ). The ~4.6 KB legend -- which IS the document on a clean tree -- was in neither term, and the 560 B constant stood in for a root tag plus a legend that are 4,900 B together. So the empty root priced 4,808 delivered bytes at 278 tokens (7.3x under, the wave-1 verifier's R2) and the budgeted root at 1.30x under (N4), both printed beside a budget_tokens="8000" a caller reads them against. --for prices exactly (9,831 B / 3,932 = 2.50); --pr-context now uses the SAME conversion, tokensForEmittedBytes at kBytesPerTokenDefault, not a second rate. HOW. prLegendText / prAnchorNoteText / prRootOpenText RETURN the bytes they used to fprintf, so the emitter can measure what it is about to write; prPriceDocument( PrPriceCtx, body, level, truncated, window ) converts envelope + this root's own start tag + body, converging over its own est_tokens digits in <=4 passes exactly as serialize.h's pricedRootAttr does. pickPrTrimLevel is handed that price instead of computing one, and the empty root is handed the same number instead of modelling its own. kPrHeaderOverheadBytes and prEstTokens are gone -- there is one estimator, and it measures. BUDGET ARITHMETIC -- the ceiling now cuts where it always should have, and at the DEFAULT it cuts at the same content. On prbudgetcheck's 3-file fixture, before -> after est_tokens: --max-tokens=100000 1123 -> 3138 trim_level 0, all 3 files (unchanged content) --max-tokens=20000 1123 -> 3138 trim_level 0, all 3 files (unchanged content) --max-tokens=8000 1123 -> 3137 trim_level 0, all 3 files (unchanged content) --max-tokens=4000 1123 -> 3137 trim_level 0, all 3 files (unchanged content) --max-tokens=2000 1123 -> 2490 trim_level 4 + files windowed shown=1 of 3, floor-exceeded Only the 2000 rung moves, and it moves because the legend alone is ~4,610 B = ~1,844 tokens: a 2,000-token ceiling has ~156 tokens left for the root tag and every changed file's structural floor, so the floor genuinely does not fit. The old number said it fit and then delivered 7,651 B (3,060 tokens) -- 1.53x the budget, silently. The new run says budget-floor-exceeded and discloses the file window. No gate assertion was loosened: prbudgetcheck #3's "est <= budget unless the floor overflows and says so" and #6's monotonicity both hold on the new numbers. Legend re-pinned to the contract: est_tokens= now reads "prices the WHOLE document this bundle emits, this legend included, at the map's markup rate of 2.50 bytes per token, and IS the number the ladder fits, so recounting the delivered bytes reproduces it". Arm #9's last row asserts that sentence is on both roots -- a definition true of one root and false of the other is the drift prRootOpenText exists to prevent. Re-pinned: test/floormarkcheck.sh (10)'s source guard r'"<pr-context%s' -> r'"<pr-context" \+', the same emitter under its new concatenation spelling; the guard is presence-checked, so leaving it stale would have read as a pass. --edit-check: pickPrTrimLevel params 2->4 and prEmptyRootTail 3->4, both contract-change, incompatible="0" (one caller each, updated in this commit). Determinism x2 + xmllint clean on --pr-context and --pr-context=HEAD~1; ASan clean on 7 pr-context paths incl. the budget ladder, the file window and --situ. quality-delta regressions=0 gating=0 (12 rows acked, all this lane's own, with the reason recorded in .ripwire_quality_acks). Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…ways promised --grep's span-tier legend says, verbatim, "tier_unclassified= hits in files nothing classified — always EMITTED, never suppressed", and grepTierAttrs()/grepTierKeys() guarded the field behind `unclassifiedHits > 0`. A comment-only match whose hit files were all classified therefore printed the promise and withheld the field: ripwire <corpus> --grep=ZEROMARK_probe <grep ... complete="1" tier="comment" tier_parsed="2" ...> # no tier_unclassified= The reading is why the promise was made. tier_unclassified= is the qualifier on the tier= LABEL — "0" is the PROOF that the label was elected over EVERY hit, and absence is the one state a reader cannot interpret: it reads equally as "none" and as "the emitter dropped it". Non-negotiable #3, on the tool's own most-quoted honesty clause. Same defect on the MCP twin, where it is worse: an MCP-only agent has no CLI to re-ask from before trusting tier=. Fix: emit it unconditionally, exactly like tier_parsed= two lines above it, inside the block hasDisclosure() already gates. tier_partial= stays conditional — its legend documents its own absence ("its absence beside a tier= means the label is a fact"). Cost, measured (G4): +22 B, and only on a --grep answer that already carries a tier disclosure. Over 15 representative queries on this repo: 7 of 15 gained the attribute, 480,529 B -> 480,683 B = +0.032%. Zero bytes on every other verb and on the default map (test/golden.xml still byte-identical). Gate: test/emittertruthcheck.sh P2.Z — a FAMILY arm, not an instance one. (Z1) the roster of consumer-visible unconditionality claims is DERIVED from src/ (every string literal claiming "always emitted"/"never omitted"/"printed even at zero"), keyed by file+phrase so it survives line moves, and diffed against the expected 11. A new or reworded claim fails until someone decides whether it is true and gives it a probe. Its first question is which member is missing. (Z2) each claim that names an attribute is proven true AT ZERO, on every surface: tier_unclassified (CLI + MCP), register-macro-excluded (XML + JSON), sub_windows, confidence/margin_pct (XML + JSON), lego/compose/routes_total. Presence guards on every probe: an extractor that finds nothing, a corpus that stops electing a non-code tier, or a legend that stops printing the clause all FAIL rather than going quietly inert. RED before the fix on (Z2a) and (Z2b) only; the other four claims already held. No other gate in the tree can see this class: legendcoveragecheck enumerates FROM the emitted document, so a suppressed attribute is structurally invisible to it, and truncvocabcheck's rules are all "if shown= then ...". --edit-check: both symbols status="unchanged" (no contract moved). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…t it can read
THREE POINTS ARE NOT A REGION, AND MIN_HULL_MEMBERS COUNTS POINTS. The only geometry gate on a module
outline was a node count, so three members that happen to sit near a line passed it and drew a hull with
almost no area — a thin coloured streak across the picture that reads as a scratch on the lens rather than
as containment. Three of them were in the django/db/migrations figure at rrf/top-k=120:
`resolve_model_field_relations` (a 107x0 px line, hull area 4 px^2), `add_operation` (178x8) and
`reload_model` (137x13). The first was being dropped SILENTLY — a fully collinear group makes convexHull
return a two-point ring and the caller skipped it without counting it anywhere, so no number on the page
ever moved.
The test is the isoperimetric ratio 4*pi*A/P^2 of the raw hull ring. The hull is convex by construction and
on a convex ring that ratio IS thinness: 1.0 for a circle, 0.785 for a square, 0.605 for an equilateral
triangle — the roundest a three-member hull can be — collapsing to pi/(2*aspect) for a thin one, which
reproduces all three measurements above to three decimals. Chosen over the two alternatives on their own
terms. An AREA floor is the wrong predicate twice: it is not scale-invariant, and it would drop a small
ROUND module (`varint`, 23x18 px, perfectly legible once HULL_PAD_PX has stood the outline off it) while
the streaks it is aimed at are long enough to survive one. An OBB aspect ratio measures the same thing and
needs rotating calipers, O(ring^2), where this is one O(ring) pass over vertices draw() already walks — and
it runs BEFORE the O(N) purity scan, so a group that cannot be a region never costs a walk over every node.
0.20 IS MEASURED. Over both corpora at the figure argv — 23 groups with 3+ members in view, this
repository at top-k=200 and django/db/migrations at rrf/top-k=120 — the ratios form a low cluster
{0.0011, 0.0718, 0.1493} and a body from 0.2879 up. The two widest adjacent gaps in the whole distribution
(2.08x and 1.93x) bracket exactly that band and 0.20 is its geometric midpoint; in shape terms, a triangle
about 8:1 base-to-height. Measured on the RAW ring, not the padded one drawn, so it is scale-invariant:
HULL_PAD_PX is a screen quantity and testing the padded shape would make an outline appear and disappear as
the reader zooms, taking the caption's count with it. On the figure corpus this takes 11 outlines to 8, and
all three that go are the streaks by name.
THE CAPTION SPLITS, because provenance and methodology travel to different places. The stamped strip had
grown to three dense lines carrying arrow semantics, the dash threshold, the label rule, the shape key and
the hull rule on top of the provenance — and the stamp font is FITTED to the bitmap width against an 8 px
floor, so every clause added to it made the provenance it exists to carry physically smaller. Method is
what a caption UNDERNEATH a figure says once, in prose; provenance is what has to survive the picture being
lifted out of the page. So renderProv builds two halves: FACTS (root, ranker, top-k, the counts, the colour
metric, and any state that makes the picture provisional) goes into the bitmap, and METHOD goes to the page
and to a companion .txt the same export writes, where a README author can lift the wording verbatim
instead of paraphrasing it. Non-negotiable #3 is why the last FACTS line exists at all: it names the .txt
AND every channel whose rule moved into it. A trimmed bitmap that said nothing about the trim would be a
picture drawing arrowheads, dashes, shapes and outlines with no way to read any of them — the quiet
omission this block was added to prevent, re-introduced by the fix for it. The hull line now states all
three truncations with a count and a reason each — the cap, the shape drop, the purity drop — rather than
naming one rule and pooling the rest.
AND THE STAMP DESCRIBES THE FRAME IT IS UNDER. Found cutting the README's crop figures: the caption's
counts are the LOADED subset and the camera is free to sit anywhere inside it, so a picture exported after
a zoom stamped "120 nodes / 183 edges" across a frame holding fifteen — a bitmap overstating its own
contents, in the one artifact the caption exists to travel in. draw() now counts node centres inside the
canvas rect (the reading that errs toward the larger number), the stamped half states it whenever it is
less than the view's own total, and draw() re-renders the caption when that number or a hull count moves,
guarded on a change so the settle loop does not rebuild innerHTML sixty times a second. Without that guard
the numbers renderProv states — this count and the three hull counts — are all computed one frame after the
caption that reports them, which is invisible on a settled page and is the previous view's numbers on an
export taken straight after a zoom.
Gate arms in htmlrendercheck, all observed RED on the pre-fix page and green after. (U13)-(U19): the
threshold is a named constant, the ratio is isoperimetric with the perimeter load-bearing (so an area floor
cannot satisfy it), the constant sits inside the band the two corpora measured, both drop counts reach the
caption separately, the uncounted collinear skip is GONE, and the cheap test precedes the O(N) scan.
(W1)-(W14): the stamp reads the FACTS half and no longer the whole caption, the methodology is still on the
page, the export writes the .txt, both files share one basename, the bitmap names the companion and
enumerates what moved, and the control PAIR (W8)/(W9) proves each moved clause LEFT the stamped half and
LANDED in the other rather than being deleted — either arm alone passes over a deletion. (W10) holds the
stamped half inside STAMP_MAX_LINES, counted with grep -o rather than -c after its own control caught two
pushes sharing a line reading as one.
(S8b) and (R11) name a different SURFACE now and say so rather than being weakened: the shape key and the
arrow semantics are still built by renderProv and still travel with the exported picture, in the .txt
instead of the bitmap. (W9a)/(W9b) hold their new home; (W8a)/(W8b) hold that they left the old one.
test/htmlrendercheck.sh declares RIPWIRE_TEST_DEPS: src/htmlexport.h. It greps a page it asks the binary to
emit and never names the source that emits it, so the script_literal rule could not see the link and
--test-gate answered a change to this file with tests="0" — naming no test at all for the one file it has
116 mutation controls about.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…tie, and every join says how much it explains Registered at 1ae72526; RED gate at 52756446. Both parts are additive (G5) and deterministic. PART 1 — informativeness before id. connectSubgraph()'s BFS recorded the FIRST-discovered parent, so among equally-short joins the winner was the one whose FILE sorted first. It now relaxes on ties: a candidate parent at the SAME distance replaces the incumbent iff ( connects, id ) is smaller. Because BFS expands in non-decreasing distance order, every parent at the minimal distance is offered exactly once, so the winner is a minimum over a total order and does not depend on expansion order — the rule REMOVES an order dependence rather than adding one, and the id is still the total final tie-break, so output stays byte-identical run to run. Distance is untouched, so path length, node count and edge count are untouched: this changes WHICH equally-short answer is served and nothing else about its shape. Prim's (dist,minId,maxId) rule over terminal PAIRS is deliberately unchanged — the defect is a parent choice, and on a 2-terminal --connect Prim has no choice to make. The two CSR halves are now ONE relax() step. They differed only in the discovery channel; the tie-break would have made them the same eight lines twice, and a rule with a comparison in it written down twice is how two halves drift. That fold is why --quality-delta reports connectSubgraph complexity 170->173 rather than the 194 the inlined version measured. PART 2 — disclosure (non-negotiable redhat-et#3). Every Steiner <s> row carries connects="N", the number of distinct symbols that intermediary joins in the undirected view the search walks (callers + callees, O(1) off the CSRs). hub="1" marks a row at or above hub_floor= on the root. Both are facts, so the --max-tokens trim never drops them. An honest "the only join found connects 764 other things" beats a confident bare `empty`. hub_floor is DERIVED, not tuned: the smallest D with D(D-1)/2 > edges — the degree at which one node alone joins more distinct symbol pairs than the whole call graph has edges. Integer bisection over one number the map header already prints; 186 here. kConnectRootBytes 200 -> 260: 200 was already short of the widest start tag it claims to bound (228 B measured attribute-by-attribute, 246 with hub_floor=), and short is the one direction that constant may not be. Legends: the full --connect legend defines all three and states the derivation; the compact dialect gets one hub_floor= term that defines hub="1" in the same clause (two terms did not fit the 400 B ceiling, which is the point of that dialect) — connect compact is 392 B. CLI and MCP share the emitter, so both surfaces move together. Worked example, before -> after: --connect=computeQualityDelta,gitFileCommitCountsInDayWindow before <s n="empty" p="src/notes.h:431"/> both call something named empty (764 callers) after <s n="computeDelta" p="src/quality.h:5284" connects="41"/> computeQualityDelta -> computeDelta -> gitFileCommitCountsInDayWindow nodes="3" edges="2" in both: the size did not move, the answer did. Battery 557 gates: pass=554 skip=3 fail=0 (baseline at 52756446: pass=552 fail=2 — this gate RED plus a stale-binary versioncheck). Determinism x2 byte-identical; xmllint clean; connectcorecheck green under ASan/UBSan. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
`ripwire <dir> --around=writeHtml --html=/tmp/h.html` exited 0, wrote no file, and said nothing on stderr. The same argv without --around writes 59 KB. writeHtml() is called from exactly ONE place -- runDefaultMap -- and every navigation and report verb pre-empts the default map in main()'s dispatch chain, so the flag was accepted and dropped. Six verbs were named in the report. A sweep of the whole flag universe (the new gate, below) found SEVENTY-FOUR flag x --html combinations in that state. This is the "accepted and silently ignored" class this repo has closed three times over, and non-negotiable redhat-et#3 is explicit: a caller must be able to tell a no-op from a typo. The MIRROR of this guard already existed and was loud -- --color-by without --html refuses and names --html (cli.h validateModifierGuards). This is the other half. REFUSE, NOT HONOUR -- and why Honouring would mean rendering the verb's scoped node set as the page, and that is a worse answer than it sounds. The page's three views are an overview of Louvain MODULES, a module subgraph, and a depth-bounded ego graph, all computed client-side over the whole selected map. Over a 20-node --around slice the module overview is empty and the ego graph is a re-derivation of the slice itself, so `--around=X --html=F` would produce a page answering a different question from the one --around answers, with no tell -- trading a silent no-op for a silent substitution. The page already does what that caller wants, and since the routing commit it does it BY NAME, so the refusal has somewhere real to point: g.html#node/SYM/2 is the depth-2 neighbourhood. Refusing is also the smaller change -- one predicate, no new emit path. DERIVED, NOT ENUMERATED The 74 verbs are not listed anywhere. htmlPreemptedBy() reuses firstFlagOutside() -- the walk over kBoolFlags/kViewFlags, the same rows parseArgs matched -- against kMapShapingFlags plus a short kHtmlRideAlong list, so the rule is "anything not on the compose list" and a verb added tomorrow refuses tomorrow with nobody editing a table. Every previous closure of this family was done by enumerating members and re-opened on the member nobody enumerated: jsonUnsupportedVerb's 77-arm chain missed 12, the shaping-flag guards missed 18. The one hand-written arm the flag tables cannot see (--export=cc.json) is named explicitly, the same honest cost jsonUnsupportedVerb pays at its own top -- and it was found by the gate, as the last silent row after the walk closed the other 73. The residual risk runs the other way: a new MAP-SHAPING flag would refuse until it is added to kHtmlRideAlong. That direction is the safe one -- a loud refusal on a legal combination is noticed the first time someone types it; a silent drop is not. Arm (B) pins the compose side so the guard cannot over-refuse its own host unnoticed. It lands in refuseInertMainModifiers rather than cli.h's validateModifierGuards for the reason that function's own header already gives: judging "alone" needs main's knowledge (the flag universe walk), which is why --no-redact and --refetch are refused there too. It still runs before any dispatch. THE GATE test/htmlhostcheck.sh is BEHAVIOURAL, not textual, and shell-only. For every flag in the universe derived from src/cli.h (test/flaguniverse.py, the same derivation shapingflagcheck arm F uses), `<verb> --html=OUT` must end in exactly one of two buckets: WRITE (OUT is non-empty) or REFUSE (non-zero exit). Exit 0 with no OUT is the silent class and fails BY NAME. Arm (D) is its mutation control: a stub binary that exits 0 writing nothing must be classified SILENT, and one that exits non-zero must be classified REFUSE, so a zero in arm (C) means "the defect is gone" rather than "the arm cannot see it". Arm (B) is the over-refusal control. Currently write=31 refuse=174 silent=0. --help now scopes --html to the default map and points at the #node/SYM/2 route; docs/COMMANDS.md regenerated from it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…e the named file's own decl/def partner TWO defects the round-C head-to-head found in --situ, both about a relationship the tool already knows and does not say. F2 — the composed co-change zero conflated two different facts. --cochange refuses loudly on an empty window (commits="0" window="18mo@HEAD", exit 1) and --rank-by=churn stamps "(no churn evidence)", but every surface that COMPOSES co-change rendered a bare (0) — the CLI report adding "(none, or no git history)", which means neither "none found" nor "none exists". Non-negotiable redhat-et#3 forbids exactly that. The window and the commit count now travel ON SituationFacts, so the four composing surfaces cannot disclose it four ways or three of them forget it: * --situ [3] — window="…" commits="…" on the header, and the empty case now says which of the two zeros it is * situational_awareness (MCP) — cochange_window / cochange_commits beside `forgotten`, the surface where an empty array reads most like an answer * --handoff <heuristic> — cochange_window= / cochange_commits= (appended AFTER n=/candidates=/capped=; `<heuristic n="…"` is a shape gates read positionally) * --pr-context <cochange> — had commits= and no window to read it against F3 — TAKEN FROM COCOINDEX. `--situ=db/wal_manager.cc` on RocksDB spent 3,104 B and never named the gold `db/wal_manager.h` anywhere, while an embedding search returned it as result redhat-et#1 in 55 bytes; same shape on table/get_context.cc; S2 recall 2/14. Not a retrieval failure: section [1] ranks by dependent-symbol count, which surfaces the biggest test files and can NEVER surface the header, because a header does not depend on the source that implements it. It is a different relationship, and one ripwire already holds. declDefPartners() — another file defining symbols under the SAME (scope, name) identity. Stated that way it is not a C++ special case: it covers a .h/.cc pair, an ObjC .h/.m, a C# partial class, a .pyi stub and a .d.ts, without reading one file extension or comparing one basename. The guard is a MAJORITY test rather than a number pulled from the air — 2*shared >= min(symbols(a), symbols(b)) — so two large files sharing one common free-function name are refused, and the row publishes `shared` so the reader audits it. No embedding mode was added and none is coming (G3). The blast-radius counts are unchanged: the partner is printed above them as its own labelled fact, never merged into them. Measured after: --situ=db/wal_manager.cc -> db/wal_manager.h (10 shared) as row one; table/get_context.cc -> table/get_context.h (16); db/version_set.cc -> db/version_set.h (138). Gates: sincewindowcheck arm 4 now 18 assertions, all GREEN, including a NEGATIVE CONTROL that two 13-symbol files sharing one name are refused (the majority test is live, not decorative) and an arm that a git repo with no commits stamps window="18mo" WITHOUT @Head rather than claiming an anchor it does not have. deckcheck: six foreign tool flags quoted in lane C2's head-to-head write-up (gortex --dataset/--detach/ --entry-point/--kind, uv --python, rg --sort) allowlisted — that gate was red on this lane's base. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…t contain
An adversarial review of the section-granular round found that its ancestor rule fails on
exactly the queries CLAUDE.md tells agents to write. Reproduced on a scratch tree holding
only docs/ARCHITECTURE.md:
--recall="crawl order sorted before ids assigned" -> lines="353-369"
--recall="how does the crawl order get sorted before the ids are assigned" -> lines="1-11,12-20"
Same two stubs at 800, 1200 and 1600 tokens; the answer first appears at 2400, third. The
landed comment called the residual "bounded and self-correcting … the worst case is a
heading-plus-intro admitted one slot early". That claim was false and is deleted with the
reasoning it rested on. This ranker has no query-side stopword list — the SIRA corpus
margin (src/lexical.h marginBits) ships disarmed — so "the", "does", "are" are live query
terms, an ancestor's SUBTREE contains every one of them everywhere, and tf is additive.
THE FIX IS AT THE RANK, AND BOTH PROPOSED ADMISSION FIXES WERE MEASURED FIRST.
- a term margin (keep the ancestor only on a query term with corpus df <= S/2) leaves the
reproduction BIT-IDENTICAL. Built and run: of the three ancestors it would have to drop,
`# ripwire — architecture` survives on "how" (df 2 of 13 own-prose units) and
`## 1. The pipeline` on "order" (df 4 of 13). Both terms separate, and both are genuinely
in the ancestor's own prose — `## 1. The pipeline` really does say "in that order".
- "keep it only if its own-prose SCORE is positive" is inert, provably: BM25's idf is
ln((S-n+0.5)/(n+0.5)+1) > 0 for every n <= S, so a positive own-prose score IS "own prose
carries a query term", which is the rule already shipping.
Admission was right to keep those units. What was wrong is that they OUTRANKED the answer,
on evidence held by the seven subsections nested inside one of them. So admission is
untouched — for a fixed heading population the pick SET is what it always was — and the rank
key becomes attributed:
rank = scorer score x (this unit's own-prose evidence / the strongest own-prose evidence
anywhere in its subtree)
Own-prose evidence is BM25 over a collection of THIS document's own-prose units: the same
idf, the same resolveBm25Params, the same forEachLexSubtoken tokenizer and the same
name(x3)/body(x1) field split the scorer uses. The collection is local because the question
is local — which section of this document the query is about — and a word in every section
of a document separates none of them, so "the" loses its idf without a stopword list. Only
ratios of two values from one call are ever read, and nothing prints them. A leaf's subtree
IS its own prose, so its maximum is itself and its key is the scorer's number bit-for-bit;
an ancestor that is the best evidence under itself is likewise untouched.
MEASURED, final binary vs the pre-change binary, two harnesses with constructed ground truth
(no hand labels). H1: 72 unique-term queries whose content words occur in one section only.
H2: 456 heading-as-query trials over docs/. The NL form of each adds ONLY function words to
the same content words, so any difference is attributable to how they are treated.
H1 @800 NL served 88->89 NL first 81->85 agree 79->83 KW first 97->97 KW bytes 100% same
H1 @1600 NL served 88->89 NL first 43->79 agree 42->78 KW first 97->97 KW bytes 100% same
H2 @800 NL served 76->80 NL first 73->79 agree 75->80 KW first 90->92 KW bytes 98% same
H2 @1600 NL served 77->81 NL first 57->74 agree 69->77 KW first 74->86 KW bytes 85% same
No metric regresses in any cell. A full own-prose re-rank (discard the scorer's number
entirely) beats this on H2 — which queries HEADINGS, the field it weights x3, so it grades
its own bias — and is a ranking replacement, not a defect fix: it moves 9% of keyword answers
for no measured keyword gain and belongs to a pre-registered ranking round. The x3/x1 split
is not a tuned constant: rerunning both harnesses with the heading at x1 is byte-identical on
every trial, because a ratio between two spans of one document is insensitive to it.
THE DEGENERATE CUT NAMED LINES IT DID NOT SERVE. findRecallBoundaryCut lands immediately
AFTER a newline, so the last kept byte opens the first line NOT served; counting every '\n'
in the kept prefix therefore disclosed a line whose bytes never reached stdout. Reproduced: a
section at line 3 followed by 400 two-line paragraphs, at --max-tokens=400, claimed
lines="3-10" with source line 9 the last byte served; now 3-9. The count runs over
kept[0, n-1), which is composeRecallUnits' own rule (it derives lineHi from ownEndByte - 1),
so the two paths compute the same thing the same way. Non-negotiable redhat-et#3.
BODY-LESS HEADINGS ARE HEADINGS. The selection dropped any heading whose section body is
empty — a heading immediately followed, with no blank line, by a same-or-shallower one — so
it was neither counted in N nor treated as a unit boundary, and --help's "up to the next
heading of any depth" was false twice over. A 4-heading fixture reported `3 of 3` and served
`5-9`, a range containing the `### Empty child` line; it now reports `3 of 4` and serves
`5-8`. On docs/EVALS.md N goes 197 -> 198, which is exactly the ATX heading count computed
independently by a line scan. --help needed no change, which is also why this repair was
chosen over rewording it. The whole-file node — the real target of the old filter — is now
excluded by NAMING it (starts at byte 0, no signature/body split, ends no earlier than any
heading in the file), measured against the file's own symbols rather than its size on disk,
so a file that grew after indexing cannot promote it into a unit that overlaps every other.
Also from the review: truncateRecallBody's out-parameter becomes a structured-binding return
(CONTRIBUTING §3) — it made the byte count optional, and the caller that skipped it is the
one whose disclosure would go stale first. And `dropped_by_budget=D` deliberately carries no
`next=`, with the reason recorded at formatRecallSectionNote: the only knob that could
restore a dropped unit is --max-tokens, which is not invertible from inside one document
(§C4 water-filling divides one budget across the set, and cross-document monotonicity is
explicitly not claimed), so a computed N would be a guess wearing a disclosure's clothes.
--quality-delta found the new scanner had cloned ownProseCarriesQueryTerm's match loop; both
now share recallMatchedTermIndex, and the boolean keeps its early exit.
Gates: 573 run, 571 pass, 2 skip (both want a reference binary), 0 fail. The twelve recall
gates plus mdsectioncheck, mcprobustcheck and redactcheck green by name. Both fixtures the
old rule was built for verified directly: guide.md answers `zqcachewarmbody` with
`1 of 7 … lines="12-26"` and no zqorientbody or zqdeepappendix; notes.md keeps
`# Geometry Fixture` and serves it first. Determinism x2 byte-identical on the map and on
--recall, warm == cold on both and on the degenerate cut; xmllint clean. ASan/UBSan/LSan over
20 arms of the attribution path, the degenerate cut at five ceilings, the empty-heading
fixture, both named fixtures and the operator's memory directory: no report.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ourth is refuted with numbers METHODOLOGY §9 #3 says never cut silently and #6 says a ceiling attribute names the ceiling ACTUALLY applied. Three caps did neither. Each fix is disclosed only when the cap actually bit — the pr_converged shape (src/prconverge.h): silence means untruncated, presence means truncated. G4 pays nothing otherwise. 1. kGrepMatchedLineMaxBytes (512 B, src/search.h) cut the matched line of EVERY --grep and --verify hit with no ellipsis and no attribute, on the only content those answers carry: a 512 B source line and a truncated 50 KB minified line printed byte-identical payloads. The row now carries line_bytes="N", the WHOLE line's byte length. The NUMBER rather than a bare capped="1" because it is what decides the reader's next move — the row already carries the deterministic follow-up (p=/l=, the root's next=), so what was missing is whether following it is worth a read. NO ellipsis here, deliberately, and this is where it differs from (2): a grep payload is raw file bytes by contract — the boolean --and/--not filter reads the same line and the --at= follow-up is expected to reproduce it — so the fact goes in an attribute (§9 #4) rather than into the bytes. lineBytes joins grepGroupByFile's fold key: two long lines can share a 512 B prefix and differ in true length, and folding those would print one line_bytes= for sites it does not describe; untruncated rows carry 0, so the key is byte-identical wherever the cap did not fire. The MCP grep verb is the third caller of grepEnrich and serves NO matched text at all, so it has no cut to disclose — asserted rather than assumed, arm A5. 2. cleanSig's kMaxSig (240 B, src/serialize.h) hard-broke every emitted signature — --pack-signatures, --for's <sigs>, <calls> callee rows, --lego — mid-token, with no marker. It now goes through truncateUtf8WithEllipsis, the tool's ONE truncator, exactly as the three other signature cuts (kForTailSigBytes, kForCapTailSigBytes, packtask.h's tail sig) already do. In-band here because a signature is already a RENDERING, not raw bytes: the body is stripped and whitespace runs collapse, so matching its three siblings is what consistency means. The visible prefix is unchanged at 240 B; only the "…" is new. The loop collects one byte past the cap so the shared truncator can do its own codepoint back-off, and a separate flag carries the fact because the trailing-space trim can pull a cut string back under the cap and a cut that trims back under is still a cut. 3. A DEFAULT --for enforced kForPayloadBudgetBytes (7500 B) on every run and named no ceiling: budget_tokens= rode only an EXPLICIT --token-budget, so a trimmed default bundle disclosed THAT it was cut (<sigs shown= total= capped="1">) while the number that cut it appeared nowhere. Compare --pack-task, whose default lands on its root as budget_tokens="6000" — same class of ceiling, two honesty outcomes. The <ctx> root now carries budget_bytes="7500", and the JSON dialect ",budget_bytes":7500. THE UNIT IS BYTES, and that is the decision the finding asked for. What the default applies IS a byte constant — bundleBudget is kForPayloadBudgetBytes verbatim and the ladder compares rendered bytes against it. A token spelling would have to divide by a rate, and the two rates in play disagree on purpose: est_tokens prices at kBytesPerTokenDefault (2.50) while a ceiling is SIZED at the conservative kMinBytesPerToken. Any token number printed here would be one no ladder ever applied, which is the exact failure §9 #6 names. DEFAULT REGIME ONLY: the explicit regime already names the caller's own ceiling, so exactly one of the two rides a trimmed bundle — never both, never neither (gate arms C8/C10). It rides the ROOT and is spliced AFTER the sigs render, with dropped_positive=/bundle=: the first cut of this fix put it on the <sigs> open tag with the ladder's own per-run share as the value, which is more precise and wrong here — forbudgetmonotoncheck pins the sig section byte-identical between the default ceiling and any explicit ceiling above it, and a per-run number breaks that identity while every served row stays the same. The late splice leaves the ladder's input untouched, so <sigs> is byte-for-byte what it was. The clause defining the attribute is spliced on the same condition (kForOverCeilingLegend's precedent) so an untrimmed bundle pays for neither. packSignatures gains cappedOut, the ladder verdict its JSON twin has always returned, so the caller reads a boolean instead of re-parsing rendered bytes. THE MCP `for` VERB CARRIES IT TOO (§P8: one element name, one attribute order, both surfaces). That dialect is budgeted by default the same way (mcpverbs.h forBudgetBytes) and had the same silence; fixing the CLI alone would have replaced one honesty gap with a disagreement between two surfaces an agent can reach for the same answer. Same two conditions, same splice point, same sentence — arm C14 pins the sentence byte-identical across the two. 4. kSliceRdMaxIter (64, src/slice.h) is REFUTED, not fixed: it is unreachable, so it cuts nothing and has nothing to disclose. Structurally, the reaching-definition lattice has no cross-slot flow — a use reads its own slot, a def assigns it a singleton (SliceRdWalker::unit) — so each slot's loop transfer is X -> const ∪ (X if a path passes through), whose ascending chain from `entry` stabilises on the SECOND round regardless of program shape. Measured: instrumenting the converged break and running --slice over 3,026 symbols of this tree's own src/ gives 4,528 loop fixpoints, max iter = 1 (508 at 0, 4,020 at 1, none higher); an adversarial C fixture (while > for > do-while nesting around a 40-case fallthrough switch with continue/break) gives max iter = 1, and an adversarial Python one (nested loops with try/except/finally and continue) the same. The bound fires at iter >= 63. Its comment already says it "only guards a broken lattice", and that is what it is: an XML attribute that can provably never appear is ungateable by construction (non-negotiable #1) and would be decoration. DEGRADED_PATH_ALERT stays the right instrument for it (guardrail #4). GATE FIRST (non-negotiable #1): test/capdisclosurecheck.sh was written and shown RED before any src/ edit. Every arm asserts three things, because a disclosure gate that only greps for its own attribute proves nothing: CROSSING (the fixture's emitted payload really is shorter than the source, or the element really is capped), DISCLOSURE, and SILENCE (the same verb on an uncrossed fixture carries nothing). MUTATION CONTROL, run against a binary built from the parent commit: every DISCLOSURE arm fails (A2, A4, B2, B4, C2, C5, C7) while every CROSSING and SILENCE arm still passes. Gate count 573 -> 574, derived from THIS tip's own `for _g in` loop after the rebase (main landed optremarkshotcheck while this lane was green; the union of the two loops is the number), across all eight published sites. Only test/regression.sh conflicted — README/EVALS/deck auto-merged clean at the stale 573, which is the failure mode, so all three were reset from origin/main and re-bumped from the merged loop. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…-uses on Elixir defs Three of the six actionable review findings on redhat-et#81 reproduced; this is those three. 1. src/ingest_model.h — expandElixirImplementationReferences appended a multi-target defimpl's cloned references to the TAIL of ing.references, whose contract is (fileId, startByte, name, role, isInherit) order. A clone carries its original's coordinates, so the tail was full of earlier bytes and earlier files. Two consumers read that order rather than re-deriving it: graph.h's chaUpDeclared records a derived type's direct bases in source order (Elixir's Extends refs from defimpl ARE cloned), and editpreview.h splices a re-parsed file in by partitioning on fileId, which only reproduces a real re-ingest while each file's refs are one ascending run. Each clone is now spliced in beside the reference it came from — the same position a re-sort would give it, at no sorting cost. 2. src/mcpverbs.h — the MCP uses verb gated the Elixir resolver path on `defs`, which resolveAllByName fills for EVERY language, while the CLI gates on the same list filtered to Lang::Elixir. An Elixir reference whose calleeName matched a definition in another language therefore took the resolver path in MCP, matched nothing, and vanished — while the CLI still reported it through the name filter. That is the divergence mcpclidiffcheck exists to prevent. Now filtered the same way, off the raw selector, exactly as resolveUsesSelector does. 3. docs/ARCHITECTURE.md — the Elixir prose claimed resolution the tree does not do. Four gaps reproduce on this tip and are now disclosed in Static limits rather than left to contradict the PR notes: a later `import M, except:` replaces an earlier `only:` selection instead of subtracting; a dotted nested `defmodule Inner.Deep` registers no implicit prefix alias, so `Inner.Deep.target()` resolves to nothing; `alias __MODULE__, as: Current` in a multi-target defimpl binds every implementation to the FIRST target; and `&_seed/0` is dropped by the underscore filter. Honesty in output is a feature (CLAUDE.md non-negotiable redhat-et#3) — these are floors, and the document now says so. The three remaining findings did not reproduce and are unchanged: reachesAny's exact name compare is correct because no non-Call Elixir reference carries an arity suffix (they are module names and @attributes, matched against symbols whose scope is empty or the enclosing module), and eliximportcheck.sh sets `set -u`, not `set -e`, so its FAIL branch and fail=1 accounting do run. Verified: elixircheck, eliximportcheck, elixirsemanticcheck, mcpclidiffcheck, editpreviewcheck, editroundtripcheck, qschemetripcheck and qextractionkeycheck pass; multi-target defimpl output is unchanged; ASan/LSan clean on the Elixir fixtures and on the repo; repeat maps byte-identical. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_013VUC4dGmaT2bFJvPJtsjvf
…ourth is refuted with numbers METHODOLOGY §9 #3 says never cut silently and #6 says a ceiling attribute names the ceiling ACTUALLY applied. Three caps did neither. Each fix is disclosed only when the cap actually bit — the pr_converged shape (src/prconverge.h): silence means untruncated, presence means truncated. G4 pays nothing otherwise. 1. kGrepMatchedLineMaxBytes (512 B, src/search.h) cut the matched line of EVERY --grep and --verify hit with no ellipsis and no attribute, on the only content those answers carry: a 512 B source line and a truncated 50 KB minified line printed byte-identical payloads. The row now carries line_bytes="N", the WHOLE line's byte length. The NUMBER rather than a bare capped="1" because it is what decides the reader's next move — the row already carries the deterministic follow-up (p=/l=, the root's next=), so what was missing is whether following it is worth a read. NO ellipsis here, deliberately, and this is where it differs from (2): a grep payload is raw file bytes by contract — the boolean --and/--not filter reads the same line and the --at= follow-up is expected to reproduce it — so the fact goes in an attribute (§9 #4) rather than into the bytes. lineBytes joins grepGroupByFile's fold key: two long lines can share a 512 B prefix and differ in true length, and folding those would print one line_bytes= for sites it does not describe; untruncated rows carry 0, so the key is byte-identical wherever the cap did not fire. The MCP grep verb is the third caller of grepEnrich and serves NO matched text at all, so it has no cut to disclose — asserted rather than assumed, arm A5. 2. cleanSig's kMaxSig (240 B, src/serialize.h) hard-broke every emitted signature — --pack-signatures, --for's <sigs>, <calls> callee rows, --lego — mid-token, with no marker. It now goes through truncateUtf8WithEllipsis, the tool's ONE truncator, exactly as the three other signature cuts (kForTailSigBytes, kForCapTailSigBytes, packtask.h's tail sig) already do. In-band here because a signature is already a RENDERING, not raw bytes: the body is stripped and whitespace runs collapse, so matching its three siblings is what consistency means. The visible prefix is unchanged at 240 B; only the "…" is new. The loop collects one byte past the cap so the shared truncator can do its own codepoint back-off, and a separate flag carries the fact because the trailing-space trim can pull a cut string back under the cap and a cut that trims back under is still a cut. 3. A DEFAULT --for enforced kForPayloadBudgetBytes (7500 B) on every run and named no ceiling: budget_tokens= rode only an EXPLICIT --token-budget, so a trimmed default bundle disclosed THAT it was cut (<sigs shown= total= capped="1">) while the number that cut it appeared nowhere. Compare --pack-task, whose default lands on its root as budget_tokens="6000" — same class of ceiling, two honesty outcomes. The <ctx> root now carries budget_bytes="7500", and the JSON dialect ",budget_bytes":7500. THE UNIT IS BYTES, and that is the decision the finding asked for. What the default applies IS a byte constant — bundleBudget is kForPayloadBudgetBytes verbatim and the ladder compares rendered bytes against it. A token spelling would have to divide by a rate, and the two rates in play disagree on purpose: est_tokens prices at kBytesPerTokenDefault (2.50) while a ceiling is SIZED at the conservative kMinBytesPerToken. Any token number printed here would be one no ladder ever applied, which is the exact failure §9 #6 names. DEFAULT REGIME ONLY: the explicit regime already names the caller's own ceiling, so exactly one of the two rides a trimmed bundle — never both, never neither (gate arms C8/C10). It rides the ROOT and is spliced AFTER the sigs render, with dropped_positive=/bundle=: the first cut of this fix put it on the <sigs> open tag with the ladder's own per-run share as the value, which is more precise and wrong here — forbudgetmonotoncheck pins the sig section byte-identical between the default ceiling and any explicit ceiling above it, and a per-run number breaks that identity while every served row stays the same. The late splice leaves the ladder's input untouched, so <sigs> is byte-for-byte what it was. The clause defining the attribute is spliced on the same condition (kForOverCeilingLegend's precedent) so an untrimmed bundle pays for neither. packSignatures gains cappedOut, the ladder verdict its JSON twin has always returned, so the caller reads a boolean instead of re-parsing rendered bytes. THE MCP `for` VERB CARRIES IT TOO (§P8: one element name, one attribute order, both surfaces). That dialect is budgeted by default the same way (mcpverbs.h forBudgetBytes) and had the same silence; fixing the CLI alone would have replaced one honesty gap with a disagreement between two surfaces an agent can reach for the same answer. Same two conditions, same splice point, same sentence — arm C14 pins the sentence byte-identical across the two. 4. kSliceRdMaxIter (64, src/slice.h) is REFUTED, not fixed: it is unreachable, so it cuts nothing and has nothing to disclose. Structurally, the reaching-definition lattice has no cross-slot flow — a use reads its own slot, a def assigns it a singleton (SliceRdWalker::unit) — so each slot's loop transfer is X -> const ∪ (X if a path passes through), whose ascending chain from `entry` stabilises on the SECOND round regardless of program shape. Measured: instrumenting the converged break and running --slice over 3,026 symbols of this tree's own src/ gives 4,528 loop fixpoints, max iter = 1 (508 at 0, 4,020 at 1, none higher); an adversarial C fixture (while > for > do-while nesting around a 40-case fallthrough switch with continue/break) gives max iter = 1, and an adversarial Python one (nested loops with try/except/finally and continue) the same. The bound fires at iter >= 63. Its comment already says it "only guards a broken lattice", and that is what it is: an XML attribute that can provably never appear is ungateable by construction (non-negotiable #1) and would be decoration. DEGRADED_PATH_ALERT stays the right instrument for it (guardrail #4). GATE FIRST (non-negotiable #1): test/capdisclosurecheck.sh was written and shown RED before any src/ edit. Every arm asserts three things, because a disclosure gate that only greps for its own attribute proves nothing: CROSSING (the fixture's emitted payload really is shorter than the source, or the element really is capped), DISCLOSURE, and SILENCE (the same verb on an uncrossed fixture carries nothing). MUTATION CONTROL, run against a binary built from the parent commit: every DISCLOSURE arm fails (A2, A4, B2, B4, C2, C5, C7) while every CROSSING and SILENCE arm still passes. RE-PIN: test/docdemotegolden_for.xml 5505 -> 5517 B and test/docdemotegolden_noroute.xml 9556 -> 9568 B, +12 B each = FOUR three-byte U+2026 markers from (2), plus the est_tokens recount. Verified before re-pinning: with est_tokens= normalised and the four new ellipses deleted, live and previous goldens are byte-identical on both fixtures — no ranking, demotion, route or budget byte moved, and neither root is capped, so (3) is absent from both. That fixture is where (2) shows up at all: a markdown section's "signature" is prose, so four of those rows were over the 240 B cap and had been cut invisibly for the life of the golden. The cause is recorded in the gate header beside the five earlier re-pins. test/printf_parity.manifest needed NO re-pin: none of its 14 pinned verbs reaches a changed emitter. Gate count 574 -> 575, derived from THIS tip's own `for _g in` loop after the LAST of three rebases (main landed optremarkshotcheck and then callsrankordercheck while this lane was in flight; the union of the two loops is the number, every time). Only test/regression.sh ever conflicted — README/EVALS/deck auto-merged clean at the stale number on all three, which is exactly the failure mode, so all three files were reset from origin/main and re-bumped from the merged loop each time. The last rebase crossed a lane that changed serialize.h and verbs_for.h too (the <calls> rank-order fix), so it was followed by a --clean-first rebuild before anything was re-measured. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ll 120 caps become visible
A cap is a ROUTING decision — it decides what an agent can and cannot find — and this tree had 120 of
them across 51 files with nothing listing them together. That is how kMaxExpandSibs came to sit at 8.
WHAT 8 COST, MEASURED. Symbols-per-file on this repo: median 4, p90 18, p99 85, max 562. At 8 the cap
fired on 68.5% of bodies and hid 89.3% of every sibling name — 10.7% of the file-context signal
survived. It looked harmless because the MEDIAN file has 4 symbols and never trips it; but symbols
concentrate in big files, so symbol-weighted it fired on two bodies in three.
WHAT IT WAS SUPPOSED TO SAVE. P16's stated cost was "582 B of a 3,067 B --expand answer (19%); ~3.5 KB
per --pack-task bundle". The second half is not reproducible: --pack-task emits NO sibs= at all, and
never did — measured against the PRE-P16 binary (2026-08-30), which also emits none. Two arms of
test/expandsibscheck.sh already asserted exactly that ("--pack-task carries no sibs=/inc="), and were
green the whole time. sibs= reaches exactly one verb: --expand.
WHY 100. It clears p99 (85), so it guards the pathological tail instead of cutting through the typical
case: 15.8% of bodies truncated instead of 68.5%, sibling names visible 10.7% -> 56.2%. Cost is +36% on
a single-symbol --expand answer and NOTHING on --for or --pack-task, which are byte-identical at every
cap value tested (9,829 B and 12,590 B at 8, 80 and 100 alike).
THE NUMBER MOVES, AND THAT IS AN EFFECT, NOT THE REASON. top-50 goes 71.0% -> 81.8%. --pack-signatures
elides exactly what it always did; --expand simply stopped hiding context, and the body side is this
ratio's DENOMINATOR. The cap was chosen on recall grounds and the figure re-derived afterwards — doing
it the other way round would be tuning a default to flatter a headline. Re-derived on THIS build and
propagated to all six artifacts that carry it (band, --help, caption, capture, COMMANDS.md, README,
EVALS), with the parity `help` label re-pinned under UPDATE_GOLDEN_EXPECT="help" — moved={help}, 39
unchanged.
THE FIXTURE HAD TO GROW, AND THAT IS THE POINT. test/expandsibsfix/manyfn.c had 45 functions: one over
the OLD cap of 40, still over 8, but UNDER 100 — so at cap=100 the gate would have passed while testing
nothing. Regrown to 145 (144 siblings), one over the new cap, by construction. Its header comment still
claimed kMaxExpandSibs=40, stale since P16 and never caught, because nothing compared it to the source.
SO THE CAPS GET AN INVENTORY, GENERATED AND GATED. docs/LIMITS.md lists every cap in src/ with its
value, site and whether its file discloses a truncation when it fires. It is a build product of
docs/limits_build.py, and test/limitstablecheck.sh fails when it drifts — 4 arms, including a
can-go-red that adds a synthetic cap to a TEMP tree via --root, never to src/, because a probe file
dropped into src/ perturbs the crawl other gates measure.
The inventory's first finding, not acted on here: 54 of the 120 caps sit in files that disclose nothing
when they fire. src/mention.h holds 7 of them, and they bound which docs can ever surface in --for. A
silent cut reads to the caller as "none exists", which non-negotiable redhat-et#3 forbids.
Gate count 570 -> 571, at all eight spellings (two in README, three in EVALS, three in the deck) — and
NOT at `1.570`, a confidence interval, or inside a CI run id.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…id nothing capdisclosurecheck closed three caps on EMISSION — a row found and then trimmed, on an element that already had a shown=/total= vocabulary. The ten caps in src/mention.h and src/gitmine.h's co-boost family are one level up: INDEXING caps, deciding what --for's ranker is allowed to LIFT. Nothing downstream can tell that a mention was dropped, and no flag, budget or paging brings it back — the answer simply does not contain the thing the task named. That is non-negotiable redhat-et#3 ("a zero means none found, never none exists") on a surface where the caller has no recourse at all. docs/LIMITS.md said `Discloses: none` for both files. THE RULE, which decides which caps got an attribute and which got a comment: DISCLOSE WHEN THE CUT CONTENT IS NOT OTHERWISE VISIBLE IN THE ANSWER. Seven disclose, each conditional, each carrying an exact `*_total=` wherever the true count is computable without changing what the cap keeps: mention_tokens_capped= / _total= kMentionMaxRawTokens (16) — mention tokens past the extraction window. extractMentions now scans the whole task and only the KEEPING stops at 16, so the total is exact and the scan costs nothing (task text is a sentence). mention_files_capped= kMentionMaxFiles (4). FACT, no total: counting the files the scan never reached means running it unbounded, and that WOULD move `matchedFile` (a full list leaves it false and routes the mention to the symbol path). The honest number available is the fact of the cut. mention_syms_capped= / _total= kMentionMaxDirectSymbols (8) — Scope.name matches. The scan runs to the end for the census; pushing is still gated on the same room test, so the kept set is unmoved. doc_mentions_capped= kDocMentionMaxDocsPerAnchor (2) and kDocMentionMaxDocsTotal (6). Set ONLY where a refusal is provable — a doc below its anchor's lift target that a cap turned away. "There might be more" never sets it. coboost_partners_capped= / _total= kCoBoostMaxPartnerFiles (8); the true count was already in hand one line above the resize. coboost_commits_capped= / _total= kCoBoostMaxFilesPerCommit (30) — whole commits thrown away before the boost ever reads them, so a partner that only moves inside big refactor commits is silently unreachable. Three do NOT, and each carries the measurement at its declaration rather than an attribute: kMentionMaxSymbolsPerFile and kCoBoostMaxSymbolsPerFile trim symbols out of a file every surviving row NAMES with p=, so the caller can see it and page in; kDocMentionMaxAnchors is a window over the top of a ranked list the bundle prints in full. WHERE THE DISCLOSURE RIDES, and why not where its prose belongs. A --for header composed before the signature ladder runs is charged against kForPayloadBudgetBytes, so bytes spent in it are bytes taken from ranked rows. The first cut of this lane folded the clause into mentionNote/docMentionNote and measured the cost: three bundles fell from 25 ranked rows to 24, and one lost 204 B of evidence. The clause was trimmed to ~50 B AND moved to the post-render splice budget_bytes= uses, so <sigs> is byte-identical and the disclosure is paid for in bytes, never in evidence. The clause carries the attributes verbatim, so it is self-defining where the reader meets them (legendcoveragecheck's `defined` predicate). HOW OFTEN THEY FIRE, over 39 realistic --for tasks on this tree: kMentionMaxRawTokens 1/39, kMentionMaxFiles 3/39, kMentionMaxDirectSymbols 0/39, the doc caps 6/39. All three mention firings came from one shape — a task pasting SEVERAL paths, exactly the multi-file case B8 exists for. The 0/39 is a property of a C++ tree with unique method names, not of the cap; the class is gated on a fixture that does cross it, and a silent cap costs nothing, so the zero is recorded rather than used as an argument to leave the cut unsaid. NON-DEGRADATION, BASE_BIN (633a1d2) vs this binary, same frozen corpus, identical flags: # base B lane B delta sigs rows invocation 1 2301 2301 0 - > - --for="cache invalidation" --format=candidates --top-k=5 2 10040 10040 0 18 > 18 --for="incremental cache invalidation when a file content hash changes" 3 14521 14627 +106 25 > 25 --for="pagerank power iteration" --detail=2 4 11055 11161 +106 25 > 25 --for="pagerank power iteration" --with-graph 5 10064 10064 0 26 > 26 --for="quality delta acks ledger rubber stamp" 6 10075 10075 0 27 > 27 --for="quality delta acks ledger rubber stamp" --no-doc-mention 7 5528 5528 0 - > - --for="rankGraphTeleport" 8 15975 16081 +106 20 > 20 --for="rankGraphTeleport" --no-route 9 2850 2850 0 - > - --for="rankGraphTeleport" --signatures-only 10 9962 9962 0 29 > 29 --for="tree-sitter parse of a source file" --adaptive 11 15958 15958 0 29 > 29 --for="tree-sitter parse of a source file" --auto-bodies 12 9575 9575 0 31 > 31 --for="tree-sitter parse of a source file" --legend=compact 13 10125 10231 +106 26 > 26 --for="why does src/lexical.h chooseForRanker pick name-exact BM25" 14 0 0 0 - > - (a truncated flag in the sweep; refusal on both) 15 9607 9607 0 - > - --pack-task="add a new output format flag to the CLI" 16 25256 25256 0 4+8+6 > --pack-task="add a new output format flag to the CLI" --partition=3 4+8+6 map 24585 24585 0 - > - the flagless map, byte-identical Twelve of sixteen are byte-identical. The four that moved are exactly the four where doc_mentions_capped="1" fired, each +106 B (the attribute, the self-defining clause, and the est_tokens digits that now price them), and NO bundle lost a ranked row. Two further real tasks, measured because they are the ones that reach the B8 caps: a five-path task +214 B (doc_mentions_capped + mention_files_capped) and a seventeen-path task +318 B (those two plus mention_tokens_capped="1" mention_tokens_total="17"), both at unchanged row counts. GATE FIRST. test/mentioncapcheck.sh, written before the code, six arms x three assertions: CROSSING, DISCLOSURE, SILENCE. Each crossing is proved WITHOUT reading the new attribute — arm A is a same-token-multiset word-ORDER contrast (every lexical signal here is order-blind, so only the extraction window can explain the difference); B/C/D read the pre-existing prose note's count against a fixture size the gate computes; E reads the partner list out of --cochange, whose cochangePartners path shares no code with applyCoChangeBoost's; F reads git itself. Arm G pins the legend definition, --json parity, zero-cost silence on an ordinary query, determinism and well-formedness. MUTATION CONTROL, run: against BASE_BIN all 8 DISCLOSURE assertions FAIL and all 15 crossing/silence assertions PASS. Against this binary, ALL PASS. docs/LIMITS.md regenerated: silent caps 54 -> 44, and both files now read `Discloses:` instead of `Discloses: none`. One defect found by READING the regenerated table rather than trusting it: limits_build.py harvests `[a-z_]+_capped` out of the source file, so the placeholder `name_capped="1"` in CapDisclosure's own doc comment was published as a mention.h attribute that does not exist. The placeholder is now spelled `<x>_capped`, which the regex cannot match, and the comment says why so the next placeholder is not spelled back. A generated doc is only as honest as its input. G1: asan+ubsan build clean on the flagless map, on the four --for shapes this lane moves (including the seventeen-path task that trips all three B8 caps) and on --pack-task; the new gate is ALL PASS under the sanitizer binary too. g1freshcheck green. Also green: capdisclosure, limitstable, docmention, mention, cochangeboost, forlens, legendcoverage, mcpattrparity, mcpforparity, fornotesbudget, forbudgetmonoton, estcharge, fornotesjson, packtask, packtaskmonoton, tokenbudget, jsonparity, formatgate, mentionsverb, docdemote, droppedpositive, xmlwellformed, qschemetrip, churnjoin, loopconservation, fixedbufsweep, binoverride, rootrel, testrowrun. Determinism x3 byte-identical, xmllint clean on the map and on a capped bundle. manifestcheck and readmedrift fail on their gate-COUNT arm only (577 -> 578); the eight published count sites are deliberately untouched this round. --quality-delta is NOT clean, and the residual is stated rather than hidden: 20 gating findings, all small deltas on already-large symbols — 5 api-surface (a census out-parameter on three miners, an out-parameter on two mention_detail helpers), 4 complexity, 7 verbosity, 4 short-horizon-churn (touching files that churn). A first pass cut it from 25: the two commit-census out-parameters were collapsed into one CommitWindowCensus struct, which removed the `params` class entirely; recordCommitFileSet was lifted out of gitLogFileSets (that function's complexity now falls below where it started); docLiftWasRefused was lifted out of applyDocMentionBoost (+17 complexity -> +7); and the in-function comment blocks were trimmed, with the durable rationale moved to file scope in mention.h where it costs no symbol its verbosity. What remains is intrinsic: six conditional disclosures across three surfaces are branches and lines that did not exist before. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…check.sh Adds .kt support (kParserVer 85) via fwcd/tree-sitter-kotlin, pinned past v0.3.8 for its scanner segfault fix. Zero ASan/UBSan/LSan findings across 501 real .kt files (8 local Android/JVM repos) despite that history. Extraction covers classes/objects/companion objects/functions/calls/imports, inheritance clauses (captureBases recognizes `delegation_specifier`), and complexity scoring (isDecisionType/cc_isNestingControl recognize `when_entry`/ `when_expression`/`do_while_statement`/Kotlin's own `catch_block` spelling — `lang == Lang::Kotlin` checked first so the other 21 languages short-circuit on a byte compare rather than a strcmp; `catch_block` is a different, unrelated node in the vendored Swift/Elixir grammars, verified against all 22 vendored parsers before scoping it). Canonical ids are scope-qualified (kotlinEnclosingScopeOf) and parameter counts are real (countParams counts `parameter` children specifically for Kotlin's parameter list — a naive "every named child" count over-counts, since a parameter's own modifiers and default value are its SIBLINGS in this grammar, not nested inside it; verified against a real parse before writing the fix). Import edges are captured (Kotlin joins Java's DepDialect, since a mixed module's imports genuinely cross the language boundary). Kotlin's function_body/class_body are positional children, not fields — the generic field-based body lookup silently returns null for every Kotlin definition, which is not cosmetic: graph.h's decl/def collapse reads a null body as "bodyless forward declaration" and deletes that candidate from resolution. Kotlin rides ingest_sidecap.h's existing ObjC positional-body fallback (widened, `class_body`/`enum_class_body` added) rather than a new mechanism — the same fallback ObjC has needed since ITS grammar exposes no `body:` field either. `enum class`'s body nests under `enum_class_body`, a DIFFERENT positional child than a plain class's `class_body`; verified via a standalone tree-sitter probe against the vendored grammar before fixing. Getting this right is what makes a same-name Kotlin/Java collision correctly surface as `ambiguous=` instead of silently resolving to whichever side happens to have a working body lookup. The JVM bridge (graph.h langCompatible) resolves Kotlin<->Java calls in both directions cleanly on real code (Nanidroid: unresolved 1425->545, ambiguous 1364->1496 — the rise is real collisions surfacing honestly via the fix above, not a regression). Disclosed, not fixed (a real new feature, not a bug): Kotlin navigation-expression receivers (`A.f()`) aren't recognized by ingest_binds.h's classifier, so an explicit receiver that would narrow a same-name collision between two Kotlin classes currently doesn't. Fixed the highest-severity language-dispatch finding: verbs_navigate.h's kExtSurfaceLangSlots was hand-sized off Lang::Elixir+1, so an unrecognized language's external-surface references silently clamped into slot 0 (Lang::Cpp). Every array in src/ hand-sized off a last-enumerator literal (verbs_navigate.h, verbs_report.h, lintrules.h, clones.h, htmlexport.h, tsprobe.cpp) now routes through model.h's kLangCount instead, so this class of bug can't recur on the next language append. Also deferred and disclosed rather than silently dropped: .kts (Gradle DSL needs its own trailing-lambda parse-quality probe), --slice support, and essential-complexity (ev=) scoring (Kotlin's when/elvis/!! vocabulary isn't measured against isDecisionType's containment claim yet). test/kotlincheck.sh: a real fixture (Kotlin<->Java bidirectional calls, a constructed same-name collision pair including an `enum class`/plain-class pair, cross-file calls, imports), every number pinned from an actual run, mutation arms proving each assertion non-tautological — including that a crash on the mutated fixture can't silently read as a passed assertion (checks binary exit status before trusting absent output). Verified: kotlincheck.sh ALL PASS on build/ripwire and a --clean-first asan/ripwire; zero ASan/UBSan/LSan findings across 501 real .kt files (8 local repos) plus the fixture; manifestcheck.sh and every gate whose golden value this change legitimately moved (kCatalogLangs, the doctor grammar count 22->23, README's --callers example, the vendored-scanner serialize() safety classification, the dependency-capable-set ledger) re-verified clean. Independently re-tested afterward on 5 larger real-world Kotlin corpora not part of the gated suite (android/nowinandroid, android/compose-samples, android/architecture-samples, ktorio/ktor, signalapp/Signal-Android — ~9,500 additional .kt files, up to 6,131 files/64,899 symbols in one repo): zero crashes, zero ASan/UBSan/LSan findings, deterministic, well-formed on every one. The grammar's one identified parsing gap — `@Composable () -> Unit`-shaped annotated function-type parameters — affects ~22.8% of Composable-heavy files on the largest sample measured, localized to a few bytes each with the enclosing declaration intact (the same shape as this repo's existing Metal precedent). Also validated against a hand-written adversarial corpus (41 files: merge conflict markers, mid-edit fragments, ObjC/Java/Kotlin interop, unicode identifiers/emoji/RTL, deep nesting, syntax edge cases, deliberately broken input including invalid-UTF-8 binary garbage) — zero crashes, zero ASan findings, byte-identical determinism. That adversarial pass surfaced one genuine disclosure gap, unrelated to Kotlin: measureFileHealth (ingest_crawl.h) validated the leading sample for well-formed UTF-8 nowhere in the pipeline, so a file with invalid UTF-8 but no tree-sitter ERROR/MISSING node reported no degraded-parse signal at all — `--skipped`'s disclosure contract silently had a hole. Fixed by reusing jsonesc::utf8SeqLen (no new UTF-8 validator) to scan the same sample already walked for whitespace, recording bad-sequence byte positions, and counting only the ones NOT already inside a recorded ERROR span into err/err_ratio — first-cut version double-counted (badUtf8 added unconditionally, before cross-checking ERROR spans), pushing err_ratio above its documented <=1.000 ceiling on `\xff\xff\xff@#$%^&*(`-shaped input; caught by two independent reviews (a backgrounded codex second opinion and a separately user-launched Opus /code-review), both of which also caught the missing kParserVer bump (84->85, ingest_cache.h + its quality.h mirror) this behavior change needed. parsehealthcheck.sh gained two arms (bad-UTF-8-flagged, good-UTF-8-negative- control) verifying both the fix and the absence of a false positive. Two independent correctness reviews ran against the original port before it landed (/code-review --fix on the fable model, and a separate codex second opinion), then a four-angle /simplify pass (reuse, simplification, efficiency, altitude); this follow-up round added a second codex pass plus a separately user-launched Opus /code-review, both run against the adversarial-corpus fixes — every fix from all reviews re-verified against 8+ corpora and re-gated clean before landing. Follow-up review round: a four-angle /simplify pass (reuse, simplification, efficiency, altitude) plus an independent /code-review high --fix pass, both run against this diff before it lands. /simplify eliminated measureFileHealth's errorSpans vector and its O(n*m) rescan in favor of a std::lower_bound dedup over sorted bad-UTF-8 positions, consolidated repeated lang==Kotlin tests in isDecisionType/cc_isNestingControl, and merged langCompatible's two return branches. The independent review then found two genuine, pre-existing test-coverage gaps the port had shipped without: captureBases's Kotlin delegation_specifier handling (both the wrapped-call `Shape()` and bare-interface `Labeled` shapes) and countParams for multi-parameter Kotlin functions had zero fixture coverage — no prior fixture exercised a base-class clause or more than one parameter, so neither fix had anything pinning it. Both are now covered (kotlincheck.sh §9/§10, with non-tautological mutation arms 7e/7f). The review also reverted the langCompatible merge back to an early-return shape (langCompatible runs per candidate on a hot resolution path; the merged form paid for the JVM-bridge comparisons even on the common C-family- only case) and refined measureFileHealth's dedup to never hold a transiently over-counted value mid-walk. Re-verified: kotlincheck.sh and parsehealthcheck.sh ALL PASS (including the new coverage arms and a bad-UTF-8-inside-an-ERROR-span dedup test), ASan clean, full gate suite clean (gates=572 pass=565 skip=4 fail=3, all three failures pre-existing/environmental and untouched by this diff). Rebased onto origin/main a second time (186 commits: upstream independently landed Dart as language redhat-et#21 while this branch was in flight). Kotlin now appends AFTER Dart in Lang (index 22, kLangCount anchored on Lang::Kotlin). kParserVer collided a second time: upstream independently claimed 86 (Ruby receiver), 87 (markdown scanner counter saturation) and 88 (Dart) while this branch still held 86/87 from the first rebase — an exact collision with this branch's own prior use of those numbers. Resolved the same way as the first collision: re-bumped to the next free values (89, 90) rather than keeping either side's number, full history preserved in the ingest_cache.h comment. Six gate failures surfaced after the merge, each investigated individually rather than assumed to be either all-expected or all-flaky: g1configcheck.sh and dartcheck.sh/elixircheck.sh each carried their own stale hardcoded doctor-probe/fuzz-target/seed counts (each language's own gate only bumps its own copy, a drift-prone hand-maintained pattern — fixed by updating all three plus a second stale fuzz-executable-name set in g1configcheck.sh that was masked locally by a Clang-only SKIP but would have failed CI's Clang legs). ingest_names.h's kotlinEnclosingScopeOf had three literal std::strcmp calls never converted to the kindIs() house style introduced by upstream's own per-node strcmp-to-kindIs refactor, since this file predated that refactor and wasn't part of the original conflict set — converted, fourth (runtime-variable) strcmp call correctly left unconverted. ingest_crawl.h's kLangTable auto-merged cleanly (both .dart and .kt rows present) but its array extent literal silently stayed at 47 — bumped to 48 after an exact row count. limitstablecheck.sh needed docs/LIMITS.md regenerated via its own generator script. docs/COMMANDS.md, its newest showcase capture, and every <!-- gatecount -->-marked reference were regenerated/re-derived rather than hand-edited, per upstream's newly introduced generated-artifact pattern for these documents. mcpattrparitycheck.sh and (in the final pre-push run) fillordercheck.sh each failed once under -j 6 parallel contention and passed cleanly on repeated standalone re-runs — confirmed transient, not code issues. Re-verified after all of the above: full gate suite clean (gates=602 pass=593-598 skip=4 fail=4, only the 4 pre-existing/environmental failures: w3fixlegendcheck.sh's root-spelling-dependent --dead-code path check, showcasecapturecheck.sh's pre-existing --stray-content=lane caption- generation bug, optremarkscheck.sh's Clang-only PGO arms, and nongitqmetricscheck.sh's non-git-corpus environmental check); ASan/UBSan clean on both src and test/kotlinfix under a --clean-first rebuild. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_016aPNTmY4XV8bHEkuhdtZnE PR redhat-et#126 review round (CodeRabbit automated review, 2026-09-10): three findings, all fixed. 1. Bodyless Kotlin class/interface definitions were silently deleted from the graph whenever a same-named Java definition existed elsewhere. Kotlin has no forward-declaration syntax for types — `interface Taggable` with no braces at all is the type's sole, complete definition, unlike a bodyless Kotlin FUNCTION (an interface member signature) or a C header prototype. The existing ingest_sidecap.h positional-body fallback (added for §5/§8's already-bodied collision fixtures) has no body child to find for a truly bodyless declaration, so bodyByte stayed 0 and graph.h's decl/def collapse (`hasBody`, used identically in buildGraph's §1e collapse and declToDefFollowThrough) read it as a droppable forward decl. Fixed by widening both `hasBody` copies with a narrow, GATED clause: `sym.lang == Lang::Kotlin && sym.kind == SymKind::Class`. New fixture (kotlincheck.sh §11): a bodyless `interface Taggable` colliding with a same-named Java class, with a mutation arm proving the assertion is not a tautology. 2. `third_party/deps/kotlin/src/scanner.c`'s `stack_push` called `abort()` outright when nested interpolated strings drove the delimiter stack to `TREE_SITTER_SERIALIZATION_BUFFER_SIZE` — a DELIBERATE process termination, not UB, so no sanitizer flagged it and it was invisible to every ASan/fuzz sweep this session ran (the fuzz seeds never happened to nest ~512 string opens deep). Violates this repo's own contract: a bad file must degrade, never crash the process. Fixed as a proper vendor patch (third_party/patches/kotlin/001-stack-push-no-abort.patch, following the yaml/001 and markdown/001 precedent for this exact defect shape): `stack_push` now returns `bool`, and `scan_string_start` propagates a `false` through the standard tree-sitter "no external token matched" recovery path (an ERROR node) instead of aborting. No kParserVer change — this only changes behavior on previously-ABORTING input, never on anything that used to parse and cache successfully. 3. `scan_string_content`'s escaped-dollar-at-string-end handling treated the FIRST quote after an escaped `\$` as the complete closing delimiter even inside a TRIPLE-quoted string, whose delimiter is three quotes — `"""a\$"""` lost its first closing quote and the remainder mis-tokenized. Fixed as a vendor patch (kotlin/002-triple-dollar-escape.patch): defers to the existing triple-quote-aware end_char handling a few lines down (the same is_triple deferral the plain-backslash case two branches below already uses). Real parse-output change on real input, so kParserVer bumped 90 -> 91 (ingest_cache.h + its quality.h mirror). Both vendor patches gained live tripwires in vendorpatchcheck.sh (new arm J): a generated 700-deep nested-interpolation fixture (the ASan-flavour crash tripwire for redhat-et#2) and a committed tripledollar.kt fixture (the plain-build exit-0-wrong-answer tripwire for redhat-et#3, checked via degraded_parse=0 and that the symbol declared right after the tricky string still extracts). Re-verified: kotlincheck.sh and vendorpatchcheck.sh ALL PASS on both build/ripwire and a --clean-first asan/ripwire; zero ASan/UBSan/LSan findings re-confirmed across all 8 local real-world Kotlin corpora (501 files) on the fixed binary.
…cs/QUALITY_DELTA_CATALOG.md The "What --quality-delta catches" placeholder (slide 29) is filled from the catalog (commit 553dbadb on docs/quality-delta-catalog-2026-09-11). Every example commit is on main 766913d; the catalog itself has not landed. Four cards, varied in kind and in what the finding led to: - redhat-et#1 new-clone-of-reused-helper: vendoredPathPrefixes re-spelled registeredMacroNames, a helper with three callers, while the round that built the per-kind dials was under way; both now call mergeBuiltinsWithConfig. - redhat-et#6 nesting (and complexity 141 -> 178): the inline per-root scan became std::any_of in a named function. - redhat-et#7 dead-code: the predicates a rewrite orphaned were deleted, not acked. The card also names the one the kind missed, namesFileNotKept, whose only caller was itself dead (IP-8). - redhat-et#3 duplication, Type-3: the shared digit-run locator became headerFieldDigits, and the card says the detector still gates a 63-token residue, which the catalog classes as a false positive (IP-4). Each card's speaker notes cite the catalog section, the base -> before -> after commits, the verbatim row(s), and a one-line reproduce command copied from the catalog's Reproduce block. All five catalog examples considered for the slide (redhat-et#1, redhat-et#3, redhat-et#6, redhat-et#7, redhat-et#8) were re-run on 2026-09-11 with a built_from=766913d02 build, and every row, gating count and exit code matched the catalog. Why four and not six: the render shows a two-column pane holds 10 lines of 39 columns at 8 pt (a 40-column line wraps). With 5 or 6 entries each pane drops to about 3 lines, too few for a before and an after. So redhat-et#5 and redhat-et#8 are left out, and redhat-et#3 carries the detector-limits card. Layout changes in qdExamples, each forced by the render: - The kind box is sized to its text. new-clone-of-reused-helper is 2.38 in at 11 pt and wrapped out of the old fixed 2.2 in box. - Card padding and pane gap are tighter, and two-column code is 8 pt (was 8.5), so a pane fits 10 x 39 instead of about 9 x 36. - Entries may carry an optional notes array, which is validated and printed under that card's line in the notes. - The kicker for a filled slide no longer reads as a placeholder. The placeholder path is unchanged. Snippets are the catalog's excerpts, trimmed to fit. "…" marks elided code, long lines are wrapped, and a comment line stands for code the catalog says was elided or deleted. The notes say so. Gates on this tree: deckcheck ALL PASS (0 bad values, 0 stale), deckclaimcheck ALL PASS (179 long flags, 34 slides), readmedriftcheck ALL PASS. All three pass with the built_from=766913d02 build, and again with the scratchpad main_096e3544 build. The pptx and PDF are regenerated (pptxgenjs 4.0.1, LibreOffice impress_pdf_Export, 34 pages). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…not list every row classifyHelpLine decides a line's role by indentation alone: four spaces is a row, five or more is the continuation prose tier 1 drops. Twelve rows in src/cli.h were written at SIX spaces, to read as sub-flags of the entry above them. Tier 1 therefore culled fifteen flags the parser accepts -- --and --not --grep-in --grep-scope --grep-context --grep-before --grep-after --ack-only --scope --edit-payload --edit-target-file --no-post-check --dry-run --apply --partition -- and v0.6.0 shipped that way. They all work; they were simply unfindable from the first screen. The sharpest edge was `--help=--and`, which refused with "`ripwire --help` lists every row" -- a false claim about our own output, which is the part non-negotiable redhat-et#3 forbids and the reason this is a bug and not a layout preference. FOUR SPACES BECOMES THE SINGLE DEFINITION, on both sides. docs/docs_commands_build.py's parse_help accepted 4..6 while the binary accepted only 4, and being the more permissive reader is exactly what hid the divergence: COMMANDS.md documented all fifteen, so nothing downstream of the generator ever showed a symptom. It is narrowed to exactly four here. The alternative -- teaching classifyHelpLine about six -- would keep the visual nesting but put a second legal row indent into a surface whose whole contract is that indentation decides, and would force helpbudgetcheck arm (C) to stop calling deep lines continuations. One indent, one meaning, is the cheaper invariant. Narrowing also lets docscommandscheck arm (H1), whose candidate detector deliberately still scans 4..6, see a future six-space row as a parser loss. THE GATE COULD NOT CATCH THIS AND REPORTED THAT AS FINE. helpbudgetcheck's `rows` extractor is the emitter's own four-space rule, so a misclassified row is absent from BOTH sides of arm (B)'s comparison: (B) printed "the split culled nothing" about exactly the rows that were culled, and its comment claimed every row was fenced. Teaching `rows` about six spaces would only move the blind spot. New arm (K) takes its population from the other side of the binary instead -- the kBoolFlags / kViewFlags / kIntFlags tables and the hand-written parseArgs arms, scraped by test/flaguniverse.py, which three other gates already share -- and asserts that every flag the tool ACCEPTS is named in tier 1. It was observed RED on the unfixed tree (14 flags) and is green after. Arm (L) is its mutation control. Eleven accepted-but-unadvertised flags sit on a committed debt list with a reason each, so that gap is on the record instead of inside an extractor's blind spot. The eight rows that had continuations now open with a summary that stands on its own, as the tier-1 contract requires and as every four-space row already did; the sentence each used to open with is kept verbatim as that entry's first continuation line, so tier 2 loses nothing. Tier 1 goes 4480 -> 4844 tokens against an unchanged 7000-token ceiling. Re-pinned deliberately: test/printf_parity.manifest, labels `help` (6a15d43c -> 3290aebf) and `help_all` (5d748a4d -> 141428b5), 40 labels unchanged; docs/COMMANDS.md regenerated (13 lines, all of them these rows' own text). Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
…ing, and the no-throw copy threw
Six findings from one review, every one of them a surface that was silently wrong rather
than loudly broken.
--pr-context PRINTED A WRONG est_tokens WITH NO DISCLOSURE. When a trim level's measurement
render fails, prRenderLevel returns an EMPTY body; pickPrTrimLevel priced that empty body, the
price fit, and the ladder broke at level 0 — while writePrContext correctly streamed the
complete untrimmed floor through emitFiles( out, kPrTrims[0], nullptr ). The only signal was
DEGRADED_PATH_ALERT, which src/infra/Diagnostics.h compiles to `do {} while (0)` under NDEBUG,
so the binary a user installs printed a modelled number with nothing at all saying so
(non-negotiable #3). The BYTES were never the defect and do not move: cutting answer rows
because a measurement buffer failed would let a cap decide the content, which is the one thing
a cap may never do. The fact goes where this class of fact already lives — truncated= now
carries ";est-unmeasured", re-priced with the label in place (the label lengthens the root tag),
and the legend defines it in budget-floor-exceeded's own voice.
THE CHARGE IS READ OFF THE LABEL, not off a boolean beside it. prPriceDocument decides the
clause from the truncated= value it is already handed, so ONE condition decides both the priced
legend and the delivered legend and they cannot drift apart; and the two conditional clauses now
arrive as a named PrLegendClauses{ runHint, estUnmeasured } rather than two bare bools, because
`prLegendText( escBase, unindexed, true, false )` says nothing at its call site about which
clause is which. Both readings came out of --quality-delta: threading a seventh parameter into
prPriceDocument and a fourth into prLegendText took the range form to gating="2" (a params row
from minor to major, and an api-surface contract change that invalidated a standing ack). The
range form is gating="0" now with NO new ack — the findings are gone rather than suppressed.
THE DEFINITION IS LABEL-GATED, which the gate found for me. Spliced unconditionally, the ~390 B
clause cost test/defaultceilingcheck.sh's fixture its entire remaining headroom: that 120-file
tree prices at 7,989 of the 8,000 default — 11 tokens spare, as E1 measured when it gated the
run-hint clause for the same reason — and went to 8,037, over budget on a document with nothing
unmeasured about it. So kPrEstUnmeasuredLegendClause rides exactly the document that carries the
label, decided by the fact the ladder recorded (PrTrimRender::rendered) and never by a search of
the rendered bytes; the pricer charges its size on the same fact, so the priced legend and the
delivered legend cannot disagree. A healthy document is byte-identical to before (est_tokens
7,989, re-measured) and prcontextcheck (F-legend) holds it that way.
THE LABEL CROSSED prBudgetTail's BUFFER. test/fixedbufsweep.sh had this buffer at 248 B of
tail[256] — "SEVEN bytes of margin ... one more attribute crosses it" — and ';est-unmeasured'
is 15 more and CAN ride beside ';budget-floor-exceeded' (a small --max-tokens puts even the
unmeasured empty-body envelope over budget). Worst case 88 lit + 90 digits + 85 label = 263 B,
so tail[320], 56 B of margin, and the sweep's row moves in this commit with the recomputed
number. rw::formatTo was not what had been saving it: it truncates SILENTLY and its return is
not read there, so an overrun would have dropped the closing quote of truncated=" and shipped a
malformed root — a G4 breach with no diagnostic.
renderToString's NO-THROW CONTRACT HAD A THROWING LAST STATEMENT. out.text.assign( buf, sz ) is
the one allocation on the success path and it sat outside the handler, so a std::bad_alloc from
it escaped a function documented to ALERT a failure and return ok == false, and jumped the
std::free( buf ) two lines below on the way out — leaking the memstream buffer. Caught in its
own handler rather than one around the whole body, because the two failures need different
cleanup (the emitter's throw owns an OPEN stream; by this point only buf is left), with its own
alert literal, and control falls THROUGH to the single free() so buf is released exactly once on
every path.
THE SHARED ROW READER'S MALFORMED-FIELD DETECTOR HAD A HOLE OF ITS OWN SPECIES.
test/testrowpaths.py found "tests_to_run" and then scanned arbitrarily far forward for a '[', so
{"tests_to_run":null,"other":[{"p":"ghost.cpp"}]} sliced the NEXT field's array and returned
ghost.cpp at exit 0 — a foreign field's paths served as this field's answer, where the docstring
already promised a TestRowParseError. The value is read adjacently now: past the key, a ':',
optional whitespace, then '[' or raise.
AND TWO PATH READERS HAD NEVER BEEN CONVERTED. The census over test/ for the four shapes the
shared reader replaced found test/affectedcheck.sh's tset() — inside the very file the reader's
docstring names among those it converted, so that claim was false — splitting EVERY row's p= on
',' including a single row's, which turns a comma-bearing path (never grouped, by testmap.h's
refusal) into two names that name nothing; and test/testgatecheck.sh's tset() matching `<t p=`
singles only, which returns the EMPTY set on a two-runner-less-test fixture where the shared
reader returns both paths. Both route through the shared reader now, and the docstring records
the census. Every other hit counts rows (listingpagingcheck, w3fixlegendcheck, testgatepagecheck
— all group-aware in place) or pins one exact row spelling with a regex that fails loudly;
deeptailcheck's `<t p=` rows are --for's tail listing, a different element sharing the tag.
Two documentation drifts beside them: skills/ripwire-mcp/SKILL.md claimed `p` for
situational_awareness, which emits `test` (src/mcp.h's kTestRowJsonShapeClause states the split
and the binary is the authority), and bench/arb/run_arb.py decoded a , the seam stopped
emitting on 2026-09-13 while decoding NONE of the entities it does emit — so a path holding '&'
was scored against a file name that does not exist. Both row shapes there share one decode now.
GATES, all red on the parent commit and green after:
* test/prcontextcheck.sh (F-legend)(F5)(F6) — est-unmeasured in truncated= on the degraded
root, the complete body still served, and the legend defining the term. RED: the degraded
root printed truncated="none" while pricing an empty body, and no legend defined the label.
* test/prcontextcheck.sh arm (G) — INFRA_FAULT_RENDER_COPY_THROW, the emitter switch's twin,
injected immediately before the assign. RED: "produced no DEGRADED_PATH_ALERT on a binary
that PROVED it can emit one". Honest in both flavours, mirroring arm (F): the switch and the
alert live only on the non-NDEBUG build, so the plain leg proves the degrade and the NDEBUG
leg asserts the verb is intact and no false disclosure appears. The est-unmeasured LEGEND
definition is asserted on EVERY flavour, which is the point of moving the disclosure off the
alert.
* test/testrowruncheck.sh arm 17 — every non-array tests_to_run value is exit 2 in both paths
and jsonlist, with a well-formed array and JSON whitespace as controls. RED: five documents
at rc=0, three of them serving ghost.cpp.
Suite: 628 gates, 626 pass, 0 fail, 2 environmental skips (argvdiffcheck and editchecknotecheck,
both wanting a reference binary), tree_writes=0. ASan/LSan clean on --pr-context healthy and on
both injected degrades.
Pins moved, two, both in test/fixedbufsweep.sh and both with the measured recomputation in the
same commit: the src/prcontext.h tail TABLE row 256 -> 320 (worst case 248 -> 263 B, margin 7 ->
56), and EXPECTED mentions 323 -> 324 — re-read from the diff, not accepted from the delta: the
one added line is the COMMENT explaining that growth, which names formatTo. calls, sites, rows
and widthforms are unchanged at 219/219/92/0. No legend or byte pin moved.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
The private-era preprint (deterministic model-free localization: the measured floor, the C++ benchmark, and the pre-registered negative result) ported into the public tree — renamed, scrubbed, and every number reconciled against docs/EVALS.md and its pinning files. Kept/updated/removed ledger in the lane record; highlights: the draft's own +4.9pp arithmetic error corrected to the published +3.9pp, EVALS §8's refused numbers absent, the shipping headline moved to the current 60.9%, the cross-paper LLM-gap subtraction dropped per the paper's own comparability rule, and "first public C++ benchmark" hedged (no gate can check a literature claim).
Also: docs/LINEAGE.md's fifth pillar now points at paper/, docs index lists it, and deckcheck scans paper/*.md as a prose surface (red-first proven; it caught a fabricated flag in the paper on first run).
🤖 Generated with Claude Code