Skip to content

paper/: the working preprint draft, ported under export discipline - #3

Merged
joyful-ii-V-I merged 5 commits into
mainfrom
paper
Aug 3, 2026
Merged

joyful-ii-V-I merged 5 commits into
mainfrom
paper

Conversation

@joyful-ii-V-I

Copy link
Copy Markdown
Collaborator

The private-era preprint (deterministic model-free localization: the measured floor, the C++ benchmark, and the pre-registered negative result) ported into the public tree — renamed, scrubbed, and every number reconciled against docs/EVALS.md and its pinning files. Kept/updated/removed ledger in the lane record; highlights: the draft's own +4.9pp arithmetic error corrected to the published +3.9pp, EVALS §8's refused numbers absent, the shipping headline moved to the current 60.9%, the cross-paper LLM-gap subtraction dropped per the paper's own comparability rule, and "first public C++ benchmark" hedged (no gate can check a literature claim).

Also: docs/LINEAGE.md's fifth pillar now points at paper/, docs index lists it, and deckcheck scans paper/*.md as a prose surface (red-first proven; it caught a fabricated flag in the paper on first run).

🤖 Generated with Claude Code

joyful-ii-V-I and others added 3 commits August 3, 2026 11:17
The draft was written pre-rename and pre-export. It is REWRITTEN here rather
than copied: the old tool name is gone from prose, titles and the author line
("The ripwire authors" — no personal names by policy), and every measured
number was reconciled against docs/EVALS.md and the bench/ files it names.

Four numbers did not survive reconciliation, and the paper says so in place
rather than quietly swapping them:

  - The headline held-out strict file@10 was 66.7%. That figure belongs to the
    BASELINE ARM of the anchor-hop experiment at that experiment's binary, and
    EVALS.md §8 refuses it in the role the draft used it for. §4.1 now leads
    with the published 60.9% (bench/locbench/README.md) and 66.7% appears only
    in §6.2, labelled as the arm it is.
  - The mention anchor's single-file stratum was +4.9pp. The machine-generated
    scoreboard gives 86.3% vs 82.4% = +3.9pp, which is also EVALS.md §4's
    figure. The draft's number was arithmetic error.
  - "LLM agents buy roughly +20pp over this floor" is removed: the subtraction
    depended on the retired 66.7%, and it compared a symbol-max-pooled held-out
    slice against a file-document full-set number across two papers — the exact
    comparison §3.1 forbids.
  - The private-corpus C++ absolutes are not restated at all (EVALS.md §8);
    §4.3 carries EVALS.md §7's framing instead, which is the publishable form.

The two-tier gate is also demoted from "policy" to what GATE_DECISION.md
actually says it is — a proposal, opt-in in the comparator — while keeping the
true and narrower claim: it was pre-registered and applied to §6's two rejects.

The negative result is the centerpiece and keeps its detail, now with the
sequel EVALS.md §7 carries (the +0.41pp / LB +0.00pp anchor candidate, also
rejected). §9 names the second contribution being folded in — the audit loop,
the two-flavour CI build, sibling-completeness, and the gated-claims machinery
that scans this very file.

Every table names the in-repo artifact that pins it; all cited paths verified
present in this tree.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
LINEAGE.md's "what is actually new" closed on a paper "in preparation" with
nothing to open; it now names the directory. Sentence shape unchanged, so
readmedriftcheck.sh's count arms are untouched — re-run anyway, ALL PASS
(E8 additionally re-verifies every repo-relative path LINEAGE cites, now 31).

docs/README.md's "Outside this directory" list gains the paper/ row, with the
status stated in the row itself so the index cannot imply a finished artifact.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A paper is the maximal fabricated-flag surface: long, written in one pass,
quoting flags in running prose rather than in command blocks, and re-read for
argument rather than for syntax — so a plausible-but-wrong flag survives longer
there than in a deck, which is the failure this gate was built for.

Follows the file's own V3 MED-1 convention rather than only bumping the total:
paper is a tagged glob FAMILY with its own paperCount and its own non-empty
assertion, so the family cannot vanish behind a total the other four still
satisfy. Total floor 38 -> 40.

RED FIRST, as the convention requires: with paper/ scratch-removed the new
assertion fires before the total floor gets a chance to —

    deckcheck: 0 paper/*.md prose source(s) found — refusing to run
    EXIT=2

and with it restored the gate scans 44 files, ALL PASS. It earned its keep on
the first green run: it caught a metasyntactic "--flag" in the paper's own §9
sentence about this very gate. Fixed in the prose, NOT allowlisted — a
placeholder is not a claim, and an allowlist entry for one would be the hiding
place this gate's rot check exists to prevent.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Aug 3, 2026

Copy link
Copy Markdown

Warning

Review limit reached

@joyful-ii-V-I, you've reached your PR review limit, so we couldn't start this review.

Next review available in: 36 minutes

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

How can I continue?

After more reviews become available, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews.

How do review limits work?

CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability.

For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window.

Please refer docs for additional details.

Review details
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: e5720788-2528-4604-a94e-07d4a2f68e46

📥 Commits

Reviewing files that changed from the base of the PR and between 2f8e282 and 3bfa853.

📒 Files selected for processing (3)
  • bench/multiswe/README.md
  • docs/EVALS.md
  • paper/PREPRINT.md
📝 Walkthrough

Summary by CodeRabbit

  • Documentation

    • Added a comprehensive preprint covering the methodology, benchmark results, comparisons, limitations, reproducibility guidance, and claim-auditing process.
    • Added documentation describing the preprint, its evaluation claims, draft status, and related artifacts.
    • Updated the project lineage documentation to link the working draft.
  • Tests

    • Expanded documentation validation to scan and require paper sources, improving coverage of research materials.

Walkthrough

The change adds a ripwire preprint covering methodology, benchmarks, results, experiments, limitations, reproducibility, and artifacts. Repository documentation links to the draft. deckcheck.sh now scans and requires paper sources.

Changes

Preprint documentation and validation

Layer / File(s) Summary
Preprint scope and evaluation framework
paper/PREPRINT.md
Defines ripwire, benchmark methodology, scoring, reproducibility checks, and acceptance gates.
Benchmark results and experiment analysis
paper/PREPRINT.md
Reports benchmark comparisons, ablations, anchor-hop designs, and rejected expansion experiments.
Reproducibility and publication record
paper/PREPRINT.md, paper/README.md
Adds limitations, reproducibility controls, claims-audit plans, references, artifact locations, and draft metadata.
Documentation links and paper validation
docs/LINEAGE.md, docs/README.md, test/deckcheck.sh
Links repository documentation to paper/ and validates paper Markdown sources and source-count requirements.

Estimated code review effort: 3 (Moderate) | ~20 minutes

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Title check ✅ Passed The title clearly identifies the addition of the working preprint draft under paper/.
Description check ✅ Passed The description directly explains the preprint addition, reconciled measurements, documentation links, and deckcheck updates.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Linked Issues check ✅ Passed Check skipped because no linked issues were found for this pull request.
Out of Scope Changes check ✅ Passed Check skipped because no linked issues were found for this pull request.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch paper

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Actionable comments posted: 4

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Inline comments:
In `@paper/PREPRINT.md`:
- Around line 24-26: Qualify the “sub-second” claim by specifying the applicable
latency statistic and cache condition, noting the 1,623.8 ms cold p95 rather
than presenting it as unconditional. Apply the same wording in paper/PREPRINT.md
lines 24-26 and paper/README.md lines 6-8 so the abstract and paper index
consistently describe the evaluated ranker.
- Around line 301-306: Remove the private-corpus numeric results and the
comparison against the public SFML figures from the paragraph beginning “The
public-versus-private divergence.” Retain only the statement that the private
corpus was removed, its results are unreproducible, and docs/EVALS.md §8 does
not publish private absolutes.
- Around line 314-327: Update the comparison description in PREPRINT.md to
reconcile the “four arms” wording with the five table rows by identifying the
aider repo-map no-persona row as an auxiliary sensitivity run, or by describing
the comparison as five arms. Explicitly state which measured rows underpin the
win/loss and latency claims.

In `@paper/README.md`:
- Around line 9-11: Update the machine-checking claim in the paper text to say
“selected publication claims” instead of “every published claim,” unless the
implementation can document and demonstrate complete binary-check coverage for
all listed benchmarks, metrics, gates, and command references.
🪄 Autofix (Beta)

Fix all unresolved CodeRabbit comments on this PR:

  • Push a commit to this branch (recommended)
  • Create a new PR with the fixes

ℹ️ Review info
⚙️ Run configuration

Configuration used: Organization UI

Review profile: CHILL

Plan: Pro Plus

Run ID: cb3f2804-2071-4be9-8894-f19d0d741aa6

📥 Commits

Reviewing files that changed from the base of the PR and between 72e43cc and 2f8e282.

📒 Files selected for processing (5)
  • docs/LINEAGE.md
  • docs/README.md
  • paper/PREPRINT.md
  • paper/README.md
  • test/deckcheck.sh

Comment thread paper/PREPRINT.md
Comment on lines +24 to +26
design space: a **deterministic, zero-runtime-dependency, sub-second, symbol-granularity** static
ranker (ripwire) that uses no model, no network, and no randomness, and whose output is
byte-identical run-to-run. We contribute: **(1)** a corrected LocBench evaluation methodology for

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟠 Major | ⚡ Quick win

Qualify the “sub-second” headline.

The paper reports a 1,623.8 ms cold p95 for an evaluated candidate. State whether “sub-second” means warm p95, median latency, or another specific path. Do not present it as an unconditional property of the evaluated ranker.

  • paper/PREPRINT.md#L24-L26: add the latency statistic and cache condition to the abstract.
  • paper/README.md#L6-L8: use the same qualified wording in the paper index.
📍 Affects 2 files
  • paper/PREPRINT.md#L24-L26 (this comment)
  • paper/README.md#L6-L8
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@paper/PREPRINT.md` around lines 24 - 26, Qualify the “sub-second” claim by
specifying the applicable latency statistic and cache condition, noting the
1,623.8 ms cold p95 rather than presenting it as unconditional. Apply the same
wording in paper/PREPRINT.md lines 24-26 and paper/README.md lines 6-8 so the
abstract and paper index consistently describe the evaluated ranker.

Comment thread paper/PREPRINT.md Outdated
Comment thread paper/PREPRINT.md
Comment on lines +314 to +327
Four local-first, no-API-key arms, identical checkouts, identical gold, and the **same metric code
imported unmodified** from the LocBench harness; zero exclusions in any arm. Versions pinned in the
report: aider-chat 0.86.2 (repo-map, fed the issue through its own ident-extraction),
codebase-memory-mcp 0.9.0 (its shipping BM25 graph-search path), graphifyy 0.9.15 (local BFS
traversal, `PYTHONHASHSEED=0` — see below). The slice is deliberately hard: 40 of 60 instances have
multi-file gold.

| arm | file@1 | file@3 | file@5 | file@10 | any@10 | median wall |
|---|---|---|---|---|---|---|
| **ripwire `--for`** | 5.0% | **18.3%** | **26.7%** | **36.7%** | **75.0%** | **0.074 s** (warm, pre-built index) |
| codebase-memory-mcp | **6.7%** | 10.0% | 16.7% | 26.7% | 66.7% | 1.14 s |
| graphify (BFS order) | 0.0% | 3.3% | 5.0% | 21.7% | 41.7% | 5.8 s |
| aider repo-map (personalized) | 0.0% | 1.7% | 6.7% | 13.3% | 33.3% | 2.5 s |
| aider repo-map (no persona) | 0.0% | 1.7% | 3.3% | 10.0% | 26.7% | — |

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win

Clarify the head-to-head arm count.

The prose says the comparison has four arms, but the table contains five measured rows because aider repo-map appears with and without a persona. Label the no-persona row as an auxiliary sensitivity run, or describe the comparison as five arms. State which row supports the win/loss and latency claims.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@paper/PREPRINT.md` around lines 314 - 327, Update the comparison description
in PREPRINT.md to reconcile the “four arms” wording with the five table rows by
identifying the aider repo-map no-persona row as an auxiliary sensitivity run,
or by describing the comparison as five arms. Explicitly state which measured
rows underpin the win/loss and latency claims.

Comment thread paper/README.md
joyful-ii-V-I and others added 2 commits August 3, 2026 11:42
…d three provenance sentences

F1: §3.2×2 + §7 claimed zstd/jq/ponyc live in bench/multiswe/dataset.lock;
the lock's languages list is ['cpp'] and its own README says the C split
needs a separate mine. Now: minable by the same harness, not in the lock.
F2: the REGISTER was wrong — EVALS §8 said 'train slice' for a number its
own pinning file books under heldout_acceptance; fixed there (n=243) and
noted in §6.2 per the port's own header rule.
F3: 'does not restate them' was false four lines under two rounded
restatements; the sentence now says exactly what is and is not carried.
F4: §1(c)'s universal hedged; C1 quotes the hedge it actually uses; the
multiswe README H1 hedged at the pin's destination.
Plus V9's two nits (§4.2 modifier, §4.5 wall/inst pointer).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…k caught it

The paper/ scanning arm flagged --languages (a mining-harness flag, not a
ripwire flag) four commits after it was added for exactly this class. The
pipe in my gate chain had swallowed deckcheck's exit code; the follow-up
commit is the penalty. Prose now names the capability, not the flag.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@joyful-ii-V-I
joyful-ii-V-I merged commit dac40d6 into main Aug 3, 2026
6 checks passed
joyful-ii-V-I added a commit that referenced this pull request Aug 7, 2026
…d now says so

Measured on this repository: 20 of the 40 default-preset panel rows were
src/main.cpp symbols, every one of them reporting the identical
`hrank=0 churn=47`. Churn is mined per PATH, so every symbol in a file
carries that file's churn=/hrank= verbatim and a symbol in a churny file
collects the historical family for free, with no property of its own.

That is the family's real unit, not a bug — ensemble.h has always ranked
it by churn alone, deliberately, so it cannot agree with structural by
construction. What was wrong is the legend a reader meets first: it said
"historical (git change frequency)" and stopped, in a report whose entire
premise is evidence families that must be able to DISAGREE. A reader
counting an inherited file fact as per-symbol corroboration is reading
the fam= column wrong, and non-negotiable #3 makes that a defect in the
output rather than a caveat for the docs.

Both legends now name the unit where they introduce the family, and both
state the consequence: file evidence INHERITED by the row, not the row's
own history. Nothing else moves — no ranking, no threshold, no family
set, no preset. The historical family's exclusion from `strict` and its
churn-alone ranking are unchanged and still measured, not re-argued.

Gate: test/qualitypanelcheck.sh arm (N), new, and in that order on
purpose — (N1) MEASURES the property from the binary's own output (both
hub.cpp rows must carry one identical historical evidence string, with
the once-committed file firing historical on no row as the control, so
(N1) pins a shared value rather than a constant one), and only then (N2)
asserts the disclosure in both the quality-panel and --ensemble legends.
A gate that pinned the sentence alone would keep passing if the family's
unit ever changed underneath it. Both (N2) arms verified red against the
pre-change binary.

Verified: qualitypanelcheck.sh PASS standalone; full suite pargates.py
-j 6 ALL PASS (362 gates, 3 standing skips); determinism double-run diff
clean; xmllint clean.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Aug 9, 2026
…call above the shadowing local no longer vanishes (kParserVer 55)

Fourth A5 iteration, and the first refutation that was a RECALL loss rather
than over-suppression: iteration 2 scoped suppression to the declaring block
but started the span at the block's opening BRACE, so a real call written
ABOVE the local was silently eaten. Verifier attack4 —

    void key() { }
    void caller() { key(); int key = 0; (void)key; }

— went base 2 (one right, one false) → postS/postSfix/postSfix2 count=0. Zero
is the wrong answer twice over: it breaks the prereg's non-regression #5 and
CLAUDE.md non-negotiable #3, and 79df919's "within-block declaration-point
ordering, disclosed" disclosed the wrong thing — a disclosed floor may
under-suppress, never delete a real site.

THE DECLARATION POINT SHIPPED IS THE END BYTE OF THE COMPLETE DECLARATOR,
which is C++ [basic.scope.pdecl] exactly (immediately after the complete
declarator, before its initializer; for a structured binding, after its
identifier-list — the outermost declarator's end byte is both). Chosen over
the end of the declaration STATEMENT because it is the only boundary that
gets a multi-declarator statement right in both directions:
`int a = probe(), probe = 0, b = probe;` keeps the call in a's initializer
AND suppresses b's read. It is a byte the grammar hands us, not an estimate;
what remains is the pre-existing name-matching floor, now merely visible
above the point too (a kept pre-declaration site is kept, never resolved).
enclosingBlockSpan becomes enclosingShadowScope and reports plainBlock, so
narrowing applies to an ordinary block ONLY: control-statement header
declarations (iteration 3) and every whole-scope shape — definition/lambda/
catch parameters, captures and init-captures, range-for variables — keep
their whole-statement/whole-body spans, because their names ARE in scope from
the start and narrowing them would re-mint iterations 1-3's false positives.

Un-masked defect, fixed in the same diff: narrowing stops covering the
declaration LINE, which exposed that isDeclSiteName never saw a `declaration`
parent at all — `int key;` leaked its own declared name as a role="read", and
`int a, key;` leaked the second one even where the parent type IS listed,
because a `declaration` carries one `declarator` FIELD PER NAME and
ts_node_child_by_field_name returns the first. Iterations 1-3 could not
observe either: the block-start span suppressed the declaration line along
with everything else.

Verifier fixtures: attack4 0→1 (the genuine call, and only it);
attack2a-c/3a_for/3a_if/3a_extent/3a_nestedwrap/3b_catch/3c_goto/
3d_genericlambda and all eleven attack5_* hold their counts exactly.
test/shadowcheck.sh grows 31→50 arms and each new arm is proven red on a
binary that fails it: 4 on the 54 binary (pre-declaration call in --uses and
in --callers, multi-declarator pre-declarator call, structured-binding
declaration), 2 on a boundary mutant that uses the statement end (b's read,
`int probe = probe;`), 2 on the un-fixed narrowing build (bare and
second-declarator name leaks), 9 on a mutant that unifies the whole-scope
shapes onto the enclosing block (param, range-for, catch, init-capture,
sibling blocks) — none was green everywhere.

Bite on src (54 → 55): --uses=key 10→10 with a byte-identical site list,
line 44→42, name 294→290, run 6→5, empty 1551→1551, count 11→11. ZERO sites
reappeared — src contains no pre-declaration-call instance of the bug class —
and all seven that vanished are declaration-site names hand-checked at
source: `std::string line;` (recall.h:400,457), `std::string name, eq, s;`
(arch.h:376), `std::string_view file, name;` (graph.h:2139,2433),
`std::string run;` (search.h:691), `std::string_view fileFilter, name;`
(layout.h:2359) — the leak this commit closes, not shadow suppression.

Span values and the reference population change → kParserVer 54→55 + quality.h
mirror + qschemetrip re-pin. EVALS gate count unchanged (no new gate script).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Aug 14, 2026
…ness into SARIF

Adversarial-verifier round on the just-merged idea-harvest wave. Five findings,
four fixed as briefed and one refuted by measurement.

F1 (MAJOR) — silent built-in subtree prunes. Three prune paths converge on
ingest.cpp's disable_recursion_pending(): a --exclude match, the committed
denylist (kCrawlSkipDirs), and the CMakeCache.txt build-output sentinel. Only the
first incremented a counter, so `--skipped` over a tree with node_modules/
reported EVERY counter zero while whole subtrees had been dropped — the honesty
contract's own failure mode one level up from the class §L1 closed. Built-in
prunes now land in their own CrawlSkips::prunedDirs, emitted as pruned_dirs= on
the <skipped> root and NOT folded into excluded_dirs=: "a rule you passed" and "a
rule this build carries" are different answers to "why is my tree missing", and
one counter cannot say both. The legend carries the same "contents UNKNOWN, not
zero, and in no count here" language for the new class. Multi-root sums it.

F2 (MAJOR) — SARIF dropped two facts the XML states, and both are the difference
between "ran and found nothing" and "never ran here". tool.driver.rules was
byte-identical whether --lint-select= kept a rule or not, and inert rules read as
ran-clean. Deselected rules now carry defaultConfiguration.enabled=false (SARIF's
own "catalogued, not enabled"; kept rules omit it and inherit the standard's
default of true); every rule carries properties.applicable, true as well as
false, because a SARIF document travels without a legend and an absent property
would be read as true; runs[0].properties mirrors the XML root's selected="K of
N" plus the raw select=/ignore=, and only under a real selection.

F3 (MINOR) — REFUTED AS BRIEFED, closed on its other half. The brief asked for
unindexed=/unindexed_exts= clauses in the DEFAULT MAP legend. Measured against a
pristine HEAD binary: on src/ at --max-tokens=500 the map's fixed floor is 1173 B
against a 1180 B allowance — 7 bytes of headroom — so any clause, conditional or
not, reds test/tokenbudgetcheck.sh arm #3; the shortest honest wording is ~150 B.
serialize.h's own note already recorded this constraint and it holds. What WAS
genuinely undefined everywhere is unindexed_exts=, the TOP-6 cap bit, and that is
now defined in the --skipped legend beside unindexed=, which is under no token
budget. The map legend is left byte-identical, the refutation is recorded at
buildUnindexedAttr with the measurement, and skipreasoncheck arm (9) now pins the
negative so a clause cannot land there silently.

F4 (MINOR) — .ripwire_quality_acks carried 8 duplicate (kind,key) rows from a
three-way merge, violating its own "a file this binary wrote never contains a
duplicate key" invariant. Deduped by quality.h::readAckRecords' D2 rule verbatim
(max(ackNow) wins, strict > so a tie keeps the first line); --quality-delta output
verified BYTE-IDENTICAL before and after. No key lost, no ratchet floor lowered.

F5 (MINOR) — --recall's per-doc lines= is the selected section range even after
the byte budget cuts that body mid-section. One clause on the header, charged only
to a run that truncated something and appended after the last attribute so the
name=value tail stays scannable. No behaviour change.

Gates, RED against a pristine HEAD binary built from `git archive HEAD` and green
after: test/skipreasoncheck.sh arms (8) pruned_dirs and (9) unindexed_exts;
test/sarifcheck.sh arms 8 (selection crossover), 9 (applicability crossover) and
10 (mutation control — the unfiltered C-family run must show the negative of both).

Every quality regression this round opened was FIXED, not acked: emitLintSarif's
complexity 11->20 / verbosity 44->69 / params and api-surface 5->6 all back to
baseline via writeSarifRuleDecl + writeSarifResult + writeSarifRunProperties and
one SarifRunProperties struct instead of a parameter per disclosure;
emitRunLintSarif 8->12 params back to 8 by taking the MainDispatch every other
handler already takes; formatRecallHeader 55->67 LOC back by hoisting the
rationale above the function. The 8 acked rows are short-horizon-churn=self only —
this change's own edit window on the verbs a disclosure gap cannot be closed
without touching.

pargates 415 gates: 412 pass, 3 skip, 0 fail. --quality-delta exit 0, gating=0.
--skipped / --lint --sarif / --recall deterministic run-to-run; XML xmllint-clean;
SARIF parses as JSON. README --callers line numbers and docs/COMMANDS.md
regenerated from the binary.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Aug 15, 2026
… fair bytes gate, conditional legend (verifier #3 #4 #5)

# Conflicts:
#	README.md
joyful-ii-V-I added a commit that referenced this pull request Aug 20, 2026
…ng (R9)

E6 / the wave-2 growth round reproduced "<bodies> ABSENT ENTIRELY" instead
of shown="0" on four of five --token-budget=3000 --for queries across four
corpora (routing note R9, exp-e6.md A5): the honesty contract ("a zero
means none found, never none exists", CONTRIBUTING.md #3) applies to
elements, not only counts, and this violated it twice over.

src/main.cpp buildForAutoBodies (the --for / MCP `for` T3 auto-bundle
path) had two branches that discarded the element instead of keeping it:
  - autoBodyIds.empty() (no positive-score candidates at all) returned
    without ever calling packBodies, so no <bodies> render existed to keep.
    Now calls packBodies with the (empty) candidate list, which renders the
    honest "<bodies shown="0" total="0" capped="0"></bodies>" shell on its
    own -- no second hand-rolled tag.
  - autoEmitted.kept.empty() (candidates existed, none fit) had ALREADY
    rendered that same honest shell via the chargeSection() call just
    above, then threw it away ("drop the empty section whole"). Now keeps
    it; only the attr= reason on <ctx> is new information.

src/packtask.h packTaskBundleText's bodiesStr had the identical gap: the
`else` of `if (!bodyIds.empty() && bodiesBudget >= kPackTaskSectionFloor)`
left bodiesStr empty, so the whole element vanished. Fixed with a bare,
byte-bounded "<bodies shown="0" total="N" capped="0|1"></bodies>" tag --
deliberately NOT routed through packBodies here: an earlier version of
this fix did call packBodies (truncateOversizedFirst=false, reusing its
per-item "body omitted (over budget)" comments), but those comments are
UNBUDGETED and pushed a --token-budget=2000 run from 5428 B allowance to
5692 B (packtaskcheck.sh's own ceiling arm caught it). The bare tag is a
small, fixed cost, same shape restatePackTaskBodiesWrapper already
hand-formats a few lines below for the same reason.

src/packtask.h::packTaskListSection (backing <callers>/<far>/<notes>/
<tests>) has the identical defect and is DELIBERATELY left alone here --
a separate, larger blast radius; a follow-up task is spawned for it.
test/packtaskcheck.sh's pre-existing "budget EXCEEDED" failure at
--token-budget=2000 is ALSO pre-existing and unrelated (verified against a
clean pre-lane baseline binary); a separate follow-up task is spawned for
that too.

Gate: test/bodiesshowncheck.sh (red-first; registered in
test/regression.sh in this commit). Six arms across all four scenarios
(--for/--pack-task x candidates-exist-but-none-fit/no-candidates-at-all):
<bodies shown="0" total=N capped=C> present and arithmetic, the existing
bundle=/reason= attributes unchanged, the non-empty case still reports a
real shown=N (proving this isn't an always-shown=0 false pass),
well-formed + deterministic output. DECISIVE ARM PROVEN: run against a
pre-fix binary (git-stash-isolated build of this commit's changes) --
4 of 6 arms FAIL there with "no <bodies shown=\"0\" ...> tag found — the
element is ABSENT (the exact R9 bug)".

Existing gates updated to match the new, more honest shape (all verified
against the pre-fix baseline to confirm they previously encoded the OLD,
buggy behavior as correct):
  - test/forautobodycheck.sh: the "no body fits whole" arm asserted
    <bodies> was WHOLLY ABSENT; now asserts shown="0" total="N" capped="1"
    is PRESENT.
  - test/packtaskcheck.sh: the tiny-budget arm asserted a bare "omitted
    (budget)" report line with no <bodies> element; now asserts <bodies
    shown="0" total="N" capped="1"> is present, and the report line reads
    "kept 0 of N" (packTaskBundleText's listStatus() derives the label
    from bodiesStr.empty(), which is no longer true on this path -- "kept
    0 of N" is a more honest label than the retired "omitted", not a
    regression).
  - test/estchargecheck.sh: the weak="1" arm assumed a flat-rate
    (bytes/2.50) document; a weak query now also carries the empty
    <bodies> shell (charged at the body rate, 3.80), so it switched from
    selfconsistent() to mixedconsistent() (already proven correct for the
    non-weak auto-bodies case just above it in the same file).

Verified: truncvocabcheck.sh (--lint payload cap), sarifcheck.sh,
lintscopecheck.sh, skippedcheck.sh, ripwirepubliccheck.sh all still ALL
PASS on this change.
joyful-ii-V-I added a commit that referenced this pull request Aug 20, 2026
…ment

Two items from the stack-graphs recon lane (#1 RefRole::Type use-sites,
#2 namespaceCompatible candidate filtering), result-free: baseline
edges=13755 ambiguous=5466 re-derived on this lane's own tree, the six
success criteria, the failure criteria that revert each item, and the
pre-fix reference binary the gates are recorded red against.

Item #3 of that report (depth-2 receiver-chain resolution) is out of
scope for this round and is not registered here.
joyful-ii-V-I added a commit that referenced this pull request Aug 22, 2026
…ing it

Result-free registration for lane J2 (item #3 of the stack-graphs recon).
Baseline re-derived on this lane's own cd30104 build against a pristine
detached checkout; success criteria #3a-#3e and the failure criteria that
revert the lane are fixed before any feature code exists.

Two recon claims re-derived and recorded as corrections: the composeEdges
hoist is unnecessary (buildFieldNarrowTables already runs ahead of the
resolve loop over the same isCompose stream), and Rule 1's bareCish path
currently WRONG-narrows `this->field.method()` as a bare unqualified call —
removing that can raise ambiguity, which is why #3b is a whole-tree
measurement rather than an assumption.
joyful-ii-V-I added a commit that referenced this pull request Aug 22, 2026
…ng it

The depth-2 lane's REJECT (lane/depth2-chains @ 21f75a9) found two real
resolver bugs and an instrument that could not see them: ambiguous= nets a
removed wrong pin against a recovered call. This round registers the
instrument the lane's finding asked for — recovered-edge and corrected-target
counts via a fixture gate (test/chainguardcheck.sh) and a two-binary
whole-tree audit (test/edgediff.py) — plus exact whole-tree predictions taken
from the lane's own measured capture-widening-only arm.

Scope re-justified from c3ebbd8 rather than restored: the widening alone IS
the fix (five recv==None guards stop misfiring); Rule 4 stays out per the
lane's finding #3.

Result-free by construction; the RESULT lands after the instrument runs.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Aug 22, 2026
…itized tier

A shipped-but-not-authored file recites a subsystem's vocabulary in a three-line document, which is
the shape BM25's length normalization rewards most — so a vendored admin script and a numbered
migration took two of five top slots on a database-backend task, and the module an agent would
actually edit never appeared. They now carry the SAME 0.35 factor decks and generated captures
already had: down-weighted, never dropped, still findable by name.

Extends the existing table rather than growing a second component list, and carries the existing
factor rather than a second uncalibrated knob.

Two rows were DELETED from the first draft of that table after checking whether they could ever
fire: vendor/, node_modules/ and dist/ are already pruned whole by the crawl, and *.min.js is
already denylisted by filename, so tier rows for them would be dead weight that reads like
coverage. test/vendoredassetcheck.sh pins that reasoning as its last arm — if a *.min.js file ever
becomes a symbol, the arm goes red and the rule becomes worth adding.

Gate is red-first on the lane base (the migration is #1, the asset #3, the real implementation
last) and green here, and it carries the two near misses a prefix match would break:
migrations/__init__.py is unnumbered and staticfiles/ is not static/.

Measured on the registered external slice: vendored slots in top-5 6/60 -> 0/60 (band [0,2]),
gold-in-top-5 6/12 -> 7/12 (guard band [-1,+4]). All simultaneous floors byte-identical.
joyful-ii-V-I added a commit that referenced this pull request Aug 24, 2026
da61bac

--test-gate=da61bac..HEAD (def-over-decl lane, finding: a ref range the verb
cannot parse) reported changed="0" at exit 0 — a silent zero on unparseable
input, breaching non-negotiable #3. New testgaterefusecheck.sh holds the verb
to --quality-delta's refusal standard (qdrefpaircheck arms (C)): exit 1, the
offending token NAMED, an adjacent probe offered — over A..B, A...B, the
half-typed --test-gate=, a no-such-file token, a mixed list, and a comma-only
list. testgatecheck arm (d), which pinned the defective silent zero, flips to
the refusal; the honest changed="0" keeps coverage via a new clean-git-tree
arm (d2). Verified red before the code: 7 of 8 arms fail on the da61bac
binary (sha256 275ecfbb…), only the valid-FILES control is green.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Aug 24, 2026
…nt changed=0

--test-gate=da61bac..HEAD (a ref range the verb has no grammar for) reported
changed="0" at exit 0 — read by a caller as 'your change touches nothing',
a non-negotiable #3 breach found by the def-over-decl lane. Now held to
--quality-delta's refusal standard (qdrefpaircheck arms (C)): exit 1, the
offending token named verbatim, an adjacent probe offered — the range shapes
(A..B, A...B) get the git diff --name-only expansion that turns the range
into the FILES the verb takes; a no-such-file token gets the --skipped
pointer; refusal is per-token, so a mixed list cannot hide a bad item behind
a good one. The half-typed --test-gate= flips EmptyValue::Meaningful→Refuse
(same ruling as --dmm=/--quality-delta=: an unset shell variable must not
silently gate a different question).

changedMaskFromList grows a checked twin (ChangedList: mask + first bad item
+ item count) and delegates to it — one loop, two surfaces, no drift; --situ
and the MCP verb keep their lenient contract. The refusal body lives in
situ.h::testGateRefusesFileList; the dispatch keeps only the branch, acked
(+3 ccx) in .ripwire_quality_acks per the parseArgs flag-arm precedent.

testgaterefusecheck.sh: 8/8 arms green (7 red at da61bac).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 4, 2026
… old walk

Makes gitignorecheck (RED at 4f6e601) green. ripwire is named for ripgrep, whose defining
default is that ignored files are not searched; the crawl walked them. On this repository's
own root that meant crawling every gitignored checkout under bench/external, which is where
the 2 GiB-cap cache thrash and the 300 s gate timeouts of several past rounds came from.

THE WALK, AND WHY IT SHELLS OUT. One `git -C <root> -c core.quotepath=false ls-files --others
--ignored --exclude-standard --directory -z` per root, not a .gitignore matcher of our own. A
native matcher has to be bug-compatible with git across nested .gitignore precedence, negation,
**, the trailing-slash directory form, core.excludesFile, .git/info/exclude and macOS case
folding — and the moment it diverges it deletes source from a corpus while claiming the
repository asked for that. Measured: 0.020 s (ugrep), 0.030 s (rocksdb), 0.150 s (duckdb),
0.032 s (this root). `--directory` is what keeps that flat — a wholly ignored directory is ONE
entry and ONE prune, never an enumeration. This is not a new dependency: gitmine.h, crossref.h,
prcontext.h, quality.h and binstale.h already popen git, and G3 is about what the binary LINKS.
Its absence degrades the same way — full walk, and --skipped says so.

WHAT IS DISCLOSED, because a corpus that shrinks silently is the defect this replaces.
ignored_files= (exact: files the rules covered) and ignored_dirs= (subtrees they pruned — the
walk stopped there, contents UNKNOWN not zero, the excluded_dirs=/pruned_dirs= contract), both
ABSENT when nothing was cut, so test/golden.xml is byte-identical and every argvdiff vector
survives. The map legend clause is charged only to a map that carries the attributes
(tokenbudgetcheck arm #3 leaves seven bytes of headroom on `src`). --skipped rows both classes
and names the mode in ignore_mode=: git / off / unavailable / root-ignored.

ORDERING IS THE CONTRACT, as it already was for the extension-vs-exclude pair. The ignore test
runs LAST — after the extension, the --exclude match and the built-in denylist — so ignored=
only ever counts a file that would OTHERWISE HAVE BEEN INDEXED (which is what lets the
accounting invariant read indexed= + oversize= + excluded= + ignored=), ignoredDirs= counts
only subtrees no existing rule had pruned, and unsupported_ext=/unindexed= mean exactly what
they meant before.

FOUR CASES KEEP THE FULL WALK. No git work tree, no git binary, --no-ignore — and the one the
registration did not name, found by building it: a root that is ITSELF inside an ignored
subtree (`ripwire build/` in a repo that ignores build/). git answers "./" — everything — and
honouring that returns an EMPTY map for a directory the user pointed at deliberately. That
answer is recognised and refused; ignore_mode="root-ignored" says why. A TRACKED file matching
a .gitignore pattern also stays indexed, because git ignores nothing it tracks.

CACHE: deliberately untouched. The blob is keyed per FILE, so a --no-ignore run writes a
SUPERSET and a default run can only read back what it crawled (gitignorecheck arm 10 pins it).
The residual is speed, not correctness, and it is in the ledger — same shape --exclude has
always had, which is why neither is in the cache key.

Files: src/ingest_crawl.h (the probe + the two extracted recorders), src/model.h (IgnoreMode
+ the two counters and row vectors), src/ingest.h + src/ingest.cpp + src/main.cpp (the flag
reaches the crawl through one trailing defaulted parameter), src/cli.h (--no-ignore, help,
flag-arm ledger 204 -> 205), src/serialize.h (buildIgnoredAttr + the conditional legend),
src/verbs_report.h (--skipped header, rows, legend), src/workspace.h (per-root merge; the
merged mode is the WEAKEST across roots, seeded from the first part).

Also: test/deckcheck_allowlist.txt retires the --no-ignore exemption in the commit that ships
the flag, and .ripwire_quality_acks records the seven gating quality-delta findings with the
reason (one deliberate contract change, three self-churn rows, and what remained of
collectSources after probeIgnoreSet/recordDirPrune were extracted: +3 ccx / +11 LOC, from
+15/+43 inline).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 4, 2026
…at 2

Battery #3 on the merged round-5 tree reddened two gates that are not about goto at all. Both pinned the
literal 2 as scaffolding for what they actually test, and the flow-sensitive slice lane added a THIRD goto
in a fixture — test/sliceflowsensfix/disclosed.cpp's cd03, the deliberately disclosed "goto falls through,
untracked" case the lane's legend now names.

- matchcapturecheck asserts that an EXPLICIT @capture is left alone (no auto_captured=) and that a trailing
  `;` comment does not defeat the scan. It now derives the population once, from the same --match it is
  validating, and fails loudly if that derivation yields nothing.
- lintbudgetcheck asserts that a built-in rule's count stays a clean total beside a saturating same-named
  USER rule. It already computed GOTO_TRUTH for the arm above; the cap arm now uses it too.

This is the fixture-addition gate footprint, the same shape as the flag-addition one: a gate that pins a
whole-tree population makes every future fixture its business. Neither gate lost coverage — both still fail
if the invariant they exist for breaks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 4, 2026
… old walk

Makes gitignorecheck (RED at 1bbd9ea) green. ripwire is named for ripgrep, whose defining
default is that ignored files are not searched; the crawl walked them. On this repository's
own root that meant crawling every gitignored checkout under bench/external, which is where
the 2 GiB-cap cache thrash and the 300 s gate timeouts of several past rounds came from.

THE WALK, AND WHY IT SHELLS OUT. One `git -C <root> -c core.quotepath=false ls-files --others
--ignored --exclude-standard --directory -z` per root, not a .gitignore matcher of our own. A
native matcher has to be bug-compatible with git across nested .gitignore precedence, negation,
**, the trailing-slash directory form, core.excludesFile, .git/info/exclude and macOS case
folding — and the moment it diverges it deletes source from a corpus while claiming the
repository asked for that. Measured: 0.020 s (ugrep), 0.030 s (rocksdb), 0.150 s (duckdb),
0.032 s (this root). `--directory` is what keeps that flat — a wholly ignored directory is ONE
entry and ONE prune, never an enumeration. This is not a new dependency: gitmine.h, crossref.h,
prcontext.h, quality.h and binstale.h already popen git, and G3 is about what the binary LINKS.
Its absence degrades the same way — full walk, and --skipped says so.

WHAT IS DISCLOSED, because a corpus that shrinks silently is the defect this replaces.
ignored_files= (exact: files the rules covered) and ignored_dirs= (subtrees they pruned — the
walk stopped there, contents UNKNOWN not zero, the excluded_dirs=/pruned_dirs= contract), both
ABSENT when nothing was cut, so test/golden.xml is byte-identical and every argvdiff vector
survives. The map legend clause is charged only to a map that carries the attributes
(tokenbudgetcheck arm #3 leaves seven bytes of headroom on `src`). --skipped rows both classes
and names the mode in ignore_mode=: git / off / unavailable / root-ignored.

ORDERING IS THE CONTRACT, as it already was for the extension-vs-exclude pair. The ignore test
runs LAST — after the extension, the --exclude match and the built-in denylist — so ignored=
only ever counts a file that would OTHERWISE HAVE BEEN INDEXED (which is what lets the
accounting invariant read indexed= + oversize= + excluded= + ignored=), ignoredDirs= counts
only subtrees no existing rule had pruned, and unsupported_ext=/unindexed= mean exactly what
they meant before.

FOUR CASES KEEP THE FULL WALK. No git work tree, no git binary, --no-ignore — and the one the
registration did not name, found by building it: a root that is ITSELF inside an ignored
subtree (`ripwire build/` in a repo that ignores build/). git answers "./" — everything — and
honouring that returns an EMPTY map for a directory the user pointed at deliberately. That
answer is recognised and refused; ignore_mode="root-ignored" says why. A TRACKED file matching
a .gitignore pattern also stays indexed, because git ignores nothing it tracks.

CACHE: deliberately untouched. The blob is keyed per FILE, so a --no-ignore run writes a
SUPERSET and a default run can only read back what it crawled (gitignorecheck arm 10 pins it).
The residual is speed, not correctness, and it is in the ledger — same shape --exclude has
always had, which is why neither is in the cache key.

Files: src/ingest_crawl.h (the probe + the two extracted recorders), src/model.h (IgnoreMode
+ the two counters and row vectors), src/ingest.h + src/ingest.cpp + src/main.cpp (the flag
reaches the crawl through one trailing defaulted parameter), src/cli.h (--no-ignore, help,
flag-arm ledger 204 -> 205), src/serialize.h (buildIgnoredAttr + the conditional legend),
src/verbs_report.h (--skipped header, rows, legend), src/workspace.h (per-root merge; the
merged mode is the WEAKEST across roots, seeded from the first part).

Also: test/deckcheck_allowlist.txt retires the --no-ignore exemption in the commit that ships
the flag, and .ripwire_quality_acks records the seven gating quality-delta findings with the
reason (one deliberate contract change, three self-churn rows, and what remained of
collectSources after probeIgnoreSet/recordDirPrune were extracted: +3 ccx / +11 LOC, from
+15/+43 inline).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 4, 2026
…at 2

Battery #3 on the merged round-5 tree reddened two gates that are not about goto at all. Both pinned the
literal 2 as scaffolding for what they actually test, and the flow-sensitive slice lane added a THIRD goto
in a fixture — test/sliceflowsensfix/disclosed.cpp's cd03, the deliberately disclosed "goto falls through,
untracked" case the lane's legend now names.

- matchcapturecheck asserts that an EXPLICIT @capture is left alone (no auto_captured=) and that a trailing
  `;` comment does not defeat the scan. It now derives the population once, from the same --match it is
  validating, and fails loudly if that derivation yields nothing.
- lintbudgetcheck asserts that a built-in rule's count stays a clean total beside a saturating same-named
  USER rule. It already computed GOTO_TRUTH for the arm above; the cap arm now uses it too.

This is the fixture-addition gate footprint, the same shape as the flag-addition one: a gate that pins a
whole-tree population makes every future fixture its business. Neither gate lost coverage — both still fail
if the invariant they exist for breaks.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 5, 2026
… — --situ joins its three siblings

H6 (capture-audit 2026-09-04, lens 6 F1). `--situ=src/nosuch.h` printed "0 changed file(s) — nothing to
analyze" at exit 0, and the MCP twin `situational_awareness{files:…}` returned all-empty arrays with a green
`_fresh: ok`, while --affected / --test-gate / --exercises refused the identical path with flag + problem +
example. An agent that misspells the file it just edited was told its edit has no blast radius and no tests
to run — the false zero non-negotiable #3 forbids, and the JSON form is the worse of the two because a
result reads as an ANSWER to every caller that only checks for an `error` key.

The refusal was not missing, it was wired to one arm: situ.h's testGateRefusesFileList already spoke the
whole triple over the SAME ChangedList and the SAME grammar. It is now parameterised by the surface's own
spelling (fileListRefusalText: `lead` + `selector`) so --situ, --test-gate and the MCP `files` field share
one text, and the MCP arm returns it in a -32602 instead of an empty report.

Second half: the path near-miss. The MCP `cochange` refusal was the only surface in the tool that suggested
the nearest indexed PATH ("did you mean './src/graph.h'?"); every CLI FILE selector had none, and --affected
ran the SYMBOL suggester over a path and answered `--affected=tow.c` with "did you mean 'TOOLS'?" (F18) with
`two.c` sitting in the file list. The candidate search moves to didyoumean.h::nearestIndexedFile — one
suggester, both surfaces — and --situ / --test-gate / --affected / --exercises / --cochange all append it.
--cochange's CLI refusal also gains the example its MCP twin had (F11).

Gate: test/fileselectorrefusecheck.sh (new, registered in test/regression.sh). RED on the pre-fix binary —
A/B fail on --situ (exit 0, no refusal), C fails on all five (no did-you-mean), E fails (MCP result with
empty arrays). GREEN after. testgaterefusecheck / affectedcheck / situdiffcheck / cochange* / didyoumeancheck
/ selectorrefusecheck / mcpcontractcheck all still pass.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 5, 2026
…(sites_l=), not just where the caller is defined

capture-audit 2026-09-04, finding P3 / M21 (lens 8 #3, lens 0). A `<c>` row's p= is the
CALLER'S DEFINITION line — `<c n="benchScores" p="bench/bench_radix_ab.cpp:133"
incompatible="1"/>` for an edit whose sites are 146 and 152. The lines an agent has to open
were printed by exactly one other verb, --uses, so "the contract moved" -> "where do I fix it"
cost a second call whose answer this document already held.

editCheckCallSites() is ONE pass over ing.references plus one sort — never a per-row rescan
(editCheckIncompatibleFlags' own rule; a widely-shared name has hundreds of caller rows) — and
it is paid for only when a row will carry the attribute. Its filters are --uses' OWN
(collectUseSites): role=="call", never a compose edge, a doc mention or a wikilink. That is
what lets the gate assert the two verbs equal without either knowing about the other. It
deliberately does NOT filter on argCountKnown: a spread call site cannot be PROVEN
incompatible but is still a line to open, and --uses prints it — the legend says so.

SPELLING: the finding named the attribute `sites=`. It is emitted as `sites_l=` instead,
because `sites` is already this tool's word for a COUNT of reference sites
(--context-ratio's <s>/<f> rows, --uses' amb_sites=/call_sites_of_name=) and a LINE LIST under
that spelling is a new same-name-different-type collision — the §P8 class test/attrvocabcheck.sh
exists to prevent, and whose stated rule is that the meaning with zero consumers moves. `_l` is
the tool's line-number suffix (--hotspots' top_l=). Reasoning recorded here, not silently.

Gate: test/editcheckcheck.sh arm (k) — the PROPERTY, not a literal. Every row flagged
incompatible="1" must carry sites_l= equal to the ascending unique set of --uses role="call"
lines whose in_id names that caller, over a fixture that exercises both the multi-site and the
single-site shape (asserted non-vacuous), plus the counterpart rule that an unflagged row
carries no sites_l=. RED on the pre-fix binary:
  caller twice: sites_l="" but --uses role=call lines are "5,6"
  caller once:  sites_l="" but --uses role=call lines are "12"
  FAIL  (k) incompatible rows do not carry the call sites --uses prints
GREEN after. MCP edit_check rides the same assembler (mcpclidiffcheck green).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 5, 2026
…oom, --tree, --external-surface; gate defaultceilingcheck

Capture-audit 2026-09-04, lane L7, plan item P4 (§5a decision 4 accepted: the first rows stay, the tail is
windowed and disclosed). Lens 8 measured the defaults on this tree: --pr-context=main~1 659,824 B (~165K tokens,
217 files), --zoom 433,867 B, --tree 187,209 B, --around=SYM 61,891 B, --external-surface 67,862 B led by sh
builtins. Usage data (§5a): --pr-context=REF 0 uses, --around 9 in 20 days.

  --pr-context   budgeted BY DEFAULT: kPrDefaultBudgetTokens = 8000 (budget_tokens= budget_default="1" on the
                 root; --token-budget/--max-tokens set it explicitly). The trim ladder runs first (detail, never
                 files); when even the structural floor of every changed file exceeds it, the FILES are windowed in
                 blast-radius order (files_shown= + the quintet, next="--pr-context[=REF] --offset=N"); --offset/
                 --limit page them (pr-context joins the paging set). PrBudget replaces the size_t budget param
                 (same arity). The shaping guard lets this one verb both page and shape by budget; --top-k on it
                 is its kShapingVerbs notice, not the paging refusal.
  --around       default depth 1 (root depth= discloses; --around-depth=2 restores): 63,471 → 6,297 B here
  --zoom         levels_shown="2" of levels= (a module AT the cut carries children=; --zoom-levels=N, 0 = all)
                 over the 40 largest top modules (shown=/capped=/total=/next_offset=, next="--zoom --offset=40"):
                 444,200 → 8,063 B here
  --tree         the 80 best files (kTreeRowCap; --limit raises): 189,209 → 11,818 B; the plan said 400 — 400 rows
                 are 54,160 B on this repo, which the same plan's ≤12,000 B gate forbids; 80 ≈ 11.5 KB
  --external-surface  100 rows (kExternalSurfaceRowCap) and the sh BUILTINS dropped and counted
                 (builtins_excluded=; externalnames.h kShellBuiltinNames — POSIX + bash builtins, never grep/sed/git,
                 which ARE the surface); --include-builtins keeps them: 68,437 → 5,569 B
Every legend defines its new attributes where the reader meets them; every cut carries next=.
Two new flags (--zoom-levels=, --include-builtins; kTotalFlagArms 207 → 209), each refusing alone naming its verb.

Gate test/defaultceilingcheck.sh (regression.sh), RED first on the wave-2 binary (30 rows): each ≤ 12,000 B at
defaults on this repo (pr-context: est_tokens ≤ 8000 on a 120-file fixture, the file window firing and --offset=N
continuing it), every cut disclosed with capped="1" + a next= that runs, the restoring flags restoring, the first 30
rows of --tree/--external-surface byte-identical to the uncapped listing's, every --around row a depth-2 row.
Re-pins, each to the new contract: prbudgetcheck #1 (default budget disclosed) / #2 (large budget: no
budget_default=) / #3 (a file cut is disclosed, never silent); treecheck (default = min(80,total) rows, "drops
nothing" on --limit=1000000); usescheck 5b (the per-language split asserted with --include-builtins; the default's
drop asserted beside it); w3fixlegendcheck §7/§11 (the default external-surface/tree roots are disclosed windows);
seedboundscheck (depth 1); mcpcontractcheck TWIN (--pr-context: no twin); the --max-tokens/--token-budget guard
sentences (pr-context named, budgetpolicycheck's parser shape kept); README flag/gate counts (179 / 537).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 5, 2026
…tet only (no files_shown= twin); paging sweep enumerates it

pageview.h rule 6: the noun-prefixed shown_/capped forms are for SECONDARY listings; the changed files are the
pr-context root's primary listing, so a cut carries shown=/capped=/total=/has_more=/next_offset=/offset=/limit= and
nothing else — P4 had emitted files_shown= beside pageDisclosure's shown= for one fact (two names, the M15 class).
Legend, --help, defaultceilingcheck (5) and prbudgetcheck #3 read shown=. pagingsweepcheck (L) table gains
--pr-context (it joined the --help HONORED-by set), so the M2 quintet rule now covers its window too.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 5, 2026
…ers a confident zero (F3)

`--affected`'s file reading is a substring PATTERN, so one item can match several files: `geo.py` also
matches `test/check_geo.py`. Every matched file's symbols became seeds, and transitiveCallers returns
"reached, MINUS the seeds" — so the very tests that reach the change were subtracted out of their own
answer. On lane-L8's sandbox (2026-09-04, found-not-fixed #1) `--affected=geo.py` reported
seeds="6" tests="0" reached="0" while `--affected=./geo.py` reported reached="1" over the same bytes.
A confidently wrong ZERO in the verb whose whole job is telling an agent which tests to run, with the
reason it was wrong nowhere in the output — non-negotiable #3's class exactly.

DECIDED: ANSWER, do not refuse (METHODOLOGY §9 #3 — a disclosed answer is terminal, a refusal costs the
agent a round trip to learn a spelling). The two facts that put a test file in the answer are different
and are now kept apart:
  - a test cannot REACH a change it is part of → its own symbols are not seeds of the caller walk;
  - the argument MATCHED it → it changed, so run it, and its row says so with seed_kind="test".
Dropping those symbols from the seed SET never hides them from the traversal — they are still found as
callers — so the answer is byte-identical for every argument that matches no test file.

The partition runs over the DEDUPED seed set and therefore over both readings: a bare symbol name that
resolves into a test file is the same fact (`--affected=check_geo.cpp` now answers tests="1" reached="0",
"run the test you changed", instead of tests="0"). On this repo, `--affected=test/affectedcheck.sh` now
names that gate with run="bash test/affectedcheck.sh" instead of reporting no tests at all.

Honesty in attributes, not prose (§9 #4): root gains seed_test_files="N" (a zero means none matched),
rows gain seed_kind="test"; both defined in the legend and in --help. Structure: the partition
(partitionAffectedSeedsByTestPath) and the answer assembly (affectedAnswer/AffectedAnswer) live in
testmap.h next to the seeding they constrain, which also removes the two complexity regressions the
inline version carried (runAffected 21→29 → back under, resolveAffectedSeeds 21→25 → back under).

Gate: test/affectedcheck.sh gains a colliding fixture pair (src/geo.cpp + test/check_geo.cpp) and 11
arms — the collision spelling must answer the same tests/reached as the unambiguous one, seeds= must
prove the pattern really matched more (non-vacuity), seed_test_files= 1 vs 0, seed_kind= on the matched
row and NOT on a merely-reached row, the test-only pattern, the legend, and the standing rule that
tests="0" while seed_test_files>0 is a failure. RED pre-fix: 8 of 11 arms fail, headed by
"--affected=geo.cpp disagrees with --affected=src/geo.cpp: tests=0/1 reached=0/2". GREEN: 33 PASS.

Verified: affectedcheck testscopecheck rootrelemitcheck testrowruncheck receiptpostcheck exercisescheck
legendcoveragecheck compactlegendcheck a9disclosurecheck floormarkcheck runhintcheck selectorchaincheck
fileselectorrefusecheck shellgateindexcheck testmacrocheck edithandlehintcheck dispatchordercheck
w3fixlegendcheck argvdiffcheck flagtablecheck — all rc=0, no gate needed re-pinning. Determinism ×2
byte-identical, xmllint clean, ASan/UBSan/LSan clean on seven --affected paths (collision, test-only,
unambiguous, multi-item, symbol reading). --edit-check: resolveAffectedSeeds and runAffected both
status="unchanged". --quality-delta regressions=0 gating=0 (7 rows acked, both reasons this lane's own).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 5, 2026
…de batch too (M1)

The capture-audit landed the compact dialect OPT-IN (§5a decision 3, P1) and registered the default flip as
follow-up #3. Opt-in made the right posture a feature an agent has to already know about: measured on this
repo, the ten-verb MCP edit loop was paying 30,839 B of legend prose per pass and 2,866 B is the same
answer. So `legend` absent now means COMPACT, `legend:"full"` restores the historic legend byte-for-byte,
and `legend:"compact"` stays accepted (a no-op spelling of the default — refusing a request for what you
already do is the worst kind of refusal).

THE BILL, this repo's src/, one call per verb, MCP, `legend:"full"` vs the new default:

  verb            legend full   legend dflt   whole doc full   whole doc dflt   doc saving
  edit_check            5,863           357           11,703            6,228        46.8%
  impact                3,660           434            7,824            4,625        40.9%
  uses                  3,807           387           12,903            9,508        26.3%
  slice                 4,500           281            4,917              724        85.3%
  whereis               4,496           294           15,831           11,657        26.4%
  flags                   854           159           17,047           16,378         3.9%
  doc_drift             5,291           199           18,271           13,209        27.7%
  lego                    931           249            1,261              604        52.1%
  exemplar                687           243            1,024              609        40.5%
  path_between            750           263            1,048              586        44.1%
  TEN-VERB EDIT LOOP   30,839         2,866                                          -90.7%
  analyze               1,349           483           43,494           42,652         1.9%
  explore               1,738           373           12,502           11,167        10.7%
  from_trace            2,002           290            5,024            3,343        33.5%
  connect               1,400           345            1,810              783        56.7%
  owners                  745           201              878              361        58.9%
  stray_content         2,267           173           17,648           15,588        11.7%
  batch                   363           142           21,224           14,435        32.0%
  ALL 17               40,703         4,873          194,409          152,457        21.6%

~7.0K tokens off a ten-verb loop, per pass, per session.

SCOPE: exactly the seventeen verbs that DECLARE `legend` (mcpVerbDeclaresLegend, read from kMcpVerbFields —
the same table the unknown-field refusal dispatches from, so a verb joins the family with one edit). The
other fourteen are untouched: their payloads are JSON or plain text with no legend, or — `for` — carry a
header that is already its own short legend whose comments are DATA (A10). A default that silently rewrote
those would be changing answers the argument was never offered for. Inside `batch` that distinction is
load-bearing, not cosmetic: batchcheck (a) asserts every sub-answer is byte-identical to its STANDALONE
verb, and compacting a sub-answer whose standalone twin is full breaks exactly that.

TWO THINGS THE FLIP FORCED, both of them real fixes rather than accommodations:

1. THE POSTURE REACHES INSIDE CDATA (applyCompactToBatchSubs, called by BOTH batch front doors). Both
   surfaces apply the dialect to the finished document and both must leave CDATA alone — a CDATA section is
   DATA and may hold JSON or plain text. So a compacted batch shipped its own legend short and up to sixteen
   sub-answers' legends long, on the one verb where legend duplication costs most. The CLI had the same hole
   under `--batch --legend=compact` and it is closed in the same commit (the owner's CLI-outranks-MCP ruling):
   measured on the fixture, uses+slice, 8,840 B full -> 8,645 B envelope-only -> 2,282 B with the sub-answers.
   `slice` takes BOTH steps, in the standalone dispatch's order — its own emitter builds the compact form
   (runBatchSub's new compactLegend parameter, which decides where schema= sits on the root) and the layer
   then replaces the legend TEXT. Doing only the first step gave a batched slice 1,542 B against the
   standalone verb's 606 B: identical rows, different legend. batchcheck (h) caught it.

2. dropped_positive= MOVED FROM PROSE TO THE ROOT (packtask.h). A2 put --pack-task's dropped_positive in the
   ledger COMMENT while its --for twin put it on the <ctx> root, and the asymmetry was invisible until
   compact became the default: the dialect strips prose, so a COMPLETENESS fact vanished from the default
   answer — measured on src/ at budget_tokens=300, dropped_positive="2" under legend:"full" and NOTHING under
   the default (droppedpositivecheck #6 went red on exactly this). METHODOLOGY §9 #4: honesty lives in
   attributes, precisely so a legend posture can never decide whether a caller is told. Same spelling as the
   --for twin (one name on both roots, mcpattrparitycheck), spliced onto BOTH root spellings including the
   ladder's route-dropped rebuild and the partition slice, and defined in the compact legend's completeness
   vocabulary. The ledger's per-section counts stay prose, which compact is entitled to drop.

WHERE AN AGENT READS IT: the `legend` field description, "a STRING legend posture: compact (the default) or
full" — one string spliced into all seventeen inputSchema stanzas, which is the per-tool sentence an MCP
client actually shows, and it names both values and the default in the SAME 54 bytes the old wording spent.
tools/list is byte-identical at 40,841 B, measured before and after. Deliberately NOT a clause added to
seventeen tool descriptions: mcpmanifestcheck's own registered rule is that the ceiling moves for a declared
argument's obliged description and never for prose, and prose there would cost ~680 B against 159 B of
headroom. --help's --legend paragraph states the split (MCP default compact, CLI default full).

GATE, WRITTEN FIRST: compactlegendcheck arm (N), the family read from src/mcprefusal.h rather than listed,
so a verb joining without a call in the gate FAILS instead of skipping. Four assertions per verb: every
posture answers; the DEFAULT is byte-identical to legend:"compact"; legend:"full" carries strictly more
legend bytes; the PAYLOAD is byte-identical between postures.
  RED  pre-flip: all seventeen — "(N) analyze: the DEFAULT is not compact — default 1161 B of legend,
       compact 290 B" and sixteen more.
  GREEN post-flip: all seventeen, e.g. "(N) edit_check: default == compact (357 B legend), legend:\"full\"
       restores 5863 B, payload byte-identical".

RE-PINS — twelve gates, each to the NEW contract, stated precisely, in this commit (rule 10b: the cost is
recorded here, it is not a reason to keep the shape). Two different re-pins, and which one a gate gets is
decided by what the gate ASSERTS, not by which was easier:
  · compact <-> compact, because the gate compares the two surfaces' DEFAULT paths and that is now compact
    on one side: mcpslicecheck (13) and sliceflowsenscheck (7) add --legend=compact to their CLI operand
    (sliceflowsenscheck keeps $FB full for arms (5)/(6) and takes a second, compact operand).
  · full <-> full, because the gate reads the FULL document by construction and compact <-> compact would
    make it compare nothing: mcpclidiffcheck (LENS 1 extracts the first `<ELEM attr="…">` tag, and a compact
    legend names its elements in bare shorthand `<iface n= p= defs=>` — zero attributes, and the gate's own
    "the probe is broken" guard would fire; LENS 3 quotes a CLI legend CLAUSE the compact dialect does not
    carry), mcpw2fixcheck, mcptranchecheck, mcpw3fixcheck, floormarkcheck (it compares disclosure TAILS),
    mcpattrparitycheck (whose slice block gains a `default (M1: compact)` row asked raw — the flip itself,
    pinned across both surfaces), mcpverbscheck (two shape assertions quote the full document's shape; the
    pack_task alias arm now also proves that `legend` survives alias resolution), editpreviewcheck (N).
  · batchcheck: three root assertions that pinned attribute ORDER (`<batch n="16" requested="20" …`) are
    re-pinned to the FACTS, asserted individually in any order — strictly stronger, since the old spelling
    would also have passed with two of those numbers swapped.
  · compactlegendcheck (U): --batch joins --for as a named exemption from the atomic-payload arm, for the
    opposite reason to --for's — its CDATA holds sub-answers whose legends legitimately moved. The property
    is not dropped: arm (N) re-asserts it on the batch after flattening the CDATA and normalising the two
    attributes leg.py already normalises on a single root (schema=, est_tokens=).

VERIFIED: all 98 --mcp-touching gates green (regression.sh excepted — it is the meta-runner). --edit-check
runBatchSub: status="contract-change" change="params" callers="2" incompatible="0". --quality-delta
regressions="0" gating="0" acked="24", every ack my own row with its reason in the ledger.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 5, 2026
…ob_sha, ONE next=; the stderr line names it and nothing else

Terminality round A (2026-09-05), lane E, E2. On a runner without a Read-before-edit policy the agent stops on the
receipt (baseline: 12/12 windows terminal). So the receipt must hold what a Read would have shown, or the stop is blind:

  region    {start,end,context,text,capped:false} — the applied lines plus context=3 each side, the bytes as they are
            ON DISK; over kReceiptRegionBudgetBytes (2048) it carries head + tail (whole lines) + elided_lines with
            capped:true — never a silently truncated text
  blob_sha  git's blob id of the written bytes (== `git hash-object FILE`; infra/gitblob.h, a hand-rolled SHA-1 as git
            IDENTITY, zero dependencies) — "what is on disk" without reading the file
  next      exactly ONE pasteable follow-up (METHODOLOGY §9 #3), read off the fold: a contract-change with broken
            callers → --uses=FILE:SYM; else the first tests_to_run run= recipe; else --test-gate=FILE; under
            --no-post-check → --edit-check=FILE:SYM (the one call that shows the state). region and blob_sha need no
            index and ride under --no-post-check too.
Same keys on the MCP twins (one engine). The stderr line now repeats the receipt's next and names no second command
(the old line prescribed --edit-check AND --affected — two calls the receipt already answered). The `ripwire wrap`
blurb no longer says "then --edit-check=SYM" after an edit (it coached the redundant check the meter counts).

Gate: test/receiptpostcheck.sh arms 10-17 (NEW, RED pre-fix on all eight): region == disk bytes, blob_sha == git
hash-object, one next= that runs and follows the rule, the stderr names only it, MCP key-set parity, the budget
head/tail/elided_lines, --no-post-check keeps the free half, the insert family, the blurb. Re-pinned in this commit
to the one-next contract: edithandlehintcheck (the printed next is the receipt's, names the RESOLVED file, never a
handle, and RUNS — asserted verbatim under --no-post-check where next IS --edit-check=FILE:SYM), rootrelemitcheck
ARM7 (next read off `next:`, its spelling == the receipt's file, and --edit-check=<receipt file>:sym runs).
postCheckJson: +1 defaulted out-param (callers 2, incompatible 0); runEditVerb unchanged.
--quality-delta regressions=0 gating=0 (22 rows acked by exact name with reasons). 17-gate sibling sweep green.
Receipt deterministic x2 (stale_index stamp aside), valid JSON.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 5, 2026
… of its body (V1 R2+N4)

Gate first: test/prbudgetcheck.sh arm #9 recounts the delivered bytes at the tool's one markup rate
against the number on the root, for BOTH --pr-context roots and every budget rung. RED on 044d4d0
with nine failing rows:

  (#9) empty-diff root (clean tree):   est_tokens=278  but the document is 4894 B = 1958 tokens  (17.60 B/tok)
  (#9) working-tree root (1 file):     est_tokens=517  but the document is 5463 B = 2185 tokens  (10.56 B/tok)
  (#9) default-budget root (8000):     est_tokens=1123 but the document is 7670 B = 3068 tokens  ( 6.82 B/tok)
  (#9) --max-tokens 100000/20000/8000/4000/2000: 1123 vs 3061/3061/3060/3060/3060

GREEN after, delta 0 on all eight roots: est_tokens == round( emitted bytes / 2.50 ) exactly.

WHY the number was wrong. prEstTokens() modelled ceil( (BODY bytes + kPrHeaderOverheadBytes 560) /
kMinBytesPerToken 2.36 ). The ~4.6 KB legend -- which IS the document on a clean tree -- was in
neither term, and the 560 B constant stood in for a root tag plus a legend that are 4,900 B together.
So the empty root priced 4,808 delivered bytes at 278 tokens (7.3x under, the wave-1 verifier's R2)
and the budgeted root at 1.30x under (N4), both printed beside a budget_tokens="8000" a caller reads
them against. --for prices exactly (9,831 B / 3,932 = 2.50); --pr-context now uses the SAME
conversion, tokensForEmittedBytes at kBytesPerTokenDefault, not a second rate.

HOW. prLegendText / prAnchorNoteText / prRootOpenText RETURN the bytes they used to fprintf, so the
emitter can measure what it is about to write; prPriceDocument( PrPriceCtx, body, level, truncated,
window ) converts envelope + this root's own start tag + body, converging over its own est_tokens
digits in <=4 passes exactly as serialize.h's pricedRootAttr does. pickPrTrimLevel is handed that
price instead of computing one, and the empty root is handed the same number instead of modelling
its own. kPrHeaderOverheadBytes and prEstTokens are gone -- there is one estimator, and it measures.

BUDGET ARITHMETIC -- the ceiling now cuts where it always should have, and at the DEFAULT it cuts at
the same content. On prbudgetcheck's 3-file fixture, before -> after est_tokens:

  --max-tokens=100000   1123 -> 3138   trim_level 0, all 3 files      (unchanged content)
  --max-tokens=20000    1123 -> 3138   trim_level 0, all 3 files      (unchanged content)
  --max-tokens=8000     1123 -> 3137   trim_level 0, all 3 files      (unchanged content)
  --max-tokens=4000     1123 -> 3137   trim_level 0, all 3 files      (unchanged content)
  --max-tokens=2000     1123 -> 2490   trim_level 4 + files windowed shown=1 of 3, floor-exceeded

Only the 2000 rung moves, and it moves because the legend alone is ~4,610 B = ~1,844 tokens: a
2,000-token ceiling has ~156 tokens left for the root tag and every changed file's structural floor,
so the floor genuinely does not fit. The old number said it fit and then delivered 7,651 B (3,060
tokens) -- 1.53x the budget, silently. The new run says budget-floor-exceeded and discloses the file
window. No gate assertion was loosened: prbudgetcheck #3's "est <= budget unless the floor overflows
and says so" and #6's monotonicity both hold on the new numbers.

Legend re-pinned to the contract: est_tokens= now reads "prices the WHOLE document this bundle emits,
this legend included, at the map's markup rate of 2.50 bytes per token, and IS the number the ladder
fits, so recounting the delivered bytes reproduces it". Arm #9's last row asserts that sentence is on
both roots -- a definition true of one root and false of the other is the drift prRootOpenText exists
to prevent.

Re-pinned: test/floormarkcheck.sh (10)'s source guard r'"<pr-context%s' -> r'"<pr-context" \+', the
same emitter under its new concatenation spelling; the guard is presence-checked, so leaving it stale
would have read as a pass.

--edit-check: pickPrTrimLevel params 2->4 and prEmptyRootTail 3->4, both contract-change,
incompatible="0" (one caller each, updated in this commit). Determinism x2 + xmllint clean on
--pr-context and --pr-context=HEAD~1; ASan clean on 7 pr-context paths incl. the budget ladder, the
file window and --situ. quality-delta regressions=0 gating=0 (12 rows acked, all this lane's own,
with the reason recorded in .ripwire_quality_acks).

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 6, 2026
…ways promised

--grep's span-tier legend says, verbatim, "tier_unclassified= hits in files nothing
classified — always EMITTED, never suppressed", and grepTierAttrs()/grepTierKeys()
guarded the field behind `unclassifiedHits > 0`. A comment-only match whose hit files
were all classified therefore printed the promise and withheld the field:

  ripwire <corpus> --grep=ZEROMARK_probe
  <grep ... complete="1" tier="comment" tier_parsed="2" ...>     # no tier_unclassified=

The reading is why the promise was made. tier_unclassified= is the qualifier on the
tier= LABEL — "0" is the PROOF that the label was elected over EVERY hit, and absence
is the one state a reader cannot interpret: it reads equally as "none" and as "the
emitter dropped it". Non-negotiable #3, on the tool's own most-quoted honesty clause.
Same defect on the MCP twin, where it is worse: an MCP-only agent has no CLI to
re-ask from before trusting tier=.

Fix: emit it unconditionally, exactly like tier_parsed= two lines above it, inside the
block hasDisclosure() already gates. tier_partial= stays conditional — its legend
documents its own absence ("its absence beside a tier= means the label is a fact").

Cost, measured (G4): +22 B, and only on a --grep answer that already carries a tier
disclosure. Over 15 representative queries on this repo: 7 of 15 gained the attribute,
480,529 B -> 480,683 B = +0.032%. Zero bytes on every other verb and on the default map
(test/golden.xml still byte-identical).

Gate: test/emittertruthcheck.sh P2.Z — a FAMILY arm, not an instance one.
  (Z1) the roster of consumer-visible unconditionality claims is DERIVED from src/
       (every string literal claiming "always emitted"/"never omitted"/"printed even at
       zero"), keyed by file+phrase so it survives line moves, and diffed against the
       expected 11. A new or reworded claim fails until someone decides whether it is
       true and gives it a probe. Its first question is which member is missing.
  (Z2) each claim that names an attribute is proven true AT ZERO, on every surface:
       tier_unclassified (CLI + MCP), register-macro-excluded (XML + JSON),
       sub_windows, confidence/margin_pct (XML + JSON), lego/compose/routes_total.
  Presence guards on every probe: an extractor that finds nothing, a corpus that stops
  electing a non-code tier, or a legend that stops printing the clause all FAIL rather
  than going quietly inert.

RED before the fix on (Z2a) and (Z2b) only; the other four claims already held. No
other gate in the tree can see this class: legendcoveragecheck enumerates FROM the
emitted document, so a suppressed attribute is structurally invisible to it, and
truncvocabcheck's rules are all "if shown= then ...".

--edit-check: both symbols status="unchanged" (no contract moved).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 6, 2026
…t it can read

THREE POINTS ARE NOT A REGION, AND MIN_HULL_MEMBERS COUNTS POINTS. The only geometry gate on a module
outline was a node count, so three members that happen to sit near a line passed it and drew a hull with
almost no area — a thin coloured streak across the picture that reads as a scratch on the lens rather than
as containment. Three of them were in the django/db/migrations figure at rrf/top-k=120:
`resolve_model_field_relations` (a 107x0 px line, hull area 4 px^2), `add_operation` (178x8) and
`reload_model` (137x13). The first was being dropped SILENTLY — a fully collinear group makes convexHull
return a two-point ring and the caller skipped it without counting it anywhere, so no number on the page
ever moved.

The test is the isoperimetric ratio 4*pi*A/P^2 of the raw hull ring. The hull is convex by construction and
on a convex ring that ratio IS thinness: 1.0 for a circle, 0.785 for a square, 0.605 for an equilateral
triangle — the roundest a three-member hull can be — collapsing to pi/(2*aspect) for a thin one, which
reproduces all three measurements above to three decimals. Chosen over the two alternatives on their own
terms. An AREA floor is the wrong predicate twice: it is not scale-invariant, and it would drop a small
ROUND module (`varint`, 23x18 px, perfectly legible once HULL_PAD_PX has stood the outline off it) while
the streaks it is aimed at are long enough to survive one. An OBB aspect ratio measures the same thing and
needs rotating calipers, O(ring^2), where this is one O(ring) pass over vertices draw() already walks — and
it runs BEFORE the O(N) purity scan, so a group that cannot be a region never costs a walk over every node.

0.20 IS MEASURED. Over both corpora at the figure argv — 23 groups with 3+ members in view, this
repository at top-k=200 and django/db/migrations at rrf/top-k=120 — the ratios form a low cluster
{0.0011, 0.0718, 0.1493} and a body from 0.2879 up. The two widest adjacent gaps in the whole distribution
(2.08x and 1.93x) bracket exactly that band and 0.20 is its geometric midpoint; in shape terms, a triangle
about 8:1 base-to-height. Measured on the RAW ring, not the padded one drawn, so it is scale-invariant:
HULL_PAD_PX is a screen quantity and testing the padded shape would make an outline appear and disappear as
the reader zooms, taking the caption's count with it. On the figure corpus this takes 11 outlines to 8, and
all three that go are the streaks by name.

THE CAPTION SPLITS, because provenance and methodology travel to different places. The stamped strip had
grown to three dense lines carrying arrow semantics, the dash threshold, the label rule, the shape key and
the hull rule on top of the provenance — and the stamp font is FITTED to the bitmap width against an 8 px
floor, so every clause added to it made the provenance it exists to carry physically smaller. Method is
what a caption UNDERNEATH a figure says once, in prose; provenance is what has to survive the picture being
lifted out of the page. So renderProv builds two halves: FACTS (root, ranker, top-k, the counts, the colour
metric, and any state that makes the picture provisional) goes into the bitmap, and METHOD goes to the page
and to a companion .txt the same export writes, where a README author can lift the wording verbatim
instead of paraphrasing it. Non-negotiable #3 is why the last FACTS line exists at all: it names the .txt
AND every channel whose rule moved into it. A trimmed bitmap that said nothing about the trim would be a
picture drawing arrowheads, dashes, shapes and outlines with no way to read any of them — the quiet
omission this block was added to prevent, re-introduced by the fix for it. The hull line now states all
three truncations with a count and a reason each — the cap, the shape drop, the purity drop — rather than
naming one rule and pooling the rest.

AND THE STAMP DESCRIBES THE FRAME IT IS UNDER. Found cutting the README's crop figures: the caption's
counts are the LOADED subset and the camera is free to sit anywhere inside it, so a picture exported after
a zoom stamped "120 nodes / 183 edges" across a frame holding fifteen — a bitmap overstating its own
contents, in the one artifact the caption exists to travel in. draw() now counts node centres inside the
canvas rect (the reading that errs toward the larger number), the stamped half states it whenever it is
less than the view's own total, and draw() re-renders the caption when that number or a hull count moves,
guarded on a change so the settle loop does not rebuild innerHTML sixty times a second. Without that guard
the numbers renderProv states — this count and the three hull counts — are all computed one frame after the
caption that reports them, which is invisible on a settled page and is the previous view's numbers on an
export taken straight after a zoom.

Gate arms in htmlrendercheck, all observed RED on the pre-fix page and green after. (U13)-(U19): the
threshold is a named constant, the ratio is isoperimetric with the perimeter load-bearing (so an area floor
cannot satisfy it), the constant sits inside the band the two corpora measured, both drop counts reach the
caption separately, the uncounted collinear skip is GONE, and the cheap test precedes the O(N) scan.
(W1)-(W14): the stamp reads the FACTS half and no longer the whole caption, the methodology is still on the
page, the export writes the .txt, both files share one basename, the bitmap names the companion and
enumerates what moved, and the control PAIR (W8)/(W9) proves each moved clause LEFT the stamped half and
LANDED in the other rather than being deleted — either arm alone passes over a deletion. (W10) holds the
stamped half inside STAMP_MAX_LINES, counted with grep -o rather than -c after its own control caught two
pushes sharing a line reading as one.

(S8b) and (R11) name a different SURFACE now and say so rather than being weakened: the shape key and the
arrow semantics are still built by renderProv and still travel with the exported picture, in the .txt
instead of the bitmap. (W9a)/(W9b) hold their new home; (W8a)/(W8b) hold that they left the old one.

test/htmlrendercheck.sh declares RIPWIRE_TEST_DEPS: src/htmlexport.h. It greps a page it asks the binary to
emit and never names the source that emits it, so the script_literal rule could not see the link and
--test-gate answered a change to this file with tests="0" — naming no test at all for the one file it has
116 mutation controls about.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
dazKind pushed a commit to dazKind/ripwire that referenced this pull request Sep 7, 2026
…tie, and every join says how much it explains

Registered at 1ae72526; RED gate at 52756446. Both parts are additive (G5) and deterministic.

PART 1 — informativeness before id. connectSubgraph()'s BFS recorded the FIRST-discovered parent, so
among equally-short joins the winner was the one whose FILE sorted first. It now relaxes on ties: a
candidate parent at the SAME distance replaces the incumbent iff ( connects, id ) is smaller. Because
BFS expands in non-decreasing distance order, every parent at the minimal distance is offered exactly
once, so the winner is a minimum over a total order and does not depend on expansion order — the rule
REMOVES an order dependence rather than adding one, and the id is still the total final tie-break, so
output stays byte-identical run to run. Distance is untouched, so path length, node count and edge
count are untouched: this changes WHICH equally-short answer is served and nothing else about its
shape. Prim's (dist,minId,maxId) rule over terminal PAIRS is deliberately unchanged — the defect is a
parent choice, and on a 2-terminal --connect Prim has no choice to make.

The two CSR halves are now ONE relax() step. They differed only in the discovery channel; the tie-break
would have made them the same eight lines twice, and a rule with a comparison in it written down twice
is how two halves drift. That fold is why --quality-delta reports connectSubgraph complexity 170->173
rather than the 194 the inlined version measured.

PART 2 — disclosure (non-negotiable redhat-et#3). Every Steiner <s> row carries connects="N", the number of
distinct symbols that intermediary joins in the undirected view the search walks (callers + callees,
O(1) off the CSRs). hub="1" marks a row at or above hub_floor= on the root. Both are facts, so the
--max-tokens trim never drops them. An honest "the only join found connects 764 other things" beats a
confident bare `empty`.

hub_floor is DERIVED, not tuned: the smallest D with D(D-1)/2 > edges — the degree at which one node
alone joins more distinct symbol pairs than the whole call graph has edges. Integer bisection over one
number the map header already prints; 186 here. kConnectRootBytes 200 -> 260: 200 was already short of
the widest start tag it claims to bound (228 B measured attribute-by-attribute, 246 with hub_floor=),
and short is the one direction that constant may not be.

Legends: the full --connect legend defines all three and states the derivation; the compact dialect
gets one hub_floor= term that defines hub="1" in the same clause (two terms did not fit the 400 B
ceiling, which is the point of that dialect) — connect compact is 392 B. CLI and MCP share the emitter,
so both surfaces move together.

Worked example, before -> after:
  --connect=computeQualityDelta,gitFileCommitCountsInDayWindow
  before  <s n="empty" p="src/notes.h:431"/>            both call something named empty (764 callers)
  after   <s n="computeDelta" p="src/quality.h:5284" connects="41"/>
          computeQualityDelta -> computeDelta -> gitFileCommitCountsInDayWindow
nodes="3" edges="2" in both: the size did not move, the answer did.

Battery 557 gates: pass=554 skip=3 fail=0 (baseline at 52756446: pass=552 fail=2 — this gate RED plus a
stale-binary versioncheck). Determinism x2 byte-identical; xmllint clean; connectcorecheck green under
ASan/UBSan.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
dazKind pushed a commit to dazKind/ripwire that referenced this pull request Sep 7, 2026
`ripwire <dir> --around=writeHtml --html=/tmp/h.html` exited 0, wrote no
file, and said nothing on stderr. The same argv without --around writes
59 KB. writeHtml() is called from exactly ONE place -- runDefaultMap --
and every navigation and report verb pre-empts the default map in main()'s
dispatch chain, so the flag was accepted and dropped.

Six verbs were named in the report. A sweep of the whole flag universe
(the new gate, below) found SEVENTY-FOUR flag x --html combinations in
that state. This is the "accepted and silently ignored" class this repo
has closed three times over, and non-negotiable redhat-et#3 is explicit: a caller
must be able to tell a no-op from a typo. The MIRROR of this guard already
existed and was loud -- --color-by without --html refuses and names --html
(cli.h validateModifierGuards). This is the other half.

REFUSE, NOT HONOUR -- and why

Honouring would mean rendering the verb's scoped node set as the page,
and that is a worse answer than it sounds. The page's three views are an
overview of Louvain MODULES, a module subgraph, and a depth-bounded ego
graph, all computed client-side over the whole selected map. Over a
20-node --around slice the module overview is empty and the ego graph is
a re-derivation of the slice itself, so `--around=X --html=F` would
produce a page answering a different question from the one --around
answers, with no tell -- trading a silent no-op for a silent
substitution. The page already does what that caller wants, and since the
routing commit it does it BY NAME, so the refusal has somewhere real to
point: g.html#node/SYM/2 is the depth-2 neighbourhood. Refusing is also
the smaller change -- one predicate, no new emit path.

DERIVED, NOT ENUMERATED

The 74 verbs are not listed anywhere. htmlPreemptedBy() reuses
firstFlagOutside() -- the walk over kBoolFlags/kViewFlags, the same rows
parseArgs matched -- against kMapShapingFlags plus a short kHtmlRideAlong
list, so the rule is "anything not on the compose list" and a verb added
tomorrow refuses tomorrow with nobody editing a table. Every previous
closure of this family was done by enumerating members and re-opened on
the member nobody enumerated: jsonUnsupportedVerb's 77-arm chain missed
12, the shaping-flag guards missed 18. The one hand-written arm the flag
tables cannot see (--export=cc.json) is named explicitly, the same honest
cost jsonUnsupportedVerb pays at its own top -- and it was found by the
gate, as the last silent row after the walk closed the other 73.

The residual risk runs the other way: a new MAP-SHAPING flag would refuse
until it is added to kHtmlRideAlong. That direction is the safe one -- a
loud refusal on a legal combination is noticed the first time someone
types it; a silent drop is not. Arm (B) pins the compose side so the
guard cannot over-refuse its own host unnoticed.

It lands in refuseInertMainModifiers rather than cli.h's
validateModifierGuards for the reason that function's own header already
gives: judging "alone" needs main's knowledge (the flag universe walk),
which is why --no-redact and --refetch are refused there too. It still
runs before any dispatch.

THE GATE

test/htmlhostcheck.sh is BEHAVIOURAL, not textual, and shell-only. For
every flag in the universe derived from src/cli.h (test/flaguniverse.py,
the same derivation shapingflagcheck arm F uses), `<verb> --html=OUT` must
end in exactly one of two buckets: WRITE (OUT is non-empty) or REFUSE
(non-zero exit). Exit 0 with no OUT is the silent class and fails BY NAME.
Arm (D) is its mutation control: a stub binary that exits 0 writing
nothing must be classified SILENT, and one that exits non-zero must be
classified REFUSE, so a zero in arm (C) means "the defect is gone" rather
than "the arm cannot see it". Arm (B) is the over-refusal control.
Currently write=31 refuse=174 silent=0.

--help now scopes --html to the default map and points at the
#node/SYM/2 route; docs/COMMANDS.md regenerated from it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
dazKind pushed a commit to dazKind/ripwire that referenced this pull request Sep 7, 2026
…e the named file's own decl/def partner

TWO defects the round-C head-to-head found in --situ, both about a relationship the tool already knows and
does not say.

F2 — the composed co-change zero conflated two different facts. --cochange refuses loudly on an empty window
(commits="0" window="18mo@HEAD", exit 1) and --rank-by=churn stamps "(no churn evidence)", but every surface
that COMPOSES co-change rendered a bare (0) — the CLI report adding "(none, or no git history)", which means
neither "none found" nor "none exists". Non-negotiable redhat-et#3 forbids exactly that. The window and the commit
count now travel ON SituationFacts, so the four composing surfaces cannot disclose it four ways or three of
them forget it:
  * --situ [3]                  — window="…" commits="…" on the header, and the empty case now says which of
                                  the two zeros it is
  * situational_awareness (MCP) — cochange_window / cochange_commits beside `forgotten`, the surface where an
                                  empty array reads most like an answer
  * --handoff <heuristic>       — cochange_window= / cochange_commits= (appended AFTER n=/candidates=/capped=;
                                  `<heuristic n="…"` is a shape gates read positionally)
  * --pr-context <cochange>     — had commits= and no window to read it against

F3 — TAKEN FROM COCOINDEX. `--situ=db/wal_manager.cc` on RocksDB spent 3,104 B and never named the gold
`db/wal_manager.h` anywhere, while an embedding search returned it as result redhat-et#1 in 55 bytes; same shape on
table/get_context.cc; S2 recall 2/14. Not a retrieval failure: section [1] ranks by dependent-symbol count,
which surfaces the biggest test files and can NEVER surface the header, because a header does not depend on
the source that implements it. It is a different relationship, and one ripwire already holds.

  declDefPartners() — another file defining symbols under the SAME (scope, name) identity. Stated that way it
  is not a C++ special case: it covers a .h/.cc pair, an ObjC .h/.m, a C# partial class, a .pyi stub and a
  .d.ts, without reading one file extension or comparing one basename. The guard is a MAJORITY test rather
  than a number pulled from the air — 2*shared >= min(symbols(a), symbols(b)) — so two large files sharing
  one common free-function name are refused, and the row publishes `shared` so the reader audits it.

No embedding mode was added and none is coming (G3). The blast-radius counts are unchanged: the partner is
printed above them as its own labelled fact, never merged into them.

Measured after: --situ=db/wal_manager.cc -> db/wal_manager.h (10 shared) as row one;
table/get_context.cc -> table/get_context.h (16); db/version_set.cc -> db/version_set.h (138).

Gates: sincewindowcheck arm 4 now 18 assertions, all GREEN, including a NEGATIVE CONTROL that two 13-symbol
files sharing one name are refused (the majority test is live, not decorative) and an arm that a git repo
with no commits stamps window="18mo" WITHOUT @Head rather than claiming an anchor it does not have.
deckcheck: six foreign tool flags quoted in lane C2's head-to-head write-up (gortex --dataset/--detach/
--entry-point/--kind, uv --python, rg --sort) allowlisted — that gate was red on this lane's base.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
AnkitArya pushed a commit to AnkitArya/ripwire that referenced this pull request Sep 9, 2026
…t contain

An adversarial review of the section-granular round found that its ancestor rule fails on
exactly the queries CLAUDE.md tells agents to write. Reproduced on a scratch tree holding
only docs/ARCHITECTURE.md:

  --recall="crawl order sorted before ids assigned"                        -> lines="353-369"
  --recall="how does the crawl order get sorted before the ids are assigned" -> lines="1-11,12-20"

Same two stubs at 800, 1200 and 1600 tokens; the answer first appears at 2400, third. The
landed comment called the residual "bounded and self-correcting … the worst case is a
heading-plus-intro admitted one slot early". That claim was false and is deleted with the
reasoning it rested on. This ranker has no query-side stopword list — the SIRA corpus
margin (src/lexical.h marginBits) ships disarmed — so "the", "does", "are" are live query
terms, an ancestor's SUBTREE contains every one of them everywhere, and tf is additive.

THE FIX IS AT THE RANK, AND BOTH PROPOSED ADMISSION FIXES WERE MEASURED FIRST.
  - a term margin (keep the ancestor only on a query term with corpus df <= S/2) leaves the
    reproduction BIT-IDENTICAL. Built and run: of the three ancestors it would have to drop,
    `# ripwire — architecture` survives on "how" (df 2 of 13 own-prose units) and
    `## 1. The pipeline` on "order" (df 4 of 13). Both terms separate, and both are genuinely
    in the ancestor's own prose — `## 1. The pipeline` really does say "in that order".
  - "keep it only if its own-prose SCORE is positive" is inert, provably: BM25's idf is
    ln((S-n+0.5)/(n+0.5)+1) > 0 for every n <= S, so a positive own-prose score IS "own prose
    carries a query term", which is the rule already shipping.
Admission was right to keep those units. What was wrong is that they OUTRANKED the answer,
on evidence held by the seven subsections nested inside one of them. So admission is
untouched — for a fixed heading population the pick SET is what it always was — and the rank
key becomes attributed:

  rank = scorer score x (this unit's own-prose evidence / the strongest own-prose evidence
                         anywhere in its subtree)

Own-prose evidence is BM25 over a collection of THIS document's own-prose units: the same
idf, the same resolveBm25Params, the same forEachLexSubtoken tokenizer and the same
name(x3)/body(x1) field split the scorer uses. The collection is local because the question
is local — which section of this document the query is about — and a word in every section
of a document separates none of them, so "the" loses its idf without a stopword list. Only
ratios of two values from one call are ever read, and nothing prints them. A leaf's subtree
IS its own prose, so its maximum is itself and its key is the scorer's number bit-for-bit;
an ancestor that is the best evidence under itself is likewise untouched.

MEASURED, final binary vs the pre-change binary, two harnesses with constructed ground truth
(no hand labels). H1: 72 unique-term queries whose content words occur in one section only.
H2: 456 heading-as-query trials over docs/. The NL form of each adds ONLY function words to
the same content words, so any difference is attributable to how they are treated.

  H1 @800   NL served 88->89  NL first 81->85  agree 79->83  KW first 97->97  KW bytes 100% same
  H1 @1600  NL served 88->89  NL first 43->79  agree 42->78  KW first 97->97  KW bytes 100% same
  H2 @800   NL served 76->80  NL first 73->79  agree 75->80  KW first 90->92  KW bytes  98% same
  H2 @1600  NL served 77->81  NL first 57->74  agree 69->77  KW first 74->86  KW bytes  85% same

No metric regresses in any cell. A full own-prose re-rank (discard the scorer's number
entirely) beats this on H2 — which queries HEADINGS, the field it weights x3, so it grades
its own bias — and is a ranking replacement, not a defect fix: it moves 9% of keyword answers
for no measured keyword gain and belongs to a pre-registered ranking round. The x3/x1 split
is not a tuned constant: rerunning both harnesses with the heading at x1 is byte-identical on
every trial, because a ratio between two spans of one document is insensitive to it.

THE DEGENERATE CUT NAMED LINES IT DID NOT SERVE. findRecallBoundaryCut lands immediately
AFTER a newline, so the last kept byte opens the first line NOT served; counting every '\n'
in the kept prefix therefore disclosed a line whose bytes never reached stdout. Reproduced: a
section at line 3 followed by 400 two-line paragraphs, at --max-tokens=400, claimed
lines="3-10" with source line 9 the last byte served; now 3-9. The count runs over
kept[0, n-1), which is composeRecallUnits' own rule (it derives lineHi from ownEndByte - 1),
so the two paths compute the same thing the same way. Non-negotiable redhat-et#3.

BODY-LESS HEADINGS ARE HEADINGS. The selection dropped any heading whose section body is
empty — a heading immediately followed, with no blank line, by a same-or-shallower one — so
it was neither counted in N nor treated as a unit boundary, and --help's "up to the next
heading of any depth" was false twice over. A 4-heading fixture reported `3 of 3` and served
`5-9`, a range containing the `### Empty child` line; it now reports `3 of 4` and serves
`5-8`. On docs/EVALS.md N goes 197 -> 198, which is exactly the ATX heading count computed
independently by a line scan. --help needed no change, which is also why this repair was
chosen over rewording it. The whole-file node — the real target of the old filter — is now
excluded by NAMING it (starts at byte 0, no signature/body split, ends no earlier than any
heading in the file), measured against the file's own symbols rather than its size on disk,
so a file that grew after indexing cannot promote it into a unit that overlaps every other.

Also from the review: truncateRecallBody's out-parameter becomes a structured-binding return
(CONTRIBUTING §3) — it made the byte count optional, and the caller that skipped it is the
one whose disclosure would go stale first. And `dropped_by_budget=D` deliberately carries no
`next=`, with the reason recorded at formatRecallSectionNote: the only knob that could
restore a dropped unit is --max-tokens, which is not invertible from inside one document
(§C4 water-filling divides one budget across the set, and cross-document monotonicity is
explicitly not claimed), so a computed N would be a guess wearing a disclosure's clothes.

--quality-delta found the new scanner had cloned ownProseCarriesQueryTerm's match loop; both
now share recallMatchedTermIndex, and the boolean keeps its early exit.

Gates: 573 run, 571 pass, 2 skip (both want a reference binary), 0 fail. The twelve recall
gates plus mdsectioncheck, mcprobustcheck and redactcheck green by name. Both fixtures the
old rule was built for verified directly: guide.md answers `zqcachewarmbody` with
`1 of 7 … lines="12-26"` and no zqorientbody or zqdeepappendix; notes.md keeps
`# Geometry Fixture` and serves it first. Determinism x2 byte-identical on the map and on
--recall, warm == cold on both and on the degenerate cut; xmllint clean. ASan/UBSan/LSan over
20 arms of the attribution path, the degenerate cut at five ceilings, the empty-heading
fixture, both named fixtures and the operator's memory directory: no report.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 10, 2026
…ourth is refuted with numbers

METHODOLOGY §9 #3 says never cut silently and #6 says a ceiling attribute names the ceiling ACTUALLY
applied. Three caps did neither. Each fix is disclosed only when the cap actually bit — the pr_converged
shape (src/prconverge.h): silence means untruncated, presence means truncated. G4 pays nothing otherwise.

1. kGrepMatchedLineMaxBytes (512 B, src/search.h) cut the matched line of EVERY --grep and --verify hit
   with no ellipsis and no attribute, on the only content those answers carry: a 512 B source line and a
   truncated 50 KB minified line printed byte-identical payloads. The row now carries line_bytes="N", the
   WHOLE line's byte length. The NUMBER rather than a bare capped="1" because it is what decides the
   reader's next move — the row already carries the deterministic follow-up (p=/l=, the root's next=), so
   what was missing is whether following it is worth a read. NO ellipsis here, deliberately, and this is
   where it differs from (2): a grep payload is raw file bytes by contract — the boolean --and/--not filter
   reads the same line and the --at= follow-up is expected to reproduce it — so the fact goes in an
   attribute (§9 #4) rather than into the bytes. lineBytes joins grepGroupByFile's fold key: two long lines
   can share a 512 B prefix and differ in true length, and folding those would print one line_bytes= for
   sites it does not describe; untruncated rows carry 0, so the key is byte-identical wherever the cap did
   not fire. The MCP grep verb is the third caller of grepEnrich and serves NO matched text at all, so it
   has no cut to disclose — asserted rather than assumed, arm A5.

2. cleanSig's kMaxSig (240 B, src/serialize.h) hard-broke every emitted signature — --pack-signatures,
   --for's <sigs>, <calls> callee rows, --lego — mid-token, with no marker. It now goes through
   truncateUtf8WithEllipsis, the tool's ONE truncator, exactly as the three other signature cuts
   (kForTailSigBytes, kForCapTailSigBytes, packtask.h's tail sig) already do. In-band here because a
   signature is already a RENDERING, not raw bytes: the body is stripped and whitespace runs collapse, so
   matching its three siblings is what consistency means. The visible prefix is unchanged at 240 B; only
   the "…" is new. The loop collects one byte past the cap so the shared truncator can do its own
   codepoint back-off, and a separate flag carries the fact because the trailing-space trim can pull a cut
   string back under the cap and a cut that trims back under is still a cut.

3. A DEFAULT --for enforced kForPayloadBudgetBytes (7500 B) on every run and named no ceiling: budget_tokens=
   rode only an EXPLICIT --token-budget, so a trimmed default bundle disclosed THAT it was cut
   (<sigs shown= total= capped="1">) while the number that cut it appeared nowhere. Compare --pack-task,
   whose default lands on its root as budget_tokens="6000" — same class of ceiling, two honesty outcomes.
   The <ctx> root now carries budget_bytes="7500", and the JSON dialect ",budget_bytes":7500.

   THE UNIT IS BYTES, and that is the decision the finding asked for. What the default applies IS a byte
   constant — bundleBudget is kForPayloadBudgetBytes verbatim and the ladder compares rendered bytes
   against it. A token spelling would have to divide by a rate, and the two rates in play disagree on
   purpose: est_tokens prices at kBytesPerTokenDefault (2.50) while a ceiling is SIZED at the conservative
   kMinBytesPerToken. Any token number printed here would be one no ladder ever applied, which is the exact
   failure §9 #6 names.

   DEFAULT REGIME ONLY: the explicit regime already names the caller's own ceiling, so exactly one of the
   two rides a trimmed bundle — never both, never neither (gate arms C8/C10). It rides the ROOT and is
   spliced AFTER the sigs render, with dropped_positive=/bundle=: the first cut of this fix put it on the
   <sigs> open tag with the ladder's own per-run share as the value, which is more precise and wrong here —
   forbudgetmonotoncheck pins the sig section byte-identical between the default ceiling and any explicit
   ceiling above it, and a per-run number breaks that identity while every served row stays the same. The
   late splice leaves the ladder's input untouched, so <sigs> is byte-for-byte what it was. The clause
   defining the attribute is spliced on the same condition (kForOverCeilingLegend's precedent) so an
   untrimmed bundle pays for neither. packSignatures gains cappedOut, the ladder verdict its JSON twin has
   always returned, so the caller reads a boolean instead of re-parsing rendered bytes.

   THE MCP `for` VERB CARRIES IT TOO (§P8: one element name, one attribute order, both surfaces). That
   dialect is budgeted by default the same way (mcpverbs.h forBudgetBytes) and had the same silence; fixing
   the CLI alone would have replaced one honesty gap with a disagreement between two surfaces an agent can
   reach for the same answer. Same two conditions, same splice point, same sentence — arm C14 pins the
   sentence byte-identical across the two.

4. kSliceRdMaxIter (64, src/slice.h) is REFUTED, not fixed: it is unreachable, so it cuts nothing and has
   nothing to disclose. Structurally, the reaching-definition lattice has no cross-slot flow — a use reads
   its own slot, a def assigns it a singleton (SliceRdWalker::unit) — so each slot's loop transfer is
   X -> const ∪ (X if a path passes through), whose ascending chain from `entry` stabilises on the SECOND
   round regardless of program shape. Measured: instrumenting the converged break and running --slice over
   3,026 symbols of this tree's own src/ gives 4,528 loop fixpoints, max iter = 1 (508 at 0, 4,020 at 1,
   none higher); an adversarial C fixture (while > for > do-while nesting around a 40-case fallthrough
   switch with continue/break) gives max iter = 1, and an adversarial Python one (nested loops with
   try/except/finally and continue) the same. The bound fires at iter >= 63. Its comment already says it
   "only guards a broken lattice", and that is what it is: an XML attribute that can provably never appear
   is ungateable by construction (non-negotiable #1) and would be decoration. DEGRADED_PATH_ALERT stays the
   right instrument for it (guardrail #4).

GATE FIRST (non-negotiable #1): test/capdisclosurecheck.sh was written and shown RED before any src/ edit.
Every arm asserts three things, because a disclosure gate that only greps for its own attribute proves
nothing: CROSSING (the fixture's emitted payload really is shorter than the source, or the element really
is capped), DISCLOSURE, and SILENCE (the same verb on an uncrossed fixture carries nothing). MUTATION
CONTROL, run against a binary built from the parent commit: every DISCLOSURE arm fails (A2, A4, B2, B4,
C2, C5, C7) while every CROSSING and SILENCE arm still passes.

Gate count 573 -> 574, derived from THIS tip's own `for _g in` loop after the rebase (main landed
optremarkshotcheck while this lane was green; the union of the two loops is the number), across all eight
published sites. Only test/regression.sh conflicted — README/EVALS/deck auto-merged clean at the stale 573,
which is the failure mode, so all three were reset from origin/main and re-bumped from the merged loop.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
henry-hz added a commit to z8-run/ripwire that referenced this pull request Sep 10, 2026
…-uses on Elixir defs

Three of the six actionable review findings on redhat-et#81 reproduced; this is those three.

1. src/ingest_model.h — expandElixirImplementationReferences appended a multi-target
   defimpl's cloned references to the TAIL of ing.references, whose contract is
   (fileId, startByte, name, role, isInherit) order. A clone carries its original's
   coordinates, so the tail was full of earlier bytes and earlier files. Two consumers
   read that order rather than re-deriving it: graph.h's chaUpDeclared records a
   derived type's direct bases in source order (Elixir's Extends refs from defimpl ARE
   cloned), and editpreview.h splices a re-parsed file in by partitioning on fileId,
   which only reproduces a real re-ingest while each file's refs are one ascending run.
   Each clone is now spliced in beside the reference it came from — the same position a
   re-sort would give it, at no sorting cost.

2. src/mcpverbs.h — the MCP uses verb gated the Elixir resolver path on `defs`, which
   resolveAllByName fills for EVERY language, while the CLI gates on the same list
   filtered to Lang::Elixir. An Elixir reference whose calleeName matched a definition
   in another language therefore took the resolver path in MCP, matched nothing, and
   vanished — while the CLI still reported it through the name filter. That is the
   divergence mcpclidiffcheck exists to prevent. Now filtered the same way, off the raw
   selector, exactly as resolveUsesSelector does.

3. docs/ARCHITECTURE.md — the Elixir prose claimed resolution the tree does not do.
   Four gaps reproduce on this tip and are now disclosed in Static limits rather than
   left to contradict the PR notes: a later `import M, except:` replaces an earlier
   `only:` selection instead of subtracting; a dotted nested `defmodule Inner.Deep`
   registers no implicit prefix alias, so `Inner.Deep.target()` resolves to nothing;
   `alias __MODULE__, as: Current` in a multi-target defimpl binds every implementation
   to the FIRST target; and `&_seed/0` is dropped by the underscore filter. Honesty in
   output is a feature (CLAUDE.md non-negotiable redhat-et#3) — these are floors, and the
   document now says so.

The three remaining findings did not reproduce and are unchanged: reachesAny's exact
name compare is correct because no non-Call Elixir reference carries an arity suffix
(they are module names and @attributes, matched against symbols whose scope is empty
or the enclosing module), and eliximportcheck.sh sets `set -u`, not `set -e`, so its
FAIL branch and fail=1 accounting do run.

Verified: elixircheck, eliximportcheck, elixirsemanticcheck, mcpclidiffcheck,
editpreviewcheck, editroundtripcheck, qschemetripcheck and qextractionkeycheck pass;
multi-target defimpl output is unchanged; ASan/LSan clean on the Elixir fixtures and
on the repo; repeat maps byte-identical.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_013VUC4dGmaT2bFJvPJtsjvf
joyful-ii-V-I added a commit that referenced this pull request Sep 10, 2026
…ourth is refuted with numbers

METHODOLOGY §9 #3 says never cut silently and #6 says a ceiling attribute names the ceiling ACTUALLY
applied. Three caps did neither. Each fix is disclosed only when the cap actually bit — the pr_converged
shape (src/prconverge.h): silence means untruncated, presence means truncated. G4 pays nothing otherwise.

1. kGrepMatchedLineMaxBytes (512 B, src/search.h) cut the matched line of EVERY --grep and --verify hit
   with no ellipsis and no attribute, on the only content those answers carry: a 512 B source line and a
   truncated 50 KB minified line printed byte-identical payloads. The row now carries line_bytes="N", the
   WHOLE line's byte length. The NUMBER rather than a bare capped="1" because it is what decides the
   reader's next move — the row already carries the deterministic follow-up (p=/l=, the root's next=), so
   what was missing is whether following it is worth a read. NO ellipsis here, deliberately, and this is
   where it differs from (2): a grep payload is raw file bytes by contract — the boolean --and/--not filter
   reads the same line and the --at= follow-up is expected to reproduce it — so the fact goes in an
   attribute (§9 #4) rather than into the bytes. lineBytes joins grepGroupByFile's fold key: two long lines
   can share a 512 B prefix and differ in true length, and folding those would print one line_bytes= for
   sites it does not describe; untruncated rows carry 0, so the key is byte-identical wherever the cap did
   not fire. The MCP grep verb is the third caller of grepEnrich and serves NO matched text at all, so it
   has no cut to disclose — asserted rather than assumed, arm A5.

2. cleanSig's kMaxSig (240 B, src/serialize.h) hard-broke every emitted signature — --pack-signatures,
   --for's <sigs>, <calls> callee rows, --lego — mid-token, with no marker. It now goes through
   truncateUtf8WithEllipsis, the tool's ONE truncator, exactly as the three other signature cuts
   (kForTailSigBytes, kForCapTailSigBytes, packtask.h's tail sig) already do. In-band here because a
   signature is already a RENDERING, not raw bytes: the body is stripped and whitespace runs collapse, so
   matching its three siblings is what consistency means. The visible prefix is unchanged at 240 B; only
   the "…" is new. The loop collects one byte past the cap so the shared truncator can do its own
   codepoint back-off, and a separate flag carries the fact because the trailing-space trim can pull a cut
   string back under the cap and a cut that trims back under is still a cut.

3. A DEFAULT --for enforced kForPayloadBudgetBytes (7500 B) on every run and named no ceiling: budget_tokens=
   rode only an EXPLICIT --token-budget, so a trimmed default bundle disclosed THAT it was cut
   (<sigs shown= total= capped="1">) while the number that cut it appeared nowhere. Compare --pack-task,
   whose default lands on its root as budget_tokens="6000" — same class of ceiling, two honesty outcomes.
   The <ctx> root now carries budget_bytes="7500", and the JSON dialect ",budget_bytes":7500.

   THE UNIT IS BYTES, and that is the decision the finding asked for. What the default applies IS a byte
   constant — bundleBudget is kForPayloadBudgetBytes verbatim and the ladder compares rendered bytes
   against it. A token spelling would have to divide by a rate, and the two rates in play disagree on
   purpose: est_tokens prices at kBytesPerTokenDefault (2.50) while a ceiling is SIZED at the conservative
   kMinBytesPerToken. Any token number printed here would be one no ladder ever applied, which is the exact
   failure §9 #6 names.

   DEFAULT REGIME ONLY: the explicit regime already names the caller's own ceiling, so exactly one of the
   two rides a trimmed bundle — never both, never neither (gate arms C8/C10). It rides the ROOT and is
   spliced AFTER the sigs render, with dropped_positive=/bundle=: the first cut of this fix put it on the
   <sigs> open tag with the ladder's own per-run share as the value, which is more precise and wrong here —
   forbudgetmonotoncheck pins the sig section byte-identical between the default ceiling and any explicit
   ceiling above it, and a per-run number breaks that identity while every served row stays the same. The
   late splice leaves the ladder's input untouched, so <sigs> is byte-for-byte what it was. The clause
   defining the attribute is spliced on the same condition (kForOverCeilingLegend's precedent) so an
   untrimmed bundle pays for neither. packSignatures gains cappedOut, the ladder verdict its JSON twin has
   always returned, so the caller reads a boolean instead of re-parsing rendered bytes.

   THE MCP `for` VERB CARRIES IT TOO (§P8: one element name, one attribute order, both surfaces). That
   dialect is budgeted by default the same way (mcpverbs.h forBudgetBytes) and had the same silence; fixing
   the CLI alone would have replaced one honesty gap with a disagreement between two surfaces an agent can
   reach for the same answer. Same two conditions, same splice point, same sentence — arm C14 pins the
   sentence byte-identical across the two.

4. kSliceRdMaxIter (64, src/slice.h) is REFUTED, not fixed: it is unreachable, so it cuts nothing and has
   nothing to disclose. Structurally, the reaching-definition lattice has no cross-slot flow — a use reads
   its own slot, a def assigns it a singleton (SliceRdWalker::unit) — so each slot's loop transfer is
   X -> const ∪ (X if a path passes through), whose ascending chain from `entry` stabilises on the SECOND
   round regardless of program shape. Measured: instrumenting the converged break and running --slice over
   3,026 symbols of this tree's own src/ gives 4,528 loop fixpoints, max iter = 1 (508 at 0, 4,020 at 1,
   none higher); an adversarial C fixture (while > for > do-while nesting around a 40-case fallthrough
   switch with continue/break) gives max iter = 1, and an adversarial Python one (nested loops with
   try/except/finally and continue) the same. The bound fires at iter >= 63. Its comment already says it
   "only guards a broken lattice", and that is what it is: an XML attribute that can provably never appear
   is ungateable by construction (non-negotiable #1) and would be decoration. DEGRADED_PATH_ALERT stays the
   right instrument for it (guardrail #4).

GATE FIRST (non-negotiable #1): test/capdisclosurecheck.sh was written and shown RED before any src/ edit.
Every arm asserts three things, because a disclosure gate that only greps for its own attribute proves
nothing: CROSSING (the fixture's emitted payload really is shorter than the source, or the element really
is capped), DISCLOSURE, and SILENCE (the same verb on an uncrossed fixture carries nothing). MUTATION
CONTROL, run against a binary built from the parent commit: every DISCLOSURE arm fails (A2, A4, B2, B4,
C2, C5, C7) while every CROSSING and SILENCE arm still passes.

RE-PIN: test/docdemotegolden_for.xml 5505 -> 5517 B and test/docdemotegolden_noroute.xml 9556 -> 9568 B,
+12 B each = FOUR three-byte U+2026 markers from (2), plus the est_tokens recount. Verified before
re-pinning: with est_tokens= normalised and the four new ellipses deleted, live and previous goldens are
byte-identical on both fixtures — no ranking, demotion, route or budget byte moved, and neither root is
capped, so (3) is absent from both. That fixture is where (2) shows up at all: a markdown section's
"signature" is prose, so four of those rows were over the 240 B cap and had been cut invisibly for the life
of the golden. The cause is recorded in the gate header beside the five earlier re-pins.
test/printf_parity.manifest needed NO re-pin: none of its 14 pinned verbs reaches a changed emitter.

Gate count 574 -> 575, derived from THIS tip's own `for _g in` loop after the LAST of three rebases (main
landed optremarkshotcheck and then callsrankordercheck while this lane was in flight; the union of the two
loops is the number, every time). Only test/regression.sh ever conflicted — README/EVALS/deck auto-merged
clean at the stale number on all three, which is exactly the failure mode, so all three files were reset
from origin/main and re-bumped from the merged loop each time. The last rebase crossed a lane that changed
serialize.h and verbs_for.h too (the <calls> rank-order fix), so it was followed by a --clean-first rebuild
before anything was re-measured.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
andriytyurnikov pushed a commit to andriytyurnikov/ripwire that referenced this pull request Sep 10, 2026
…ll 120 caps become visible

A cap is a ROUTING decision — it decides what an agent can and cannot find — and this tree had 120 of
them across 51 files with nothing listing them together. That is how kMaxExpandSibs came to sit at 8.

WHAT 8 COST, MEASURED. Symbols-per-file on this repo: median 4, p90 18, p99 85, max 562. At 8 the cap
fired on 68.5% of bodies and hid 89.3% of every sibling name — 10.7% of the file-context signal
survived. It looked harmless because the MEDIAN file has 4 symbols and never trips it; but symbols
concentrate in big files, so symbol-weighted it fired on two bodies in three.

WHAT IT WAS SUPPOSED TO SAVE. P16's stated cost was "582 B of a 3,067 B --expand answer (19%); ~3.5 KB
per --pack-task bundle". The second half is not reproducible: --pack-task emits NO sibs= at all, and
never did — measured against the PRE-P16 binary (2026-08-30), which also emits none. Two arms of
test/expandsibscheck.sh already asserted exactly that ("--pack-task carries no sibs=/inc="), and were
green the whole time. sibs= reaches exactly one verb: --expand.

WHY 100. It clears p99 (85), so it guards the pathological tail instead of cutting through the typical
case: 15.8% of bodies truncated instead of 68.5%, sibling names visible 10.7% -> 56.2%. Cost is +36% on
a single-symbol --expand answer and NOTHING on --for or --pack-task, which are byte-identical at every
cap value tested (9,829 B and 12,590 B at 8, 80 and 100 alike).

THE NUMBER MOVES, AND THAT IS AN EFFECT, NOT THE REASON. top-50 goes 71.0% -> 81.8%. --pack-signatures
elides exactly what it always did; --expand simply stopped hiding context, and the body side is this
ratio's DENOMINATOR. The cap was chosen on recall grounds and the figure re-derived afterwards — doing
it the other way round would be tuning a default to flatter a headline. Re-derived on THIS build and
propagated to all six artifacts that carry it (band, --help, caption, capture, COMMANDS.md, README,
EVALS), with the parity `help` label re-pinned under UPDATE_GOLDEN_EXPECT="help" — moved={help}, 39
unchanged.

THE FIXTURE HAD TO GROW, AND THAT IS THE POINT. test/expandsibsfix/manyfn.c had 45 functions: one over
the OLD cap of 40, still over 8, but UNDER 100 — so at cap=100 the gate would have passed while testing
nothing. Regrown to 145 (144 siblings), one over the new cap, by construction. Its header comment still
claimed kMaxExpandSibs=40, stale since P16 and never caught, because nothing compared it to the source.

SO THE CAPS GET AN INVENTORY, GENERATED AND GATED. docs/LIMITS.md lists every cap in src/ with its
value, site and whether its file discloses a truncation when it fires. It is a build product of
docs/limits_build.py, and test/limitstablecheck.sh fails when it drifts — 4 arms, including a
can-go-red that adds a synthetic cap to a TEMP tree via --root, never to src/, because a probe file
dropped into src/ perturbs the crawl other gates measure.

The inventory's first finding, not acted on here: 54 of the 120 caps sit in files that disclose nothing
when they fire. src/mention.h holds 7 of them, and they bound which docs can ever surface in --for. A
silent cut reads to the caller as "none exists", which non-negotiable redhat-et#3 forbids.

Gate count 570 -> 571, at all eight spellings (two in README, three in EVALS, three in the deck) — and
NOT at `1.570`, a confidence interval, or inside a CI run id.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
lennix1337 pushed a commit to lennix1337/ripwire that referenced this pull request Sep 11, 2026
…id nothing

capdisclosurecheck closed three caps on EMISSION — a row found and then trimmed, on an
element that already had a shown=/total= vocabulary. The ten caps in src/mention.h and
src/gitmine.h's co-boost family are one level up: INDEXING caps, deciding what --for's
ranker is allowed to LIFT. Nothing downstream can tell that a mention was dropped, and no
flag, budget or paging brings it back — the answer simply does not contain the thing the
task named. That is non-negotiable redhat-et#3 ("a zero means none found, never none exists") on a
surface where the caller has no recourse at all. docs/LIMITS.md said `Discloses: none` for
both files.

THE RULE, which decides which caps got an attribute and which got a comment: DISCLOSE WHEN
THE CUT CONTENT IS NOT OTHERWISE VISIBLE IN THE ANSWER.

Seven disclose, each conditional, each carrying an exact `*_total=` wherever the true count
is computable without changing what the cap keeps:

  mention_tokens_capped= / _total=  kMentionMaxRawTokens (16) — mention tokens past the
      extraction window. extractMentions now scans the whole task and only the KEEPING
      stops at 16, so the total is exact and the scan costs nothing (task text is a
      sentence).
  mention_files_capped=             kMentionMaxFiles (4). FACT, no total: counting the
      files the scan never reached means running it unbounded, and that WOULD move
      `matchedFile` (a full list leaves it false and routes the mention to the symbol
      path). The honest number available is the fact of the cut.
  mention_syms_capped= / _total=    kMentionMaxDirectSymbols (8) — Scope.name matches. The
      scan runs to the end for the census; pushing is still gated on the same room test,
      so the kept set is unmoved.
  doc_mentions_capped=              kDocMentionMaxDocsPerAnchor (2) and
      kDocMentionMaxDocsTotal (6). Set ONLY where a refusal is provable — a doc below its
      anchor's lift target that a cap turned away. "There might be more" never sets it.
  coboost_partners_capped= / _total=  kCoBoostMaxPartnerFiles (8); the true count was
      already in hand one line above the resize.
  coboost_commits_capped= / _total=  kCoBoostMaxFilesPerCommit (30) — whole commits thrown
      away before the boost ever reads them, so a partner that only moves inside big
      refactor commits is silently unreachable.

Three do NOT, and each carries the measurement at its declaration rather than an attribute:
kMentionMaxSymbolsPerFile and kCoBoostMaxSymbolsPerFile trim symbols out of a file every
surviving row NAMES with p=, so the caller can see it and page in; kDocMentionMaxAnchors is
a window over the top of a ranked list the bundle prints in full.

WHERE THE DISCLOSURE RIDES, and why not where its prose belongs. A --for header composed
before the signature ladder runs is charged against kForPayloadBudgetBytes, so bytes spent
in it are bytes taken from ranked rows. The first cut of this lane folded the clause into
mentionNote/docMentionNote and measured the cost: three bundles fell from 25 ranked rows to
24, and one lost 204 B of evidence. The clause was trimmed to ~50 B AND moved to the
post-render splice budget_bytes= uses, so <sigs> is byte-identical and the disclosure is
paid for in bytes, never in evidence. The clause carries the attributes verbatim, so it is
self-defining where the reader meets them (legendcoveragecheck's `defined` predicate).

HOW OFTEN THEY FIRE, over 39 realistic --for tasks on this tree: kMentionMaxRawTokens 1/39,
kMentionMaxFiles 3/39, kMentionMaxDirectSymbols 0/39, the doc caps 6/39. All three mention
firings came from one shape — a task pasting SEVERAL paths, exactly the multi-file case B8
exists for. The 0/39 is a property of a C++ tree with unique method names, not of the cap;
the class is gated on a fixture that does cross it, and a silent cap costs nothing, so the
zero is recorded rather than used as an argument to leave the cut unsaid.

NON-DEGRADATION, BASE_BIN (633a1d2) vs this binary, same frozen corpus, identical flags:

  #   base B   lane B   delta   sigs rows   invocation
  1     2301     2301       0     - > -     --for="cache invalidation" --format=candidates --top-k=5
  2    10040    10040       0    18 > 18    --for="incremental cache invalidation when a file content hash changes"
  3    14521    14627    +106    25 > 25    --for="pagerank power iteration" --detail=2
  4    11055    11161    +106    25 > 25    --for="pagerank power iteration" --with-graph
  5    10064    10064       0    26 > 26    --for="quality delta acks ledger rubber stamp"
  6    10075    10075       0    27 > 27    --for="quality delta acks ledger rubber stamp" --no-doc-mention
  7     5528     5528       0     - > -     --for="rankGraphTeleport"
  8    15975    16081    +106    20 > 20    --for="rankGraphTeleport" --no-route
  9     2850     2850       0     - > -     --for="rankGraphTeleport" --signatures-only
 10     9962     9962       0    29 > 29    --for="tree-sitter parse of a source file" --adaptive
 11    15958    15958       0    29 > 29    --for="tree-sitter parse of a source file" --auto-bodies
 12     9575     9575       0    31 > 31    --for="tree-sitter parse of a source file" --legend=compact
 13    10125    10231    +106    26 > 26    --for="why does src/lexical.h chooseForRanker pick name-exact BM25"
 14        0        0       0     - > -     (a truncated flag in the sweep; refusal on both)
 15     9607     9607       0     - > -     --pack-task="add a new output format flag to the CLI"
 16    25256    25256       0   4+8+6 >     --pack-task="add a new output format flag to the CLI" --partition=3
                                  4+8+6
 map   24585    24585       0     - > -     the flagless map, byte-identical

Twelve of sixteen are byte-identical. The four that moved are exactly the four where
doc_mentions_capped="1" fired, each +106 B (the attribute, the self-defining clause, and
the est_tokens digits that now price them), and NO bundle lost a ranked row. Two further
real tasks, measured because they are the ones that reach the B8 caps: a five-path task
+214 B (doc_mentions_capped + mention_files_capped) and a seventeen-path task +318 B
(those two plus mention_tokens_capped="1" mention_tokens_total="17"), both at unchanged
row counts.

GATE FIRST. test/mentioncapcheck.sh, written before the code, six arms x three assertions:
CROSSING, DISCLOSURE, SILENCE. Each crossing is proved WITHOUT reading the new attribute —
arm A is a same-token-multiset word-ORDER contrast (every lexical signal here is
order-blind, so only the extraction window can explain the difference); B/C/D read the
pre-existing prose note's count against a fixture size the gate computes; E reads the
partner list out of --cochange, whose cochangePartners path shares no code with
applyCoChangeBoost's; F reads git itself. Arm G pins the legend definition, --json parity,
zero-cost silence on an ordinary query, determinism and well-formedness.

MUTATION CONTROL, run: against BASE_BIN all 8 DISCLOSURE assertions FAIL and all 15
crossing/silence assertions PASS. Against this binary, ALL PASS.

docs/LIMITS.md regenerated: silent caps 54 -> 44, and both files now read `Discloses:`
instead of `Discloses: none`.

One defect found by READING the regenerated table rather than trusting it: limits_build.py
harvests `[a-z_]+_capped` out of the source file, so the placeholder `name_capped="1"` in
CapDisclosure's own doc comment was published as a mention.h attribute that does not exist.
The placeholder is now spelled `<x>_capped`, which the regex cannot match, and the comment
says why so the next placeholder is not spelled back. A generated doc is only as honest as
its input.

G1: asan+ubsan build clean on the flagless map, on the four --for shapes this lane moves
(including the seventeen-path task that trips all three B8 caps) and on --pack-task; the new
gate is ALL PASS under the sanitizer binary too. g1freshcheck green.

Also green: capdisclosure, limitstable, docmention, mention, cochangeboost, forlens,
legendcoverage, mcpattrparity, mcpforparity, fornotesbudget, forbudgetmonoton, estcharge,
fornotesjson, packtask, packtaskmonoton, tokenbudget, jsonparity, formatgate, mentionsverb,
docdemote, droppedpositive, xmlwellformed, qschemetrip, churnjoin, loopconservation,
fixedbufsweep, binoverride, rootrel, testrowrun. Determinism x3 byte-identical, xmllint
clean on the map and on a capped bundle.

manifestcheck and readmedrift fail on their gate-COUNT arm only (577 -> 578); the eight
published count sites are deliberately untouched this round.

--quality-delta is NOT clean, and the residual is stated rather than hidden: 20 gating
findings, all small deltas on already-large symbols — 5 api-surface (a census out-parameter
on three miners, an out-parameter on two mention_detail helpers), 4 complexity, 7 verbosity,
4 short-horizon-churn (touching files that churn). A first pass cut it from 25: the two
commit-census out-parameters were collapsed into one CommitWindowCensus struct, which
removed the `params` class entirely; recordCommitFileSet was lifted out of gitLogFileSets
(that function's complexity now falls below where it started); docLiftWasRefused was lifted
out of applyDocMentionBoost (+17 complexity -> +7); and the in-function comment blocks were
trimmed, with the durable rationale moved to file scope in mention.h where it costs no
symbol its verbosity. What remains is intrinsic: six conditional disclosures across three
surfaces are branches and lines that did not exist before.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
xCatG added a commit to xCatG/ripwire that referenced this pull request Sep 11, 2026
…check.sh

Adds .kt support (kParserVer 85) via fwcd/tree-sitter-kotlin, pinned past
v0.3.8 for its scanner segfault fix. Zero ASan/UBSan/LSan findings across
501 real .kt files (8 local Android/JVM repos) despite that history.

Extraction covers classes/objects/companion objects/functions/calls/imports,
inheritance clauses (captureBases recognizes `delegation_specifier`), and
complexity scoring (isDecisionType/cc_isNestingControl recognize `when_entry`/
`when_expression`/`do_while_statement`/Kotlin's own `catch_block` spelling —
`lang == Lang::Kotlin` checked first so the other 21 languages short-circuit
on a byte compare rather than a strcmp; `catch_block` is a different,
unrelated node in the vendored Swift/Elixir grammars, verified against all
22 vendored parsers before scoping it). Canonical ids are scope-qualified
(kotlinEnclosingScopeOf) and parameter counts are real (countParams counts
`parameter` children specifically for Kotlin's parameter list — a naive
"every named child" count over-counts, since a parameter's own modifiers
and default value are its SIBLINGS in this grammar, not nested inside it;
verified against a real parse before writing the fix). Import edges are
captured (Kotlin joins Java's DepDialect, since a mixed module's imports
genuinely cross the language boundary).

Kotlin's function_body/class_body are positional children, not fields — the
generic field-based body lookup silently returns null for every Kotlin
definition, which is not cosmetic: graph.h's decl/def collapse reads a null
body as "bodyless forward declaration" and deletes that candidate from
resolution. Kotlin rides ingest_sidecap.h's existing ObjC positional-body
fallback (widened, `class_body`/`enum_class_body` added) rather than a new
mechanism — the same fallback ObjC has needed since ITS grammar exposes no
`body:` field either. `enum class`'s body nests under `enum_class_body`, a
DIFFERENT positional child than a plain class's `class_body`; verified via a
standalone tree-sitter probe against the vendored grammar before fixing.
Getting this right is what makes a same-name Kotlin/Java collision correctly
surface as `ambiguous=` instead of silently resolving to whichever side
happens to have a working body lookup.

The JVM bridge (graph.h langCompatible) resolves Kotlin<->Java calls in both
directions cleanly on real code (Nanidroid: unresolved 1425->545, ambiguous
1364->1496 — the rise is real collisions surfacing honestly via the fix
above, not a regression). Disclosed, not fixed (a real new feature, not a
bug): Kotlin navigation-expression receivers (`A.f()`) aren't recognized by
ingest_binds.h's classifier, so an explicit receiver that would narrow a
same-name collision between two Kotlin classes currently doesn't.

Fixed the highest-severity language-dispatch finding: verbs_navigate.h's
kExtSurfaceLangSlots was hand-sized off Lang::Elixir+1, so an unrecognized
language's external-surface references silently clamped into slot 0
(Lang::Cpp). Every array in src/ hand-sized off a last-enumerator literal
(verbs_navigate.h, verbs_report.h, lintrules.h, clones.h, htmlexport.h,
tsprobe.cpp) now routes through model.h's kLangCount instead, so this class
of bug can't recur on the next language append.

Also deferred and disclosed rather than silently dropped: .kts (Gradle DSL
needs its own trailing-lambda parse-quality probe), --slice support, and
essential-complexity (ev=) scoring (Kotlin's when/elvis/!! vocabulary isn't
measured against isDecisionType's containment claim yet).

test/kotlincheck.sh: a real fixture (Kotlin<->Java bidirectional calls, a
constructed same-name collision pair including an `enum class`/plain-class
pair, cross-file calls, imports), every number pinned from an actual run,
mutation arms proving each assertion non-tautological — including that a
crash on the mutated fixture can't silently read as a passed assertion
(checks binary exit status before trusting absent output).

Verified: kotlincheck.sh ALL PASS on build/ripwire and a --clean-first
asan/ripwire; zero ASan/UBSan/LSan findings across 501 real .kt files (8
local repos) plus the fixture; manifestcheck.sh and every gate whose golden
value this change legitimately moved (kCatalogLangs, the doctor grammar
count 22->23, README's --callers example, the vendored-scanner serialize()
safety classification, the dependency-capable-set ledger) re-verified clean.
Independently re-tested afterward on 5 larger real-world Kotlin corpora not
part of the gated suite (android/nowinandroid, android/compose-samples,
android/architecture-samples, ktorio/ktor, signalapp/Signal-Android —
~9,500 additional .kt files, up to 6,131 files/64,899 symbols in one repo):
zero crashes, zero ASan/UBSan/LSan findings, deterministic, well-formed on
every one. The grammar's one identified parsing gap — `@Composable () ->
Unit`-shaped annotated function-type parameters — affects ~22.8% of
Composable-heavy files on the largest sample measured, localized to a few
bytes each with the enclosing declaration intact (the same shape as this
repo's existing Metal precedent).

Also validated against a hand-written adversarial corpus (41 files: merge
conflict markers, mid-edit fragments, ObjC/Java/Kotlin interop, unicode
identifiers/emoji/RTL, deep nesting, syntax edge cases, deliberately broken
input including invalid-UTF-8 binary garbage) — zero crashes, zero ASan
findings, byte-identical determinism.

That adversarial pass surfaced one genuine disclosure gap, unrelated to
Kotlin: measureFileHealth (ingest_crawl.h) validated the leading sample for
well-formed UTF-8 nowhere in the pipeline, so a file with invalid UTF-8 but
no tree-sitter ERROR/MISSING node reported no degraded-parse signal at all —
`--skipped`'s disclosure contract silently had a hole. Fixed by reusing
jsonesc::utf8SeqLen (no new UTF-8 validator) to scan the same sample already
walked for whitespace, recording bad-sequence byte positions, and counting
only the ones NOT already inside a recorded ERROR span into err/err_ratio —
first-cut version double-counted (badUtf8 added unconditionally, before
cross-checking ERROR spans), pushing err_ratio above its documented <=1.000
ceiling on `\xff\xff\xff@#$%^&*(`-shaped input; caught by two independent
reviews (a backgrounded codex second opinion and a separately user-launched
Opus /code-review), both of which also caught the missing kParserVer bump
(84->85, ingest_cache.h + its quality.h mirror) this behavior change needed.
parsehealthcheck.sh gained two arms (bad-UTF-8-flagged, good-UTF-8-negative-
control) verifying both the fix and the absence of a false positive.

Two independent correctness reviews ran against the original port before it
landed (/code-review --fix on the fable model, and a separate codex second
opinion), then a four-angle /simplify pass (reuse, simplification,
efficiency, altitude); this follow-up round added a second codex pass plus
a separately user-launched Opus /code-review, both run against the
adversarial-corpus fixes — every fix from all reviews re-verified against 8+
corpora and re-gated clean before landing.

Follow-up review round: a four-angle /simplify pass (reuse, simplification, efficiency,
altitude) plus an independent /code-review high --fix pass, both run against this diff
before it lands. /simplify eliminated measureFileHealth's errorSpans vector and its
O(n*m) rescan in favor of a std::lower_bound dedup over sorted bad-UTF-8 positions,
consolidated repeated lang==Kotlin tests in isDecisionType/cc_isNestingControl, and
merged langCompatible's two return branches. The independent review then found two
genuine, pre-existing test-coverage gaps the port had shipped without: captureBases's
Kotlin delegation_specifier handling (both the wrapped-call `Shape()` and bare-interface
`Labeled` shapes) and countParams for multi-parameter Kotlin functions had zero fixture
coverage — no prior fixture exercised a base-class clause or more than one parameter, so
neither fix had anything pinning it. Both are now covered (kotlincheck.sh §9/§10, with
non-tautological mutation arms 7e/7f). The review also reverted the langCompatible merge
back to an early-return shape (langCompatible runs per candidate on a hot resolution
path; the merged form paid for the JVM-bridge comparisons even on the common C-family-
only case) and refined measureFileHealth's dedup to never hold a transiently
over-counted value mid-walk. Re-verified: kotlincheck.sh and parsehealthcheck.sh ALL
PASS (including the new coverage arms and a bad-UTF-8-inside-an-ERROR-span dedup test),
ASan clean, full gate suite clean (gates=572 pass=565 skip=4 fail=3, all three failures
pre-existing/environmental and untouched by this diff).

Rebased onto origin/main a second time (186 commits: upstream independently
landed Dart as language redhat-et#21 while this branch was in flight). Kotlin now
appends AFTER Dart in Lang (index 22, kLangCount anchored on Lang::Kotlin).
kParserVer collided a second time: upstream independently claimed 86 (Ruby
receiver), 87 (markdown scanner counter saturation) and 88 (Dart) while this
branch still held 86/87 from the first rebase — an exact collision with this
branch's own prior use of those numbers. Resolved the same way as the first
collision: re-bumped to the next free values (89, 90) rather than keeping
either side's number, full history preserved in the ingest_cache.h comment.

Six gate failures surfaced after the merge, each investigated individually
rather than assumed to be either all-expected or all-flaky: g1configcheck.sh
and dartcheck.sh/elixircheck.sh each carried their own stale hardcoded
doctor-probe/fuzz-target/seed counts (each language's own gate only bumps
its own copy, a drift-prone hand-maintained pattern — fixed by updating all
three plus a second stale fuzz-executable-name set in g1configcheck.sh that
was masked locally by a Clang-only SKIP but would have failed CI's Clang
legs). ingest_names.h's kotlinEnclosingScopeOf had three literal std::strcmp
calls never converted to the kindIs() house style introduced by upstream's
own per-node strcmp-to-kindIs refactor, since this file predated that
refactor and wasn't part of the original conflict set — converted, fourth
(runtime-variable) strcmp call correctly left unconverted. ingest_crawl.h's
kLangTable auto-merged cleanly (both .dart and .kt rows present) but its
array extent literal silently stayed at 47 — bumped to 48 after an exact
row count. limitstablecheck.sh needed docs/LIMITS.md regenerated via its
own generator script. docs/COMMANDS.md, its newest showcase capture, and
every <!-- gatecount -->-marked reference were regenerated/re-derived rather
than hand-edited, per upstream's newly introduced generated-artifact
pattern for these documents. mcpattrparitycheck.sh and (in the final
pre-push run) fillordercheck.sh each failed once under -j 6 parallel
contention and passed cleanly on repeated standalone re-runs — confirmed
transient, not code issues.

Re-verified after all of the above: full gate suite clean (gates=602
pass=593-598 skip=4 fail=4, only the 4 pre-existing/environmental failures:
w3fixlegendcheck.sh's root-spelling-dependent --dead-code path check,
showcasecapturecheck.sh's pre-existing --stray-content=lane caption-
generation bug, optremarkscheck.sh's Clang-only PGO arms, and
nongitqmetricscheck.sh's non-git-corpus environmental check); ASan/UBSan
clean on both src and test/kotlinfix under a --clean-first rebuild.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_016aPNTmY4XV8bHEkuhdtZnE

PR redhat-et#126 review round (CodeRabbit automated review, 2026-09-10): three findings, all fixed.

1. Bodyless Kotlin class/interface definitions were silently deleted from the graph whenever a
   same-named Java definition existed elsewhere. Kotlin has no forward-declaration syntax for
   types — `interface Taggable` with no braces at all is the type's sole, complete definition,
   unlike a bodyless Kotlin FUNCTION (an interface member signature) or a C header prototype. The
   existing ingest_sidecap.h positional-body fallback (added for §5/§8's already-bodied collision
   fixtures) has no body child to find for a truly bodyless declaration, so bodyByte stayed 0 and
   graph.h's decl/def collapse (`hasBody`, used identically in buildGraph's §1e collapse and
   declToDefFollowThrough) read it as a droppable forward decl. Fixed by widening both `hasBody`
   copies with a narrow, GATED clause: `sym.lang == Lang::Kotlin && sym.kind == SymKind::Class`.
   New fixture (kotlincheck.sh §11): a bodyless `interface Taggable` colliding with a same-named
   Java class, with a mutation arm proving the assertion is not a tautology.

2. `third_party/deps/kotlin/src/scanner.c`'s `stack_push` called `abort()` outright when nested
   interpolated strings drove the delimiter stack to `TREE_SITTER_SERIALIZATION_BUFFER_SIZE` — a
   DELIBERATE process termination, not UB, so no sanitizer flagged it and it was invisible to
   every ASan/fuzz sweep this session ran (the fuzz seeds never happened to nest ~512 string opens
   deep). Violates this repo's own contract: a bad file must degrade, never crash the process.
   Fixed as a proper vendor patch (third_party/patches/kotlin/001-stack-push-no-abort.patch,
   following the yaml/001 and markdown/001 precedent for this exact defect shape): `stack_push`
   now returns `bool`, and `scan_string_start` propagates a `false` through the standard
   tree-sitter "no external token matched" recovery path (an ERROR node) instead of aborting. No
   kParserVer change — this only changes behavior on previously-ABORTING input, never on anything
   that used to parse and cache successfully.

3. `scan_string_content`'s escaped-dollar-at-string-end handling treated the FIRST quote after an
   escaped `\$` as the complete closing delimiter even inside a TRIPLE-quoted string, whose
   delimiter is three quotes — `"""a\$"""` lost its first closing quote and the remainder
   mis-tokenized. Fixed as a vendor patch (kotlin/002-triple-dollar-escape.patch): defers to the
   existing triple-quote-aware end_char handling a few lines down (the same is_triple deferral the
   plain-backslash case two branches below already uses). Real parse-output change on real input,
   so kParserVer bumped 90 -> 91 (ingest_cache.h + its quality.h mirror).

Both vendor patches gained live tripwires in vendorpatchcheck.sh (new arm J): a generated
700-deep nested-interpolation fixture (the ASan-flavour crash tripwire for redhat-et#2) and a committed
tripledollar.kt fixture (the plain-build exit-0-wrong-answer tripwire for redhat-et#3, checked via
degraded_parse=0 and that the symbol declared right after the tricky string still extracts).

Re-verified: kotlincheck.sh and vendorpatchcheck.sh ALL PASS on both build/ripwire and a
--clean-first asan/ripwire; zero ASan/UBSan/LSan findings re-confirmed across all 8 local
real-world Kotlin corpora (501 files) on the fixed binary.
joyful-ii-V-I added a commit to andriytyurnikov/ripwire that referenced this pull request Sep 12, 2026
…cs/QUALITY_DELTA_CATALOG.md

The "What --quality-delta catches" placeholder (slide 29) is filled from the catalog (commit 553dbadb on
docs/quality-delta-catalog-2026-09-11). Every example commit is on main 766913d; the catalog itself has not landed.
Four cards, varied in kind and in what the finding led to:

- redhat-et#1 new-clone-of-reused-helper: vendoredPathPrefixes re-spelled registeredMacroNames, a helper with three callers,
  while the round that built the per-kind dials was under way; both now call mergeBuiltinsWithConfig.
- redhat-et#6 nesting (and complexity 141 -> 178): the inline per-root scan became std::any_of in a named function.
- redhat-et#7 dead-code: the predicates a rewrite orphaned were deleted, not acked. The card also names the one the kind
  missed, namesFileNotKept, whose only caller was itself dead (IP-8).
- redhat-et#3 duplication, Type-3: the shared digit-run locator became headerFieldDigits, and the card says the detector
  still gates a 63-token residue, which the catalog classes as a false positive (IP-4).

Each card's speaker notes cite the catalog section, the base -> before -> after commits, the verbatim row(s), and
a one-line reproduce command copied from the catalog's Reproduce block. All five catalog examples considered for
the slide (redhat-et#1, redhat-et#3, redhat-et#6, redhat-et#7, redhat-et#8) were re-run on 2026-09-11 with a built_from=766913d02 build, and every row, gating
count and exit code matched the catalog.

Why four and not six: the render shows a two-column pane holds 10 lines of 39 columns at 8 pt (a 40-column line
wraps). With 5 or 6 entries each pane drops to about 3 lines, too few for a before and an after. So redhat-et#5 and redhat-et#8 are
left out, and redhat-et#3 carries the detector-limits card.

Layout changes in qdExamples, each forced by the render:
- The kind box is sized to its text. new-clone-of-reused-helper is 2.38 in at 11 pt and wrapped out of the old
  fixed 2.2 in box.
- Card padding and pane gap are tighter, and two-column code is 8 pt (was 8.5), so a pane fits 10 x 39 instead
  of about 9 x 36.
- Entries may carry an optional notes array, which is validated and printed under that card's line in the notes.
- The kicker for a filled slide no longer reads as a placeholder. The placeholder path is unchanged.

Snippets are the catalog's excerpts, trimmed to fit. "…" marks elided code, long lines are wrapped, and a comment
line stands for code the catalog says was elided or deleted. The notes say so.

Gates on this tree: deckcheck ALL PASS (0 bad values, 0 stale), deckclaimcheck ALL PASS (179 long flags, 34
slides), readmedriftcheck ALL PASS. All three pass with the built_from=766913d02 build, and again with the
scratchpad main_096e3544 build. The pptx and PDF are regenerated (pptxgenjs 4.0.1, LibreOffice impress_pdf_Export,
34 pages).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit to andriytyurnikov/ripwire that referenced this pull request Sep 12, 2026
…not list every row

classifyHelpLine decides a line's role by indentation alone: four spaces is a row, five or
more is the continuation prose tier 1 drops. Twelve rows in src/cli.h were written at SIX
spaces, to read as sub-flags of the entry above them. Tier 1 therefore culled fifteen flags
the parser accepts -- --and --not --grep-in --grep-scope --grep-context --grep-before
--grep-after --ack-only --scope --edit-payload --edit-target-file --no-post-check --dry-run
--apply --partition -- and v0.6.0 shipped that way. They all work; they were simply
unfindable from the first screen. The sharpest edge was `--help=--and`, which refused with
"`ripwire --help` lists every row" -- a false claim about our own output, which is the part
non-negotiable redhat-et#3 forbids and the reason this is a bug and not a layout preference.

FOUR SPACES BECOMES THE SINGLE DEFINITION, on both sides. docs/docs_commands_build.py's
parse_help accepted 4..6 while the binary accepted only 4, and being the more permissive
reader is exactly what hid the divergence: COMMANDS.md documented all fifteen, so nothing
downstream of the generator ever showed a symptom. It is narrowed to exactly four here.
The alternative -- teaching classifyHelpLine about six -- would keep the visual nesting but
put a second legal row indent into a surface whose whole contract is that indentation
decides, and would force helpbudgetcheck arm (C) to stop calling deep lines continuations.
One indent, one meaning, is the cheaper invariant. Narrowing also lets docscommandscheck
arm (H1), whose candidate detector deliberately still scans 4..6, see a future six-space
row as a parser loss.

THE GATE COULD NOT CATCH THIS AND REPORTED THAT AS FINE. helpbudgetcheck's `rows` extractor
is the emitter's own four-space rule, so a misclassified row is absent from BOTH sides of
arm (B)'s comparison: (B) printed "the split culled nothing" about exactly the rows that
were culled, and its comment claimed every row was fenced. Teaching `rows` about six spaces
would only move the blind spot. New arm (K) takes its population from the other side of the
binary instead -- the kBoolFlags / kViewFlags / kIntFlags tables and the hand-written
parseArgs arms, scraped by test/flaguniverse.py, which three other gates already share --
and asserts that every flag the tool ACCEPTS is named in tier 1. It was observed RED on the
unfixed tree (14 flags) and is green after. Arm (L) is its mutation control. Eleven
accepted-but-unadvertised flags sit on a committed debt list with a reason each, so that gap
is on the record instead of inside an extractor's blind spot.

The eight rows that had continuations now open with a summary that stands on its own, as the
tier-1 contract requires and as every four-space row already did; the sentence each used to
open with is kept verbatim as that entry's first continuation line, so tier 2 loses nothing.
Tier 1 goes 4480 -> 4844 tokens against an unchanged 7000-token ceiling.

Re-pinned deliberately: test/printf_parity.manifest, labels `help` (6a15d43c -> 3290aebf)
and `help_all` (5d748a4d -> 141428b5), 40 labels unchanged; docs/COMMANDS.md regenerated
(13 lines, all of them these rows' own text).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
joyful-ii-V-I added a commit that referenced this pull request Sep 14, 2026
…ing, and the no-throw copy threw

Six findings from one review, every one of them a surface that was silently wrong rather
than loudly broken.

--pr-context PRINTED A WRONG est_tokens WITH NO DISCLOSURE. When a trim level's measurement
render fails, prRenderLevel returns an EMPTY body; pickPrTrimLevel priced that empty body, the
price fit, and the ladder broke at level 0 — while writePrContext correctly streamed the
complete untrimmed floor through emitFiles( out, kPrTrims[0], nullptr ). The only signal was
DEGRADED_PATH_ALERT, which src/infra/Diagnostics.h compiles to `do {} while (0)` under NDEBUG,
so the binary a user installs printed a modelled number with nothing at all saying so
(non-negotiable #3). The BYTES were never the defect and do not move: cutting answer rows
because a measurement buffer failed would let a cap decide the content, which is the one thing
a cap may never do. The fact goes where this class of fact already lives — truncated= now
carries ";est-unmeasured", re-priced with the label in place (the label lengthens the root tag),
and the legend defines it in budget-floor-exceeded's own voice.

THE CHARGE IS READ OFF THE LABEL, not off a boolean beside it. prPriceDocument decides the
clause from the truncated= value it is already handed, so ONE condition decides both the priced
legend and the delivered legend and they cannot drift apart; and the two conditional clauses now
arrive as a named PrLegendClauses{ runHint, estUnmeasured } rather than two bare bools, because
`prLegendText( escBase, unindexed, true, false )` says nothing at its call site about which
clause is which. Both readings came out of --quality-delta: threading a seventh parameter into
prPriceDocument and a fourth into prLegendText took the range form to gating="2" (a params row
from minor to major, and an api-surface contract change that invalidated a standing ack). The
range form is gating="0" now with NO new ack — the findings are gone rather than suppressed.

THE DEFINITION IS LABEL-GATED, which the gate found for me. Spliced unconditionally, the ~390 B
clause cost test/defaultceilingcheck.sh's fixture its entire remaining headroom: that 120-file
tree prices at 7,989 of the 8,000 default — 11 tokens spare, as E1 measured when it gated the
run-hint clause for the same reason — and went to 8,037, over budget on a document with nothing
unmeasured about it. So kPrEstUnmeasuredLegendClause rides exactly the document that carries the
label, decided by the fact the ladder recorded (PrTrimRender::rendered) and never by a search of
the rendered bytes; the pricer charges its size on the same fact, so the priced legend and the
delivered legend cannot disagree. A healthy document is byte-identical to before (est_tokens
7,989, re-measured) and prcontextcheck (F-legend) holds it that way.

THE LABEL CROSSED prBudgetTail's BUFFER. test/fixedbufsweep.sh had this buffer at 248 B of
tail[256] — "SEVEN bytes of margin ... one more attribute crosses it" — and ';est-unmeasured'
is 15 more and CAN ride beside ';budget-floor-exceeded' (a small --max-tokens puts even the
unmeasured empty-body envelope over budget). Worst case 88 lit + 90 digits + 85 label = 263 B,
so tail[320], 56 B of margin, and the sweep's row moves in this commit with the recomputed
number. rw::formatTo was not what had been saving it: it truncates SILENTLY and its return is
not read there, so an overrun would have dropped the closing quote of truncated=" and shipped a
malformed root — a G4 breach with no diagnostic.

renderToString's NO-THROW CONTRACT HAD A THROWING LAST STATEMENT. out.text.assign( buf, sz ) is
the one allocation on the success path and it sat outside the handler, so a std::bad_alloc from
it escaped a function documented to ALERT a failure and return ok == false, and jumped the
std::free( buf ) two lines below on the way out — leaking the memstream buffer. Caught in its
own handler rather than one around the whole body, because the two failures need different
cleanup (the emitter's throw owns an OPEN stream; by this point only buf is left), with its own
alert literal, and control falls THROUGH to the single free() so buf is released exactly once on
every path.

THE SHARED ROW READER'S MALFORMED-FIELD DETECTOR HAD A HOLE OF ITS OWN SPECIES.
test/testrowpaths.py found "tests_to_run" and then scanned arbitrarily far forward for a '[', so
{"tests_to_run":null,"other":[{"p":"ghost.cpp"}]} sliced the NEXT field's array and returned
ghost.cpp at exit 0 — a foreign field's paths served as this field's answer, where the docstring
already promised a TestRowParseError. The value is read adjacently now: past the key, a ':',
optional whitespace, then '[' or raise.

AND TWO PATH READERS HAD NEVER BEEN CONVERTED. The census over test/ for the four shapes the
shared reader replaced found test/affectedcheck.sh's tset() — inside the very file the reader's
docstring names among those it converted, so that claim was false — splitting EVERY row's p= on
',' including a single row's, which turns a comma-bearing path (never grouped, by testmap.h's
refusal) into two names that name nothing; and test/testgatecheck.sh's tset() matching `<t p=`
singles only, which returns the EMPTY set on a two-runner-less-test fixture where the shared
reader returns both paths. Both route through the shared reader now, and the docstring records
the census. Every other hit counts rows (listingpagingcheck, w3fixlegendcheck, testgatepagecheck
— all group-aware in place) or pins one exact row spelling with a regex that fails loudly;
deeptailcheck's `<t p=` rows are --for's tail listing, a different element sharing the tag.

Two documentation drifts beside them: skills/ripwire-mcp/SKILL.md claimed `p` for
situational_awareness, which emits `test` (src/mcp.h's kTestRowJsonShapeClause states the split
and the binary is the authority), and bench/arb/run_arb.py decoded a &#44; the seam stopped
emitting on 2026-09-13 while decoding NONE of the entities it does emit — so a path holding '&'
was scored against a file name that does not exist. Both row shapes there share one decode now.

GATES, all red on the parent commit and green after:
  * test/prcontextcheck.sh (F-legend)(F5)(F6) — est-unmeasured in truncated= on the degraded
    root, the complete body still served, and the legend defining the term. RED: the degraded
    root printed truncated="none" while pricing an empty body, and no legend defined the label.
  * test/prcontextcheck.sh arm (G) — INFRA_FAULT_RENDER_COPY_THROW, the emitter switch's twin,
    injected immediately before the assign. RED: "produced no DEGRADED_PATH_ALERT on a binary
    that PROVED it can emit one". Honest in both flavours, mirroring arm (F): the switch and the
    alert live only on the non-NDEBUG build, so the plain leg proves the degrade and the NDEBUG
    leg asserts the verb is intact and no false disclosure appears. The est-unmeasured LEGEND
    definition is asserted on EVERY flavour, which is the point of moving the disclosure off the
    alert.
  * test/testrowruncheck.sh arm 17 — every non-array tests_to_run value is exit 2 in both paths
    and jsonlist, with a well-formed array and JSON whitespace as controls. RED: five documents
    at rc=0, three of them serving ghost.cpp.

Suite: 628 gates, 626 pass, 0 fail, 2 environmental skips (argvdiffcheck and editchecknotecheck,
both wanting a reference binary), tree_writes=0. ASan/LSan clean on --pr-context healthy and on
both injected degrades.

Pins moved, two, both in test/fixedbufsweep.sh and both with the measured recomputation in the
same commit: the src/prcontext.h tail TABLE row 256 -> 320 (worst case 248 -> 263 B, margin 7 ->
56), and EXPECTED mentions 323 -> 324 — re-read from the diff, not accepted from the delta: the
one added line is the COMMENT explaining that growth, which names formatTo. calls, sites, rows
and widthforms are unchanged at 219/219/92/0. No legend or byte pin moved.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant