Skip to content

review-pr: add T5, X8, D13, V7 and V8 from #1797 - #1817

Open
xiaofei-zheng wants to merge 3 commits into
mainfrom
feature/xiaofei/review-pr-order-rules
Open

xiaofei-zheng wants to merge 3 commits into
mainfrom
feature/xiaofei/review-pr-order-rules

Conversation

@xiaofei-zheng

@xiaofei-zheng xiaofei-zheng commented Oct 10, 2026 •

Copy link
Copy Markdown
Collaborator
  • Description: Five review-pr rules learned from fix(cli): export every --extra-env pin so all readers resolve it alike #1797, which fixed one defect across four rounds — each round finding a defect introduced by the round before it, and each one caught by the reviewer rather than by the checks. The rules name the five reasons the checks stayed silent.

    T5 (tests): the fix moved a projection relative to the ladder that resolves a pinned value, and the tests added with it called the two in the order they wanted — passing identically before and after, while the production path still ran them the wrong way round. An ordering fix asserts the order itself, or drives the real entry point.

    V7 (review method): the AST walk used to record every os.environ write in _run_optimize never listed TP/CONC/EP, because they are written one frame down inside a helper. A tool scoped to one function and a tree with no writers produce the same empty output, so a search is reported together with its scope, and a clearance written without the enumeration says SKIPPED instead.

    V8 (review method): splitting a function and moving its call earlier broke nine tests across three files the author had not opened — a Namespace that suddenly needed one more attribute, cases calling the half that no longer did the job, cases that now needed an empty environment. Tests are callers; the caller list comes from a grep on the symbol, not from the files the diff already touches. V7 enumerates the readers of a value, V8 the callers of a symbol.

    D13 (design): treating an --extra-env pin as a third source alongside the flag and the environment required five mechanisms whose only job was to keep the pin from colliding with the projection. The tree already resolved the same shape in one ladder; reusing it removed all five, 92 lines, and both defects found in the intervening rounds could not arise in that form. The count on its own is taste — the finding is the count plus the simpler form the tree already contains.

    X8 (cross-artifact): the comment block that PR added to cli/__init__.py was appended to across four rounds of review and recounted how the arrangement reached its final shape, including a docstring describing a call that had been deleted. AGENTS.md Comment below the local average already forbids narrating the change; this is the reviewer-side check. The ratio is measured per non-test file and net of removals -- count a line as comment when it is a whole line leading with # or part of a bare string-expression statement, drop blanks, net the adds against the removes, compare to net code lines -- and it is taken on the head under review rather than on the merged result: fix(cli): export every --extra-env pin so all readers resolve it alike #1797 cut its own commentary back in e8ff4867c, so its final state no longer fires, while 7a9752e4d, a head the review actually ran on, does.

  • Linked issue(s): none

  • Tests: none — rules and skill prose only, no executable code. Backtested against fix(cli): export every --extra-env pin so all readers resolve it alike #1797's own history, which is where they came from: V7's sweep fires on the exact commit that moved the TP projection past _preflight and clears on the one that moved it back; V8's git grep on _resolve_workload_knobs lists the two test files the split broke and the author had not opened; D11 (review-pr: add D11 and D12, and make complexity above 20 a gate #1811) fires on the first head with 5 bare readers of the backend in files the diff never touched; T5's anchor assertion fails on the head before the ordering fix. X8 was corrected by the backtest itself. Verified by hand that all 60 rules have unique ids (the first draft collided with the existing D12), that every rule resolves to an index row or is one of the always-on V1-V8/X2, that no row names an undefined id, and that the SKILL.md rule count and always-on list moved with them.

  • Size/complexity triggers crossed: none

  • If this simplifies or refactors: n/a

  • Observable effect: a reviewer running review-pr on a diff that reorders calls, splits a function, or defends itself with five interlocking mechanisms now has rules that fire. On fix(cli): export every --extra-env pin so all readers resolve it alike #1797 the first four rounds returned no blocking issues from the author's own review each time.

  • Breaking changes: no

  • PR addresses single concern: yes

  • Root cause is upstream: no

Five rules from one PR's review history. #1797 fixed a real defect four
times over, and each round was caught by its reviewer rather than by the
checks, for a reason the rules did not name.

T5 -- the fix moved a projection relative to a resolver, and the tests added
with it called the two in the order they wanted, passing before and after.
An ordering fix asserts the order itself or drives the real entry point.

V7 -- the AST walk recording every os.environ write in _run_optimize never
listed TP/CONC/EP: they are written one frame down. A tool scoped to one
function and a tree with no writers give the same empty output, so a search
is reported with its scope.

V8 -- splitting a function and moving its call earlier broke nine tests in
three files the author had not opened. Tests are callers; the caller list
comes from a grep on the symbol, not from the files already in the diff.

D13 -- treating a pin as a third source alongside the flag and the
environment took five mechanisms whose only job was to keep them apart. The
tree already resolved the same shape in one ladder; reusing it removed all
five, 92 lines, and both defects found in between.

X8 -- that PR added 56 lines of comment against 28 of code, most of it
recounting how the arrangement got there. AGENTS.md already forbids
narrating the change; this is the reviewer-side check, measured on the
diff's own ratio.
@xiaofei-zheng
xiaofei-zheng requested a review from a team as a code owner October 10, 2026 15:28
Backtesting the rule against #1797's own history: the whole-diff totals read
6591 code against 1216 comment and cleared at every head, because test code
dilutes the ratio and the merge of main swamped it. cli/__init__.py measured
on its own read 28 against 56 and fired from the second commit onward. A
per-commit reading inverts for the opposite reason -- a docs-only commit
adds comment by construction.
xiaofei-zheng added a commit that referenced this pull request Oct 11, 2026
…ationale

Self-review with the rules in #1817 found two things this PR had missed.

T5: the case covering the seeding calls it and the ladder in the order it
wants, so it passes with the two swapped in _run_optimize -- verified by
swapping them, where the behavioural case still passed and the defect was
live. The ordering anchor table gains the pair, resolved against the resume
branch's own ladder call rather than the fresh one, and fails on the swap.

X3: a tree-wide sweep for the claim this PR's ladder contradicts found one
more copy -- the Qwen skill told operators that "CLI defaults can otherwise
override the intended workload", which stopped being true when the default
became the bottom rung. The advice stands; the reason is now the real one.
Two further copies state the advice without a rationale and are left alone.
@jiaqiang-dot-liu

Copy link
Copy Markdown
Collaborator

PR #1817 -- review-pr: add T5, X8, D13, V7 and V8 from #1797

What it does: PR #1797 took four rounds, each round introducing the defect the next one found, and
every one caught by the reviewer rather than by a check. This adds the five review-pr rules that
name why the checks stayed silent -- T5 (an ordering fix needs an assertion on the order, not a test
that replays it), V7 (enumerate a value's writers tree-wide, not inside the diff's frame), V8 (take
a symbol's caller list from a grep, tests included), D13 (count the mechanisms a fix needs to hold
itself up) and X8 (a comment states a constraint, not the history of the change) -- places each
body at the end of its family in rules.md, adds the index rows that make them reachable, and moves
SKILL.md's count and always-on line to 60 rules and V1-V8.

Blocking issues: 1

  1. [free:x8-backtest] X8's calibration figure for fix(cli): export every --extra-env pin so all readers resolve it alike #1797 is not reproducible, and the real figure
    does not fire the rule [verified]
    Problem: X8 cites one PR as evidence and states its measurement twice --
    .claude/skills/review-pr/rules.md:267 ("cli/__init__.py gained 56 lines of comment and
    docstring against 28 lines of code") and .claude/skills/review-pr/rules.md:275-276 ("cli/__init__.py
    on its own read 28 against 56"). Measured over fix(cli): export every --extra-env pin so all readers resolve it alike #1797's own range (880c167..2f2f7bf), the net
    delta for that file is +28 code / +26 comment. The code half matches exactly; the comment half
    does not, and cannot: the file grows 66 lines of which 12 are blank, so net code plus net comment
    must equal 54 for any partition of lines, while the quoted pair sums to 84. The figure belongs to
    a mid-branch state -- at 7a9752e4d, the commit before fix(cli): export every --extra-env pin so all readers resolve it alike #1797's own cleanup
    e8ff4867c docs(cli): cut the commentary back to the constraints, the file reads +28 code /
    +49 comment, and net comment peaked at +58 earlier in the branch. The companion sentence at
    rules.md:275 has the same shape: "the whole-diff totals read 6591 code against 1216 comment"
    is not an added-line count of a 1022-line diff; it is a whole-file census of the touched files
    (5408 code / 1229 comment at head, or 6341 / 1229 charging blanks as code).
    Impact: X8's measurable tell is "net of removals, the comment and docstring lines the diff adds
    outnumber the code lines". At 26 against 28 the rule stays silent on the one diff it was learned
    from -- the evidence block, read as the rule instructs, clears it. A reviewer who checks the
    citation finds a number that does not reproduce and discounts the rule; a reviewer who does not
    carries a threshold whose only backtest is a branch state the author had already cleaned up.
    SKILL.md makes the citation load-bearing: "Add the rule body to rules.md ... with the real PR
    it was learned from quoted as evidence. A rule with no PR behind it is a hypothetical and does
    not go in."
    Action: re-measure cli/__init__.py over 880c167..2f2f7bf and replace both figures with +28 code
    / +26 comment, then either state the threshold that fires at that margin or cite the mid-branch
    head the rule actually fires on by sha. Relabel the 6591/1216 sentence as whole-file totals, or
    recompute it as added lines (337 code / 94 comment gross, 254 / 75 net).

Checked: .claude/skills/review-pr/rules.md and SKILL.md at ce8f794 and at base 880c167 (all 60
rule headings extracted and set-compared against every id on the index table -- no duplicate id, no
row naming an undefined id, no body off every row except the always-on V1-V8/X2); tree-wide
git grep at head for V1-V6, 55 rules, rules across and review-pr (only SKILL.md:87 carries
a count; .github/workflows/pr-review.yml, .github/scripts/pr-review-publish.sh and AGENTS.md:27
name no rule id); AGENTS.md:123 and :129 for the bullets X8 and D13 quote (both quoted
accurately); src/hyperloom/inference_optimizer/cli/__init__.py at 880c167, 7a9752e and 2f2f7bf;
#1797's test_cli_resume_launch_shape.py for the ordering assertion T5 describes (present, an
ast.walk over _run_optimize comparing first-call line numbers); ci.txt at head (reuse, Gitleaks,
CodeQL, pyproject, wheel, readthedocs green; e2e/resolve path-filtered out).
Ran: git show <sha>:src/.../cli/__init__.py at three shas, then a tokenize + ast classifier
over each (comment = #-leading lines plus bare string-expression statements, non-blank), both
whole-file and remapped onto every +/- line of #1797's diff; the same classifier summed over all
eight .py files #1797 touches; independently re-derived by a second reviewer who also swept 15
alternative metrics for one yielding 56 and found none.
Base: 880c167 | Head: ce8f794

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants