Skip to content

ci: code-metrics gate on touched files with baseline-free checks; fold in import-linter, PR hygiene and ruff principle rules - #1812

Open
zoroyihan7 wants to merge 38 commits into
mainfrom
ci/code-metrics-gate
Open

zoroyihan7 wants to merge 38 commits into
mainfrom
ci/code-metrics-gate

Conversation

@zoroyihan7

@zoroyihan7 zoroyihan7 commented Oct 10, 2026 •

Copy link
Copy Markdown
Contributor
  • Description: what and why

    One PR for the CI quality gates. It adds the code-metrics gate and folds in three sibling PRs, merged in with git merge (no rebase); they will be closed as superseded.

    Folded in

    code-metrics: dimensions and why each number (pyproject.toml [tool.hyperloom.code_metrics], every number commented there)

    Dimension Fails when Why Baseline entries
    Cyclomatic complexity > 20 The style guide's Complexity ceiling (review-pr: add D11 and D12, and make complexity above 20 a gate #1811, review rule D12), now measured: same number, same two cases 122
    Cognitive complexity > 30 Twice SonarSource's default of 15, the same step as the ceiling 326
    Function length > 80 physical lines def through the last line, decorators excluded, comments/docstrings included, a nested function is measured on its own, not folded into its parent; past one screen 477
    Nested block depth > 5 pylint R1702 default 29
    Module length > 1200 lines (> 800 is a report-only warning) half again the 800-line review trigger 55
    Duplication (jscpd) any clone >= 100 tokens, >= 10 lines SonarSource CPD defaults, unchanged 123
    Dead code (vulture) any finding >= 80% confidence unchanged 2

    Dropped: maintainability index (56 of 78 baselined files sat at MI 0, so worsening there was invisible; its comment term rewards adding comments, which the new comment check works against; no surveyed major Python repo gates on it), and pylint's statements, branches, returns, arguments, positional arguments, locals and public methods (function length replaces statements). radon is no longer pinned or installed.

    Baseline: scripts/code_metrics_baseline.txt, 1134 entries in 1138 lines (was 3406 entries in 8025 lines of nested JSON), one sorted line per entry, <metric> <path>::<Class.method> <value> (module metrics: the path alone), with a header stating the rule and the update command, so a baseline change is one diff line per unit. Seeded on this branch with python scripts/code_metrics.py --seed-baseline, then tightened with --update-baseline.

    Scope: touched files only. A finding is judged only when its file is changed by the PR (diff against the merge base; for a push to main, against github.event.before). Untouched files, baselined or drifted on main, never fail an unrelated PR and are listed under "Outside this change (informational)". In a touched file: a new or worsened unit fails (for function length: a function that newly exceeds 80 lines or grows while over it; editing a recorded long function without lengthening it is not a failure); a baselined unit that was fixed must leave the baseline in the same PR (out of date fails with the update command). New and worsened units are also compared with the same file at the merge base, so a unit the base already had at that value or worse is backlog, exactly D12's comparison. The "baseline may only shrink relative to the base" check stays whole-file.

    Override: the baseline-raise label. Read from the GitHub API when the gate runs (--pr-number), never from the event payload. It turns baseline growth and new/worse findings into waived findings, still printed in full in the job log (untruncated) and in the report. It waives nothing else (out-of-date entries, config loosening and the checks below still fail). Use it only with the reason in the PR description. Adding or removing a label, or editing the title or body, re-runs the job (pull_request types include labeled, unlabeled, edited).

    Checks without a baseline (same job, same sticky report, each a section; each fails the job):

    • Comments (added lines of .py files, docstrings excluded): an added run of more than 8 full-line # comments; an added comment with a PR/issue reference (#123+, PR 1234, PR-1234, GH-1234, issue #12, a /pull/ or /issues/ link) or incident narration (a dated incident or outage, "after the outage", "postmortem"). A comment starting with TODO and the line after it may link an issue (Ruff TD003). The diff is read with -M -w --text --no-textconv, so a renamed file, a moved or re-indented comment is not "added" and .gitattributes cannot hide hunks. Measured base rate on the last 80 commits: 2 commits had >8-line added blocks, 0 had PR references.
    • English only: no CJK character (ideographs, CJK symbols and punctuation, halfwidth/fullwidth forms) in any tracked text file, nor in the PR title, body or any commit message (read from the API). Backlog fixed here: scripts/ray_exec_p0_smoke.py, src/hyperloom/inference_optimizer/tests/test_inferencex_client_unit.py, src/hyperloom/orchestrator/enablement/tests/test_artifacts.py (multi-byte test data now uses the euro sign and Arabic-Indic digits). Absolute, no baseline.
    • Production code does not import test code: no non-test module under src/ or scripts/ imports a tests package or a test_* module (an AST pass, so it covers scripts/, which import-linter cannot see; importlib.import_module/__import__ with a literal name and conftest modules count). kernelforge.mcp_server.tools.test is a tool module and does not fire. Backlog fixed here: scripts/refresh_inferencex_anchor_contract.py imported hyperloom.inference_optimizer.tests.test_inferencex_anchor_contract; the shared contract helpers moved to hyperloom.orchestrator.actions.executors.inferencex_anchor_contract (public, with INFERENCEX_REPO_DEFAULT made public in preflight, so the script and the module pass A1 below), used by the script and the test. Absolute.
    • Repeated literals (diff-only, production modules the PR touches): taking a string or bytes value (3+ chars) or a number (not 0, 1, -1, 2) from fewer than 3 occurrences in a module to 3 or more fails, counted before and after; a literal already repeated at the base is backlog (a mapping table may gain a row), a renamed module is compared with its source, and a literal whose total over the touched modules did not grow only moved (a split). Dict keys, x["k"], the key of .get/.pop/.setdefault and of an in test, keyword names, docstrings, f-string text, annotations, __all__ and test files do not count. Diff-only on purpose: a whole-tree version hits 6089 literals in 479 of 832 modules.

    Design checks (diff-only, scripts/code_metrics_design.py; each its own report section, refusal message naming the rule and the fix, and how-to-fix line; tests exempt from all). A line is added when the diff adds it and the change did not remove the same text elsewhere (the comment and literal checks' -M -w machinery: moves, re-indents and renames are not additions); a function or class is new when its def/class line is added and the base file (or its rename source) had no unit of that qualified name. Replayed with the full gate over the 30 PRs merged just before (fix(orchestrator): bound the prompt sections that grow with the session #1735-fix(orchestrator): keep a PRELUDE skip_to_kernel hint from ending FRAMEWORK_AGENT #1808, merge commit vs first parent), the hit rate per check is:

    Check Fails when PRs failed of 30
    A1 private-name reach an added import or attribute access of a private name (or module under a private package) owned by another package unit; units = root packages, layers and independent modules in .importlinter, longest prefix wins, scripts/ its own unit, units read from the merge base's contracts (moving import-contracts is loosening); dunders and same-unit use exempt 1 (refactor(cli): one hyperloom <command> entry point; remove dead CLI shells and fix CLI bugs #1736: scripts/check_cli_references.py reads hyperloom.cli._COMMANDS and _SESSION_COMMANDS; true positive)
    A8 flag argument a new function whose first statement branches on a bool/two-value Literal parameter (either side of the comparison) into a branch ending in return/raise; main, click/typer commands and argparse.Namespace handlers exempt 0
    A9 tuple return a new function returning a tuple of more than tuple-return-max-elements (2) positions, by annotation (tuple, typing.Tuple, also inside Optional/Union/|) or display; tuple[X, ...] (and the displays such a function returns) and dunders exempt 0
    A10 raw vocabulary an added ==/!=/in-display comparison or case pattern with the string value of a member of exactly one Enum under vocabulary-roots (src), in a module that names that enum, against a subject that reads as the member (.value, or a name whose words end like the enum's: status for RunStatus); raw input on its way to the enum (raw == "running") passes 0 (narrowed; see below)
    A11 hidden global write an added global statement or store into globals(); nonlocal passes 0
    A12 pure forwarder a new function whose body is one return of a call passing exactly its own parameters through; bare-name builtins, classes (factories), .get lookups, methods of a constant, decorated functions (other than staticmethod/classmethod) and methods overriding a base-class method exempt 0
    A12 single implementation a new ABC/Protocol with exactly one implementation (subclass, or for a protocol a class with all its methods) in src, in the same package unit; an interface a lower unit declares for a higher one is exempt 0

    Across the 30 PRs the replay saw 89 new functions and 4 new classes (cross-checked from git blobs). Narrowed with evidence: A10 first matched any enum value in src; its one PR hit (feat(runtime-findings): scan server logs, surface findings to every role, allow verified correctness-fix KEEP #1793 entry["status"] == "unknown" vs SpecialistFailureType.UNKNOWN) was another vocabulary reusing the word, and on the whole tree 165 of 173 hits were in modules that never name the enum, so it now needs a value unique to one enum and a module that names it (whole tree 173 -> 3); review then showed two of those three were the conversion of raw input or a version sentinel, so the compared side must also read as the member (whole tree 3 -> 1: recipe_journal.py comparing getattr(config.mode, "value") with "remote"). A12 forwarders skip builtins, factories and .get (whole tree 52 -> 43: SECTION_SHAPES.get(section), json.dumps(value) were accessors, not layers). Config in [tool.hyperloom.code_metrics]: import-contracts, vocabulary-roots, tuple-return-max-elements (raising it, or dropping a vocabulary root, is refused as loosening; a missing contracts file is "could not run"). On this branch the checks found four true positives in the gate's own new code, fixed here: the private anchor-contract module and constant above, Finding.key returning tuple[str, str, str] (now a Key NamedTuple), and a violates() that only forwarded to worse(). Adversarial pass before push: the Optional/Union triple, the reversed Literal comparison, case patterns, globals()[...] stores and a delegate method named like a builtin (self._client.next(job)) got past their checks and are now caught; a functools.wraps wrapper, an adapter method overriding its ABC, ", ".join(parts), __reduce__, a display returned under tuple[int, ...], an attribute of a parameter that shadows an import and an @lru_cache function were refused and now pass; A1 no longer trusts a contracts file the PR edits. Each has a test that is red without the fix. Rerun of the design checks alone over the last 30 first-parent commits of main with the final rules: A1 2 of 30 (refactor(cli): one hyperloom <command> entry point; remove dead CLI shells and fix CLI bugs #1736 above, and explore.py in f261cde importing grid_server_args._MULTI_VALUE_FLAGS across units, also a true positive), every other check 0 of 30. Also fixed: git output is decoded with replacement, so a binary file in a PR's diff (read with --text) no longer stops every diff-only check with a decode error.

    Kept from before: the sticky PR comment, the fork workflow_run poster with the conclusion header taken from GitHub, the base branch's copy of the gate scripts judging the PR (code_metrics_checks.py and code_metrics_design.py now among them), scope-loss and pathspec refusals, best-effort posting. code-metrics still runs on every PR (no paths-ignore). The job holds a token, so nothing PR-controlled is read outside the checkout: the baseline must be a plain repository path to a regular file (a malformed line is reported by number, never quoted), a symlinked .py in scope is refused rather than read or skipped, and the report and artifact are written under RUNNER_TEMP, not in the checkout.

    Admin actions

    • Create the label baseline-raise (it does not exist yet). size-exception was created with this PR, which carries it: the diff budget counts 2,699 production lines, almost all of them docs(review-pr): add the A rule family for code shape, and the AGENTS.md principles it rests on #1815's mechanical Ruff fixes across 122 files, bundled on request.
    • Mark code-metrics and pr-hygiene as required status checks on main. Ruff and import-linter live in lint.yml, which has paths-ignore, so those two always-run jobs are the ones to require.
    • Enable require_code_owner_review and require_last_push_approval in the main ruleset: a pull_request run uses the PR's own workflow file, so only CODEOWNERS review of /.github/workflows/ stops a PR from replacing the gate steps.

    Known limits

    • The head-versus-base comparison covers the per-file dimensions (cyclomatic, cognitive, function length, nesting, module length); duplication and dead code are cross-file and are compared with the baseline only. If main ever gains an unrecorded clone (a baseline-raise merge that skipped the baseline update), the next PR touching either copy sees it as new and needs the label or the baseline entry.
    • Diff-only comment checks are heuristics: a long comment moved to a new file is excused only when the same lines are removed elsewhere in the change; #256-style counts in prose still read as a reference.
    • Non-UTF-8 text without a byte-order mark (GBK, for instance) is treated as binary by the English-only check.
    • A module is matched across a move by its file name, so renaming a module that is already over 1200 lines needs the label (functions and classes are matched by qualified name and move freely).
    • Literal data rows such as [("A", 100), ("B", 100), ("C", 100)] count as three uses of 100; name the value or carry the label.
    • Like every pull_request gate, a PR can edit its own .importlinter contracts or [tool.coverage.run].omit; CODEOWNERS review of those files (admin action above) is what closes that, not the jobs.
    • The design checks are syntactic: A1 follows attribute access only through imported names (not obj._x on an instance), A10 knows only string-valued enum members, and A12 matches a protocol's implementations by method names and an ABC's by base-class name.
    • If the API cannot be read (labels, PR metadata) the gate exits 2 ("could not run"), which is red, never a pass.
    • review-pr: add D11 and D12, and make complexity above 20 a gate #1811 recorded 124 units above 20 on 2026-10-10; the gate measures 122 today over src and scripts without tests (Ruff 0.16.2).
  • Linked issue(s): supersedes ci(lint): enforce architecture contracts with import-linter #1813, ci: PR behaviour gates -- diff coverage 80%, diff budget, PR template check #1814, docs(review-pr): add the A rule family for code shape, and the AGENTS.md principles it rests on #1815

  • Tests: added/updated? commands run?

    • scripts/tests/test_code_metrics.py, new scripts/tests/test_code_metrics_checks.py and new scripts/tests/test_code_metrics_design.py (285 tests; the design file asserts which check and line each case refuses, and that every exemption holds): each asserts which refusal fires (check name and message), including touched-file scope, out-of-date entries in touched files, the ceiling's head-versus-base comparison, the label read from the API (a forged event payload changes nothing), waived rows listed in full in the log, PR title/body/commit CJK, function-length counting, module-length warnings, and the workflow wiring. 52 guards mutated one at a time; each mutation turned its tests red for the stated reason, then restored. The design checks: 27 more guards mutated (each exemption, the added-line and new-unit rules, the wiring, the config refusal, the binary-diff decode), each turning its own named test red, restored from the commit.
    • pytest -n 4 scripts/tests: 544 passed (before the design checks). With them: 659 passed, 6 skipped. ruff check . and ruff format --check .: clean. PYTHONPATH=src lint-imports --no-cache: 7 kept, 0 broken. actionlint 1.7.12 (with shellcheck) and yamllint 1.38.0 on the changed workflows: clean.
    • Real tree, all pinned tools: python scripts/code_metrics.py --base-ref origin/main passes in 27 s locally (the previous gate took 37 s of a 53 s CI job). End to end on scratch commits in a throwaway worktree: a CC 21 function in a touched file fails (CC 20 passes, the limit is "> 20"); a CC 25 function on main in a file the PR does not touch is informational and the PR passes; a 9-line added comment block fails; a CJK character in a .py fails; a production import of a tests module fails; a literal written a third time fails; with a simulated baseline-raise label a new violation and a baseline growth pass as waived findings, printed in the log. Replays: renaming category_mapping.py and llm_attribution.py, adding a row to CATEGORY_TO_KIND, .get("model")/"model" in three times, and the relocation commits of [Relocate] Kernel agent: move the kernel tools to their owner packages and retire HYPERLOOM_KERNEL_AGENT_ROOT #1787 and [Relocate] Coordinator layer: move code to the collaborator that owns it #1794 raise no comment or literal refusal ([Relocate] Kernel agent: move the kernel tools to their owner packages and retire HYPERLOOM_KERNEL_AGENT_ROOT #1787 keeps one true positive: a third identical package string).
  • Size/complexity triggers crossed: none; the gate measures its own scripts and they have no baseline entries.

  • If this simplifies or refactors: the nested-JSON baseline becomes one line per entry; the maintainability index and the pylint count dimensions, with radon, are removed.

  • Observable effect: every PR gets the code-metrics check and one sticky report comment. A PR fails it when a file it touches gains a unit over a limit or makes a recorded one worse, leaves a fixed unit in the baseline, grows the baseline without the baseline-raise label, adds a long comment block or history in a comment, adds CJK text (files, PR title, body or commits), imports test code from production code, or adds a third copy of a literal, or adds code that reaches into another package's private names, branches on a flag argument, returns a 3+-tuple, compares a raw enum value, writes a global, only forwards its parameters, or declares a single-implementation abstraction. Violations in files the PR does not touch no longer fail it.

  • Breaking changes: no (CI only; the checks are not required until an admin marks them)

  • PR addresses single concern: no: it bundles the four CI-gate PRs on request, so they land together and the review covers one gate set.

  • Root cause is upstream: n/a

🤖 Generated with Claude Code

@zoroyihan7
zoroyihan7 requested a review from a team as a code owner October 10, 2026 12:20
@zoroyihan7 zoroyihan7 added type:feature New capability or improvement github_actions Pull requests that update GitHub Actions code labels Oct 10, 2026
@github-actions

github-actions Bot commented Oct 10, 2026 •

Copy link
Copy Markdown

Code metrics gate: PASSED

Judged: the 150 files this change touches. Files it does not touch never fail it; anything found there is listed under Outside this change for information.

Dimension Limit New Worse Baseline out of date Baseline entries vs base
Cyclomatic complexity fails > 20 0 0 0 122 n/a
Cognitive complexity fails > 30 0 0 0 326 n/a
Function length (lines) fails > 80 0 0 0 477 n/a
Nested block depth fails > 5 0 0 0 29 n/a
Module length (lines) warns > 800, fails > 1200 0 0 0 55 n/a
Duplicated lines per file fails > 0 0 0 0 123 n/a
Dead code (vulture) fails > 0 0 0 0 2 n/a
Check Refusals
Comments 0
English only 0
Production code imports test code 0
Repeated literals 0
A1 Private name of another package 0
A8 Flag argument 0
A9 Tuple return 0
A10 Raw vocabulary comparison 0
A11 Hidden global write 0
A12 Pure forwarder 0
A12 Single-implementation abstraction 0

Module length warning: over 800 lines, within the 1200-line failure limit (not failing): 8

Module Lines Limit
scripts/code_metrics_design.py 840 warns > 800
src/hyperloom/inference_optimizer/trace/langfuse_emitter.py 1200 warns > 800
src/hyperloom/orchestrator/actions/executors/_ray_serving.py 1003 warns > 800
src/hyperloom/orchestrator/actions/executors/_server_patcher.py 1184 warns > 800
src/hyperloom/orchestrator/kernel/attempt_summary.py 1013 warns > 800
src/hyperloom/orchestrator/specialists/dispatch.py 996 warns > 800
src/hyperloom_kb/http_service.py 824 warns > 800
src/kernelforge/agent_backends/codex.py 871 warns > 800

Already over at the base (backlog, not this change's): 4

Unit Dimension Base Now
src/hyperloom/orchestrator/actions/executors/_workload_envs.py:1339 materialize_config_with_envs Cognitive complexity 291 291
src/hyperloom/orchestrator/actions/executors/_workload_envs.py:1339 materialize_config_with_envs Cyclomatic complexity 111 111
src/hyperloom/orchestrator/actions/executors/_workload_envs.py:1339 materialize_config_with_envs Function length (lines) 929 929
src/hyperloom/orchestrator/actions/executors/_workload_envs.py:1 Module length (lines) 2273 2273
8 findings in files this change does not touch

Outside this change (informational): 8

Unit Dimension Baseline Now
src/hyperloom/inference_optimizer/cli/__init__.py Module length (lines) 2569 2635
src/hyperloom/orchestrator/loop/intent_router.py Module length (lines) 1548 1556
src/hyperloom/orchestrator/phases/machine_state.py Module length (lines) 2275 2289
src/hyperloom/orchestrator/phases/machine.py MachinePhase.advance_phase_if_needed Cognitive complexity 32 within the limit, or removed
src/hyperloom/inference_optimizer/cli/__init__.py _run_optimize Cognitive complexity 238 234
src/hyperloom/inference_optimizer/cli/__init__.py _run_optimize Cyclomatic complexity 99 97
src/hyperloom/inference_optimizer/cli/__init__.py _run_optimize Function length (lines) 940 939
src/hyperloom/orchestrator/phases/machine.py MachinePhase.advance_phase_if_needed Function length (lines) 193 192

HEAD^1 has no code-metrics config yet; growth and loosening checks start once it does.

Thresholds and sources: [tool.hyperloom.code_metrics] in pyproject.toml. Tools: complexipy 8.0.1, jscpd 5.4.1, ruff 0.16.2, vulture 2.16.

@zoroyihan7
zoroyihan7 force-pushed the ci/code-metrics-gate branch from 5048fb4 to bf0caf7 Compare October 10, 2026 12:23
Comment thread scripts/tests/test_code_metrics.py Fixed
Comment thread scripts/tests/test_code_metrics.py Fixed
zoroyihan7 and others added 5 commits October 10, 2026 12:39
Adds the `code-metrics` CI job. It measures cyclomatic and cognitive
complexity, statements/branches/returns/arguments/locals/nesting per
function, public methods per class, module length, maintainability index,
duplication and dead code at each tool's industry-default threshold, over
src/hyperloom, src/kernelforge and scripts.

Units already over a threshold are recorded at their current value in
scripts/code_metrics_baseline.json and may only go down: a new violation
fails, a worse baselined unit fails, an improved or removed one fails until
`python scripts/code_metrics.py --update-baseline` tightens the baseline, and
the baseline and config may not grow or loosen relative to the base branch.

The report goes to the job summary and to one sticky PR comment; fork PRs
are commented on by a separate workflow_run workflow that treats the
uploaded report as data.

Co-Authored-By: Claude <noreply@anthropic.com>
A move was matched on the last name segment and greedily, so a worse
unit could take over the allowance of a same-named one (B.run moving
next to an improved A.run, or two main functions). Match on the
qualified name (file name for module metrics) and only when exactly one
removed and one added key share it.

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Claude <noreply@anthropic.com>
Without a base ref, a push straight to main (an admin bypassing PRs)
could raise a baseline value or loosen a threshold unnoticed.

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Claude <noreply@anthropic.com>
@zoroyihan7
zoroyihan7 force-pushed the ci/code-metrics-gate branch from bf0caf7 to 49ec9c1 Compare October 10, 2026 12:41
zoroyihan7 and others added 5 commits October 10, 2026 13:15
Resolve the docs conflicts with #1811's complexity ceiling: the review
ceiling stays, and the code-metrics gate is described as stricter inside
its scope.

Co-Authored-By: Claude <noreply@anthropic.com>
- roots/exclude must be plain repository paths, git reads them with
  --literal-pathspecs, and loosening is judged by the files each config
  measures in the head tree, so a ':(exclude)' root can no longer drop
  a file from the gate while reading as a widened scope.
- CI runs the base branch's copy of the gate scripts (and comment
  poster) over the change, so editing scripts/code_metrics*.py does not
  change the verdict on the same PR; the report lists edits to the gate
  and its tool pins under "Gate implementation changed".
- vulture and jscpd read copies with '# noqa' and 'jscpd:ignore' markers
  defused, so the documented "suppression comments do not hide a unit"
  holds for dead code and duplication too.
- The scope is all of src (hyperloom_kb was outside it) and scripts; the
  baseline is re-seeded on the merged tree: +45 entries, 39 for
  hyperloom_kb and 6 that '# noqa' had hidden from vulture; no existing
  entry changed.

Co-Authored-By: Claude <noreply@anthropic.com>
Add .importlinter with seven contracts: hyperloom top-level layers,
kernelforge depends on hyperloom.common only, hyperloom.common is a
leaf, agent packages and kernel backends are independent, only the
orchestrator and hyperloom.cli import kernelforge directly, and no
cycles between sibling packages.

Existing violations are recorded as exact ignore_imports lines (56
upward imports of hyperloom.inference_optimizer.cli, 155 cycle
breakers) with unmatched_ignore_imports_alerting = error, so the
baselines can only shrink.

Run it as a hard gate in lint.yml (import-linter==2.15 pinned) and as
a local pre-commit hook; document the rules in the style guide.

Co-Authored-By: Claude <noreply@anthropic.com>
Make G, TD001/TD003-TD007, FIX001/FIX003/FIX004, PGH, RSE, ERA, PIE,
RET, DTZ, LOG, A and C4 hard gates via [tool.ruff.lint] extend-select,
after clearing the current findings:

- safe autofixes (PIE807/PIE790/PIE808, C420, RET501/RET502/RET505,
  RSE102), then reviewed unsafe fixes for C401/C405/C416/C408,
  PIE810 and RET504 (all behaviour-preserving rewrites);
- ERA001: every hit was a prose comment that parses as Python; the
  comments are reworded, no code is deleted;
- PGH003: name the ignored code (import-not-found), as elsewhere;
- LOG014: pass the exception in scope instead of exc_info=True outside
  a handler;
- documented per-file ignores where the old shape is behavioural:
  naive persisted timestamps (DTZ), stdlib-mirroring keyword names (A002),
  root-logger calls (LOG015), Sphinx `copyright` (A001); tests may use
  dict(...) and builtin-named fake kwargs. `aiter` (AMD's library) is
  allowlisted for A.

Also align the pre-commit ruff hook with CI's pinned 0.16.2, drop the
CODEOWNERS entry for /setup.py (removed in #808; root *.py is already
covered by /*.py), and update the style guide's lint section.

I (isort) and UP stay off pending a team decision on landing the mass
rewrite.

Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
- tests-coverage.yml: on pull requests, the Python 3.10 coverage job
  writes coverage.xml and runs diff-cover 9.7.1 against the merge
  commit's base parent with --fail-under=80; the markdown report goes to
  the job summary.
- pr-hygiene.yml (new, always runs, read-only token): diff budget over
  production Python lines (src/, scripts/, tests excluded; warn > 400,
  fail > 1000 unless a writer added the size-exception label), the
  agent-doc CLI reference check moved from docs.yml, and an advisory PR
  template completeness check that reads the body only as data.
- scripts/pr_hygiene.py + tests pinning each refusal message.

Co-Authored-By: Claude <noreply@anthropic.com>
zoroyihan7 and others added 3 commits October 10, 2026 13:37
A pure rename has zero numstat lines, so moving a 1001-line test file to
a production path passed the diff budget at 0 lines.

Co-Authored-By: Claude <noreply@anthropic.com>
… text diffs

check_cli_references.py prints PR-controlled file paths; a crafted path
could start a workflow command. A PR's .gitattributes marking *.py as
-diff hid its lines from both the budget and diff-cover; .git/info/
attributes takes precedence over it.

Co-Authored-By: Claude <noreply@anthropic.com>
zoroyihan7 and others added 5 commits October 10, 2026 14:05
…result

- Dead code: vulture reads the real files again, so a `# noqa: F401`
  side-effect import or re-export (Ruff's sanctioned form, with RUF100
  flagging unused ones) is not forced into an importlib rewrite. jscpd
  still reads copies with `jscpd:ignore` defused. Baseline tightened with
  --update-baseline: 3412 -> 3406 entries (6 dead-code entries removed,
  none added or raised).
- Fork poster: the comment opens with the conclusion, head SHA and run link
  taken from github.event.workflow_run, above the artifact text, so a
  fork-uploaded "PASSED" cannot misstate the result. Cancelled/skipped runs
  are not posted, a run with no artifact is a notice, and one poster runs
  per fork branch.
- A refused comment (403, 5xx, rate limit) or artifact upload is a warning;
  the job result comes from the gate step alone and a gate crash stays red.

Co-Authored-By: Claude <noreply@anthropic.com>
…seline

Rework the code-metrics gate:

- Dimensions: cyclomatic complexity > 20 (the style guide's complexity
  ceiling, now measured), cognitive complexity > 30, function length > 80
  physical lines, nested blocks > 5, module length > 1200 lines (> 800 is a
  report-only warning), duplication and dead code unchanged. Drop the
  maintainability index and the pylint count dimensions.
- Scope: only files the change touches are judged; findings elsewhere are
  informational. A unit the base already had at that value or worse is
  backlog, as the complexity ceiling defines it. Baseline growth stays a
  whole-file check.
- The baseline-raise PR label, read from the API at run time, waives new,
  worse and growth findings, which stay listed in the log and the report.
- New checks without a baseline: added comment blocks over 8 lines and
  PR/issue/incident history in added comments, CJK characters in tracked
  files and in the PR's title, body and commits, production imports of
  test code, and added repeated literals.
- Baseline is one sorted line per entry in scripts/code_metrics_baseline.txt.
- Fix the three files that carried CJK characters, and move the InferenceX
  anchor-contract helpers the refresh script imported from a test module
  into a production module both use.

Co-Authored-By: Claude <noreply@anthropic.com>
zoroyihan7 and others added 4 commits October 10, 2026 17:51
Swapping an entry to a same-named new unit while the recorded one is still
over the limit was matched as a move, so the baseline stayed the same size
while the tree gained a second violation.

Co-Authored-By: Claude <noreply@anthropic.com>
…seline line

The baseline path comes from the PR's pyproject.toml and the job holds a
token; a baseline set to /proc/self/environ (or a symlink to it) made the
parse error quote the environment into the PR comment. The baseline must be
a plain repository path and a regular file, pyproject.toml must not be a
symlink, measured files skip symlinks, and a malformed baseline line is
reported by number only.

Co-Authored-By: Claude <noreply@anthropic.com>
…Python

A PR could replace code-metrics/ with a symlink into the runner temp
directory and have upload-artifact follow it, uploading the persisted git
credential. The report and the PR number now live under RUNNER_TEMP. A
symlinked .py in scope was skipped, so its code escaped measurement; it is
now refused (exit 2) without being read.

Co-Authored-By: Claude <noreply@anthropic.com>
@zoroyihan7 zoroyihan7 changed the title ci(metrics): gate white-box code metrics with a ratcheting baseline and a PR report ci: code-metrics gate on touched files with baseline-free checks; fold in import-linter, PR hygiene and ruff principle rules Oct 10, 2026
@zoroyihan7 zoroyihan7 added the size-exception Waives the PR diff budget; added by a maintainer with a reason in the PR description label Oct 10, 2026
Comment thread scripts/code_metrics_report.py Fixed
Comment thread scripts/tests/test_code_metrics.py Fixed
Comment thread scripts/code_metrics_report.py Fixed
Comment thread scripts/tests/test_code_metrics.py Fixed
…ract

The helpers moved out of the test module into a production module, where
diff coverage now grades them; pin build_record's record and each refusal,
and fetch_pinned_file's fallbacks.

Co-Authored-By: Claude <noreply@anthropic.com>
zoroyihan7 and others added 2 commits October 10, 2026 19:05
The module-length rule warns above module-lines-warning and fails above
module-lines, but the report printed only "fails > 1200" in the dimension
table and in the Limit column of the warning section, so the warnings
read as failures. The dimension row now reads "warns > 800, fails > 1200",
the warning section title names both lines, and its Limit column shows the
warning threshold; both numbers come from the config.

Co-Authored-By: Claude <noreply@anthropic.com>
… and concatenation

The module-length warning applies to lengths over module-lines-warning
and at most module-lines, so a warning value at or above the failing
length can never warn; the dimension row then states only the failing
level. Wrap the multi-line how-to-fix bullets in parentheses so the
string concatenation is explicit, and import each module one way per
test file (code_metrics_collect, _inferencex_anchor_contract), as the
code-quality scan asked.

Co-Authored-By: Claude <noreply@anthropic.com>
zoroyihan7 added a commit that referenced this pull request Oct 11, 2026
A1-A14 were replayed on 30 recent merged PRs and every fire was judged
by an independent adversarial judge. Severity now follows the measured
precision: a rule or delimited sub-case is blocking at precision >= 0.8
on at least two real findings with an AGENTS.md sentence behind it.

- A1: demoted to advisory as a whole (5/10); the added non-test import
  sub-case stays blocking (5/5). Tests and unchanged lines are carved out.
- A2 stays blocking (2/2); A8(a) keeps blocking (no fire).
- A12 pure forwarder (2/2) and A13 redundant re-check (2/2) promoted.
- A4, A5, A9, A10, A11 get Not-a-finding clauses for each replayed
  false positive; every Severity line records its count.
- Rules whose shape the code-metrics job (#1812) mechanizes say so.

Co-Authored-By: Claude <noreply@anthropic.com>
zoroyihan7 added a commit that referenced this pull request Oct 11, 2026
…lause to V1

The code-metrics job in #1812 (head 0962cff) checks none of the A
shapes, so the Mechanized notes cited a gate that does not exist. The
A11 clause no longer exempts an in-place resolve_* by convention; it
covers only a write that moved unchanged from the parent (V1).

Co-Authored-By: Claude <noreply@anthropic.com>
zoroyihan7 and others added 8 commits October 11, 2026 07:21
New scripts/code_metrics_design.py, run by the gate next to the comment and
literal checks. Each judges only what the change adds (a moved, re-indented or
renamed line is not added; a function or class is new only when the base file
had no unit of that qualified name), skips test code, and has its own report
section, refusal message and how-to-fix line:

- A1 private-name reach: an added import or attribute access of a private name
  owned by another package unit; the units are the root packages, layers and
  independent modules of .importlinter.
- A8 flag argument: a new function whose first statement branches on a bool or
  two-value Literal parameter into two bodies (CLI entry points exempt).
- A9 tuple return: a new function returning a tuple of more than
  tuple-return-max-elements (2) positions, by annotation or by display.
- A10 raw vocabulary: an added comparison with the string value of a member of
  exactly one Enum under vocabulary-roots, in a module that names that enum.
- A11 hidden global write: an added global statement.
- A12 forwarder and single implementation: a new function that only passes its
  parameters to another call; a new ABC or Protocol with one implementation in
  its own package unit.

Config: import-contracts, vocabulary-roots and tuple-return-max-elements in
[tool.hyperloom.code_metrics]; the last may not be raised and a vocabulary root
may not be dropped. The gate's own tree now passes them: the anchor-contract
helper module and its repository default become public, Finding.key returns a
Key NamedTuple, and the violates() forwarder is gone. git output is decoded
with replacement, so a binary file in the diff no longer stops the checks.

Co-Authored-By: Claude <noreply@anthropic.com>
…d a missing contracts file

Co-Authored-By: Claude <noreply@anthropic.com>
…in the design checks

Co-Authored-By: Claude <noreply@anthropic.com>
Co-Authored-By: Claude <noreply@anthropic.com>
A8 reads a flag comparison with the constant on the left. A9 sees a
triple inside Optional/Union and leaves dunders, whose shape the
protocol fixes. A10 reads case patterns. A11 also catches a store into
globals(). A12 matches builtins only by bare name (a method named like
one is still a forwarder), and passes a method of a constant, a
functools.wraps wrapper and a method overriding its base class.

Co-Authored-By: Claude <noreply@anthropic.com>
…n the design checks

A9 no longer counts a returned display when the annotation admits any
length (tuple[int, ...]). A1 no longer resolves an attribute through an
import its enclosing function shadows with a parameter or an assignment.

Co-Authored-By: Claude <noreply@anthropic.com>
… their own shapes

A1 reads the package units from the merge base's contracts file, and
moving import-contracts is reported as loosening, so a change cannot
redraw the units it is judged by. A10 judges a comparison only when the
other side reads as the member (its .value, or a name whose words end
like the enum's): the code converting raw input to the member has to
compare the string. One enum defined in two modules is one vocabulary.
A12 leaves decorated functions (a cache, a registration) alone.

Co-Authored-By: Claude <noreply@anthropic.com>
Against the current main (#1793, #1805) apply_agentx_switch is 104 lines
and loop/coordinator.py 1429; the gate reports both entries out of date.

Co-Authored-By: Claude <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

github_actions Pull requests that update GitHub Actions code size-exception Waives the PR diff budget; added by a maintainer with a reason in the PR description type:feature New capability or improvement

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants