Skip to content

fix(cli): export every --extra-env pin so all readers resolve it alike - #1797

Merged
jiaqiang-dot-liu merged 15 commits into
mainfrom
feature/xiaofei/agentx-extra-env-spec
Oct 10, 2026
Merged

jiaqiang-dot-liu merged 15 commits into
mainfrom
feature/xiaofei/agentx-extra-env-spec

Conversation

@xiaofei-zheng

@xiaofei-zheng xiaofei-zheng commented Oct 9, 2026 •

Copy link
Copy Markdown
Collaborator
  • Description: The CLI serialized the operator's --extra-env pins into INFERENCE_OPTIMIZER_EXTRA_ENV alone and never exported the individual names. Every Hyperloom control variable is read with a bare os.environ.get, so none of them could see a pin, and AgentX compensated by re-merging the blob inside agentx_env_for_conc — one reader resolving a knob differently from all the others. A pinned HYPERLOOM_AGENTIC_BACKEND=mlperf therefore selected mlperf_agentic_client.sh while the seeded grading axis, state.agentx_backend and resolve_benchmark_timeouts() still said aiperf: the MLPerf client ran while the session graded it on an interactivity series MLPerf does not publish, so every round REVERTed.

    The fix is at the producer. _export_operator_launch_shape exports every pin under its own name, withholding none, so a pin and an operator's own export of the same name become the same thing. The local merge layer is gone and the three copies of the blob parser collapse into one operator_extra_env() in env_safety; the blob itself stays, because it is what reaches state.json and what names the previous launch's pins so a resume that drops one can unset it.

    Knobs Hyperloom resolves for itself then take the environment as one rung of their own ladder — explicit flag, then $NAME, then the resumed session's value, then the default — which is the shape --max-model-len / $MAX_MODEL_LEN already had in _resolve_run_max_model_len_inner. args becomes the single answer and the projection writes it back, so there is no second writer to collide with. _resolve_precision splits off from _resolve_workload_knobs because it reads the checkpoint and cannot run until the model is resolved, while the five numeric knobs must be settled earlier: _preflight's check_gpu_visibility compares $TP against the visible GPU count. The resulting fresh-branch order is launch-shape export, ladder, projection, topology gates, _preflight; the gates become _enforce_topology_gates so the same two checks see the resolved shape. On the resume branch the pins are restored ahead of agentx_state_is_stale, which resolves the AgentX backend from the environment and would otherwise refuse a session whose identity came from a pin.

  • Linked issue(s): none

  • Tests: test_cli_resume_launch_shape.py adds twelve cases, each failing on the commit before the one that fixes it. A pin reaching the environment under its own name; agentx_client_script and seed_grading agreeing once it does; a pin and an export of the same name resolving identically; the ladder's four rungs including a non-positive value falling through; a full _run_optimize resume of a session seeded with a pinned backend, which exits 2 before the guard fix; a resume dropping a pin unsetting it; the MAX_MODEL_LEN rungs; and a table of ordering anchors asserted against _run_optimize's source — launch-shape export before the ladder, ladder before the projection, projection before _preflight, ladder before the gates. The anchor table is the one that pins the ordering defects: the behavioural case beside it passes on the broken revisions because it calls the functions in the order it wants. test_agentx_switch.py's _pin_extra_env now pins the way the CLI does. Every test file touching INFERENCE_OPTIMIZER_EXTRA_ENV / agentx_env_for_conc / the pin parsers runs; remaining failures reproduce on the merge base and are pre-existing Windows-platform ones.

  • Size/complexity triggers crossed: none

  • If this simplifies or refactors: removes the {**os.environ, **_operator_extra_env()} merge in agentx_env_for_conc, two duplicate blob parsers, and — after a round where the pin was treated as a third source alongside the flag and the environment — the exclusion set, the export filter, the matching exclusion in the unset loop, the pin-reading int helper and a resume-only MAX_MODEL_LEN resolver that existed only to keep those two apart. The contract preserved is that benchmark.envs still carries every pin to the benchmark subprocess, covered by test_agentx_switch.py and test_kernel_integrate_and_report.py's recipe case.

  • Observable effect: a pin means the same thing in every reader and on both branches. Before: --extra-env HYPERLOOM_AGENTIC_BACKEND=mlperf ran the MLPerf client under aiperf grading and could not be resumed; --extra-env ISL=4096 measured 4096 while recording 1024 in the comparability fingerprint and the manifest, then reported 4096 after a resume; --extra-env TP=4 launched at TP=1 while state.json said 4; check_gpu_visibility could not warn about --tp 8 on a 2-GPU box; a resume re-passing MAX_MODEL_LEN was silently ignored.

  • Breaking changes: no. A pin that names a Hyperloom control variable now takes effect where it previously did not, and an explicit flag for the same knob outranks it; the docs state the ladder.

  • PR addresses single concern: yes

  • Root cause is upstream: no

@xiaofei-zheng
xiaofei-zheng requested a review from a team as a code owner October 9, 2026 10:07
@xiaofei-zheng

Copy link
Copy Markdown
Collaborator Author

PR #1797 -- fix(agentx): derive client knobs and workload_spec from --extra-env pins

What it does: the CLI serializes --extra-env pins into the single INFERENCE_OPTIMIZER_EXTRA_ENV JSON var and never exports the individual names, so apply_agentx_switch -- which derived every AgentX client knob from os.environ -- could not see them, while materialize_config_with_envs merged the raw pins into benchmark.envs afterwards and overwrote the values the switch had derived. The result was a published workload_spec that disagreed with the environment the client actually ran under, and a pinned warmup grace that was either passed through unscaled or replaced by the 1800 s default. The PR makes agentx_env_for_conc layer the operator pins over the process environment, routes both the switch's forwarding loop and the later extra-env merge through one _is_agentx_client_key predicate, and has that merge skip the knobs the switch already settled.

Blocking issues: none

Checked: agentx_env_for_conc and its four call sites; _is_agentx_client_key against the inline condition it replaces (same key set, same mlperf gating); apply_agentx_switch forwarding loop, _raw_grace source, and build_agentx_workload_spec's client_knob/served_knob precedence; materialize_config_with_envs ordering -- the filter runs on the operator pins only, before extra_envs is merged, so variant-proposed envs still pass through filter_untrusted_env_mapping; every reader of safe_extra_envs after the filter (envs, _sync_repo_aliases prefer_dir, the reference unset_envs guard); _benchmark_runtime.apply_runtime_benchmark_overrides for the rebuild path; agentx_warmup_grace_sec scaling arithmetic; and a grep for any other non-test os.environ reader of AGENTX_* / AIPERF_BIN / WEKA_LOADER_OVERRIDE (_recipe_script.py:41 reads envs first, so it sees the pin the switch wrote).

Cases run, through the real materialize_config_with_envs, at head and at the merge base:

  • AgentX on, --extra-env AGENTX_CANONICAL_DATASET=...062126 AGENTX_WARMUP_GRACE_PERIOD=600 AGENTX_WARMUP_GRACE_CONC=4 at CONC=8. Base: envs carried the full corpus and the raw 600 while the spec published _256k and 1800. Head: envs and spec agree on the full corpus, and both carry the scaled 1200.
  • AgentX on, no pins: byte-identical at both revisions (_256k, 1800).
  • AgentX off, same pins: byte-identical at both revisions; the pins still land in envs and no workload_spec is emitted. A non-AgentX pin (MY_TUNING_KNOB) survives the filter in every case.

Ran: pytest test_agentx_switch.py at head (27 passed) and against the merge base, where both added cases fail with the exact values above ('1800' != '1200', spec corpus _256k).

Base: 1c14e79 | Head: 7546c7c

LGTM

@jiaqiang-dot-liu

Copy link
Copy Markdown
Collaborator

PR #1797 -- fix(agentx): derive client knobs and workload_spec from --extra-env pins

What it does: the CLI serializes --extra-env into the single INFERENCE_OPTIMIZER_EXTRA_ENV JSON var and never exports the individual names, so apply_agentx_switch -- which derived every AgentX client knob from os.environ -- could not see an operator pin, while materialize_config_with_envs merged the raw pin into benchmark.envs afterwards. The published workload_spec therefore disagreed with the environment the client actually ran under, and a pinned warmup grace was either passed through unscaled or replaced by the 1800 s default. The PR has agentx_env_for_conc layer the operator pins over the process environment, routes the switch's forwarding loop and the later extra-env merge through one _is_agentx_client_key predicate, and has that merge skip the knobs the switch already settled.

Blocking issues: 1

  1. [D8] A pinned HYPERLOOM_AGENTIC_BACKEND now swaps the whole workload, while every session-level reader of it still says aiperf [verified]
    Problem: agentx_env_for_conc makes --extra-env outrank os.environ, and HYPERLOOM_AGENTIC_BACKEND is BACKEND_ENV, so agentx_client_script(_agentx_env) at _workload_envs.py:541 now dispatches on the pin. Nothing rejects the pin on the way in -- parse_operator_extra_env (cli/bootstrap.py:42) does a bare NAME=VALUE split with no key validation, _export_operator_launch_shape (cli/__init__.py:1204) writes only the JSON blob and never exports the individual name, and _preflight_agentx_backend (cli/__init__.py:977) does not look at this variable. Meanwhile every session-level reader of it still reads bare os.environ. src/hyperloom/orchestrator/actions/executors/_workload_envs.py:247
    Impact: with HYPERLOOM_AGENTX=1 and --extra-env HYPERLOOM_AGENTIC_BACKEND=mlperf, head materializes benchmark_script: mlperf_agentic_client.sh, workload_spec.kind: agentx_mlperf_agentic, PORT: 30000 (base: aiperf_client.sh, agentx_trace_replay, no PORT), but agentic_backend() at common/agentx_workload.py:49 returns aiperf in both. So the MLPerf client runs while: intvty_grading_enabled() (common/perf_metric.py:105) returns True and graded_metric_key() returns e2e_norm_intvty_p90, an interactivity series the MLPerf harness does not publish; _accuracy_gate.py:429 gates the inline scores.json path on the same bare is_mlperf_backend() and skips it, so every KEEP is decided without the harness score; resolve_benchmark_timeouts() (_subprocess_kill.py:291, all ten production callers pass no env) hands back the 7800 s default instead of the trajectory-sized MLPerf cap; and cli/bootstrap.py:340 persists state.agentx_backend = "aiperf", which is what agentx_state_is_stale() compares a later resume against -- so the guard that exists to stop two measurement sets sharing one ledger cannot see the mismatch.
    Action: resolve BACKEND_ENV from one place. Either exclude it from the keys agentx_env_for_conc lets the pins override and reject such a pin in parse_operator_extra_env with a message telling the operator to export it, or point agentic_backend() and its seed/grading/timeout callers at the same resolved mapping.

Note on severity, stated because it is the one thing that weakens this: docs/reference/environment-variables.md:948 documents HYPERLOOM_AGENTIC_BACKEND as a process env var, not an --extra-env knob, and the merge base was already half-inconsistent here (it shipped the pin into benchmark.envs while still running aiperf_client.sh). This PR promotes that half-honoured pin into a full workload swap with no corresponding move in the session identity. Nothing rejects the pin at any layer, and this PR's own premise is that --extra-env pins are the operator's environment, so an operator has every reason to set it this way.

Checked: agentx_env_for_conc and both call sites; _is_agentx_client_key against the inline condition it replaces (same prefixes, same tuple, same mlperf gating -- the two ifs becoming one or cannot change the written value); the apply_agentx_switch forwarding loop, _raw_grace source and build_agentx_workload_spec client_knob/served_knob precedence; materialize_config_with_envs ordering -- the strip runs on the operator pins only, before extra_envs is merged, so variant-proposed envs still pass filter_untrusted_env_mapping (probed: a variant AGENTX_KEEP_ALIVE_ENV is still dropped at head); _agentx_timeouts.agentx_warmup_grace_sec arithmetic (600*8//4 = 1200); _benchmark_runtime.apply_runtime_benchmark_overrides for the rebuild path; the five materialize_config_with_envs callers and which of them pass a state-derived agentx_mode; every non-test reader of BACKEND_ENV / agentic_backend; BLOCKED_VARIANT_ENV_NAMES against the forwarded key set.

Ran: pytest src/hyperloom/inference_optimizer/tests/test_agentx_switch.py at head (27 passed) and with the same file against the merge base (2 failed, 25 passed -- '1800' != '1200' and spec corpus _256k, the exact values the description claims). Cases run through the real materialize_config_with_envs at head and at the merge base: AgentX on with corpus + grace pins at CONC 8; AgentX off with the same pins; --extra-env HYPERLOOM_AGENTIC_BACKEND=mlperf; --extra-env AGENTX_KEEP_ALIVE_ENV=LD_PRELOAD on both the operator and the variant channel; a pinned aiperf benchmark_script with the switch inactive.

Base: 1c14e79 | Head: 7546c7c

The CLI serialized the operator's pins into INFERENCE_OPTIMIZER_EXTRA_ENV
alone, so a bare os.environ.get could not see them and the AgentX switch had
to re-merge the blob locally. That split let one reader resolve a knob
differently from the rest: a pinned HYPERLOOM_AGENTIC_BACKEND selected the
MLPerf client while the seeded grading axis, the persisted backend identity
and the benchmark timeout all still read aiperf.

Pins now become real environment variables, applied after the flag-derived
workload knobs so a pinned TP/CONC/EP wins, and dropped on a resume that no
longer passes them. The blob stays for the state roundtrip and to name what
the previous launch exported. Nothing is withheld from the export: an
operator who can pass --extra-env can equally export the same name.
@xiaofei-zheng xiaofei-zheng changed the title fix(agentx): derive client knobs and workload_spec from --extra-env pins fix(cli): export each --extra-env pin so every reader resolves it alike Oct 10, 2026
@xiaofei-zheng

Copy link
Copy Markdown
Collaborator Author

Response to the blocking finding on HYPERLOOM_AGENTIC_BACKEND

The finding holds and is fixed, but at the root rather than at the symptom.

The two remedies it offered were (a) exclude BACKEND_ENV from what the pins may override and reject such a pin, or (b) point agentic_backend() and its callers at one resolved mapping. Both treat BACKEND_ENV as a special case. The actual root cause is one layer down: the CLI serialized pins into INFERENCE_OPTIMIZER_EXTRA_ENV alone and never exported the individual names, so every bare os.environ.get was blind to them and agentx_env_for_conc had to re-merge the blob locally. BACKEND_ENV was simply the first knob where that split became observable; any other Hyperloom control variable a pin names had the same shape.

So the pins now become real environment variables in _export_operator_launch_shape, and the local merge in agentx_env_for_conc is gone. With one resolution point there is nothing left to keep in sync: agentic_backend(), seed_grading(), state.agentx_backend, resolve_benchmark_timeouts() and the accuracy gate all read the same value without any of them changing. No exception list — an operator who can pass --extra-env can equally export the same name, so withholding a name would buy no safety and would reintroduce the split this removes.

Title and description have been rewritten to match; the change is no longer AgentX-scoped.

Author self-review of the new commit

Blocking issues: none.

Checked: _export_operator_launch_shape and both call sites in _run_optimize — 1652 (after _export_workload_envs_for_optimize, so a pinned TP/CONC/EP outranks the flag-derived value) and 1910 on the resume branch, confirming the clear-then-restore order leaves the right final state and is idempotent when --extra-env is re-passed; the unset loop against the case where a name is both an inherited variable and a previous pin — reaching it needs the internal handoff blob exported by hand, and grep found no path that re-enters optimize as a subprocess (rebench is a separate entry point), so no inherited-blob window exists; redact_env in trace/langfuse_mapping.py — a separately exported pin is matched by BENCHMARK_SECRET_ENV_NAMES, is_secret_shaped_env_name and every SENSITIVE_ENV_NAME_MARKERS substring, which is strictly stricter than the _is_credential_env_name path blob entries take (TOKEN there is segment-matched), so the export widens no trace exposure; filter_untrusted_env_mapping in materialize_config_with_envs — pins it drops now reach the benchmark subprocess through build_benchmark_env's os.environ.copy() base, with scrub_benchmark_process_env still removing BLOCKED_CHILD_ENV_NAMES, and the recipe_overwritten entries are names the recipe re-exports unconditionally so the subprocess value is unchanged; the consolidated operator_extra_env() against both parsers it replaces; _operator_extra_env and _operator_pinned_envs still have live call sites (_workload_envs.py:871, :1908, apply_aiter_moe_pin_filter), so neither is left as a writer without a reader.

Ran: a differential harness driving operator_extra_env() and verbatim copies of both replaced parsers over 15 blobs built through the real parse_operator_extra_env — spaced keys, a key containing a space, empty value, blank name, no =, plus malformed, non-object and null raw values — all three agree on every case, 0 differences, and the unparseable and non-object branches log rather than swallow. pytest test_cli_resume_launch_shape.py test_agentx_switch.py (38 passed); the three added cases each fail on the parent commit. All 13 test files touching INFERENCE_OPTIMIZER_EXTRA_ENV / agentx_env_for_conc / the pin parsers run locally; every failure was reproduced on the parent commit and is a pre-existing Windows-platform failure (symlink privileges, POSIX paths), none new. CI at this head: 30 checks pass, 0 fail, all 12 test shards green on Linux.

SKIPPED: independent second reader — the agent assigned to it terminated on an API gateway error, so that axis was not independently checked; the attack list given to it was worked through directly instead, and this review is the author's own.

Base: 1c14e79 | Head: 63c66e7

@jiaqiang-dot-liu

Copy link
Copy Markdown
Collaborator

PR #1797 (re-review at 63c66e7) -- fix(cli): export each --extra-env pin so every reader resolves it alike

What it does: the CLI serialized --extra-env into INFERENCE_OPTIMIZER_EXTRA_ENV alone and never exported the individual names, so every Hyperloom control variable -- all of which are read with a bare os.environ.get -- was blind to a pin. _export_operator_launch_shape now exports each pin under its own name, applied after the flag-derived workload knobs so a pinned TP/CONC/EP wins, and unsets a name the previous launch exported but this one no longer passes. The local {**os.environ, **_operator_extra_env()} merge in agentx_env_for_conc is gone, and the three copies of the blob parser collapse into one operator_extra_env() in env_safety.

The previous head's finding is resolved at the root, verified: with HYPERLOOM_AGENTIC_BACKEND=mlperf pinned, agentic_backend(), agentx_client_script(), intvty_grading_enabled(), graded_metric_key() and resolve_benchmark_timeouts() now all agree (mlperf / mlperf_agentic_client.sh / False / output_throughput / 9600.0), against aiperf / aiperf_client.sh / True / e2e_norm_intvty_p90 / 7800.0 at the previous head. Removing the merge layer rather than special-casing BACKEND_ENV is the right call.

Blocking issues: 1

  1. [free:resume-clears-pins-before-the-guard] A session whose AgentX identity came from a --extra-env pin cannot be resumed [verified]
    Problem: _export_operator_launch_shape is called twice on a resume. The first call at src/hyperloom/inference_optimizer/cli/__init__.py:1652 is unconditional and passes parse_operator_extra_env(args), which is {} on a resume that does not re-pass --extra-env -- so the pinned names are not in os.environ. agentx_state_is_stale(state) then runs at :1809 and reads them with a bare os.environ.get (cli/bootstrap.py:115 for HYPERLOOM_AGENTX, :135 for agentic_backend()). The persisted pins are restored only at :1909-1910, a hundred lines after the guard has already hard-exited.
    Impact: a session launched with --extra-env HYPERLOOM_AGENTIC_BACKEND=mlperf seeds state.agentx_backend = "mlperf" (cli/bootstrap.py:340, now correctly, because the fresh-launch export at :1652 precedes it) and persists the pin in state.operator_extra_env (cli/bootstrap.py:391). Resuming it, with HYPERLOOM_AGENTX exported in the shell as usual, gives had_backend='mlperf' against want_backend='aiperf' and sys.exit(2): "session was measured with agentic backend 'mlperf' but this run is 'aiperf'". Simulated at this head with the function body verbatim: the verdict is "agentic backend 'mlperf' vs this run 'aiperf'" at :1809 and "" once :1910 has run. --extra-env HYPERLOOM_AGENTX=1 fails the same way one check earlier, on the benchmark_mode comparison. This is new: before this PR the pin never reached agentic_backend(), so the seed recorded aiperf and the guard matched. The docs this PR updates promise the opposite -- "On --resume the persisted pins are re-exported, so they only need re-passing when you want to change them".
    Action: restore the persisted pins before the guard reads them -- load state.operator_extra_env and export it ahead of agentx_state_is_stale at :1809, or skip the :1652 call on the resume branch so it does not clear names the resume block is about to restore. A regression test resuming a session seeded with a pinned HYPERLOOM_AGENTIC_BACKEND would pin the discrimination; test_export_operator_launch_shape_unsets_pins_dropped_on_resume covers the unset half but not the ordering against the guard.

Checked: the move of _export_operator_launch_shape from :1627 to :1652 and everything between the two positions (only _export_workload_envs_for_optimize and the multi-node exports -- nothing there reads the pins, so the TP/CONC/EP precedence claim holds); both call sites on the resume branch and every reader between them (_preflight_agentx_backend:977 and _apply_agentx_budget_profile:1009 gate on _agentx_enabled() but read no backend; the resolve_benchmark_timeouts() at :1743 discards its return, so only agentx_state_is_stale is consequential in that window); the new operator_extra_env() in env_safety against the three parsers it replaces (same empty/unparseable/non-object handling, now with a log line on the non-object branch that _workload_envs lacked and _grid_variant_filter swallowed); agentx_env_for_conc now returning dict(os.environ) unconditionally -- both callers treat it read-only; the unset loop reading the previous blob before it is rewritten; filter_untrusted_env_mapping on the variant channel, still unaffected (a variant-proposed AGENTX_KEEP_ALIVE_ENV is still dropped).

Ran: pytest src/hyperloom/inference_optimizer/tests/test_agentx_switch.py at this head -- 27 passed. test_cli_resume_launch_shape.py cannot be collected on Windows (cli/kb.py:9 imports fcntl), pre-existing and unrelated; its three added cases were read rather than run, and CI covers them. Reader agreement under a pinned backend compared at this head and at 7546c7c. The resume ordering was simulated with the _export_operator_launch_shape body copied verbatim from :1205.

Base: 1c14e79 | Head: 63c66e7 | CI: all 37 checks success at this head

@jiaqiang-dot-liu

Copy link
Copy Markdown
Collaborator

PR #1797 -- fix(cli): export each --extra-env pin so every reader resolves it alike

What it does: the CLI serialized the operator's --extra-env pins into the single INFERENCE_OPTIMIZER_EXTRA_ENV JSON blob and never exported the individual names, so every Hyperloom control variable -- all of which are read with a bare os.environ.get -- was blind to them, and AgentX had to re-merge the blob locally in agentx_env_for_conc. That local layer made one reader resolve a knob differently from every other: a pinned HYPERLOOM_AGENTIC_BACKEND=mlperf selected the MLPerf client while the seeded grading axis, state.agentx_backend and resolve_benchmark_timeouts() all still said aiperf. The PR exports each pin as a real environment variable in _export_operator_launch_shape, moves that call after the flag-derived workload knobs so a pinned TP/CONC/EP wins, unsets a pin a resume no longer passes, deletes the local merge layer, and collapses the three copies of the blob parser into one operator_extra_env() in env_safety.

Blocking issues: 1

  1. [free:pin-order-fresh-vs-resume] The fresh-launch branch overwrites a pinned ISL/OSL/MAX_MODEL_LEN/PRECISION/SKIP_VARIANTS/MODEL_PATH/FRAMEWORK/GPU_TYPE; the resume branch of the same session honours it [verified]
    Problem: the pin export moves to src/hyperloom/inference_optimizer/cli/__init__.py:1652, which on the fresh branch is followed by eight unconditional writes of names a pin may carry -- SKIP_VARIANTS (:1658), MODEL_PATH (:2088), FRAMEWORK (:2100), GPU_TYPE (:2137), MAX_MODEL_LEN (:2157), ISL (:2158), OSL (:2159), PRECISION (:2164) -- and _resolve_workload_knobs (:1156) fills those values from flag > state > default without reading os.environ, so the pin cannot come back through it. On the resume branch the same names are written at :1839-:1902 and the export sits after them at :1910, so there the pin wins. state.operator_extra_env is persisted at cli/bootstrap.py:391 and restored at cli/__init__.py:1909, so a bare --resume-from with no --extra-env is enough to flip the answer. src/hyperloom/inference_optimizer/cli/init.py:2158
    Impact: --extra-env ISL=4096 launches a session whose process environment says ISL=1024 while benchmark.envs says 4096, then resumes it with both at 4096. The in-process readers record the difference: canonical_fingerprint.py:118-119 builds the comparability fingerprint from os.environ["ISL"]/["OSL"], session/manifest.py:245-246 writes the manifest workload from the same, cli/__init__.py:1057 derives MAX_MODEL_LEN from os.environ["MAX_MODEL_LEN"], and bypass_runner.py:151-154 prefers os.environ.get("ISL") over bench_envs.get("ISL"). One session, one ledger, two workload identities across the resume -- which is the failure mode this PR exists to close. It also contradicts what the PR writes: docs/how-to/optimize-custom-workload.md:124-126 now says a pin "also sets any Hyperloom control variable of the same name -- there is no exception list", :139-143 that "a pinned TP/CONC/EP outranks --tp and friends" (--isl/--osl are those friends), and the description that "a pin now means the same thing everywhere".
    Action: give the pins one resolution point that dominates both branches -- move the fresh-branch export below :2164 (and SKIP_VARIANTS at :1658), or route those eight writes through a helper that will not overwrite a pinned name -- and add a case that pins ISL through the real fresh-launch ordering rather than calling _export_operator_launch_shape in isolation.

Checked: _export_operator_launch_shape and both call sites (:1652, :1910) with an AST walk of _run_optimize recording the enclosing condition of every os.environ[...] write; _resolve_workload_knobs (:1156-1190) for an env read; parse_operator_extra_env (cli/bootstrap.py:42) for key validation; state.operator_extra_env seed (bootstrap.py:391) and restore (:1909); env_safety.operator_extra_env against both parser bodies it replaces (same strip, coercion, blank-key drop and isinstance gate, plus a warning the _grid_variant_filter copy lacked); _is_agentx_client_key against the inline condition it replaces; agentx_env_for_conc and both call sites; the new extra-env strip in materialize_config_with_envs against apply_agentx_switch's early returns and every non-test writer of aiperf_client.sh / mlperf_agentic_client.sh; BLOCKED_VARIANT_ENV_NAMES / BLOCKED_CHILD_ENV_NAMES / build_benchmark_env for what a now-global pin reaches that the variant filter used to drop; head-tree occurrence counts for every added identifier; non-test bare readers of ISL / OSL / MAX_MODEL_LEN / SKIP_VARIANTS. | Ran: the real _export_operator_launch_shape and the real materialize_config_with_envs over pins {ISL=4096, OSL=512, TP=4} in the exact statement order of each branch, in two worktrees -- base: fresh os.environ 1024/128/1, resume 1024/128/1 (identical); head: fresh 1024/128/4, resume 4096/512/4. benchmark.envs is 4096/512/4 in all four. | Base: 1c14e79 | Head: 63c66e7

The guard resolves the AgentX backend and benchmark mode with a bare
os.environ.get, but the persisted pins were re-exported a hundred lines
later. A session seeded with --extra-env HYPERLOOM_AGENTIC_BACKEND=mlperf
therefore recorded agentx_backend="mlperf" and was then refused on resume
as "measured with agentic backend 'mlperf' but this run is 'aiperf'";
--extra-env HYPERLOOM_AGENTX=1 failed one check earlier on benchmark_mode.
Both became reachable only once the pins started reaching those readers.

The restore moves ahead of the guard; the re-export block below keeps the
state writeback and the operator-facing print.
Which value a pinned ISL/OSL/CONC/TP/EP/PRECISION ended up with was decided
by write order, and the two branches write in different orders: the fresh
branch projects args back into the environment after the pin export, so
--extra-env ISL=4096 ran at the default 1024 while benchmark.envs said 4096;
the resume branch exported later and the pin survived. One session, two
workload identities across a resume.

_resolve_workload_knobs now takes these pins as a rung of its ladder --
flag, pin, resumed state, default -- so args is the single answer and every
later projection writes it. An explicit flag outranks a pin, matching the
ladder _resolve_run_max_model_len_inner already uses. Those names are no
longer exported directly, because feeding both the export and the ladder is
what let a pin beat a flag on one branch and lose to it on the other; the
blob still carries them, which is what the ladder and state.json read.

MAX_MODEL_LEN and FRAMEWORK stay exported: their ladders read the
environment, so the pin has to reach it. SKIP_VARIANTS is not pinnable --
it is the policy for one run, not part of the measurement contract.
@xiaofei-zheng

Copy link
Copy Markdown
Collaborator Author

Both findings confirmed and fixed. Reproduced each through the production entry point before changing anything.

free:resume-clears-pins-before-the-guard (f95df524). Confirmed: a _run_optimize resume of a session seeded with --extra-env HYPERLOOM_AGENTIC_BACKEND=mlperf exits 2 with exactly the message quoted — "session was measured with agentic backend 'mlperf' but this run is 'aiperf'". The restore of state.operator_extra_env now sits directly after SharedState.load_or_init and ahead of agentx_state_is_stale; the block below keeps the state writeback and the operator-facing print. The regression test drives the real resume path rather than the helper in isolation, so it pins the ordering against the guard rather than the unset half.

free:pin-order-fresh-vs-resume (4a48ac4b). Confirmed for ISL/OSL/PRECISION, with one correction: MAX_MODEL_LEN does not belong on the list. _resolve_run_max_model_len_inner:1057 reads os.environ["MAX_MODEL_LEN"] as its second rung, so that pin was already reaching it; the same is true of FRAMEWORK at :2129 (args.framework or os.environ.get("FRAMEWORK", "")). The names that genuinely had no env rung are the ones _resolve_workload_knobs fills.

The fix is not a reorder. Moving the export below the projections only relocates the collision, because the two branches run it at different points and a pin that both exports itself and feeds a ladder has to beat an explicit flag on one branch and lose to it on the other. So ISL/OSL/CONC/TP/EP/PRECISION now enter through _resolve_workload_knobs as a rung — flag > pin > resumed state > default, the ladder _resolve_run_max_model_len_inner already uses — and are excluded from the direct export. args becomes the single answer and every later projection writes it, so where the export sits no longer matters. The blob still carries them: it is what the ladder reads and what state.json persists. The unset loop excludes them on both sides, so it cannot clear a projection it never wrote.

This reverses something the earlier description claimed: a pinned TP/CONC/EP no longer outranks --tp. An explicit flag wins. docs/how-to/optimize-custom-workload.md and the PR description now state the ladder instead of "there is no exception list".

SKIP_VARIANTS is left unpinnable rather than given a rung: it selects the grid policy for one invocation, not part of the session's measurement contract, so a resume should take it from the flag it was passed rather than from what the original launch pinned. Called out in case you read it the other way.

Added cases, each failing on the commit before its fix: a full _run_optimize resume of a session with a pinned backend; the four ladder rungs including a malformed pin falling through to the default rather than raising; a ladder-resolved pin staying out of the export but present in the blob; and the unset loop leaving a projected value alone.

…ze through

_run_optimize writes SKIP_VARIANTS, PD_MODE and INFERENCE_OPTIMIZER_NODES
directly, so monkeypatch had nothing recorded to undo and the next test in
the worker graded on what this one left behind: shard 1/6 failed on
geak_metric_axis and the graded-comparison axes.
@jiaqiang-dot-liu

Copy link
Copy Markdown
Collaborator

PR #1797 -- fix(cli): export each --extra-env pin so every reader resolves it alike

What it does: the CLI serialized the operator's --extra-env pins into the single INFERENCE_OPTIMIZER_EXTRA_ENV JSON blob and never exported the individual names, so every Hyperloom control variable -- all read with a bare os.environ.get -- was blind to them, and AgentX had to re-merge the blob locally in agentx_env_for_conc, which is how a pinned HYPERLOOM_AGENTIC_BACKEND=mlperf could select the MLPerf client while the grading axis, state.agentx_backend and resolve_benchmark_timeouts() all still said aiperf. Pins now become real environment variables; the six knobs Hyperloom resolves for itself are withheld from that export and enter _resolve_workload_knobs as a rung instead (flag > pin > resumed state > default); the resume branch restores the persisted pins ahead of agentx_state_is_stale so a session seeded with a pinned backend is resumable; the local merge layer is gone and the three blob parsers are one operator_extra_env() in env_safety.

Blocking issues: 2

  1. [free:ladder-knobs-not-reprojected-on-fresh] On a fresh launch the ladder's answer for TP/CONC/EP never reaches the environment, because their only projection runs before the ladder [verified]
    Problem: _export_workload_envs_for_optimize is the sole non-test writer of os.environ["TP"|"CONC"|"EP"] in the tree (cli/__init__.py:1221-1223) and is called once, at :1686, from tp_resolved/ep_resolved computed at :1636-1637 from args alone -- about 500 lines before _resolve_workload_knobs fills the pin at :2194. The re-projection that follows the fresh ladder covers MAX_MODEL_LEN/ISL/OSL/PRECISION (:2198-2205) and omits these three; the resume branch projects all of them after its ladder (:1930-1942), so the pin lands there. src/hyperloom/inference_optimizer/cli/init.py:1686
    Impact: --extra-env TP=4 EP=2 CONC=63 with no flags gives, on a fresh launch, args and state.json at 4/2/63 (_seed_shared_state runs at :2284, after the ladder) while os.environ stays at the defaults 1/1/64 -- and the environment is what the server actually launches from. _workload_envs.py:1487/:1213/:567 build the Magpie YAML from it, cli/preflight.py:1513 gates the visible-GPU count on int(os.environ.get("TP","1")), session/manifest.py:220/:247 records conc and tp from the environment only, and canonical_fingerprint.py:117/:121 builds the comparability fingerprint from it. So the session runs TP=1 while its state, its ledger and a later resume of it all say TP=4. The two fail-fast topology gates at :1648/:1660 also compare against the pre-ladder values, so a pinned TP=16 on 8 GPUs passes them. This reverses 63c66e7d9, where the fresh environment did carry 4/2/63, and the merge base was at least self-consistent at 1/1/64 everywhere.
    Action: project TP/CONC/EP from args after _resolve_workload_knobs on the fresh branch -- move the _export_workload_envs_for_optimize call below :2194 and feed it args.tp/args.ep rather than the pre-ladder tp_resolved/ep_resolved, and recompute the topology gates from the resolved values. Add a case that pins TP through the real fresh-branch ordering and asserts os.environ["TP"]; all four ladder cases currently call _resolve_workload_knobs in isolation and use ISL/OSL/CONC only, which is why CI is green.

  2. [free:max-model-len-pin-lost-on-resume] A MAX_MODEL_LEN pin re-passed on a resume to change it is silently discarded [verified]
    Problem: MAX_MODEL_LEN is deliberately not a ladder name, so the :1852 export places the pin in the environment; _resume_max_model_len = getattr(args,"max_model_len",None) or getattr(state,"max_model_len",0) at :1929 never reads the environment, and :1936 writes that value back over the pin. _resolve_run_max_model_len, whose second rung is $MAX_MODEL_LEN (:1057-1060), has exactly one caller, :2196, on the fresh branch; nothing on the resume branch seeds args.max_model_len from the pin. src/hyperloom/inference_optimizer/cli/init.py:1936
    Impact: with state.max_model_len = 32768, a resume passing --extra-env MAX_MODEL_LEN=65536 runs at 32768 with no message -- reproduced as 65536 after :1852 and 32768 after the loop, against (65536, '$MAX_MODEL_LEN') on the fresh branch. docs/how-to/optimize-custom-workload.md tells the operator pins "only need re-passing when you want to change them", and 63c66e7d9 honoured this one because its resume export sat after the loop; moving it up to :1852 for the staleness guard left MAX_MODEL_LEN behind.
    Action: feed $MAX_MODEL_LEN into the resume rung -- make _resume_max_model_len consult operator_extra_env() (or call _resolve_run_max_model_len on both branches) so flag > pin > state holds there too.

Checked: _resolve_workload_knobs and its new pin rung (:1157-1198), _pinned_int's fall-through; LADDER_RESOLVED_PIN_NAMES (:1226) against the set the ladder fills, and both its uses (:1240 filter, :1243 unset loop); _export_operator_launch_shape and both call sites (:1678, :1852) with an AST walk of _run_optimize recording the enclosing condition of every os.environ[...] write, variable-key writes included; a tree-wide sweep for every non-test writer of TP/CONC/EP/MAX_MODEL_LEN; tp_resolved/ep_resolved at :1636-1637 and the topology gates they feed; _seed_shared_state (bootstrap.py:368-372, called at :2284) for what state.json records; every caller of _resolve_run_max_model_len_inner; the relocated resume restore against agentx_state_is_stale at :1856 and the state writeback at :1952; env_safety.operator_extra_env against both parser bodies it replaces; _is_agentx_client_key against the inline condition it replaces; non-test bare readers of TP/CONC/EP/MAX_MODEL_LEN; head-tree occurrence counts for every added identifier; the eight added test cases against the branch each exercises. | Ran: the real _export_operator_launch_shape, _export_workload_envs_for_optimize and _resolve_workload_knobs in each branch's statement order with pins {TP=4, EP=2, CONC=63} -- fresh args 4/2/63 against os.environ 1/1/64, resume 4/2/63 against 4/2/63; the same through the real materialize_config_with_envs; and the resume order with state.max_model_len=32768 against a pinned 65536. Compared against 63c66e7d9 and 1c14e79e1, where neither split appears. | Base: 1c14e79 | Head: 786fe72

…n change MAX_MODEL_LEN

Two regressions from withholding the ladder-resolved names from the export.

_export_workload_envs_for_optimize is the only non-test writer of
os.environ["TP"|"CONC"|"EP"], and it ran about 500 lines before
_resolve_workload_knobs filled the pin, from flag-derived values. The fresh
re-projection after the ladder covered MAX_MODEL_LEN/ISL/OSL/PRECISION and
omitted these three, so --extra-env TP=4 EP=2 CONC=63 gave args and
state.json 4/2/63 against an environment still at 1/1/64 -- and the
environment is what the server launches from, what the Magpie YAML and the
comparability fingerprint are built from, and what the preflight GPU-count
gate reads. The projection moves below the ladder. The topology gates are now
a function called twice: once on the flags, so a bad --tp still fails before
the slow preflight work, and again on the resolved values, so a pinned TP
cannot pass a gate the shape it launches with would fail.

On the resume branch _resume_max_model_len read flag-or-state and wrote it
back over the pin the restore had just placed, so re-passing --extra-env
MAX_MODEL_LEN to change it was silently discarded -- 63c66e7 honoured it
because its export sat after that loop. It now goes through
_resolve_resume_max_model_len: flag, pin, recorded value.
@xiaofei-zheng

Copy link
Copy Markdown
Collaborator Author

Both confirmed and fixed in 6ac02d0d2. Reproduced each in the branch's own statement order before changing anything; the numbers match yours exactly -- fresh args 4/2/63 against os.environ 1/1/64, and the resume MAX_MODEL_LEN going 65536 after the restore to 32768 after the loop.

free:ladder-knobs-not-reprojected-on-fresh. Correct, and my mistake was mechanical: the AST walk I used to check the write order only descended _run_optimize, and os.environ["TP"] is written one frame down inside _export_workload_envs_for_optimize, so the three names never appeared in my own ordering table. The projection now sits below _resolve_workload_knobs and takes args.tp/args.ep.

On the gates: I did not move them, because they also run on the resume branch, where args.tp is only filled by that branch's own ladder much later -- moving them would have left a resume with no topology check at all. They are now _enforce_topology_gates(...), called twice on a fresh launch. The first call keeps the early fail on a bad --tp, before any of the slow preflight work; the second runs after the ladder, so a pinned TP=16 is refused by the gate it would otherwise have walked past. Same two gates, same messages.

free:max-model-len-pin-lost-on-resume. Correct. MAX_MODEL_LEN goes through _resolve_resume_max_model_len(args, pins, state) -- flag, pin, recorded value -- rather than an inline flag-or-state read. I kept it a separate rung instead of calling _resolve_run_max_model_len on both branches: that function's AgentX path derives from the model's native context length, which is a fresh-launch decision, and running it on a resume would change what a resumed AgentX session uses for a reason unrelated to this fix.

On the tests, you were right that the four ladder cases prove nothing about ordering. The case that pins TP through the function sequence would pass on the previous head too, because it calls them in the order it wants. So the one that actually pins the defect asserts against the source: every _export_workload_envs_for_optimize call in _run_optimize must appear after the first _resolve_workload_knobs. It fails on 786fe72f with the two line numbers in the message. The MAX_MODEL_LEN case covers all four rungs including a malformed pin.

One thing I did not change, in case you read it differently: SKIP_VARIANTS stays unpinnable, and the GPU_TYPE pin is still overwritten by the probe/flag resolution at :2166. GPU_TYPE has --gpu-type plus hardware autodetection behind it, and the env name carries the Magpie runner type rather than the probed value, so folding a pin into that ladder would make a pinned value reappear as something else. If you want it pinnable I would rather do it as its own change.

CodeQL flagged the file for importing hyperloom.inference_optimizer.cli with
both 'import' and 'import from'. monkeypatch.setattr takes a dotted string
target, so the module object is not needed to patch _preflight and friends.
…e enumeration

Two blocking defects in this PR survived a green CI and my own verification,
each for a reason the rules did not name.

T5: the fix moved a projection relative to a resolver, and the tests added
with it called the two in the order they wanted -- passing both before and
after, while the production path still ran them the wrong way round. An
ordering fix has to assert the order itself, or drive the real entry point.

V7: the AST walk I used to record every os.environ write in _run_optimize
never listed TP/CONC/EP, because they are written one frame down inside
_export_workload_envs_for_optimize. A tool scoped to one function and a tree
with no writers produce the same empty output, so the scope of the search
has to be stated with its result.
@jiaqiang-dot-liu

Copy link
Copy Markdown
Collaborator

PR #1797 -- fix(cli): export each --extra-env pin so every reader resolves it alike

What it does: the CLI serialized the operator's --extra-env pins into the single INFERENCE_OPTIMIZER_EXTRA_ENV JSON blob and never exported the individual names, so every Hyperloom control variable -- all read with a bare os.environ.get -- was blind to them, and AgentX had to re-merge the blob locally in agentx_env_for_conc. Pins now become real environment variables; the six knobs Hyperloom resolves for itself are withheld from that export and enter _resolve_workload_knobs as a rung (flag > pin > resumed state > default); the fresh branch's TP/CONC/EP projection moves below that ladder and the topology gates become _enforce_topology_gates, run once on the flags and again on the resolved shape; the resume branch restores the pins ahead of agentx_state_is_stale and resolves MAX_MODEL_LEN through _resolve_resume_max_model_len; the local merge layer is gone and the three blob parsers are one operator_extra_env() in env_safety. It also adds the T5 and V7 review rules to .claude/skills/review-pr/.

Blocking issues: 1

  1. [free:tp-projected-after-preflight] Moving the TP/CONC/EP projection below the ladder carries it past _preflight, so the check_gpu_visibility step no longer sees the run's TP [verified]
    Problem: _export_workload_envs_for_optimize is the only non-test writer of os.environ["TP"|"CONC"|"EP"] in the tree (cli/__init__.py:1240-1242) and has one call site. At the merge base 880c1672a it ran at :1634, before _preflight at :1697; at head it runs at :2236, after _preflight at :1782. _preflight (cli/preflight.py:2161-2709) registers the check_gpu_visibility step at :2328, whose action _check_gpu_visibility (:1457) reads wanted = int(os.environ.get("TP", "1") or "1") at :1499. _export_operator_launch_shape deliberately withholds TP from the environment at :1308 because it is in LADDER_RESOLVED_PIN_NAMES, so on a fresh launch from a shell that does not itself export TP the name is unset when the step runs. src/hyperloom/inference_optimizer/cli/init.py:2236
    Impact: wanted is 1 for every fresh launch, so if wanted > visible is unreachable for any non-zero GPU count and the warning "TP=N but rocm-smi sees M GPU(s); sglang/vllm may fail to load weights" can no longer fire -- --tp carries default=None (cli/parser.py:349-351), so this is the explicit --tp 8-on-a-2-GPU-box case the check exists for, not only the pinned one. _enforce_topology_gates does not cover it: that returns at nodes < 2 and this is the single-node check. The step's tp_requested detail is also recorded as 1 regardless of --tp, and it is durable -- _run_install_step (preflight.py:2063-2085) folds the detail into event["ext"]["steps"] and _persist_install_event writes it into the session at cli/__init__.py:1644, :1862 and :2281. The description says the projection "moves below the ladder" and does not say it also moves below preflight.
    Action: project TP/CONC/EP before _preflight as well as after the ladder, or pass the resolved TP into the check instead of reading it from the environment -- _check_gpu_visibility already returns tp_requested in its detail, so giving _run_install_step the value is the smaller change. Extend the ordering assertion the PR adds: it pins _export_workload_envs_for_optimize against _resolve_workload_knobs and says nothing about _preflight, which is why CI is green.

Checked: _resolve_workload_knobs and its pin rung (:1184-1218) with _pinned_positive_int (:1156); _resolve_resume_max_model_len (:1169, called at :1960) against the inline read it replaces; LADDER_RESOLVED_PIN_NAMES (:1248) against the set the ladder fills, and both uses (:1308 filter, :1313 unset loop); _enforce_topology_gates (:1251) against the inline gates it replaces -- same two gates, same messages, same nodes < 2 guard -- and both call sites (:1697 on the flags, :2228 on the resolved values); the full fresh and resume orderings by AST walk with the enclosing condition of every os.environ[...] write (fresh: ladder :2225, gates :2228, TP/CONC/EP :2236, MAX_MODEL_LEN/ISL/OSL/PRECISION :2240-2247, manifest :2280, seed :2326; resume: export :1883, guard :1887, ladder :1959, MAX_MODEL_LEN :1960, loop :1970, PRECISION :1980, writeback :1986-1987); a tree-wide sweep -- not scoped to _run_optimize -- for every non-test writer and reader of TP / CONC / EP / MAX_MODEL_LEN, which returned one writer and placed preflight.py:1499 as the only reader inside the window the projection moved across; the _run_install_step -> _record_install_step -> _persist_install_event chain; head-tree occurrence counts for every added identifier; T5 and V7 in .claude/skills/review-pr/rules.md against the "Adding a new rule" contract and the index rows they are listed on. | Ran: AST comparison of the _export_workload_envs_for_optimize / _preflight / _resolve_workload_knobs call lines in the merge-base and head trees -- base 1634 before 1697, head 2236 after 1782, one projection call site in each. | Base: 880c167 | Head: a6dc206

Moving the TP/CONC/EP projection below the ladder carried it past
_preflight, and check_gpu_visibility reads $TP to warn when a run asks for
more shards than rocm-smi can see. TP is withheld from the pin export
because the ladder owns it, and --tp defaults to None, so on a fresh launch
the name was simply unset there: wanted fell back to 1, the warning became
unreachable for any non-zero GPU count, and the step recorded tp_requested=1
into the session regardless of --tp.

_resolve_workload_knobs now fills the five numeric knobs only, and runs
before the gates, the projection and _preflight, so all four see the same
TP. Precision splits off into _resolve_precision, which stays after the
model is resolved because it reads the checkpoint -- running it early would
detect nothing and settle on the default with no later call able to correct
it. The second _enforce_topology_gates call goes away: with the ladder ahead
of the gates, the first call already sees the resolved shape.

The ordering test becomes a table of anchors rather than one pair. Each time
this order was corrected for one anchor it broke for another, which a single
assertion could not catch.
@xiaofei-zheng

Copy link
Copy Markdown
Collaborator Author

Confirmed and fixed in cbdbf4de4. The AST comparison reproduces exactly what you reported -- base: projection 1634 before _preflight 1697; head a6dc206f7: projection 2236 after _preflight 1782 -- and --tp carrying default=None with TP withheld from the export means the name was simply absent, so this was not only the pinned case.

I did not add a second projection. The projection moved late because the ladder was late, so the ladder moves instead: _resolve_workload_knobs now fills the five numeric knobs only and runs right after the launch-shape export, before the gates, the projection and _preflight. All four see the same TP, and the second _enforce_topology_gates call added in the previous round is gone -- with the ladder ahead of them, the first call already sees the resolved shape, which is simpler than what I had.

Precision could not come along: it reads the checkpoint's config.json, and running it before the model is resolved would detect nothing, settle on the default, and leave the later call with args.precision already set and nothing to correct. It splits into _resolve_precision, called where the old combined ladder used to sit on the fresh branch and alongside the knob ladder on the resume branch.

The ordering test is now a table of anchors rather than one pair, which is the actual lesson -- each time I corrected this order for one anchor it broke for another, and a single assertion could not see that:

_export_operator_launch_shape  <  _resolve_workload_knobs        (the ladder reads the pins from the environment)
_resolve_workload_knobs        <  _export_workload_envs_for_optimize  (else the default is published over the pin)
_export_workload_envs_for_optimize  <  _preflight                 (check_gpu_visibility reads $TP)
_resolve_workload_knobs        <  _enforce_topology_gates         (a pinned TP faces the same gate as --tp)

On a6dc206f7 it fails on the third row with 553 < 99 and the reason in the message.

I also ran the sweep in the direction that would have caught this: for every name _preflight reads, which does _run_optimize write only after calling it -- collecting writes from the helpers it calls, not just its own body. At head that returns FRAMEWORK alone, which is also true on the merge base (cli:1836, read at preflight.py:331), so it is pre-existing and out of scope here; TP is back on the before side. A second sweep confirms each ladder-resolved name still has exactly one writer in the tree, and that no json.loads of INFERENCE_OPTIMIZER_EXTRA_ENV survives outside env_safety.operator_extra_env -- the blob is now read in one place, plus the Langfuse redactor that masks it by name.

Worth stating plainly: that sweep is the one I should have run before moving the projection, and not running it is why this round existed. It is now V7 in .claude/skills/review-pr/rules.md, together with T5 for the ordering-assertion half; both ride along in this PR at the maintainer's request, and the PR body marks the single-concern answer as no because of it.

_enforce_topology_gates still said it was called twice, and the launch-shape
export still justified withholding the ladder names by where it runs on each
branch. Both describe the previous arrangement; the gate runs once, and the
reason those names stay out is that the projection is their single writer.
…ird source

The pin machinery had grown five mechanisms -- an exclusion set, a filter in
the export, a matching exclusion in the unset loop, a pin-reading int helper
and a resume-only MAX_MODEL_LEN resolver -- all of them there to keep a pin
from colliding with the projection of the ladder's answer.

They existed because I treated a pin as a third source alongside the flag and
the environment. It is not: _export_operator_launch_shape puts it in the
environment, which MAX_MODEL_LEN has always read as its second rung. Export
every pin, have the ladders read $NAME, and the collision cannot arise -- the
export runs before the ladder, the ladder writes args, the projection writes
args. An explicit flag still outranks both, and a pin and an operator's own
export now mean exactly the same thing, which is what this PR claims.

Net 92 lines removed from cli/__init__.py against the previous commit, and one
concept fewer to explain.
@xiaofei-zheng xiaofei-zheng changed the title fix(cli): export each --extra-env pin so every reader resolves it alike fix(cli): export every --extra-env pin so all readers resolve it alike Oct 10, 2026
Three sets of tests outside the files I had been running, which is how CI
caught them and I did not.

The early ladder call reached args.resume_from on a Namespace that the
topology-gate tests build with four attributes, since nothing before the
gates used to need it; it goes through getattr like every other attribute
read in that stretch.

_resolve_workload_knobs no longer fills precision, so the cases asserting it
call _resolve_precision -- renamed to say so -- or both halves through a
local helper, the way _run_optimize calls them.

The ladder now has an environment rung, so a test asserting what the flags,
the state and the defaults produce has to start from an empty environment
rather than whatever the worker ran before it. The autouse fixture in
test_cli_workload_envs.py clears the set on the way in as well as restoring
it on the way out.
The docstrings and comments added by this PR ran to twice the lines of the
code they sat on, most of it recounting how the arrangement got here rather
than what holds it in place. Keep the constraints -- the model has to be
resolved before precision is detected, the pins have to be exported before
the guard reads the backend, SKIP_VARIANTS is not part of the measurement
contract -- and drop the rest. One orphaned comment block left over from
moving the projection goes with it.
They are review-process rules with no runtime code, and carrying them here
made this a two-concern PR. They go out on their own branch instead.
xiaofei-zheng added a commit that referenced this pull request Oct 10, 2026
Backtesting the rule against #1797's own history: the whole-diff totals read
6591 code against 1216 comment and cleared at every head, because test code
dilutes the ratio and the merge of main swamped it. cli/__init__.py measured
on its own read 28 against 56 and fired from the second commit onward. A
per-commit reading inverts for the opposite reason -- a docs-only commit
adds comment by construction.
@jiaqiang-dot-liu

Copy link
Copy Markdown
Collaborator

PR #1797 -- fix(cli): export every --extra-env pin so all readers resolve it alike

What it does: the CLI serialized the operator's --extra-env pins into the single INFERENCE_OPTIMIZER_EXTRA_ENV JSON blob and never exported the individual names, so every Hyperloom control variable -- all read with a bare os.environ.get -- was blind to them, and AgentX compensated by re-merging the blob inside agentx_env_for_conc, which is how a pinned HYPERLOOM_AGENTIC_BACKEND=mlperf could run the MLPerf client while the grading axis, state.agentx_backend and resolve_benchmark_timeouts() all still said aiperf. The fix is at the producer: every pin is exported under its own name with no exclusions, the local merge layer is gone and the three blob parsers collapse into one operator_extra_env() in env_safety. The knobs Hyperloom resolves for itself then take the environment as one rung of their own ladder -- flag, $NAME, resumed state, default -- so args is the single answer and the projections write it back; _resolve_precision splits out because it reads the checkpoint, the topology gates become _enforce_topology_gates on the resolved shape, and the resume branch restores the pins ahead of agentx_state_is_stale.

Blocking issues: 1

  1. [free:env-rung-reverses-the-reference-docs] The reference docs still say these six environment variables are ignored, in three places, none of which the PR moved [verified]
    Problem: _resolve_workload_knobs (cli/__init__.py:1165) and _resolve_precision (:1190) now read $ISL/$OSL/$CONC/$TP/$EP/$PRECISION as the rung below an explicit flag, via _positive_env_int (:1156). docs/reference/environment-variables.md:80-82 still reads "Set with CLI flags, not env vars. Pre-set ISL / OSL / CONC / PRECISION / TP / EP env vars are ignored and overwritten (GPU_TYPE is a fallback when --gpu-type is omitted)" -- that file is touched by this PR, at line 948 only. docs/reference/multi-node.md:154 still reads "These must be flags -- the corresponding env vars are ignored and overwritten", and :155 "Flags are authoritative; env vars are not" for --gpu-type, --precision, which is now half true. The how-to the PR does rewrite states the new ladder, so the tree contradicts itself on the page an operator consults to learn exactly this. docs/reference/environment-variables.md:80
    Impact: ran _resolve_workload_knobs in both trees with ISL=4096 CONC=256 TP=4 PRECISION=fp8 pre-set and no flags -- merge base returns the defaults (1024, 1024, 64, 1, 1)/bf16 with state=None and the state's (3333, 777, 99, 8, 2)/int4 with a state; head returns (4096, …, 256, 4, …)/fp8 in both, so the pre-set shell beats the resumed session's recorded workload too. The trigger is a workflow Hyperloom generates: reference_script.py:420-425 writes export TP= / export MAX_MODEL_LEN= / export GPU_TYPE= into the launch-recipe script, orchestrator/kernel/request_handlers.py:999-1003 writes export TP/CONC/ISL/OSL into the kernel-agent repro script, and examples/hyperloom-custom-advanced/SKILL.md:301-304 has operators export the same names. An operator who sources one of those and then runs optimize -- or --resume-from -- in the same shell now measures a different workload than the reference page promises, silently, and canonical_fingerprint.py:117-121 and session/manifest.py:218-247 record the shell's numbers as the session's.
    Action: rewrite docs/reference/environment-variables.md:80-82 and docs/reference/multi-node.md:154-155 to the ladder the how-to now states (--precision moves off the "env vars are not" row; GPU_TYPE stays on it). Decide in the same edit whether the environment should outrank a resumed session's recorded value -- the ladder puts it there, and a stale export in the shell is indistinguishable from a deliberate one.

Checked: _resolve_workload_knobs (:1165) and its environment rung via _positive_env_int (:1156); _resolve_precision (:1190) against the half of the old combined ladder it was cut from -- same persisted rung, same detection rung, same WARN branch; _export_operator_launch_shape (:1259) with every exclusion now gone and its unset loop at :1278; _enforce_topology_gates (:1229) against the inline block it replaces, same two gates and messages, now called once at :1678 on the resolved shape; the full fresh and resume orderings by AST walk with the enclosing condition of every call and every os.environ[...] write (fresh: export :1655, ladder :1663, projection :1676, gates :1678, _preflight :1753, precision :2202, MAX_MODEL_LEN :2204, ISL/OSL :2207-2208, PRECISION :2213, manifest :2246, seed :2292; resume: export :1853, guard :1857, ladder :1929, precision :1930, MAX_MODEL_LEN rung :1934, loop :1939, writeback :1963-1964) -- all four of the PR's ordering anchors hold, and last round's _preflight finding is closed; a tree-wide sweep for every non-test writer and reader of the six names plus MAX_MODEL_LEN, which found one writer and no reader in any window the reordering crossed; whether the ladder running at :1663 instead of :2136 exposes a None-sensitive consumer (preflight.py reads none of them, model_gate.py:1161/1457 is reached at :2103/:2299/:2309, bootstrap.py:370 at :2292); the five helpers deleted this round for dangling references; both existing test files the ladder split touched, whose assertions are kept and whose fixture now pops the seven rung names; a grep confirming no json.loads of the blob survives outside env_safety.operator_extra_env; and a tree-wide sweep for other copies of the "ignored and overwritten" claim, which is how the third one was found. | Ran: _resolve_workload_knobs / _resolve_precision in the merge-base and head worktrees with the six names pre-set, with and without a resumed state and with --isl 512, plus the AST ordering comparison of both trees. | Base: 880c167 | Head: 2f2f7bf

@jiaqiang-dot-liu
jiaqiang-dot-liu merged commit 20d4309 into main Oct 10, 2026
37 checks passed
@jiaqiang-dot-liu
jiaqiang-dot-liu deleted the feature/xiaofei/agentx-extra-env-spec branch October 10, 2026 23:54
xiaofei-zheng added a commit that referenced this pull request Oct 11, 2026
On #1797 the grep on the split symbol reaches two of the three files; the third
drives the entry point the moved call sits inside and names the symbol nowhere.
The evidence line now asks for that grep too.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants