Repository navigation
fix(cli): let a resumed session's recorded workload outrank the shell - #1818
xiaofei-zheng wants to merge 2 commits into
Conversation
Giving the knobs an environment rung also gave a stale export one. Hyperloom writes export TP/CONC/ISL/OSL into the repro scripts it generates and the custom-workload SKILL has operators export the same names, so sourcing one and then resuming re-measured the session at the shell's shape and recorded those numbers in the fingerprint and the manifest. Reproduced: shell TP=4 against a session recorded at TP=8 resolved to 4. _seed_env_from_resumed_state puts the recorded values in the environment before the ladder reads it, skipping any name this resume re-pins, so the workload only changes through a flag or a re-passed --extra-env. The reference docs said the opposite of the new behaviour in two places the PR had not touched -- environment-variables.md called these names "ignored and overwritten", multi-node.md "env vars are not" authoritative. Both now state the ladder, with --gpu-type kept on the flag-only row.
…ationale Self-review with the rules in #1817 found two things this PR had missed. T5: the case covering the seeding calls it and the ladder in the order it wants, so it passes with the two swapped in _run_optimize -- verified by swapping them, where the behavioural case still passed and the defect was live. The ordering anchor table gains the pair, resolved against the resume branch's own ladder call rather than the fresh one, and fails on the swap. X3: a tree-wide sweep for the claim this PR's ladder contradicts found one more copy -- the Qwen skill told operators that "CLI defaults can otherwise override the intended workload", which stopped being true when the default became the bottom rung. The advice stands; the reason is now the real one. Two further copies state the advice without a rationale and are left alone.
PR #1818 -- fix(cli): let a resumed session's recorded workload outrank the shellWhat it does: #1797 gave Blocking issues: 2
Checked: src/hyperloom/inference_optimizer/cli/init.py (_seed_env_from_resumed_state, |
Description: fix(cli): export every --extra-env pin so all readers resolve it alike #1797 gave
ISL/OSL/CONC/TP/EP/PRECISIONan environment rung — flag,$NAME, resumed state, default — which also gave a stale shell export one. Hyperloom writesexport TP/CONC/ISL/OSLinto the repro scripts it generates (reference_script.py:420-425,orchestrator/kernel/request_handlers.py:999-1003) andexamples/hyperloom-custom-advanced/SKILL.md:301-304has operators export the same names, so sourcing one and then running--resume-fromin that shell re-measured the session at the shell's shape, withcanonical_fingerprint.py:117-121andsession/manifest.py:218-247recording those numbers as the session's. Reproduced on the merge base: a shell atTP=4 CONC=256 ISL=4096against a session recorded attp=8 conc=99 isl=3333resolved to the shell's values._seed_env_from_resumed_stateputs the recorded workload into the environment before the ladder reads it, skipping any name this resume re-pins. The recorded workload therefore outranks the shell for everything resolved from the ladder, and changing it stays an explicit act — a flag, or--extra-env NAME=VALUEpassed again on the resume. The fresh branch is unchanged. One reader is deliberately outside this:_preflightruns before the seeding, socheck_gpu_visibilitystill compares$TPfrom the shell. That is pre-existing — the resume branch has never projected TP before preflight — and it only shapes a warning, so it is left for its own change.The reference docs also still stated the pre-fix(cli): export every --extra-env pin so all readers resolve it alike #1797 behaviour in two places that PR did not touch:
environment-variables.mdcalled these names "ignored and overwritten" andmulti-node.mdsaid "env vars are not" authoritative. Both now state the ladder.--gpu-typestays on the flag-only row, sinceGPU_TYPEremains a fallback rather than a rung.Linked issue(s): follows fix(cli): export every --extra-env pin so all readers resolve it alike #1797
Tests:
test_cli_resume_launch_shape.pyadds a case driving_seed_env_from_resumed_stateand_resolve_workload_knobsin the resume branch's order — a shell carryingTP=4 CONC=256 ISL=4096, a session recorded attp=8 isl=3333, and--extra-env CONC=16re-pinned — asserting the recorded values survive while the re-pinned one changes. Without the seeding the shell'sTP=4wins and the case fails.test_cli_resume_launch_shape.pyandtest_cli_workload_envs.pypass (47).Size/complexity triggers crossed: none
If this simplifies or refactors: n/a
Observable effect: resuming a session from a shell that has
TP/CONC/ISL/OSLleft in it — which is what sourcing a generated repro script does — now continues at the workload the session was measured at, instead of silently switching to the shell's and recording that in the fingerprint and the manifest. Re-passing--extra-envstill changes it.Breaking changes: no. It restores the pre-fix(cli): export every --extra-env pin so all readers resolve it alike #1797 outcome for a resume whose shell carries these names, while keeping the rung fix(cli): export every --extra-env pin so all readers resolve it alike #1797 added for a fresh launch and for an explicitly re-passed pin.
PR addresses single concern: yes
Root cause is upstream: no