Repository navigation
Repair a local Ray head without serving_slot; stop early on unrecoverable baseline failures - #1816
Open
zoroyihan7 wants to merge 6 commits into
Open
zoroyihan7 wants to merge 6 commits into
zoroyihan7 wants to merge 6 commits into
Conversation
…ne failures A Ray head started before Hyperloom without the serving_slot resource was reused as-is, so every serving lease failed its feasibility check and the baseline failed three times before the run ended as baseline_failed. - When the connected head lacks serving_slot and it is a lone, idle head on this host (no explicit RAY_ADDRESS, not multi-node, one live node that is the local raylet, no resources in use), restart it the way the installer does: ray stop --force, then ray start --head with the resource. Any other cluster keeps the existing error, now saying why it was not restarted. - Tag RayInfeasibleError messages with a stable marker. Two consecutive subprocess_nonzero baselines with the same output (modulo numbers) now stop the run at two; an infeasible Ray cluster stops at two even while the enablement lane is open and is not handed to it as a launch log. - Docs: the manual ray start commands declare serving_slot. Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
…whole failure text - ray stop --force stops every head on the host, so require exactly one visible gcs_server, and treat live actors (including resourceless ones) and other live drivers as work that forbids the restart. - RAY_ADDRESS=local starts a new cluster on ray.init, so it no longer qualifies for the repair. - Leases now connect, check feasibility and create their actor under the same lock the repair holds. - The baseline failure signature hashes the whole normalised text instead of its tail. Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
…ore GPUs - The only gcs_server on the host must be the connected cluster's (matched by GCS port), every node record counts (a dead worker is still a worker), and another driver is identified by job id alone (pids repeat across PID namespaces). - Check serving_slot before the GPU count so a hand-started head short of both reaches the repair. Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
…e host ray stop --force also stops a worker node of an unrelated cluster, which runs a raylet but no GCS server. Require the only raylet on the host to be the connected node (by --node_id) as well as the only GCS server. Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
…e actor Actor creation is asynchronous, so a specialist submitted a moment ago can be missing from the GCS actor table and hold no resources yet. Count lease actor submissions under the startup lock and refuse the restart while any exist. Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
Waiting for a second one was not bounded: the first failure opens the enablement lane, and with no launch log to author against that lane never schedules another baseline. What the runtime could repair it already did, so stop on the first, like the AgentX preflight. Co-Authored-By: Claude <noreply@anthropic.com> Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description: what and why
A Ray head that already exists when Hyperloom starts (for example one started by hand from the troubleshooting doc) is reused by
ensure_ray_clusteras long asray statusanswers. If it lacks theserving_slotcustom resource, every serving lease fails_assert_cluster_feasiblewithexisting Ray head has no serving_slot resource, each baseline attempt fails assubprocess_nonzero, and the session only ends after the retry budget, with long idle stretches in between (the first failure opens the enablement lane, which has no launch log to author against and never schedules another baseline). The installer already restarts such a head; the runtime path did not._ray_serving._ensure_cluster_feasible,_ray_runtime.local_head_restartable/restart_local_head_with_serving_slot): when the connected head lacksserving_slot, restart it the way the installer does (ray stop --force, thenray start --head ... --resources={"serving_slot": 1}), reconnect, and re-check feasibility. Becauseray stop --forcestops every Ray process on the host, the restart is refused (and the existing error kept, with the reason appended) unless all of these hold: no explicitRAY_ADDRESS(only unset/auto;localwould start a new cluster on reconnect), not a multi-node run, exactly one node record in the cluster and it is this host's live raylet, the onlygcs_serverandrayletvisible on the host are the connected cluster's (by--gcs_server_port/--node_id), no resource held, no actor that is not dead, no other live driver (by job id), and this process has not submitted a lease actor yet (actor creation is asynchronous). Leases now connect, check feasibility and create their actor under one lock that the repair also holds. The slot is checked before the GPU count so a hand-started head short of both still reaches the repair.loop/writeback.py): infeasibility errors now carry a stableray_cluster_infeasiblemarker. A baseline that fails with it stops the run asbaseline_failedon the first occurrence (also in enablement) and is not stashed as an enablement launch log; whatever the runtime could repair it already has. Separately, two consecutivesubprocess_nonzerofailures whose whole error text matches after folding numbers/hex ids stop at two instead of three (outside enablement only; other classes carry fixed sentences and keep the three-strike rule). New state fieldbaseline_last_failure_signature.ray start --headcommands introubleshooting.mdandupgrade.mdnow declareserving_slot, with a note on the automatic repair and when it is refused.Linked issue(s): none
Tests: added/updated? commands run?
orchestrator/actions/executors/tests/test_ray_serving_slot_repair.py: lone idle local head without the slot is restarted then feasible (also end to end throughServingLease.ensure, and when it is also short of GPUs); head with the slot, or a lease that needs no slot, is never restarted; explicitRAY_ADDRESS(incl.local), multi-node, a second live or dead node record, a head on another host, held resources, a resourceless live actor, another live driver (same pid, different job id), an uninspectable cluster, a second head / a foreign raylet / a non-matching head on the host, and a lease actor already submitted all keep the error; actor creation happens under the startup lock;/procparsing; the restart runs stop then a start declaring the slot.test_coordinator_async_methods_coverage_unit.py: same failure twice stops at two; different failures, a shared long traceback tail, a repeated non-subprocess_nonzeroclass, and a repeat inside enablement keep the old budget; an infeasible cluster stops on the first failure, also in enablement, with no launch log stashed.ruff check .,ruff format --check .(0.16.2), andpytestontest_ray_*.py,test_ray_backend_unit.py,test_coordinator_async_methods_coverage_unit.py,test_dead_holder_failure_accounting.py,test_dispatched_task_policy.py,test_bringup_round_scenario.py,test_enablement_*,test_sbd_v6_enablement_wiring.py,test_phase_state_machine.py,test_shared_state_evolution.py,test_backend_gating.py,test_gaps_field.py: all pass. Ray itself is faked in the unit tests (the privateray._private.statereads and/proclayout were checked against the Ray 2.44.1 sources).Size/complexity triggers crossed: the diff is mostly tests (~470 of ~850 added lines); production code is split into small helpers (
_visible_ray_daemons,_cluster_activity,local_head_restartable,_ensure_cluster_feasible,_baseline_failure_signature).If this simplifies or refactors: no;
force_restart_local_clustergained an optionalreasonand a generic failure message, behaviour otherwise unchanged.Observable effect: a session started on a host where an idle, lone, hand-started Ray head lacks
serving_slotnow logsRay head has no serving_slot resource; restarting the local head with itand proceeds. Where the restart is not safe, the existing error now says why (for example... not restarting it automatically: RAY_ADDRESS='...' names an explicit cluster) and the session ends asbaseline_failedon the first baseline instead of after three attempts and hours of idling. Two identicalsubprocess_nonzerobaseline failures end the session one attempt earlier.Breaking changes: no.
PR addresses single concern: yes (a pre-existing Ray head without
serving_slotand the retries it caused).Root cause is upstream: no.
🤖 Generated with Claude Code