Skip to content

Repair a local Ray head without serving_slot; stop early on unrecoverable baseline failures - #1816

Open
zoroyihan7 wants to merge 6 commits into
mainfrom
fix/ray-head-without-serving-slot
Open

zoroyihan7 wants to merge 6 commits into
mainfrom
fix/ray-head-without-serving-slot

Conversation

@zoroyihan7

Copy link
Copy Markdown
Contributor
  • Description: what and why

    A Ray head that already exists when Hyperloom starts (for example one started by hand from the troubleshooting doc) is reused by ensure_ray_cluster as long as ray status answers. If it lacks the serving_slot custom resource, every serving lease fails _assert_cluster_feasible with existing Ray head has no serving_slot resource, each baseline attempt fails as subprocess_nonzero, and the session only ends after the retry budget, with long idle stretches in between (the first failure opens the enablement lane, which has no launch log to author against and never schedules another baseline). The installer already restarts such a head; the runtime path did not.

    1. Runtime repair (_ray_serving._ensure_cluster_feasible, _ray_runtime.local_head_restartable / restart_local_head_with_serving_slot): when the connected head lacks serving_slot, restart it the way the installer does (ray stop --force, then ray start --head ... --resources={"serving_slot": 1}), reconnect, and re-check feasibility. Because ray stop --force stops every Ray process on the host, the restart is refused (and the existing error kept, with the reason appended) unless all of these hold: no explicit RAY_ADDRESS (only unset/auto; local would start a new cluster on reconnect), not a multi-node run, exactly one node record in the cluster and it is this host's live raylet, the only gcs_server and raylet visible on the host are the connected cluster's (by --gcs_server_port / --node_id), no resource held, no actor that is not dead, no other live driver (by job id), and this process has not submitted a lease actor yet (actor creation is asynchronous). Leases now connect, check feasibility and create their actor under one lock that the repair also holds. The slot is checked before the GPU count so a hand-started head short of both still reaches the repair.
    2. Baseline stop (loop/writeback.py): infeasibility errors now carry a stable ray_cluster_infeasible marker. A baseline that fails with it stops the run as baseline_failed on the first occurrence (also in enablement) and is not stashed as an enablement launch log; whatever the runtime could repair it already has. Separately, two consecutive subprocess_nonzero failures whose whole error text matches after folding numbers/hex ids stop at two instead of three (outside enablement only; other classes carry fixed sentences and keep the three-strike rule). New state field baseline_last_failure_signature.
    3. Docs: the manual ray start --head commands in troubleshooting.md and upgrade.md now declare serving_slot, with a note on the automatic repair and when it is refused.
  • Linked issue(s): none

  • Tests: added/updated? commands run?

    • New orchestrator/actions/executors/tests/test_ray_serving_slot_repair.py: lone idle local head without the slot is restarted then feasible (also end to end through ServingLease.ensure, and when it is also short of GPUs); head with the slot, or a lease that needs no slot, is never restarted; explicit RAY_ADDRESS (incl. local), multi-node, a second live or dead node record, a head on another host, held resources, a resourceless live actor, another live driver (same pid, different job id), an uninspectable cluster, a second head / a foreign raylet / a non-matching head on the host, and a lease actor already submitted all keep the error; actor creation happens under the startup lock; /proc parsing; the restart runs stop then a start declaring the slot.
    • test_coordinator_async_methods_coverage_unit.py: same failure twice stops at two; different failures, a shared long traceback tail, a repeated non-subprocess_nonzero class, and a repeat inside enablement keep the old budget; an infeasible cluster stops on the first failure, also in enablement, with no launch log stashed.
    • Every new guard was mutated (removed or weakened) and at least one test failed for each; restored afterwards.
    • Ran: ruff check ., ruff format --check . (0.16.2), and pytest on test_ray_*.py, test_ray_backend_unit.py, test_coordinator_async_methods_coverage_unit.py, test_dead_holder_failure_accounting.py, test_dispatched_task_policy.py, test_bringup_round_scenario.py, test_enablement_*, test_sbd_v6_enablement_wiring.py, test_phase_state_machine.py, test_shared_state_evolution.py, test_backend_gating.py, test_gaps_field.py: all pass. Ray itself is faked in the unit tests (the private ray._private.state reads and /proc layout were checked against the Ray 2.44.1 sources).
  • Size/complexity triggers crossed: the diff is mostly tests (~470 of ~850 added lines); production code is split into small helpers (_visible_ray_daemons, _cluster_activity, local_head_restartable, _ensure_cluster_feasible, _baseline_failure_signature).

  • If this simplifies or refactors: no; force_restart_local_cluster gained an optional reason and a generic failure message, behaviour otherwise unchanged.

  • Observable effect: a session started on a host where an idle, lone, hand-started Ray head lacks serving_slot now logs Ray head has no serving_slot resource; restarting the local head with it and proceeds. Where the restart is not safe, the existing error now says why (for example ... not restarting it automatically: RAY_ADDRESS='...' names an explicit cluster) and the session ends as baseline_failed on the first baseline instead of after three attempts and hours of idling. Two identical subprocess_nonzero baseline failures end the session one attempt earlier.

  • Breaking changes: no.

  • PR addresses single concern: yes (a pre-existing Ray head without serving_slot and the retries it caused).

  • Root cause is upstream: no.

🤖 Generated with Claude Code

zoroyihan7 and others added 6 commits October 10, 2026 14:20
…ne failures

A Ray head started before Hyperloom without the serving_slot resource was
reused as-is, so every serving lease failed its feasibility check and the
baseline failed three times before the run ended as baseline_failed.

- When the connected head lacks serving_slot and it is a lone, idle head on
  this host (no explicit RAY_ADDRESS, not multi-node, one live node that is
  the local raylet, no resources in use), restart it the way the installer
  does: ray stop --force, then ray start --head with the resource. Any other
  cluster keeps the existing error, now saying why it was not restarted.
- Tag RayInfeasibleError messages with a stable marker. Two consecutive
  subprocess_nonzero baselines with the same output (modulo numbers) now
  stop the run at two; an infeasible Ray cluster stops at two even while the
  enablement lane is open and is not handed to it as a launch log.
- Docs: the manual ray start commands declare serving_slot.

Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
…whole failure text

- ray stop --force stops every head on the host, so require exactly one
  visible gcs_server, and treat live actors (including resourceless ones)
  and other live drivers as work that forbids the restart.
- RAY_ADDRESS=local starts a new cluster on ray.init, so it no longer
  qualifies for the repair.
- Leases now connect, check feasibility and create their actor under the
  same lock the repair holds.
- The baseline failure signature hashes the whole normalised text instead
  of its tail.

Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
…ore GPUs

- The only gcs_server on the host must be the connected cluster's (matched by
  GCS port), every node record counts (a dead worker is still a worker), and
  another driver is identified by job id alone (pids repeat across PID
  namespaces).
- Check serving_slot before the GPU count so a hand-started head short of
  both reaches the repair.

Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
…e host

ray stop --force also stops a worker node of an unrelated cluster, which runs
a raylet but no GCS server. Require the only raylet on the host to be the
connected node (by --node_id) as well as the only GCS server.

Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
…e actor

Actor creation is asynchronous, so a specialist submitted a moment ago can be
missing from the GCS actor table and hold no resources yet. Count lease actor
submissions under the startup lock and refuse the restart while any exist.

Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
Waiting for a second one was not bounded: the first failure opens the
enablement lane, and with no launch log to author against that lane never
schedules another baseline. What the runtime could repair it already did, so
stop on the first, like the AgentX preflight.

Co-Authored-By: Claude <noreply@anthropic.com>
Signed-off-by: zoroyihan7 <Yihan.Wang@amd.com>
@zoroyihan7
zoroyihan7 requested a review from a team as a code owner October 10, 2026 14:49
@zoroyihan7 zoroyihan7 added the type:bug Something is broken or not working as expected label Oct 10, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

type:bug Something is broken or not working as expected

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant