Repository navigation
fix(libero): fail loudly on tasks with empty init-state sets - #267
Open
akushonkamen wants to merge 1 commit into
Open
akushonkamen wants to merge 1 commit into
akushonkamen wants to merge 1 commit into
Conversation
Shipped LIBERO-PRO perturbation suites contain tasks whose pruned_init files hold zero states (libero_10_task#2, libero_spatial_task#3/RLinf#7). make_env() indexed straight into those sets: seed % 0 died on a cryptic ZeroDivisionError, and any aggregation path that tolerates the crash silently shrinks the evaluation denominator (100 -> 90/80), making results incomparable. Guard make_env(): when the requested task ships 0 init states, raise a RuntimeError naming the suite, the task id and every empty task in the suite (best-effort scan), pointing at the HF-snapshot restore path in pro_hybrid_guide.md section 2.3. Document the currently known empty pruned_init files there. Offline unit tests cover the guard with a fake suite (no dataset needed) and pin the healthy-task offset/reset-id math. Fixes RLinf#246
Codex Review SummaryThis comment shows the latest Codex review activity on this pull request.
ℹ️ About Codex in GitHubYour team has set up Codex to review pull requests in this repo. Reviews are triggered when you
Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #246
Description
make_env()inrobots/libero/env_server.py:115-117indexed straight intoget_task_init_states(...)with no guard for tasks whose pruned_init file ships zero states:seed % len(...)degenerated toseed % 0, surfacing only as a context-freeZeroDivisionErroratenv_server.py:117;Fix:
make_env(): when the requested task has 0 init states, raise aRuntimeErrornaming the suite, the requested task id, and every empty task in the suite (best-effort scan viaget_num_tasks(); API signature verified againstRLinf/LIBERO-PRO@rpent,liberopro/liberopro/benchmark/__init__.py:269-295).libero_10_task#2,libero_spatial_task#3/#7) plus the HF-snapshot restore procedure inpro_hybrid_guide.md§2.3.tests/unit_tests/robots/libero/test_env_server_init_states.py.Data side (BDDL consistency check + pruned_init regeneration) belongs to the upstream LIBERO-PRO simulation pipeline per the issue triage and was not done locally; a
timeout 60probe found no LIBERO entries under~/.cache/huggingface, so real-dataset reproduction was not executed here (the issue-sanctioned fake-suite path covers the guard).Testing
venvs/rlinf1662/bin/python -m pytest tests/unit_tests/robots/libero/test_env_server_init_states.py -q→ 2 failed (test_make_env_fails_loudly_on_empty_init_states,test_make_env_error_lists_every_empty_task_in_suite, bothZeroDivisionErroratenv_server.py:117), 1 passed.timeout 300 venvs/rlinf1662/bin/python -m pytest tests/unit_tests -k "env_server or init_state" -q --continue-on-collection-errors→ 4 passed, 1 skipped (the flag is needed because of a pre-existing, unrelated collection error intests/unit_tests/rpent/planner/test_codex_contracts.py/openai_codex). Re-confirmed right before submission with the same numbers.timeout 590 venvs/rlinf1662/bin/python -m pytest tests/unit_tests -q --continue-on-collection-errors→ 14 failed, 1147 passed, 7 skipped.main+ new test file): 16 failed = the same 14 pre-existing failures + the 2 guard tests red → zero new failures introduced by the fix.timeout 120 ruff check robots/libero→ All checks passed!;ruff format --check robots/libero→ 18 files already formatted;tests/unit_tests/robots/liberopasses both as well./tmp/issue246-repro/repro_zero_states.py→RuntimeErrorlisting suite, requested task, and the full empty-task inventory for the suite.Manual verification
Not needed — offline tests fully cover the change. The guard's runtime message was exercised end-to-end via the fake-suite repro script above (no GPU/simulator/model service involved). Real-dataset reproduction was not possible locally (no HF LIBERO cache; see Description).
Checklist
mainare unchanged by this PR; see baseline diff above.)