Skip to content

Migrate hyperloom-workload-optimizer skill to AMD-AGI/Hyperloom product repo - #1482

Closed
danielholanda wants to merge 1 commit into
AMD-AGI:mainfrom
danielholanda:dholanda/federate_skill
Closed

danielholanda wants to merge 1 commit into
AMD-AGI:mainfrom
danielholanda:dholanda/federate_skill

Conversation

@danielholanda

@danielholanda danielholanda commented Sep 10, 2026 •

Copy link
Copy Markdown

Description

This PR migrates hyperloom-workload-optimizer skill from amd/skills to AMD-AGI/Hyperloom product repo. The skill will continue to be part of the amd/skills catalog.

This enables AMD-AGI/Hyperloom to be the single source of truth (SSOT) for the hyperloom-workload-optimizer skill.

From now on amd/skills will simply mirror any changes made to this skill instead of requiring code to be changed at amd/skills.

@danielholanda
danielholanda requested a review from a team as a code owner September 10, 2026 14:45
@danielholanda

Copy link
Copy Markdown
Author

Looks like your linter is failing on the skill your team created on amd/skills. Feel free to make changes to this PR directly.

@johnl-amd

johnl-amd commented Sep 11, 2026 •

Copy link
Copy Markdown

The only red check is ruff, and it is formatting rather than anything
functional: two files want lines rejoined under this repo's 120-char limit.

uvx ruff@0.15.20 format skills/hyperloom-workload-optimizer/

That is 5 insertions and 16 deletions across scripts/preflight.py and
scripts/tests/test_preflight.py, and ruff check --force-exclude passes
cleanly afterwards. Verified on this branch.

Worth using ruff directly rather than pre-commit run here: the hook set needs
python3.11 and fails to build its virtualenv on a machine without it, whereas
ruff is a standalone binary and picks up [tool.ruff] from pyproject.toml
either way.

Everything else is green, and the skill's own checks pass: Structural, Discover
and Results. Behavioral is skipping, which is expected given machine.yml asks
for [mi300x, gpu, rocm] and that lane is dispatch-only.

lishuoshuo-amd added a commit that referenced this pull request Sep 17, 2026
The workflow and its test are named for the relationship they guard --
this folder is federated out of here -- rather than for the catalog on
the other end of it.

Both keep their own name rather than #1482's `AMD Skills Checks`, which
belongs to a workflow that really does call the catalog's skillscope
harness. This one asserts four rules itself, so borrowing that name would
show a green check for a harness that never ran.
lishuoshuo-amd added a commit that referenced this pull request Sep 21, 2026
The workflow and its test are named for the relationship they guard --
this folder is federated out of here -- rather than for the catalog on
the other end of it.

Both keep their own name rather than #1482's `AMD Skills Checks`, which
belongs to a workflow that really does call the catalog's skillscope
harness. This one asserts four rules itself, so borrowing that name would
show a green check for a harness that never ran.
xiaofei-zheng pushed a commit that referenced this pull request Sep 21, 2026
)

* feat(skills): add the catalog entry skill for Hyperloom bootstrap

amd/skills now federates every catalog skill from the product repo that
owns it, so `hyperloom-workload-optimizer` has to live here and the
catalog vendors a copy nightly.

What it holds is only the bootstrap: confirm the workspace, install the
wheel, run `/hyperloom-setup`, then hand the run to the skill that owns
it. Everything after setup already ships with the runtime --
`hyperloom-setup` for credentials and run mode, the demo skills for a
workload preset, `inference_optimizer` for the launcher gates, resume and
monitoring -- so they stay in step with the installed version by
construction.

This is the agent-facing form of examples/README.md, which stays as the
human quickstart.

The catalog copy of this skill was written against an empty workspace and
carries its own launch, resume and GPU-preflight scripts. Those are not
imported: a second launch path in the product repo would drift from the
CLI it wraps. The skill says so explicitly rather than leaving it to the
reader.

No packaging change. The entry point earns its keep before the wheel is
installed, so shipping it in the wheel would only overwrite the copy the
user installed from the catalog.

* refactor(skills): keep the catalog entry skill beside the examples it hands off to

The skill's whole job is to reach the demo skills in examples/, and its
prose is the agent-facing form of examples/README.md, so it reads better
next to both than at the repository root.

It stays out of pyproject's data-files on purpose, unlike the four demo
skills one level up: this entry point is what a user follows before the
wheel exists, so shipping it would only overwrite the copy they installed
from the catalog.

* docs(skills): require the user's go-ahead on the launch plan

The run skills report the plan before starting; the walkthrough in
amd/skills already promises the user is asked, and an unattended run that
holds the GPU for hours should not begin on a plan nobody accepted.

* test(skills): hold the federated skill to the rules the catalog validates

amd/skills imports this folder nightly and validates it there, so until
now a broken edit here would surface as a red bot pull request in that
repo, where nobody on this side is watching.

Nothing in this repo would have caught it either: the skill is markdown,
and lint.yml and tests-coverage.yml both ignore **/*.md, so the one file
that ships to the catalog was the one file no job read. The new workflow
carries no paths-ignore for exactly the reason packaging.yml carries
none.

The four assertions are the rules the catalog enforces, and the
description is the one with no room left: at 943 of 1024 characters, a
single added trigger sentence takes it to 1098. Each assertion was
checked against the break it exists for -- an over-long description, a
name that no longer matches the directory, and a moved folder, which is
the case that also needs a federation.json pull request upstream.

* rename: say federated skill, not catalog skill

The workflow and its test are named for the relationship they guard --
this folder is federated out of here -- rather than for the catalog on
the other end of it.

Both keep their own name rather than #1482's `AMD Skills Checks`, which
belongs to a workflow that really does call the catalog's skillscope
harness. This one asserts four rules itself, so borrowing that name would
show a green check for a harness that never ran.

* fix(ci): point the renamed workflow at the renamed test

The rename commit carried the git mv but not the edits inside the file,
so the workflow still named itself Catalog skill and ran a test path that
no longer existed.
@danielholanda

Copy link
Copy Markdown
Author

Closing as most of this has been addressed here: #1515

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants