Skip to content

Latest commit

 

History

History
44 lines (32 loc) · 1.52 KB

File metadata and controls

44 lines (32 loc) · 1.52 KB

Local test rig — no GPU required

gpuwaste synth writes CSVs directly, so it tests the analyzer but not the extract path. This rig runs a fake dcgm-exporter behind a real Prometheus, using the same label spellings a live cluster produces — so the query, label-mapping and join code all get exercised for real.

Runs on any machine, Apple Silicon included. No NVIDIA hardware involved.

cd dev
docker compose up -d

# let it accumulate a few minutes of scrapes.
# time is compressed: 10 real seconds = 1 simulated hour, so a full
# day/night cycle passes every 4 minutes.
sleep 300

gpuwaste extract --prometheus http://localhost:9090 --days 1 --step 15 --out ./local
gpuwaste analyze --data ./local --step 15 --cluster local-test

Expect the same seven findings synth produces. If extract returns them, the Prometheus path works and the only thing left untested is your cluster's specific label schema.

Check the raw metrics directly:

curl -s localhost:9400/metrics | head -30
open http://localhost:9090        # Prometheus UI

Tear down with docker compose down.

What this does and doesn't prove

Does: the range queries are correct, DCGM label variants resolve, the utilization/allocation join works, fb_total is derived properly, and uniform sampling is preserved end to end.

Doesn't: your cluster's actual label spellings, whether DCGM profiling is enabled there, or how kube-state-metrics is configured. Those only surface against the real thing — which is the point of running it internally.