Skip to content

Latest commit

 

History

History
125 lines (92 loc) · 3.19 KB

File metadata and controls

125 lines (92 loc) · 3.19 KB

l9gpu-collector

Pre-built OpenTelemetry Collector distribution bundling:

  • Standard OTLP receivers/exporters
  • batch, memorylimiter, k8sattributes, resource processors
  • k8sprocessor — per-GPU pod attribution (this repo)
  • slurmprocessor — Slurm job enrichment (this repo)
  • healthcheck + zpages extensions
  • prometheus receiver (for scraping DCGM exporter, NIM, vLLM, etc.)

Shipped as a single container image and a set of signed tarballs.


Install

Docker

docker run --rm \
  -v $PWD/config.yaml:/etc/l9gpu/config.yaml:ro \
  -p 4317:4317 -p 4318:4318 \
  ghcr.io/last9/l9gpu-collector:latest

Use :latest for tracking the newest release, or pin to a version:

docker pull ghcr.io/last9/l9gpu-collector:v0.1.0

The docker image is published for linux/amd64 only. For linux/arm64, use the binary tarball below (a native arm64 image will be added once an arm64 build runner is available).

Kubernetes

Deploy as a Deployment or DaemonSet. Point the image at ghcr.io/last9/l9gpu-collector and mount your config from a ConfigMap at /etc/l9gpu/config.yaml (or wherever --config points).

Binary (Linux / macOS)

Tarballs are attached to every GitHub release:

VERSION=v0.1.0
OS=linux       # or darwin
ARCH=amd64     # or arm64

curl -sSLO https://github.com/last9/gpu-telemetry/releases/download/collector-${VERSION}/l9gpu-collector_${VERSION#v}_${OS}_${ARCH}.tar.gz
tar -xzf l9gpu-collector_${VERSION#v}_${OS}_${ARCH}.tar.gz
./l9gpu-collector --config=config.yaml

Config

A complete working example is in deploy/collector/config.example.yaml.

Minimum viable config:

receivers:
  otlp:
    protocols:
      grpc: {endpoint: 0.0.0.0:4317}
      http: {endpoint: 0.0.0.0:4318}

processors:
  k8sattributes: {}
  k8s: {}
  batch: {}

exporters:
  otlp:
    endpoint: ${env:OTEL_EXPORTER_OTLP_ENDPOINT}
    headers:
      Authorization: ${env:OTEL_EXPORTER_OTLP_HEADERS}

service:
  pipelines:
    metrics:
      receivers: [otlp]
      processors: [k8sattributes, k8s, batch]
      exporters: [otlp]
    logs:
      receivers: [otlp]
      processors: [k8sattributes, k8s, batch]
      exporters: [otlp]

Building from source

If you need a different component set, edit deploy/collector/builder-config.yaml and rebuild with ocb:

go install go.opentelemetry.io/collector/cmd/builder@v0.126.0
builder --config deploy/collector/builder-config.yaml
./_build/l9gpu-collector --config=config.yaml

Why ship a pre-built distribution?

The OTel-standard way to get custom processors into a collector is for every user to run ocb with their own manifest. That's correct for advanced users but a poor first experience:

  • Requires the Go toolchain
  • Requires understanding ocb semantics
  • Produces an untested binary

Shipping a pre-built distribution — same pattern as Grafana Alloy, Splunk OTel Collector, AWS ADOT — gives new users a docker run on-ramp while advanced users can still roll their own with the builder-config in this repo.