This document belongs to the deployment layer added by the
TylrDn/tensorizer fork; the tensorizer
library itself is upstream
CoreWeave's. See
About this fork for the split.
The layer demonstrates a CoreWeave-aligned stack for high performance model serving. It combines Slurm on Kubernetes (SUNK), Tensorizer, vLLM, and CoreWeave observability to provide fast, reproducible, and secure deployments.
Want the whole path in one page? Read the full-stack example — serialize, publish, deploy, serve, verify — including a values table for the Helm chart, a troubleshooting table, and the layer's known gaps.
# clone and enter the repo
git clone https://github.com/TylrDn/tensorizer
cd tensorizer
# run the Tensorizer demo
python examples/tensorizer/serialize_and_load.py --local-only- Deploy SUNK to a Kubernetes cluster and schedule a pod from a Slurm job.
- Serialize and host a model using Tensorizer and CoreWeave Object Storage.
- Launch vLLM pointing at the tensorized weights via the provided Helm chart.
- Observe GPU and network metrics in CoreWeave Grafana dashboards.
The following documents provide more detail: