Skip to content

Latest commit

 

History

History
47 lines (37 loc) · 1.67 KB

File metadata and controls

47 lines (37 loc) · 1.67 KB

Overview

This document belongs to the deployment layer added by the TylrDn/tensorizer fork; the tensorizer library itself is upstream CoreWeave's. See About this fork for the split.

The layer demonstrates a CoreWeave-aligned stack for high performance model serving. It combines Slurm on Kubernetes (SUNK), Tensorizer, vLLM, and CoreWeave observability to provide fast, reproducible, and secure deployments.

Want the whole path in one page? Read the full-stack example — serialize, publish, deploy, serve, verify — including a values table for the Helm chart, a troubleshooting table, and the layer's known gaps.

Five‑Minute Quickstart

# clone and enter the repo
git clone https://github.com/TylrDn/tensorizer
cd tensorizer

# run the Tensorizer demo
python examples/tensorizer/serialize_and_load.py --local-only

Fifteen‑Minute Deep Dive

  1. Deploy SUNK to a Kubernetes cluster and schedule a pod from a Slurm job.
  2. Serialize and host a model using Tensorizer and CoreWeave Object Storage.
  3. Launch vLLM pointing at the tensorized weights via the provided Helm chart.
  4. Observe GPU and network metrics in CoreWeave Grafana dashboards.

The following documents provide more detail: