Tensorizer converts PyTorch modules into a single .tensors file that can be
streamed from HTTP or S3 at wire speed.
python examples/tensorizer/serialize_and_load.py --local-onlyThe script serializes a tiny GPT‑2 model, serves it over HTTP, and lazily loads it back into a fresh module.
TensorSerializer.write_modulecreatestiny-gpt2.tensors.upload_to_s3(optional) pushes the file to CoreWeave Object Storage.TensorDeserializer(..., device=..., lazy_load=True, num_readers=8)streams the model directly to CPU or GPU memory.- KNative/KServe benefit from faster cold starts because weights are fetched on demand rather than baked into the container.
Deserialization is network-bound, so throughput tracks your link speed. The
often-quoted ~5GB/s on 40GbE (GPT-J, 20GB) and the letter-value plot in the
README are CoreWeave's published benchmarks for upstream
tensorizer (release 2.5.0 methodology, examples/benchmark_buffer_size). They
have not been independently reproduced in this fork, and this deployment
layer ships no benchmark of its own. Measure your own cluster before quoting a
number — see observability.md for what to watch.