Long video generation and world-action modeling research from NVIDIA. Each project has its own directory, code, documentation and model weights.
| Directory | What it is | Use it when you want to | Venue |
|---|---|---|---|
Long-WAM/ |
Long-context world-action models | Train and evaluate robot policies, generate robot videos, or deploy on YAM, Franka and Unitree G1 | arXiv |
LongLive-Plug/ |
Once-for-all distillation for video generation | Distill a capability once on a base model and reuse it across downstream models, without retraining | arXiv |
LongLive2.0/ |
An NVFP4 parallel infrastructure for long video generation | Train or serve long-video models fast, with NVFP4 quantization and sequence parallelism | arXiv |
LongLive1.0/ |
Real-time interactive long video generation | Type prompts and watch a long video appear in real time, steered as you go | ICLR 2026 |
Each project is self-contained. Download only the project you need using partial clone and sparse checkout; the example below selects Long-WAM.
git clone --filter=blob:none --sparse --single-branch --branch main --depth 1 https://github.com/NVlabs/LongLive.git
cd LongLive
git sparse-checkout set Long-WAM
cd Long-WAMFor another project, replace Long-WAM in the last two commands with
LongLive-Plug, LongLive2.0 or LongLive1.0. Only the selected project and
shared root files are downloaded; no other project is required to run it.
Click a thumbnail to watch on YouTube.
- π₯ [2026.10.08] The Long-WAM paper is available on arXiv.
- π₯ [2026.10.07] We release Long-WAM: long-context world-action models, robot-video pretraining, benchmark training/evaluation and real-robot deployment demos. β
Long-WAM/Β· Project page - π₯ [2026.09.28] We release LongLive-Plug training and inference code with four recipes covering CFG and few-step distillation. β
LongLive-Plug/ - π₯ [2026.07.08] LongLive 2.0 supports FP8 inference. Please refer to here.
- π₯ [2026.06.01] We released LongLive-RAG, a general retrieval-augmented framework for long video gen.
- π₯ [2026.05.30] LongLive 2.0 now supports I2V AR teacher-forcing training and I2V DMD distillation for Wan2.2-TI2V-5B.
- β‘ [2026.05.25] We optimized the NVFP4 inference path with fused Triton RoPE/adaLN kernels, reduced KV-cache synchronization overhead, in-place quantized KV-cache updates, faster FP4 KV dequantization, pinned VAE transfers, and safer LoRA-before-quantization setup, improving overall throughput by 18.6%.
- π₯ [2026.05.13] We release LongLive 2.0, infra with NVFP4, parallelism and multi-shot for AR training, DMD distillation, and inference (β‘45.7 FPS).
- π₯ [2026.04.12] LongLive supports kv cache compression with TriAttention, with 50% KV reduction and no quality drop. Check it here
- π [2026.01.27] LongLive is accepted by ICLR-2026.
- π₯ [2026.01.11] LongLive supports adapting LongLive's original RoPE into KV-cache relative RoPE and generates infinite long videos!
- π₯ [2025.11.03] We implement LongLive on linear attention model SANA-Video! Now SANA-Video can generate 60s interactive videos in real-time.
- π₯ [2025.09.29] We release Paper, this GitHub repo LongLive with all training and inference code, the model weight LongLive-1.3B, and demo page Website.
Public checkpoints on Efficient-Large-Model. See download and inference instructions.
| Model family | Released variants | Use |
|---|---|---|
| Long-WAM LIBERO | IDM Β· COD | LIBERO policy inference |
| Long-WAM RoboTwin 2.0 | IDM Β· COD | RoboTwin 2.0 policy inference |
| Long-WAM RoboCasa GR1 | 0 s Β· 2.4 s Β· 4.8 s Β· 9.6 s Β· 19.2 s | Context-length scaling |
| Long-WAM RoboCasa365 | RoboCasa365 | Generalist kitchen policy |
| Long-WAM YAM | YAM | Bimanual pretraining on ABC 130K |
| LongLive 2.0 Robot | Robot-S Β· Robot-M | Autoregressive robot-video generation |
LongLive-Plug model collection
| LongLive-Plug | Directory | Supported Models |
|---|---|---|
| LongLive-Plug-MiniMax-H3-few-step | LongLive-Plug/ |
H3-World, Code World Model, SolarWM-H3, Fun ControlNet-Union, LineartAnime, ... |
| LongLive-Plug-MiniMax-H3-cfg | LongLive-Plug/ |
MiniMax-H3 (base model), SolarWM-H3 |
| LongLive-Plug-Wan2.1-T2V-14B-few-step | LongLive-Plug/ |
FantasyWorld, DreamZero, Fun Control, Wan-Move, MagicTryOn, ... |
| LongLive-Plug-Wan2.1-T2V-14B-cfg | LongLive-Plug/ |
FantasyWorld, DreamZero, Fun Control, Wan-Move, MagicTryOn, ... |
| LongLive-Plug-Wan2.2-TI2V-5B-few-step | LongLive-Plug/ |
Matrix-Game 3.0, SCOPE, Fast-WAM LIBERO, Kiwi-Edit, Ovi, ... |
| LongLive-Plug-Wan2.2-TI2V-5B-cfg | LongLive-Plug/ |
Matrix-Game 3.0, SCOPE, Fast-WAM LIBERO, Kiwi-Edit, Ovi, ... |
Model examples are from the paper appendix, Complete Transfer Coverage and Additional Cases; each row lists up to five examples. Wan coverage includes both few-step and CFG transfer. MiniMax-H3 few-step and CFG adapters are used separately.
| Model | Directory | FPS β | Params | VBench β | Multi-shot |
|---|---|---|---|---|---|
| LongLive-2.0-5B | LongLive2.0/ |
24.8 | 5B | 85.06 | β |
| LongLive-2.0-5B-NVFP4-4Step | LongLive2.0/ |
29.7 | 5B | 84.51 | β |
| LongLive-2.0-5B-NVFP4-2Step | LongLive2.0/ |
45.7 | 5B | 83.14 | β |
| LongLive-1.3B | LongLive1.0/ |
20.7 | 1.3B | 84.87 |
FPS counts generated output frames per second of diffusion time and excludes VAE decoding, following the LongLive 1.0 and Self Forcing protocol; the LongLive-2.0 paper reports end-to-end times separately. The LongLive-2.0 rows were measured on one GB200 (180 GB), as in Table 3 of the LongLive-2.0 paper, and the NVFP4 rows include KV-cache quantization (inference.kv_quant: true, as in the released NVFP4 config). The LongLive-1.3B row was measured on one H100.
Long-context robot policies with streaming visual memory, built on LongLive 2.0 Robot video pretraining.
- Benchmarks: LIBERO, RoboTwin 2.0, Domino, RoboCasa GR1 and RoboCasa365.
- Inference: streaming IDM/COD, configurable Agentic API and image-to-video demos.
- Deployment: RTX 5090, DGX Spark and AGX Thor runtime profiles; YAM, Franka and Unitree G1 interfaces.
| Generated robot video | YAM long-horizon task | Unitree G1 cup stacking |
|---|---|---|
![]() |
![]() |
![]() |
β Paper Β· Code and quick start Β· Checkpoints Β· Project page Β· Video
Separates reusable capabilities from task-specific customization: train functional LoRAs on a base model, then attach them to compatible downstream models while keeping those models' task-specific weights. Supports Wan2.1-14B, Wan2.2-TI2V-5B and MiniMax-H3.
β Code and documentation in LongLive-Plug/
Training and inference infrastructure built around NVFP4 quantization and sequence parallelism.
- For training, it supports
- Balanced sequence parallel for T2V/I2V AR training (teacher-forcing).
- T2V/I2V AR training on multi-shot (or single-shot) videos.
- NVFP4 (or BF16) for both AR training and few-step distillation.
- For inference, it supports
- NVFP4 inference (W4A4) and NVFP4 KV Cache.
- TorchAO FP8 PTQ inference (W8A8) from the BF16 checkpoint.
- Multi-shot attention sink.
- Sequence parallel inference.
- Async decoding.
β Code and documentation in LongLive2.0/
Accepts sequential user prompts and generates the corresponding video in real time, so a person can steer a long video while it is being produced. The key ideas are the attention sink, KV-recache, and streaming long tuning.
β Code and documentation in LongLive1.0/
- DreamForge-World 0.1: Adapts the LongLive AR video stack with a residual action pathway for low-compute real-time controllable world modeling.
- DreamX-World 1.0: Follows LongLive by adapting the model on long sequences with long rollouts and local temporal windows for stable long-horizon AR world generation.
- SANA-Video: Combines SANA-Video with LongLive to build LongSANA, a real-time minute-long video generation variant with constant-memory KV cache.
- Daydream Scope: Wraps LongLive as a streaming AR video diffusion pipeline for interactive text-to-video and video-to-video workflows.
- MemFlow: Builds on the LongLive codebase and adds adaptive memory retrieval for more consistent long narrative video generation.
- ShotStream: Builds on LongLive's distillation procedure for real-time streaming multi-shot AR video generation.
- Stream-T1: Builds on LongLive's codebase and algorithm, adding test-time scaling with noise propagation, reward pruning, and memory sinking.
- KVPO: Builds on LongLive and related AR video codebases to perform GRPO-style alignment through historical KV semantic exploration.
- LoL: Builds on LongLive to study and mitigate sink-collapse for ultra-long AR streaming video generation.
- TriAttention: Integrates trigonometric KV-cache compression into LongLive's causal inference pipeline, reducing KV memory inside LongLive's local-attention window.
- StreamEdit: Provides a
LongLive_StreamEditimplementation for training-free streaming video editing built on the LongLive 1.0 codebase. - LongLive-RAG: A general retrieval-augmented framework for long video generation.
- Streaming Autoregressive Video Generation via Diagonal Distillation: Builds on the LongLive codebase and supports direct initialization from
LongLive-1.3Bcheckpoints for streaming AR video distillation. - Forcing-KV: Adds hybrid KV-cache compression to LongLive, including LongLive inference and interactive-generation scripts.
- Dummy Forcing: Unifies Self-Forcing, LongLive, and Causal-Forcing pipelines with LongLive inference, VBench, and interactive-generation configs.
- MemRoPE: Uses LongLive as a supported base model for training-free infinite video generation with evolving memory tokens.
- Astrolabe: Supports LongLive as a distilled autoregressive video backbone with LongLive-specific RL configs and LoRA initialization.
- OPSD-V: Post-trains LongLive with cache-aware on-policy self-distillation to improve long-horizon visual quality and motion dynamics while preserving its few-step autoregressive inference pipeline.
Released under the Apache License 2.0. Each directory also carries its own copy of the license and, where applicable, its own third-party notices.
Please consider citing our work if you find it useful:
Long-WAM: Scaling the Context of World-Action Models
@misc{huang2026longwamscalingcontextworldaction,
title={Long-WAM: Scaling the Context of World-Action Models},
author={Wei Huang and Bohan Zhang and Chenzhi Liu and Isabella Liu and
Shuai Yang and Weian Mao and Luozhou Wang and Yicheng Xiao and
Weifeng Lin and Qixin Hu and Bryan Chu and Sifei Liu and
Linxi Fan and Xiaojuan Qi and Song Han and Yukang Chen},
year={2026},
eprint={2610.10528},
archivePrefix={arXiv},
primaryClass={cs.RO},
url={https://arxiv.org/abs/2610.10528},
}@misc{yang2026longliveplug,
title = {LongLive-Plug: Once-for-All Distillation for Video Generation},
author = {Shuai Yang and Luozhou Wang and Wei Huang and ZhiFei Chen and
Bohan Zhang and Xiao Fu and Qianli Ma and Chen-Hsuan Lin and
Weian Mao and Bryan Chu and Song Han and Yukang Chen},
year = {2026}
}@article{longlive_2.0,
title={LongLive2.0: An NVFP4 Parallel Infrastructure for Long Video Generation},
author={Chen, Yukang and Wang, Luozhou and Huang, Wei and Yang, Shuai and Zhang, Bohan and Xiao, Yicheng and Chu, Ruihang and Mao, Weian and Hu, Qixin and Liu, Shaoteng and Zhao, Yuyang and Mao, Huizi and Chen, Ying-Cong and Xie, Enze and Qi, Xiaojuan and Han, Song},
journal={arXiv preprint arXiv: 2605.18739},
year={2026}
}@inproceedings{longlive,
title={Longlive: Real-time interactive long video generation},
author={Yang, Shuai and Huang, Wei and Chu, Ruihang and Xiao, Yicheng and Zhao, Yuyang and Wang, Xianbang and Li, Muyang and Xie, Enze and Chen, Yingcong and Lu, Yao and others},
booktitle={ICLR},
year={2026},
}Related project:
@article{longlive_rag,
title = {LongLive-RAG: A General Retrieval-Augmented Framework for Long Video Generation},
author = {Hu, Qixin and Yang, Shuai and Huang, Wei and Han, Song and Chen, Yukang},
journal = {arXiv preprint arXiv:2606.02553},
year = {2026}
}- Self-Forcing: the AR training codebase and formulation we build upon.
- Wan2.1: the video diffusion backbone used in LongLive 1.0 and LongLive-Plug.
- Wan2.2: the base video diffusion model components used in this release.
- MiniMax-H3: the audio-video generation backbone used in LongLive-Plug.










