A curated list of awesome tools, guides, playbooks, and resources for the NVIDIA DGX Spark, the GB10 Grace Blackwell personal AI supercomputer.
DGX Spark is a desktop machine built on the GB10 Grace Blackwell Superchip (SM 12.1 / sm_121), with 128 GB of unified CPU+GPU memory. You can link two units over 200 Gb/s networking to run larger models. This list collects community projects for setting it up, serving models, fine-tuning, benchmarking, and day-to-day operation.
Platform essentials: aarch64 · CUDA 13.x · sm_121 · 128 GB unified memory · 200 Gb/s ConnectX-7
- Official
- Setup & Configuration
- Inference & Serving
- Fine-tuning
- Quantization & NVFP4
- Models & Benchmarks
- Multi-node
- Image & Media Generation
- Audio & Speech
- Science & HPC
- Remote Access & Desktop
- Tools & Monitoring
- Operating Systems & Containers
- Community & Resource Collections
- NVIDIA/dgx-spark-playbooks - Step-by-step DGX Spark playbooks spanning vLLM, SGLang, llama.cpp, NVFP4 quantization, speculative decoding, cuTile kernels, NCCL, and two- and three-node clustering.
- a1exus/sparky - Self-hosted DGX Spark LLM stack, vLLM, Ollama, and a llama.cpp router serving every cached GGUF behind Traefik, with mDNS, Cloudflare Tunnel, and Tailscale ingress.
- Albatross1382/onnxruntime-aarch64-cuda-blackwell - ONNX Runtime 1.24.4 CUDA shared libraries for sm_121 on aarch64, loaded through the Rust ort crate or dlopen.
- botAGI/AGmind - One-command private RAG stack for DGX Spark, Dify with vLLM, Weaviate, RAGFlow, and Docling across 30+ containers, plus two-node clustering over 200 Gb/s QSFP.
- christopherowen/dgx-spark-memory-saver - Patch for nvidia-uvm that packs GPU page tables on 64 KiB DGX Spark kernels, 3.03 GiB recovered versus stock 64 KiB and about 1.8 GiB per node over 4 KiB.
- Chrizz-lab/GB10-Agentig-Coding-Framework - Agentic coding stack for DGX Spark with dual-vLLM Qwen3 and CrewAI orchestration.
- csabakecskemeti/dgx-spark-community-playbooks - Community playbook collection for DGX Spark, covering dual-Spark RDMA inference, heterogeneous RoCE clustering, and local Claude Code.
- Entrpi/dgx-spark-serving-mode - Three-state serving-mode script for DGX Spark that pares GNOME and maintenance timers down to multi-user.target for 10-15 GB more unified memory.
- Fulton-Engineering-Services/dgx-spark-wheels - Wheel index for GB10 covering packages with no upstream aarch64 build: flash-attn, sageattention, flashinfer, and torch 2.13 on CUDA 13.3, each kernel-launch verified.
- getainode/ainode - Browser-UI AI appliance for GB10 doing inference and LoRA fine-tuning, with UDP-discovered multi-node clustering, TP=2 on the current 0.5.x and TP=4 only on 0.4.x.
- GuigsEvt/dgx_spark_config - Source-build guide for LLVM, Triton, and PyTorch 2.9.1 against sm_121, with release wheels and ~1.5x on 8192 FP16 GEMM versus the stock cu130 build.
- HeKun-NVIDIA/dgx-spark-openclaw - Two-script deploy of a local LLM plus OpenClaw frontend, Qwen3.5-35B-A3B and MiniMax-M2.5-REAP-NVFP4 on a GB10 NVFP4-kernel vLLM image, or GLM-4.7-Flash on Ollama.
- HendrikSchoettle/ragflow-dgx-spark - Build and deploy pipeline for RAGFlow v0.24.0 on DGX Spark aarch64, with a source-built onnxruntime-gpu wheel for sm_121 and multilingual OCR.
- install-safe-press/gb10-playbooks - Chinese-language walkthrough of NVIDIA's official GB10 playbooks covering the basics and AI-agent sections, with hardware, DAC cabling, and Dell switch-config notes of its own.
- JetBrains-Hardware/spark-setup - Remote deploy scripts for Qwen, GPT-OSS 120B, Nemotron 3 NVFP4, and Gemma 4 on vLLM, with MTP speculative decoding holding 14.4 tok/s at 200k context.
- jschmied/dgx-spark-setup-guide - End-to-end DGX Spark guide: llama-server router mode, key-only SSH, DCGM monitoring, SWE-bench evals, and vLLM appendices where MTP-3 takes Qwen3.6-35B-A3B NVFP4 from 73 to 102 tok/s single-stream.
- m9h/neurocontainers-arm - Prebuilt causal-conv1d wheel plus four published neuroimaging containers built against NGC PyTorch CUDA 13, with locally built FreeSurfer 8.2.0 packages inside two of them.
- mARTin-B78/dgx-spark_lite-llm_llama-swap_vllm_llama-cpp_ollama - Multi-engine LLM stack for DGX Spark with llama-swap idle eviction behind a LiteLLM gateway, plus a harness pairing llama-benchy with a 69-scenario tool-calling bench.
- natolambert/dgx-spark-setup - Training setup for CUDA 13.x on aarch64, cu130 vLLM wheels, SDPA over flash-attn, swap-off OOM guards, and profiled SFT, DPO, and LoRA batch ceilings.
- seitzbg/onnxruntime-gpu-sm121-aarch64 - Prebuilt onnxruntime-gpu 1.27.1 aarch64 wheel, CUDA 13.x execution provider for sm_121, 7.3x over CPU on GB10.
- Sggin1/DGX-SPARK - Dated GB10 lab notes, TurboQuant 3-bit KV cache at 240K context, dual-Spark 195 Gb/s RDMA, and sm_121a FP4 SASS evidence.
- sjug/dgx-spark-ethernet-patch - Binary patch for the DGX Spark OOBE ethernet-detection bug, an 8-byte aarch64 HasInternet edit for FastOS 1.120.38.
- Th0rgal/cuda-blackwell-carry-bug - Repro for a PTXAS bug that drops the carry flag between separate inline-asm blocks on GB10 sm_121, with a Docker harness and a 128-bit reference path to compare against.
- timothystewart6/ubuntu-gb10 - Ubuntu 24.04 setup for GB10 in place of DGX OS, Ansible roles for NVIDIA driver, CUDA 13.x, DOCA-OFED, dual-node NCCL, and a 33-check read-only verify playbook.
- tonyd2wild/DGX-Spark-Hard-Poweroff-Fix - Diagnosis of GB10 log-less hard power-offs as an embedded-controller cut, fixed by a 2200 MHz clock cap and page-cache drops at 5% decode cost.
- 0xSero/deepseek-v4-flash-0731-spark-sparkinfer - DeepSeek-V4-Flash-0731 on one DGX Spark via EXL3 and SparkInfer sparse MLA, 38.1 tok/s median c1 code decode at a 262K limit, 432-byte NVFP4 KV disabled.
- AEON-7/vllm-ultimate-dgx-spark - DGX Spark vLLM 0.29.0 image compiled for sm_121a with DSpark quantized Markov heads, DFlash 2 at 3.39x single-stream on Qwen3.8-27B, and Triton NVFP4 KV cache.
- airawatraj/dgx-spark-nemotron-super-agent - Nemotron-3-Super-120B agentic stack on DGX Spark scoring 93 of 100 on tool-eval-bench, with spark-arena 23.7 tok/s.
- albond/DenseSpark-Qwen3.8-27B - Qwen3.8-27B INT4 AutoRound on one DGX Spark with vLLM 0.27.1 and sm_121 kernels, MTP drafting, 49.1 tok/s at one request and 260.2 at 16.
- albond/SingleSpark-Qwen3.8-Flash-Next - Qwen3.8-Flash-Next on one DGX Spark from a 101 GiB four-bit checkpoint with the n-gram table kept resident, 43.34 tok/s median single stream at 65,536 context.
- Anemll/dspark-vllm-gx10 - Two-node GB10 port of DeepSeek-V4-Flash DSpark to vLLM 0.25.1 with nvfp4_ds_mla KV format and a b12x MXFP4 MoE backend, 48.5 tok/s decode at TP=2.
- atcuality2021/vllm-gb10-gemma4 - Vendored ManthanQuant 3-bit Lloyd-Max KV-cache compression patched into vLLM's attention backends for Gemma 4 on GB10, 5.12x smaller KV at ~0.978 cosine, quantized in NumPy on the Grace CPU.
- Avarok-Cybersecurity/dgx-vllm - vLLM image for DGX Spark, tag v22 from February 2026, NVFP4 through a software E2M1 conversion and Marlin backends, ~42 tok/s on Qwen3-Next-80B-A3B-Instruct-NVFP4, 20% over AWQ INT4.
- bjk110/spark_vllm_docker - vLLM serving from one DGX Spark at TP=1 to two over 200 Gb/s RoCE at TP=2, 42 presets, pinned DeepSeek-V4-Flash-0731 and Solar-Open2-250B production images with rollbacks.
- blazux/qwen3.8-Flash-DGX - Qwen3.8-Flash-Next NVFP4 on one GB10 with the 48 GiB PLE table mmapped from NVMe, fixing a prefix-cache hit that restored an all-zero Mamba state and a non-deterministic sparse-attention top-k.
- dolf3131/qwen3.8-flash-next-dgx-spark - NVIDIA's Qwen3.8-Flash-Next NVFP4 on one DGX Spark, 47.7 GiB n-gram table paged to SSD swap, 33.0 tok/s single-stream at 524K context, plus the PLE-offload hang at TP=1.
- EmilHaase/DGX-Spark-VLLM-Hydra-Manager - vLLM manager for DGX Spark with sm_121a source builds and UMA KV-cache limits for multi-model launch.
- Entrpi/ds4-spark-vllm - One-command 2-bit DeepSeek-V4-Flash vLLM install on a single DGX Spark, 85 GiB IQ2_XXS plus Q2_K checkpoint validated against antirez/ds4 at 1.75 tok/s under enforce-eager.
- eugr/spark-vllm-docker - vLLM Docker for one to eight DGX Sparks, native PyTorch distributed by default with Ray opt-in, and 42 run-recipe.sh YAML recipes, Qwen3.8-Flash-Next-NVFP4 solo to GLM-5.2-NVFP4 on eight nodes.
- gitcommit90/glm-5.3-one-spark - GLM-5.3-Flash EXL3 2.05 bpw on one DGX Spark under vLLM TP1 with DFlash2 K5, 40.1 tok/s on code and 29.9 on prose at 262K context.
- jordanovski/overdrive - Web console and CLI for launching concurrent vLLM containers on DGX Spark, with preflight GPU-memory admission control and a SWE-bench page comparing resolution rates.
- mark-ramsey-ri/vllm-dgx-spark - Run vLLM on 1-to-N DGX Spark servers (single Spark, 2 via direct cable, or 3+ via switched fabric) to serve or benchmark LLMs.
- MiaAI-Lab/Nemotron-Labs-3-Puzzle-75B-DGX-Spark - Nemotron-Labs-3-Puzzle-75B-A9B NVFP4 hybrid Mamba MoE on one DGX Spark, vLLM 0.24 launcher with aarch64 NCCL and FlashInfer cuda_ipc patches, 256K context, MTP k=3.
- MiaAI-Lab/Ornith-1.5-35B-A3B-DGX-Spark - Ornith-1.5-35B-A3B NVFP4 with in-checkpoint MTP on one DGX Spark, 86.3 to 440 tok/s at 24 streams, plus two b12x patches for CUDA-graph capture.
- MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark - Qwen3.8-Flash-Next NVFP4 on one DGX Spark under vLLM with the PLE table packed and memory-mapped, 48.7 tok/s prose single stream and 162.9 at eight on 512k YaRN.
- mouwp2026/qwen3.8-flash-next-gx10-mtp-hashk - Qwen3.8-Flash-Next on one GB10 with the PLE table compressed to a 12.8 GB HashK build, FP8 dense cast, and a shrunk-vocabulary MTP head, 104 tok/s at six streams.
- mrexodia/Kolibri-1-vLLM-DGX-Spark - Kolibri-1 FP8 on one DGX Spark under vLLM 0.29 with Aleph Alpha's inference plugin, 262,144-token context, 42.0 tok/s generation and 97% prefix-cache hits in a pi coding-agent run.
- omnia-projetcs/spark-dgx - Interactive vLLM Docker launcher for DGX Spark, 22 preset model configs from single-node NVFP4 to TP=4 Ray clusters, 10 with measured TTFT and concurrency tables.
- Sapid-Labs/vllm-spark-arena - Crowd-optimization arena for vLLM on sm_121, scoring sitecustomize.py patches over a pinned wheel as paired ratios, gated on byte-identical output and a held-out timed speedup.
- sayyidfareed/qwen3.8-flash-next-dgx-spark-1m - Qwen3.8-Flash-Next NVFP4 at a validated 989,801-token request on one ASUS GX10, clean PLE pages released by MADV_DONTNEED watermark, 5/5 needles at 26.7 tok/s single stream.
- spark-arena/sparkrun - One-command launcher for vLLM, SGLang, and llama.cpp on one or more DGX Sparks, where --tp 2 means two hosts over auto-detected RDMA, plus git-based recipe registries.
- sudoingX/dgx-spark-ling - Official Ling-3.0-flash INT4 on one DGX Spark under vLLM with its MTP layer, 38.7 tok/s against 35.2 for the community GGUF, but 7.9 against 33.6 at 45K context.
- timothystewart6/vllm-gb10 - Prebuilt GHCR vLLM image for GB10 (sm_121a) with NCCL and FlashInfer built from source, pinned by SHA, digest or version, and
latestpromoted only after a four-model verification gate. - tonyd2wild/Qwen3.8-Flash-Next-NVFP4-DGX-Spark - NVIDIA's own Qwen3.8-Flash-Next NVFP4 checkpoint byte for byte on one DGX Spark, the 47.68 GiB PLE table left on NVMe and gathered 16 rows per token, 43.9 tok/s median.
- 0xBakeer/qwen38-flash-next-spark - Qwen3.8-Flash-Next on one DGX Spark with the 51B lookup table on SSD, in two profiles that differ by 2x: 88 tok/s rewriting a file against 32 on prose.
- cahlen/glm-5.3-flash-GGUF-1bit-dgx-spark - GLM-5.3-Flash as a 1-bit UD-IQ1_S GGUF on one DGX Spark for agentic coding, where leaving
--reasoning-budgetunset returns nothing on 35% of turns against 0 of 80 with it. - croll83/llama.cpp-dgx - Deprecated llama.cpp fork for DGX Spark, kept for its TurboQuant TQ3_0 KV cache and weight kernels, which upstream does not expose.
- gitcommit90/angelslim-hy3-iq1m-mtp-dgx-spark - AngelSlim Hy3 IQ1_M GGUF with MTP drafting on one DGX Spark through patched llama.cpp, 100K context, 18-19 tok/s on prose and 26-30 on code and JSON.
- marknx/flash-next-gguf-tools - Qwen3.8-Flash-Next GGUF split across one DGX Spark and an RTX 5090 over 10 GbE, 140 tok/s aggregate at eight streams that OOM-kill the single box, plus converter fixes.
- phuongncn/qwen3.6-27b-speedhack-gx10-dgx-spark - DFlash block-diffusion spec-decode llama.cpp fork for Qwen3.6-27B on GB10, 7-11 to 38-40 tok/s coding via a p_min drafting threshold, and 60-66 to 113 on a 35B-A3B MoE.
- Sapid-Labs/llamacpp-spark-arena - Crowd-optimization arena for llama.cpp CUDA kernels on sm_121, with a thermal gate, alternating baseline and candidate runs, and referee-verified held-out speedup.
- shamily/gemma4-llama-dgx-spark - Dockerized llama.cpp for all four Gemma 4 models on GB10, benchmarked at 69.9 tok/s tg128 for the 26B-A4B MoE against 11.0 for the dense 31B.
- sxuff/ternary-bonsai-2-27b-gx10 - Ternary-Bonsai-2-27B on one DGX Spark under the PrismML llama.cpp fork, PTQ1_0 (1.75 bpw) at 34.21 tok/s tg128 in 5.53 GiB, build pinned to commit 1a07bfa.
- Weschera/GLM-5.3-Flash-Unsloth-1x-DGX-Spark - GLM-5.3-Flash UD-IQ3_XXS GGUF on one DGX Spark under Unsloth's llama.cpp fork, MTP n=2 at 20.8 tok/s against 15.5 without, 64K context.
- 0xBakeer/ling3-flash-spark - Ling-3.0-flash MXFP4 on one DGX Spark pairing the Humming MoE backend and an online FP8 LM head with DSpark drafting, with the cookbook and notebook configs selectable for comparison.
- 0xWhiteMage/qwen3.8-27b-kearuga-sglang-dgx-spark-dflash2 - Qwen3.8-27B Kearuga target with a distilled DFlash 2 drafter on SGLang for one DGX Spark, 35.33 tok/s at C1 and 108.90 at C4, +14.5% and +10.8% over stock.
- BTankut/dgx-spark-sglang-moe-configs - Tuned Triton MoE configs for GB10's 101,376-byte shared memory limit, where SGLang defaults need 147,456 and EAGLE crashes, GLM-4.7-FP8 at 20-27 tok/s on four nodes.
- hasso5703/dgx-spark-qwen38 - One-command SGLang service on GB10 for seven switchable Qwen3.8 targets, Qwen3.8-27B NVFP4 at 71.4 tok/s greedy single-stream median via DFlash2 and Qwen3.8-Flash-Next 176B at 262K context.
- InquiringMinds-AI/longcat-next-multimodal - LongCat-Next 75B-A3B any-to-any multimodal through one SGLang process on a single GB10, image generation and voice-clone TTS on OpenAI endpoints at w8a8_int8 after 4-bit collapsed both.
- mark-ramsey-ri/sglang-dgx-spark - Run SGLang on 1-to-N DGX Spark servers (single Spark, 2 via direct cable, or 3+ via switched fabric) to serve or benchmark LLMs.
- MiaAI-Lab/Nemotron3.5-Lightning-DGX-Spark-RTX-5090-6000-PRO - Nemotron 3.5 Lightning 30B-A3B NVFP4 with its DSpark draft model on one DGX Spark via SGLang, 4.93M-token KV pool, 48 max concurrent, up to 1M per request.
- MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark - Qwen3.8-27B NVFP4 on one DGX Spark with EAGLE/MTP, DSpark and DFlash2 as measured swap-in modes, GDN in bf16 and the scheduler pinned to the X5 cores.
- pangoleen/qwen3.8-27b-dgx-spark-dflash2 - Qwen3.8-27B NVFP4 on one DGX Spark with a DFlash2 draft budget of 16, 64-78 tok/s single stream on code against ~30 on chat, 387 aggregate at 32 streams.
- robbiemu/dgx-spark-inference - SGLang services on one DGX Spark under systemd with per-role memory budgets, refusing a launch that will not fit, on a digest-pinned v0.5.14-cu130 runtime.
- scottgl9/sglang-spark-gb10-optimizations - SGLang fork that routes NVFP4 through Marlin FP4 where CUTLASS returns zeros on sm_121, Qwen3.5-122B-A10B at 43-45 tok/s with ~90% MTP acceptance.
- ubehera/sglang-spark - SGLang patch stack for DGX Spark with an sm_121a-only sgl-kernel 0.4.4 wheel, model-agnostic TP=2 launch recipes at NEXTN k=3, and a systemd watcher that stops wedged nodes.
- Weschera/Qwen3.8-27B-NVFP4-DFlash2-DGX-Spark - Qwen3.8-27B NVFP4 with DFlash2 on one DGX Spark, digest-pinned, where a boot-time fp8_gemm autotune race decides between 42 and 33 tok/s for the process lifetime.
- 0xBakeer/deepseek-v41-flash-spark - DeepSeek-V4.1-Flash on one DGX Spark in a plain-PyTorch engine with routing-ranked expert keep-sets, 24.3-36.6 tok/s by workload at 44% keep, above the reliable 39% default.
- 0xBakeer/TandemLLM - Inference engine for Qwen3.8-27B on one DGX Spark with StairCut draft-tree sizing and custom NVFP4 kernels, 49.89 tok/s single request with the lookup store off against 13.69 plain greedy.
- antirez/ds4 - DwarfStar C inference engine for DeepSeek, GLM and Qwen3.8 Flash Next with a
make cuda-sparktarget, DeepSeek V4 Flash Q2 prefill above 820 t/s through 65K, decode 18.1 to 13.8. - ashhart/TensorFold - Speculative-decoding LLM server for Apple Silicon and NVIDIA with replies equal to serial decoding, CUDA engines for one or two DGX Spark, GLM-5.3-Flash at 256k tokens on two.
- Avarok-Cybersecurity/atlas - Pure Rust and CUDA inference engine in one 75 MB binary with GB10 as verified target, ahead of vLLM 0.27.1 at every C=1 to C=128 rung on Qwen3.8-27B NVFP4.
- Baekpica/ds4-dfm-rs - Rust-host continuation of ds4 with DGX Spark as release reference and recorded gates per model artifact, Qwen3.8 Flash Next Q5 at 1,323 prefill and 28.93 decode tok/s with MTP 2.
- blake-snc/sm121-kernels - Hand-written PTX kernel library for sm_121 in 259 files, covering flash attention, GEMM, Gated DeltaNet, and MoE, driver-only via cudarc with FP8 attention at ~108 TFLOPS.
- calico88x/DGX-Model-Manager - Control plane for managing Ollama, SGLang, vLLM, llama.cpp, LocalAI, and ComfyUI on DGX Spark, with roles, API tokens, and Hugging Face cache inventory.
- HawkBearPig/dgpp - C++/CUDA inference engine for one, two, or four DGX Spark with tensor parallelism over RoCE, 29 benchmarked configurations across GLM, Qwen, DeepSeek, and MiMo models.
- jdaln/dgx-spark-inference-stack - Docker serving stack for a single DGX Spark with on-demand model loading, automatic idle shutdown, and a unified API gateway.
- joshhu/meetaclawtaipei - Three concurrent NVFP4 vLLM models on one DGX Spark with a 3-LLM voice-clone roommate demo.
- kshetrajna12/sparkstation - Headless control plane for DGX Spark model fleets with profile-driven placement over vLLM and SGLang, unified-memory admission control, and LiteLLM gateway sync.
- lrozewicz/vLLM-Moet-GB10 - vLLM-Moet fork for one GB10 running GLM-5.3-Flash at ~29 tok/s on code and DeepSeek-V4-Flash-0731 at ~21, via 2-bit expert planes plus an FP4 delta tier.
- mark-ramsey-ri/trt-dgx-spark - TensorRT-LLM serving for 1 to N DGX Spark with the aarch64 nvcr 1.2.1 container and TP set by node count, verified on 1 and 2 nodes.
- MiaAI-Lab/exllamav3 - ExLlamaV3 fork that builds on aarch64 GB10, with NVFP4 and FP8 paged-attention KV at about 3.5x the context per GB of fp16, plus DFlash2, DSpark, and MTP drafting.
- MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold - Qwen3.8-Flash-Next top-5-expert INT4-AutoRound checkpoint on one DGX Spark under TensorFold's Zig engine, 64.4 tok/s prose at c1 and 200.9 at c8, about 10% lower capability.
- r0b0tlab/glm53-flash-exl3-exllamav3-gb10 - ExLlamaV3 and TabbyAPI runtime for GLM-5.3-Flash EXL3 2.25 bpw with DFlash2 on one GB10, a fail-closed memory guard, and 259,993-token retrieval at 262,144 context.
- rdaum/eider - Rust and CUDA inference server for sm_121 NVFP4 with no tensor runtime, Qwen3.8-Flash-Next with its 51B PLE table paged from NVMe at ~300 tok/s uncached prefill, ~13 MTP decode.
- rdoiron/mimo-mods-for-dgx-spark - Ten vLLM runtime patches for MiMo-V2.5 on sm_121a, with a CUTLASS block-FP8 bypass and a backported tool-call corruption fix (PR #42969).
- sf-stav/veloGB10 - Rust and CUDA inference engine built only for GB10 with bitwise-lossless MTP, Qwen3.8-Flash-Next EXL3 at 137 tok/s on one machine and 221 on four, pure-code decode.
- sixteen-miles-labs/sparklab - GB10-only inference runtime for one DGX Spark with NVMe-backed MoE expert banks, GLM-5.3 753B at 1.29 decode tok/s and DeepSeek V4.1 Flash 552B at 1.05, both experimental.
- Th0rgal/dgx-spark-router - Stdlib-only OpenAI-compatible router that swaps six llama.cpp GGUF and vLLM NVFP4 models in and out of 128 GB unified memory, with Marlin GEMM defaults for GB10.
- vcruz305/DeepSeek-V4.1-Flash-EXL3-DGX-Spark-recipe - DeepSeek-V4.1-Flash EXL3 1.59 bpw on one DGX Spark in native ExLlamaV3 with the DSpark drafter aliased from mmap over ATS, 17.5 tok/s median on fresh prompts.
- xangel82/DS4-GB10-GX10-DSpark-CUDA - DS4 fork for DeepSeek-V4-Flash on one GB10, lossless DSpark and HybridLC decode at 24-26 t/s on tool calls against the original 13, 900-953 t/s prefill.
- albond/DGX_Spark_Unsloth_Lossless_Speedup - Triton kernels for sm_121a that keep the loss curve bit-identical at 7.67x LoRA and 8.35x full fine-tune over stock Unsloth on Qwen3.5-2B.
- alicankiraz1/DGX-Spark-Asus-Ascent-Nvidia-GB10-SFT-Finetuner - No-code SFT fine-tuning tool for DGX Spark.
- haven-jeon/unsloth-vllm-gb10 - Unsloth training and vLLM inference Docker image for DGX Spark GB10 with source-built xformers and Triton.
- JanneckGit/slm-agentic-training-pipeline - LoRA SFT then GRPO/verl for a Qwen3-4B tool agent on one GB10, verifier-gated synthetic trajectories over 48 task templates, with sm_121 flags where FA2 fails the actor forward.
- kreuzhofer/dgx-spark-unsloth-qwen3.5-training - BF16 LoRA on Qwen3.5-35B-A3B without quantization, eager per-shard CUDA loads and page-cache eviction that cut peak unified memory from 134 GB to 72 GB.
- NvMayMay/nvfp4-lora-spark - LoRA trained on the served NVFP4 weights of a 100B+ MoE on one GB10, with a serve-time logprob check that catches no-op adapters.
- ubehera/finetune - Unsloth LoRA fine-tune run on GB10 with five idempotent in-venv hotfixes and a committed perplexity table, 30.3 base to 14.3 after five epochs.
- waybarrios/dgx-spark-finetune-llm - LLM fine-tuning with LoRA + NVFP4/MXFP8 on DGX Spark.
GB10's Blackwell architecture supports NVFP4 (4-bit floating point) in hardware. It runs faster than INT4 at similar quality.
- 0xBakeer/Qwen3.8-27B-4-bit-on-a-single-DGX-Spark - Qwen3.8-27B NVFP4 on one DGX Spark, 75 tok/s edit-heavy and 29.6 fresh with DSpark at k=14, its k=7 edge over FP8 falling from 27% at c1 to 0.2% at c16.
- AEON-7/Gemma-4-26B-A4B-it-Uncensored-NVFP4 - NVFP4 Gemma 4 26B MoE on DGX Spark with DFlash speculative decoding, 49.8 tok/s on prose to 202.4 on extraction, and 1,937 tok/s aggregate at 64 concurrent.
- AEON-7/Gemma-4-31B-Uncensored-NVFP4-DFlash - Prebuilt vLLM image for Gemma 4 31B Deckard Heretic with z-lab DFlash k=15 and CUTLASS NVFP4, decode 11 to 38.82 tok/s at c=1.
- AEON-7/Nemotron-3-Nano-Omni-AEON-Ultimate-Uncensored - Source-built vLLM image with sm_121a patches for abliterated multimodal Nemotron-3-Nano-Omni NVFP4, refusals down from 99/100 to 16/100 with thinking off.
- AEON-7/Qwen3.6-35B-A3B-heretic-NVFP4-DFlash - Source-built vLLM image serving NVFP4 Qwen3.6-35B-A3B with DFlash speculative decoding, averaging 97 tok/s single-stream across six prompt categories.
- AEON-7/supergemma4-26b-abliterated-multimodal-nvfp4 - Plain-NVFP4 SuperGemma4-26B abliterated multimodal for DGX Spark, as a prebuilt vLLM container, one loader patch down from three in the AWQ release.
- BioInfo/turboquant-dgx - TurboQuant KV-cache quantization on GB10 with 3.88x compression and 8.4x Triton kernel speedup.
- drowzeys/keys-vLLm.0.27-Qwen3.8-27B-ADay777Ablit-NVFP4-A4Q-NVFP4-KV-4M-KV-token-pool-MTP3-Single-DGX-Spark - Abliterated Qwen3.8-27B NVFP4 on one DGX Spark with FA2 NVFP4-KV back-ported to sm_121, a 4.34M-token pool for 4.14x concurrency at 1M context, A4Q prefill +8-10% at 48-96K.
- drowzeys/vLLm-0.24-optimized-NVIDIA-Nemotron-Lab-Puzzle-75B-A9B-A4Q-MTP3-NVFP4-KV-2.7M-Pool-Single-DGX-Spark - Nemotron-Labs-3 Puzzle-75B-A9B on a single DGX Spark with an NVFP4 KV cache holding 2.71M tokens, MTP k=3, 25.8 tok/s decode and 220.3 aggregate at 32 streams.
- jethac/vllm-gemma4-nvfp4-kv-repro - Repro and wheels for Gemma-4 garbage output under
--kv-cache-dtype nvfp4on sm_120 and sm_121, traced to three open vLLM fixes: FULL cudagraph capture, KV-sharing scales, V-scale writer. - jethac/vllm-nvfp4-kv-consumer-blackwell-repro - Model-free one-command repro for the vLLM NVFP4 KV V-scale swizzle bug on sm_120 and sm_121, V rel-L2 0.667 stock versus 0.095 fixed, plus 3.556x KV capacity math.
- Libertai/vllm-sparse-mla-blackwell - Hand-written NoPE sparse-MLA kernel and vLLM plugin for sm_121 and sm_120, tiles sized to the 101,376-byte shared-memory ceiling, serving GLM-5.3-Flash at 24.04 tok/s c1 with MTP on two GB10.
- localai-org/apex-quant - MoE-aware mixed-precision GGUF recipe, Q8_0 perplexity at 21.3 GB instead of 34.4 GB and 62.3 against 52.5 t/s on Qwen3.5-35B-A3B, measured on GB10.
- Logos-Flux/optimized-CUDA-GB10 - Vectorized sm_121 RMSNorm kernel, 2.59x average over PyTorch BF16, on the Hugging Face Kernel Hub.
- mitkox/sparser-faster-llms - GB10 sm_121 CUDA-core TwELL sparse-kernel port of SakanaAI's sparser-faster-llms for DGX Spark builds without Hopper WGMMA.
- r0b0tlab/gemma4-26b-a4b-nvfp4-gb10-native-cutlass - Gemma-4-26B-A4B NVFP4 for GB10 via native VLLM_CUTLASS MoE backend on CUDA-13 nightly, 260 tok/s at concurrency 8.
- r0b0tlab/gemma4-31b-it-nvfp4-gb10 - Gemma-4-31B-IT NVFP4 reproducibility pack for GB10 with native FlashInfer/CUTLASS FP4 GEMM, 54 tok/s at concurrency 8.
- r0b0tlab/ling30vl-nvfp4-mp-sm121 - Ling-3.0-flash-VL quantized to mixed NVFP4 and FP8 for one GB10, 249.7 GB of BF16 down to 72.62 GB, served by the inclusionAI vLLM fork in an sm_121 container.
- r0b0tlab/nemotron3-super-120b-a12b-nvfp4-gb10-native-mtp - Nemotron-3-Super-120B-A12B NVFP4 for GB10 on SGLang native MTP, 21.64 tok/s and +45.8% over baseline.
- r0b0tlab/qwen36-35b-a3b-nvfp4-fast-sm121-vllm - Qwen3.6-35B-A3B NVFP4-Fast on one GB10 via sm_121-native vLLM, 80.6 tok/s single-stream to 344 tok/s at concurrency 32, with GSM8K 86.73% and 86.33% MTP acceptance over 474K draft tokens.
- r0b0tlab/qwen36-35b-a3b-nvfp4-sm121-vllm - NVIDIA Qwen3.6-35B-A3B-NVFP4 on one GB10 under vLLM 0.25.0, W4A16 targets routed to native W4A4, 93 tok/s single stream, GSM8K 86.5%, 262,144-token context qualified.
- r0b0tlab/qwen38-27b-nvfp4-sm121-vllm - Qwen3.8-27B NVFP4 on one GB10 via an sm_121 vLLM image plus DFlash2 K=8 overlay, 67.1 to 279.2 tok/s aggregate at c1 to c6, NIAH 3/3 at 262K.
- RobTand/gridbook - Product-codebook weight format with a vLLM plugin whose decoded tiles are native NVFP4 or FP8, a 295B MoE at 2.9 bpp on one DGX Spark, 2.6x GGUF prefill, slower decode.
- secYOUre/nvfp4bench - NVFP4 peak-throughput CLI for GB10 sm_121a, 1022 TFLOPS sparse and 511 dense on packed mxf4nvf4 MMA, exactly half on byte-padded mxf8f6f4.
- spped2000/thaillm-nvfp4-dgx-spark - ThaiLLM-30B quantized to NVFP4 on DGX Spark, 3.4x smaller at 63 tok/s decode, Thai accuracy flat under paired McNemar tests, MMLU down 1.4 points.
- sudoingX/dgx-spark-laguna - Laguna S 2.1 NVFP4 on DGX Spark with DFlash, 25-30 tok/s at 128K, plus hard-hang protection.
- vcruz305/glm53-flash-nvfp4-one-spark-quantize - GLM-5.3-Flash quantized from a 598 GiB BF16 source to a 177 GiB NVFP4 pack on one DGX Spark via ModelOpt layerwise PTQ and Accelerate disk offload.
- vcruz305/vllm-exl3 - Out-of-tree vLLM plugin for EXL3 trellis MoE packs serving GLM-5.3-Flash, DeepSeek-V4.1-Flash, and Qwen3.8-Flash-Next, the last at 47.6 tok/s with MTP k=2 at 65,536-token context on one GB10.
- VincentKaufmann/fp4-cuda-kernel - FP4 tensor-core GEMM as a one-line Python call on sm_121 through CUTLASS, pre-quantized weight cache at 85-129 TFLOPS and 1.0-2.4x BF16
F.linearup to M=2048, 0.7x at M=4096. - vladimir-voinea/gb10-laguna-s-2.1-w4a16-moe - W4A16 grouped-MoE kernels and GB10-tuned Triton configs for sm_121a, reaching 157-230 GB/s achieved weight bandwidth against 109-142 for stock Triton on a 117B MoE.
- albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4 - Qwen3.5-122B-A10B on DGX Spark, tuned from 28.3 to 52 tok/s (+82%) with a hybrid INT4+FP8 checkpoint and an INT8 LM head.
- Avarok-Cybersecurity/atlas-recipes - Recipe corpus for Atlas on GB10, run by the atlasctl launcher that replaces sparkrun, with per-model KV dtype and MTP width.
- Blackwellboy/laguna-s21-lab - Laguna S 2.1 NVFP4 testing lab on one DGX Spark, 20-cell tuning sweep with its losing cells, 12-hour soak of 3,096 turns, and 450-turn thinking-gate study.
- colonel-otto/3x-dgx-spark-mesh-deepseek-v4-0731-flash - DeepSeek-V4-Flash at TP=3 on three DGX Spark, decode 6.7 to 20.2% faster than two nodes at n=30 on the retired Anemll engine, 84.7 tok/s c1 on eugr's b12x image.
- DanTup/spark-evals - Harbor-run leaderboard of evo eval, aider polyglot, and swe-rebench scores under Codex and Terminus-2 agents for four model configurations that fit on one DGX Spark.
- DG1001/local-agentic-coding-128gb - Coding-agent benchmark of 14 local models on one ASUS Ascent GX10 under three harnesses, 86 hidden tests, with Nemotron scoring 47 to 85 across 13 runs.
- elsung/dgx-spark-deepseek-v4-flash - DeepSeek-V4-Flash official FP8 across two DGX Spark under vLLM TP=2 over a 200 Gb/s QSFP56 RoCE cable, 41 tok/s single stream and ~350 aggregate at c=32.
- Entrpi/ds4-on-spark - DGX Spark fork of antirez/ds4 at 2.4-3.3x upstream prefill and 1.33-1.47x decode, 59 tok/s over 12 concurrent requests, and 2.26M tokens of active context at shipped defaults.
- Entrpi/qwen3.5-122B-A10B-on-spark - Qwen3.5-122B-A10B on a single DGX Spark via DFlash block-diffusion spec-decode, 81 tok/s on agent traffic.
- evanwtf/local-llm - Coding-agent benchmark of model, engine and harness stacks on a two-node DGX Spark cluster, ranked by suite pass rate then wall time, 29 configurations under OpenCode.
- GaelicThunder/colibri-gb10-attention-cliff - Fixed-size kernel guard at 8192 context tokens in colibri on GB10, silent CPU fallback at 0.45 versus 0.86 tok/s, plus per-branch counters and a sweep harness.
- GaelicThunder/DeepSeek-V4-Flash-Vision-One-DGX-Spark - DeepSeek-V4-Flash-Vision-Exp in EXL3 on one DGX Spark at 245,760 context, 37.6 tok/s on code with images in the same process, keeping 90% of the original token probability.
- GaelicThunder/gb10-uma-inference-notes - Five measured GB10 unified-memory properties that hold across models and engines, including host-to-device copies as pure waste worth 63% once removed.
- GaelicThunder/moe-offload-findings - Nine findings on MoE decode with experts streamed from NVMe, mostly negative: batching does not amortize expert reads and prefetching cannot beat a static pin.
- gitcommit90/qwen38-27b-dgx-spark - Qwen3.8-27B NVFP4 on one GB10 with Inco DFlash 2 speculative decoding, 44.46 tok/s versus DSpark 31.95 at k=7, plus a BF16 lm_head guard fix.
- hebo1221/motif3-dgx-spark - Motif-3 315B as mixed IQ2_XXS on one DGX Spark, 83.56 GiB at 16.49 tok/s, published with the author's verdict that the quant misses BF16 quality retention.
- jeremy-newhouse/dgx-spark-nemotron-super-bench - Single-stream decode benchmark of Nemotron-3-Super-120B-A12B-NVFP4 on one GB10, ~26-27 tok/s realistic with MTP vs ~33.6 microbench.
- jiayuqi7813/DeepSeek-V4-Flash-0731-CRACK-2x-DGX-Spark - Rank-1 Householder refusal edit of DeepSeek-V4-Flash-0731 in native FP8 UE8M0 blocks on two DGX Spark, compliance 3.53 to 90.59% with HumanEval 148 against 150.
- jvr0x/dgx-spark-bench - Closed-loop concurrency sweeps on DGX Spark with pinned recipes and a GitHub Pages dashboard, per-session and aggregate tok/s across 43 published runs.
- k3net/docai-evals - Evidence repository for 19 Hungarian document-AI and GB10 serving experiments, where Marlin beats the publisher-recommended MoE backend by up to 3.3x and NVFP4 roughly doubles counterparty-role errors.
- Kleybrink/dgx-spark-bench - Ollama benchmarking framework for DGX Spark measuring throughput, latency, memory, and answer quality with an LLM-as-a-judge pipeline over 21 prompts in 9 categories.
- marksunner/dgx-spark-single-stack - Single-box agent stack on DGX Spark, Hermes runtime and Honcho memory on CPU beside vLLM serving Qwen3.5-122B hybrid INT4+FP8 with MTP at 41-47 tok/s.
- marksunner/dgx-spark-step37-flash - StepFun's Step 3.7 Flash (198B MoE) on a single DGX Spark with llama.cpp at ~27 tok/s and 96K context on a q8_0 KV cache, reduced from 128K for CUDA-graph stability.
- martimramos/dgx-spark-ml-guide - Troubleshooting playbook for PyTorch on GB10, 16 numbered failures with an error-to-fix table, from cu128 nightly wheels to aarch64 mmcv builds and NVRTC symlinks.
- Memoriant/dgx-spark-kv-cache-benchmark - KV cache quantization on GB10: q4_0 costs 37% of generation throughput at 110K context, TurboQuant turbo3 up to 23.6% at 32K, and prompt processing is untouched.
- mneha05/gb10-attn - Context-split FlashDecoding paged-attention decode kernel at 82-85% of GB10's 273 GB/s on large working sets, fp16 gpt2 geometry, after correcting a 2x roofline and an L2-cached benchmark.
- msuiche/weightless - Serve-time refusal steering with GGUF layer-projection vectors instead of redistributed weights, validated on Qwen3.8-27B on one DGX Spark and GLM-5.3-Flash on four.
- nabe2030/dense-27b-31b-dgx-spark - Dense Qwen 3.5/3.6-27B and Gemma 4-31B on llama.cpp b8922, Q4_K_M at 10-12 tok/s versus 3.8-4.5 BF16, JCommonsenseQA cost 0.2-0.6 points.
- nabe2030/gemma4-vs-qwen35-dgx-spark - Gemma 4 26B-A4B versus Qwen 3.5/3.6-35B-A3B MoE on llama.cpp, F16 thinking-mode bug isolated, 26.5 against 58 tok/s decode and 0.35 point JCommonsenseQA spread.
- OscarActual/gb10-llm-benchmark - Ollama benchmarks for GB10 across 11 LLMs and 10 embedders: decode tok/s, TTFT, and Czech RAG recall.
- pendakwahteknologi/gx10-benchmarks - Benchmark roster for the ASUS Ascent GX10 (GB10) with nine published runs across inference, training, efficiency, and generation, each carrying timestamped CSV and log artifacts.
- r0b0tlab/deepseek-v4-flash-nvfp4-gb10-benchmark - DeepSeek-V4-Flash FP8 on two DGX Spark at TP=2 with MTP over RoCE, 38.4 tok/s c1 and 144.6 aggregate at c16 on driver 580.142, down about 3.5x on 580.159.03.
- r0b0tlab/diffusiongemma-26b-nvfp4-sm121-vllm - DiffusionGemma 26B-A4B NVFP4 under vLLM on GB10 with the FlashInfer CUTLASS FP4 MoE path, 146.3 tok/s at c1 and 242.9 at c16, no Marlin fallback.
- r0b0tlab/laguna-s-2.1-nvfp4-sm121-vllm - Laguna S 2.1 NVFP4 on GB10 under vLLM 0.25.1, DFlash K=7 at 22.2 tok/s c1 with thinking off on revision 0761412, plus an 8,620-case scorecard for the earlier 216d1f1.
- r0b0tlab/nex-n2-mini-nvfp4 - NVFP4 vLLM container for Nex-N2-mini (Qwen3.5-MoE-35B) on GB10, 185 tok/s aggregate at concurrency 8.
- r0b0tlab/step37-flash-nvfp4-sm121-vllm-docker - vLLM container for StepFun's Step 3.7 Flash NVFP4 (198B MoE VLM) on dual GB10 TP=2, with verified native-CUTLASS sm_121 execution at 16.49 tok/s.
- ramsred/llm-engine-benchmark - Long-context comparison of vLLM, SGLang, and TensorRT-LLM on GB10 under one client, plus a prefill-budget study cutting P95 TTFT 26.9% by moving 8192 to 2048.
- styles01/sparkrun-recipes - Source-pinned SparkRun recipes for nine exclusive lanes on one GB10, led by a Qwen3.8-Flash-Next EXL3 native-MTP daily driver at 54.3 tok/s single-stream decode.
- ThinkCode/glm53-flash-2x-gb10-bench - GLM-5.3-Flash NVFP4 against EXL3 on two GB10, with NVFP4 ahead cold at every concurrency and EXL3 ahead 1.6x at warm C4 on prefix-cache reuse NVFP4 never gets.
- VincentMarquez/glm52-gb10-colibri - GLM-5.2 744B on one DGX Spark at full top-8 via a colibri engine branch, 11.1 tok/s on a synthetic prompt with PLD, 5-7 on real chat.
- Weschera/spark-bench - LLM benchmark for DGX Spark across 80 scenarios in 13 domains, with 12 multi-turn agentic workflows, 4 machine-graded long-generation builds, and a TrueScore weighting speed at 5%.
You can connect two DGX Spark units directly over 200 Gb/s QSFP for double the memory and compute.
- 0xdfi/GLM-5.2-1M-4x-DGX-Spark - Profile index for unpruned GLM-5.2 744B on 4x DGX Spark with NVFP4 DS-MLA KV, O14 Fast at 250K total KV, 819 tok/s prefill, 42.3 peak decode.
- 0xdfi/GLM-5.2-R9-Adaptive-MTP-FULL-CUDA-4x-DGX-Spark - GLM-5.2 on four DGX Spark nodes with FULL CUDA graphs for every adaptive MTP depth K2/K4/K5, 83.4 tok/s at C4 measured before the 420K retune, 520K balanced profile.
- 0xSero/glm-5.3-flash-sglang-sm121 - GLM-5.3-Flash NVFP4 under SGLang on two or four DGX Spark, 22.69 tok/s single stream and 82.37 at eight on TP=4, digest-pinned with an acceptance checklist.
- ajensenwaud/Kolibri-1-2x-DGX-Spark-TensorFold - Kolibri-1 on two DGX Spark under TensorFold with 17 patches over PR #328, 83.7 tok/s single-stream decode at FP8 against 47.1 on one Spark.
- alexellis/glm-5.3-flash-4x-dgx-spark-switchless - GLM-5.3-Flash NVFP4 at TP4 across four DGX Spark cabled as a switchless RoCE ring, ~45 tok/s on agentic traffic over patched NCCL 2.30.7 that skips the wedging tree connect.
- ArgentAIOS/dgx-spark-cluster - Two-node DGX Spark guide with DMA-BUF NCCL settings for GB10 where nvidia-peermem GPU Direct fails, 93.5% DDP scaling on a 200 Gb/s RoCEv2 link.
- bertholomus/deepseek-v4.1-tensorfold-tp2-2xgb10 - DeepSeek-V4.1-Flash EXL3 2.9 bpw on two DGX Spark under a TensorFold TP2 engine with DSpark drafting, 1,048,576-token window at four streams, 100.5 tok/s greedy single-stream code.
- bird/GLM-spark - GLM-5.2 across four DGX Spark nodes, 1M context at 24.7 tok/s on unpruned Int4-Int8Mix weights, TP4 plus DCP4 KV sharding and lossless speculative decoding.
- botAGI/dspark-0731-gb10 - DeepSeek-V4-Flash-0731 on two DGX Spark under vLLM 0.25.2, flashinfer_b12x MoE backend worth 9-12% step rate to 256K, 28 tok/s single stream at 1,042,600 tokens.
- bumasoft/Qwen3.8-Flash-Next-TP3-DGX-Spark-VLLM - Qwen3.8-Flash-Next FP8 across three DGX Spark at tensor-parallel 3, checkpoint padded in place at 60.9 tok/s, where padding key heads alone silently breaks the 1:3 GDN pairing.
- chadhurley25075-png/pd-bridge - DeepSeek-V4-Flash prefilled on two DGX Spark, decoded on a Mac Studio over plain 10 GbE with no shared cache format, 3.7x vs the Mac alone on a 241K cold prompt.
- chishiki37/dgx-spark-fabric - Switched MikroTik CRS812 fabric replacing the direct DAC between two DGX Spark nodes, fixing broken RoCE for 26-33% more decode, with 100 and 200 Gb/s breakouts measured identical.
- christopherowen/dgx-spark-networking - Switchless RoCE collectives library for DGX Spark rings, one-shot BF16 all-reduce at 16.9 µs against NCCL's 90.6 to 92.4 µs for 10 KiB on four nodes.
- christopherowen/spark-ds41f - DeepSeek-V4.1-Flash on three switchless DGX Spark at TP3 with TileLang sm_121 kernels, 52.2 prose and 63.0 code tok/s single stream at 512K context.
- ciprianveg/gb10-vllm - Model-agnostic sm_121 vLLM platform, KIMI-K3 v5 at TP16 plus DCP8 with RoCEnante collectives on 16 GB10, 29.81 tok/s at C1 and 87.12 aggregate at C8, GLM-5.3 on GLM-5.2's v19 image.
- CosmicRaisins/glm-5.2-gb10 - Unpruned GLM-5.2 Int4-Int8Mix on four GB10 nodes with DCP KV sharding, 320K context at 598 t/s prefill or 640K at 430, decode flat near 22 t/s.
- CosmicRaisins/minimax-m3-awq-gb10 - MiniMax-M3-AWQ-INT4 vLLM serve recipe for 4x GB10, FP8 KV cache, EAGLE3 spec-decode, and indexer-corruption fix.
- digchick/dgx-spark-200g-link-fix - Troubleshooting playbook for the 200G ConnectX-7 link failing to train between two Sparks (CX7 hotplug power-saving), with the fix and NCCL/RoCE verification.
- drowzeys/keys-1M-CTX-Inkling-Small-NVFP4-Dspark-NVFP4-KV-Cache-SGlang-SM121-optimized-on-Two-DGX-Sparks - NVFP4 KV cache in the SGLang triton backend for Inkling-Small on two DGX Spark, 3.12x pool for 1M context, 23.9 tok/s on a pooled open-ended mix.
- drowzeys/keys-MiMo-V2.6-Pro-MOPD-C3-Jarrelscy-ARVQ-4-DGX-Sparks-1M-Context - MiMo-V2.6-Pro-MOPD in Jarrelscy's ARVQ and NVFP4 hybrid on four DGX Spark at TP4 and 1M context, single-stream 32.2 tok/s on code and 24.2 on prose with MTP k=2.
- drowzeys/Keys-Setup-Autonomous-Self-Improving-Local-Inference-Stack - Mixture-of-Agents stack for four DGX Spark nodes with a DeepSeek-V4-Flash router and an NVFP4 Two-Tower consolidated onto one node at 29 tok/s.
- drowzeys/keys-TensorFold-GLM-5.3-TP4-4x-DGX-Spark - GLM-5.3 753B at EXL3 2.75 bpw on four DGX Spark under TensorFold TP4 with MTP 2, 40.2 tok/s prose with thinking on, needle found at 905,182 tokens.
- Enntity/sparkglm - GLM-5.3-Flash NVFP4 on two DGX Spark under the Atlas engine with DFlash2 and prefix caching, 0.67 s median first token on turns 2+ of 30-48K conversations.
- epappas/gx10-cluster - Ansible provisioning for a GB10 cluster over ConnectX-7 with diagnostic tools, a recipe catalogue, and measured-or-forum-read labels on hardware claims, 22.7 GB/s two-node NCCL busbw.
- FujitsuPolycom/sparkring - One-command installer and vLLM/SGLang serving for switchless two- or four-Spark GB10 rings, using RoCEnante RDMA collectives for decode and NCCL 2.32.3 for larger ones.
- HeNryous/mimo-v25-dflash-dgx-spark - MiMo-V2.5 (309B MoE, NVFP4 plus a required MXFP8 o_proj overlay) on two DGX Spark nodes, DFlash drafting ~54 tok/s on structured content, 1.67M-token fp8 KV pool at 500K.
- idonati/spark-vllm-docker-festr2 - vLLM patches for festr2 MiMo-V2.5-Pro NVFP4/MXFP8 on an 8-node sm_121 cluster, fused-QKV fix for Q mis-slotted as K/V on 7 of 8 ranks, 114 tok/s aggregate at 20 concurrent.
- jakejharris/jspark3 - GLM-5.3-Flash on three DGX Sparks with a TensorFold fork and DFlash2 drafting, RigMark estimates at reasoning effort low of 91.3 tok/s code, 51.6 prose, 2,124 tok/s cold 64K prefill.
- jayleaton/glm53-tensorfold-spark - GLM-5.3-Flash abliterated EXL3 4-bit on two DGX Spark under TensorFold plus 77 patches, RigMark 57.5 tok/s at C1 and 95.1 at C4.
- joesinvestments/DeepSeek-V4-Flash-0731-TP4-4x-DGX-Spark - DeepSeek-V4-Flash-0731 at TP=4 on four DGX Spark under vLLM 0.25.2 with DSpark k=7, 123.13 tok/s single stream at 2K tokens, 57.9 at ~150K agentic prompts.
- joesinvestments/glm52-spark-kit - GLM-5.2 on four DGX Spark as source overlays over stock vLLM, fused nvfp4_ds_mla KV writer at 51 times the torch reference, sparse indexer unblocked above DCP 1.
- joeynyc/Hy3-295B-NVFP4-2x-DGX-Spark - Tencent Hunyuan 3 295B MoE in NVFP4 at TP2 on two DGX Spark, 26 tok/s single stream at 262K with TurboQuant k8v4 KV, MTP measured 20% slower.
- joeynyc/MiniMax-H3-2x-DGX-Spark - Cross-host diffusion executor for one MiniMax H3 video across two DGX Spark nodes, Ulysses sequence parallelism over RoCEv2, 68.8 s against 155.0 s single-box.
- josephdrose/joe-spark-patches - Out-of-tree vLLM patches for DeepSeek on four DGX Spark, Engram kept on NVMe because
cpu_offloadfrees nothing on GB10, plus a DSpark proposer that survives concurrent requests. - josephdrose/nccl-spark-switchless - NCCL 2.30.7 patches for a switchless 4-node GB10 RoCE ring, tree-skip plus 2-hop relay over the uncabled diagonals, MiniMax-M3 229B NVFP4 at ~24 tok/s.
- karolpalys/glm52-triple-spark-tuning - Tuning log and negative results for GLM-5.2 753B on three DGX Spark nodes, with the measurement methodology and the evidence that killed each rejected change.
- kindlingai/glm-5.3-flash-gx10 - GLM-5.3-Flash NVFP4 on two, three, four or six GB10 under vLLM with mentat in place of Ray, 117.3 tok/s code decode at TP=4 with DFlash2.
- kindlingai/glm-5.3-full-exl3-tp6 - Full GLM-5.3 EXL3 3.25 bpw on six DGX Spark at TP6 by re-fragmenting experts without requantizing, 33.8 tok/s prose at MTP k=2, 958 tok/s prefill at 32K.
- kingjones30/GLM-5.3-Flash-2x-DGX-Spark - GLM-5.3-Flash NVFP4 on two DGX Spark, NoPE latent padded to the rope width sm_121 kernels expect, 24.7 tok/s on code under MTP-5 on the stock image.
- knapcio/DeepSeek-V4.1-Flash-4x-DGX-Spark-TP4 - DeepSeek-V4.1-Flash TP4 on four DGX Spark under SGLang with RoCEnante all-reduce, 89.7 tok/s prose and 131.9 code at c1 with the SM clock capped at 2200 MHz.
- knapcio/GLM-5.3-Flash-4x-DGX-Spark-TP4 - GLM-5.3-Flash NVFP4 at TP4 on four DGX Spark with 8-bit dense layers and DFlash2, 90.3 tok/s prose single stream, 3,470 tok/s cold prefill at 32k.
- Libertai/glm53-flash-vllm-gb10 - GLM-5.3-Flash under vLLM on two DGX Spark at 24.2 tok/s through a hand-written sparse-MLA kernel, with degenerate output traced to an uninitialized MoE activation scale.
- magicbear/DeepSeek-V4.1-Flash-SGLang-DGX-Spark - DeepSeek-V4.1-Flash under SGLang on four against eight DGX Spark, c1 code decode 57.9 against 67.0 tok/s, 1.63-1.76x prefill at eight, while its own vLLM cross-check reaches 75.5 on four.
- makiisthenes/dgx-spark-multinode-vllm-ray - Dual-DGX Spark vLLM deployment with NVIDIA vLLM 26.04, Ray, and 200 GbE QSFP.
- maliubiao/dgx-spark-2-deepseek-flash-0731 - Bilingual manual for DeepSeek-V4-Flash-0731 and Vision-Exp on two DGX Spark, with LMCache v0.5.3 built on aarch64 for NVMe cold KV backup, 1 GiB L1, 500K context.
- MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark - Two-node vLLM TP=2 recipe for DeepSeek-V4-Flash-Vision-Exp with native image input, DSpark speculative decode and nvfp4_ds_mla KV at a 1M-token ceiling.
- MiaAI-Lab/DeepSeek-v4.1-Flash-DGX-Sparks - DeepSeek-V4.1-Flash under SGLang on three or four DGX Spark with NVMe-resident Engram tables, TP4 prose c1 87.7 tok/s on a switch, 1M-context needle passed at 1,011,084 tokens.
- MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks - DeepSeek-V4.1-Flash at 2.9 bpw EXL3 on two DGX Spark with vLLM TP2 and file-backed Engram, 31.6 tok/s single stream with DSpark k=3, 600K context.
- MiaAI-Lab/GLM-5.2-NVFP4-AQLM-Triple-DGX-Sparks - GLM-5.2 NVFP4 plus AQLM 2-bit hybrid on three DGX Spark at TP3, 21 tok/s structured at 348K vision context with top-4 routing, 25-26 on fp8 KV at 235K.
- MiaAI-Lab/GLM-5.3-EXL3-3x-DGX-Sparks-TensorFold - GLM-5.3 EXL3 2.75 bpw on three DGX Spark under TensorFold, a 499,712-token window by context parallelism, 41.7 tok/s code, a 94K prompt resumed from NVMe in 3.0 s.
- MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold - GLM-5.3-Flash EXL3 on two DGX Spark under TensorFold v0.6.0 with 96 patches, 1,048,576-token window at 4 concurrent requests, 60.4 tok/s prose and 114.7 structured at c1.
- MiaAI-Lab/GLM-5.3-Flash-NVFP4-Dual-DGX-Spark - GLM-5.3-Flash NVFP4 on two DGX Spark over Ray TP=2 with image and video input, MTP at four draft tokens, FP8 KV cache and 262K context.
- MiaAI-Lab/Inkling-Small-NVFP4-Dual-DGX-Sparks - Inkling-Small-NVFP4 across two DGX Spark on SGLang at a 1,142,712-token KV pool, 33.9 tok/s single stream, page size 1 required for the triton DSpark verify path.
- MiaAI-Lab/MiMo-V2.6-Flash-2x-DGX-Sparks - MiMo-V2.6-Flash MXFP4 under SGLang on two DGX Spark, DFlash at 69.8 tok/s on code and EAGLE MTP at 35.0 on prose, single stream.
- MiaAI-Lab/Qwen3.8-Flash-Dual-DGX-Sparks-TensorFold - Qwen3.8-Flash-Next NVFP4 on two DGX Spark under TensorFold's Zig engine at TP2 with FP8 KV, 1,048,576-token window, 67.4 tok/s prose at c1 and 390.1 at c16.
- MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks - Qwen3.8-Flash-Next NVFP4 on two DGX Spark under vLLM TP2 with expert parallel and MTP-3, 56.8 tok/s prose at one stream and 225 aggregate at eight with 47k draft vocabulary.
- nabe2030/dgx-spark-2node-rpc - GLM-5.2 GGUF at 228.5 GB split across two DGX Spark nodes over llama.cpp RPC, CX7 measured as two ~100 Gb/s PCIe Gen5 x4 paths, RDMA vs TCP A/B.
- nacyot/vllm-ds4f-gb10 - vLLM fork for DeepSeek-V4.1-Flash on four GB10 at TP=4 with disk KV offload, a 493K-token session restored from SSD in 7.8 s against a 392 s cold prefill.
- neko-legends/spark-bench - Four DGX Sparks as one TP4 cluster over six dated model lanes, the live one DeepSeek-V4.1-Flash uncensored on TensorFold at 66.5 tok/s prose and 104.1 code, 1k-160k prompts.
- ondigo-winder/SWITCHLESS-NCCL-RING-4-GB10 - NCCL 2.30.7 patch for a switchless four-node GB10 ring using both PCIe halves of each QSFP cable, 193.1 Gb/s all-reduce bus bandwidth against 111.6.
- OsakaTX/qwen3.8-flash-next-vllm-dgx-spark - Qwen3.8-Flash-Next NVFP4 on two DGX Spark under vLLM TP2 with MTP n=3, 41-44 tok/s c1 and 162 aggregate at eight, RDMA passthrough worth 40-45%.
- pfn/spark-vllm-compose - Head and worker Docker Compose files that run vLLM across DGX Spark nodes with native --nnodes/--node-rank instead of Ray, shipped services for Qwen3.5-397B-A17B-int4 and MiniMax-M2.7.
- r0b0tlab/DeepSeek-V4-Flash-DSpark-v026-SM121 - Dual-GB10 DeepSeek-V4-Flash-0731 on vLLM 0.26 with a digest-pinned GHCR image and runtime gate, NIAH PASS at 1M, BFCL multi_turn_base 0.755, ~79 tok/s c1 decode.
- r0b0tlab/qwen38-flash-next-w4a16-sm121-sglang - Qwen3.8-Flash-Next self-quantized to NVFP4 W4A16 on two GB10 under SGLang TP=2 with NEXTN MTP, 63.84 tok/s at c1 and 132.64 at c4 on 1,024-token prompts.
- rajsinghtechbot/dgx-spark-vllm-k8s - Kubernetes cookbook for DeepSeek-V4-Flash on dual DGX Spark, with Multus/Spiderpool RDMA over RoCEv2, UMA-aware container memory limits, and Prometheus monitoring.
- raullenchai/twinspark - Self-healing two-node vLLM cluster for DeepSeek-V4-Flash-0731 on DGX Spark at 74.8 tok/s, needle sweep cutting long-context failures from 8/25 to 2/25 by trading nvfp4_ds_mla KV for fp8_ds_mla.
- Reederey87/dgx-spark-2x-deepseek-v4-flash - DeepSeek-V4-Flash-0731 on two DGX Spark from a full-source vLLM main build, four gated one-file upstream fix layers promoted on Welch non-inferiority, 3,027,217-token KV pool.
- Reederey87/glm53-flash-exl3-2x-dgx-spark - GLM-5.3-Flash EXL3 on two DGX Spark at 1M context, 110k replay prefix-cache hits lifted from 0 to 97-98% by 3,584-token batches plus a drafter-prune patch.
- RustRunner/DGX-Llama-Cluster - Three DGX Spark nodes as a switchless ConnectX-7 RDMA star running llama.cpp RPC, 384 GB pooled unified memory for 400B+ models in 4-bit, NFS model share.
- sfxnz/DeepSeek-V4-Flash-Vision-Exp-vLLM-2x-DGX-Spark - DeepSeek-V4-Flash-Vision-Exp on two DGX Spark at TP=2, ViT loaded as a vLLM plugin that preserves DSpark-6 draft acceptance, 26.2 tok/s prose single stream at 1M context.
- strusty/GLM-5.3-Flash-NVFP4-3x-DGX-Sparks-kindling-DGXOS - GLM-5.3-Flash NVFP4 at TP=3 on three DGX Spark in a switchless triangle on stock DGX OS, 3,054,135-token KV pool, about 3,500 tok/s cold prefill at 64k.
- tomsti/guides - GB10 cluster guide for DGX Spark over ConnectX-7 RoCE, covering NCCL rail pinning, the duplicate-MAC workaround, and MikroTik 400G switching.
- tonyd2wild/DeepSeek-v4-Flash-Vision-Exp-DSpark-1M-NVFP4-KV-2x-DGX-Spark - DeepSeek-V4-Flash-Vision-Exp ported into vLLM on two or four DGX Spark at 1M context, 53 tok/s real-prompt decode single stream and a 2.79M-token NVFP4 KV pool at TP2.
- tonyd2wild/DeepSeek-V4.1-Flash-vLLM-DGX-Spark - DeepSeek-V4.1-Flash on four DGX Spark at TP4 with Engram rows read from NVMe, 3.5 bpw EXL3 serving lane at 95.3 tok/s idle code and 500K context.
- tonyd2wild/DS4-H3-Video-Gen-Factory - DeepSeek-V4-Flash at 1M context beside two MiniMax H3 video renders on two DGX Spark, 35% of idle throughput kept, and a load order set by H3's 50 GB footprint.
- tonyd2wild/GLM-5.2-655K-MTP-4x-DGX-Spark---25-32tok-s - GLM-5.2 744B unpruned (QuantTrio Int4-Int8Mix) at 655,360-token context across four DGX Spark nodes via decode-context-parallelism, 23.0 tok/s single and 47.9 aggregate at 4 concurrent, MTP k=3.
- tonyd2wild/GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s - Speed-shape recipe for unpruned GLM-5.2 at 200K context on four DGX Spark nodes, 36 tok/s peak, 75 aggregate at 6 concurrent, 63% of each step overhead.
- tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark - GLM-5.3-Flash NVFP4 on two DGX Spark with knapcio's TP4 stack ported to TP2 and DFlash2, single-stream code decode 68.9 tok/s against 44.7 for the previous recipe at 262K context.
- tonyd2wild/GLM-5.3-Int4-Int8Mix-TP4-4x-DGX-Spark - GLM-5.3 743B Int4-Int8Mix at TP4 on four DGX Spark, DFlash2 at 53.32 tok/s structured output (1.98x MTP-4, level on prose), NVFP4 KV pool of 293,447 tokens at 270K.
- tonyd2wild/MiMo-V2.5-TP2-1M-NVFP4-KV-2xDGX-Spark - MiMo-V2.5 Omni at tensor-parallel 2 on two DGX Spark, NVFP4 4-bit KV cache for a 2.17M-token pool at 1M context, thinking-OFF eval 97.8 against 90.6.
- tonyd2wild/MiMo-V2.5-TP3-NVFP4-KV-3xDGX-Spark - MiMo V2.5 Omni (310B MoE, text/image/video/audio) at tensor-parallel 3 across three DGX Sparks, with 4-bit NVFP4 KV cache for a ~10.6M-token KV pool at 1M context.
- tonyd2wild/MiMo-V2.6-Flash-DGX-Spark-Recipe - MiMo-V2.6-Flash-RL on two DGX Spark under vLLM TP2 with DFlash and fp8 KV, 53.3 tok/s single stream under a 300K limit, after four patches to the stock image.
- tonyd2wild/MiniMax-M3-2x-DGX-Spark-36-tok-s - Reproduction of a3refaat's MiniMax-M3 428B unpruned stack on two DGX Spark, W4A16 GPTQ with NVFP4 KV and EAGLE-3 at 36.6 tok/s single-stream JSON and 31.8 code, 196K context.
- tonyd2wild/Minimax-M3-NVFP-3x-DGX-Sparks-TP-3 - MiniMax-M3 NVFP4 428B-A23B at tensor-parallel 3 across three DGX Sparks, 10.5 tok/s at 200K context over a switchless 200 Gb/s RoCE mesh.
- tonyd2wild/nfs-model-weights - NFS recipe for sharing one checkpoint across N DGX Spark nodes, taking a 4-node model library from 6.8 TB to 1.7 TB with no per-node copies.
- tonyliu312/GLM-5.3-Flash-1M-Context-4x-DGX-Spark - GLM-5.3-Flash at 1,048,576-token context on four DGX Spark, 2.67x decode from dropping --enforce-eager and 1M concurrency from 1.35x to 4.20x by raising --kv-cache-memory, which silently overrides --gpu-memory-utilization.
- tpurtell/glm-5.2-4x-spark-1x-rtx6k-96gb - Rust attention-FFN disaggregation engine for GLM-5.2 on one RTX PRO 6000 coordinator and four DGX Spark expert nodes, 26.6 NVFP4 and 28.5 EXL3 K3 tok/s weighted decode, 1,807 prefill.
- tsarihan/qwen3.8-flash-next-nvfp4-2x-dgx-spark-playbook - Qwen3.8-Flash-Next NVFP4 against FP8 on the same two DGX Spark, 3.67x the KV pool and 66% more decode at 245K, needles 5/5 on both lanes.
- urbanspr1nter/dgx-spark-bare-metal - Four-node DGX Spark Ray and vLLM cluster on a Mikrotik switch, with an 8-node guide on 400 Gb/s breakout cables, plus GLM-5.2 NVFP4, Kimi K2.7 Code and DeepSeek-V4-Flash launchers.
- ursuciprian/qwen3.8-flash-next-dgx-spark-tp-2 - Qwen3.8-Flash-Next NVFP4 on two DGX Spark under vLLM V2 with b12x kernels, where one replica per node behind a router finishes synthetic agent replays 29-37% sooner than TP=2.
- vroomfondel/dgxarley - Ansible playbooks for a K3s cluster of four DGX Spark nodes and an x86 control plane, SGLang over SR-IOV RoCE at 9.78 GB/s NCCL bus bandwidth.
- Weschera/Ling-3.0-flash-VL-2x-DGX-Sparks - Ling-3.0-flash-VL FP8 on two DGX Spark under SGLang TP2 at 131,072 tokens, 26.0 tok/s single-stream on prose and 18.5 on code, with no MTP head.
- www-ai-rs/gb10-deepseek-v4-flash - Two-node GB10 operator tooling and executed-code test harness for DeepSeek-V4-Flash 304B at 1M context, GPU rail power sampled at 2 Hz.
- yunwei37/dgx-spark-4-ring-no-switch - Serving profiles and results log for four DGX Spark on a switchless ConnectX ring, DeepSeek-V4.1-Flash TP4 at 1M context 50.14 tok/s C1, gmu 0.88 recorded as unsafe.
- ZD-AI-Lab/Triple-GB10 - Three-node GB10 QSFP ring with three /30 subnets and forwarding routes, pooling about 300 GB for Ray and vLLM pipeline-parallel across 3 Sparks.
- zorost/sparkduet - Model-swap layer for two DGX Spark, four checkpoints on disk behind one OpenAI port, DeepSeek-V4-Flash c1 from 72.2 tok/s on math to 33.6 on prose by DSpark draft acceptance.
- AEON-7/comfyui-aeon-spark - ComfyUI Docker for DGX Spark with SageAttention v3 compiled for sm_121a, CUDA 13, NVFP4, and Flux 2 / LTX 2.3 pre-bundled.
- alexhegit/h3-spark.c - CUDA port of antirez/h3.c for MiniMax-H3 on GB10, 15.6 s warm for a 512x512 22-frame clip at layers 45 and reuse 2.
- bjarkebolding/spark-comfyui - Single-script containerized ComfyUI for DGX Spark, sm_121 SageAttention and sha256-verified workflow recipes, with a 2100 MHz clock cap for overcurrent reboots.
- CoconutMacaroon/blender-arm64 - Blender build for GB10 aarch64 with CUDA, OptiX, and Vulkan, shipping a prebuilt DGX Spark binary release.
- dr-vij/ComfyUI-DGX-Spark-Docker-opinionated - ComfyUI Docker for DGX Spark with self-built aarch64 wheels for decord NVDEC, flash-attn, onnxruntime-gpu and SageAttention 2.2.0, plus a SAM3 blacklist for WanVAE hangs.
- dr-vij/Hunyuan3D-2.1-DGX-Spark-Docker - Hunyuan3D-2.1 3D generation on DGX Spark via Docker Compose, building custom_rasterizer and DifferentiableRenderer CUDA components on-box.
- dr-vij/Trellis2-DGX-Spark-Docker - Dockerized TRELLIS.2 3D generation for DGX Spark, with nvdiffrast, CuMesh, FlexGEMM, and torchvision built from source for sm_121 on CUDA 12.9.
- drowzeys/keys-SM121-Optimized-MiniMax-H3-Nvidia-Sol-Engine-Kijai-SolAttn_Triton-Single-DGX-Spark - MiniMax H3 in ComfyUI on one DGX Spark with flex_attention forced to Triton where sm_121 miscompiles it, plus Sol-Engine FirstBlockCache and batched VAE ports at 1.54x end to end.
- Hitheshkaranth/LTX-2.5_Video_DGX_Spark_Setup - LTX-2.5 NVFP4 text and image-to-video with audio on DGX Spark via a three-line sm_121a kernel patch, 768x512 4 s clip in 61 s.
- jayleaton/localrouter - MCP and OpenAI-style server that loads Qwen-Image 2.1 and MiniMax H3 on demand with Zig engines on DGX Spark, 1024x1024 NVFP4 in 12.7 s warm.
- joeynyc/cosmos-locateanything-dgx - Two-stage DGX Spark pipeline: Cosmos 3 video generation, then NVIDIA LocateAnything object grounding.
- joeynyc/MiniMax-H3-DGX-Spark - MiniMax H3 FL2VA video generation on one DGX Spark via vLLM-Omni and online FP8, sm_121 loader and AdaLN patch, 111 s warm request or 80 s with Cache-DiT.
- kabilankb/cosmos3-nano-gb10 - Cosmos3-Nano 16B video and image generation on a single GB10 instead of NVIDIA's recommended 8x H100, 480p 57-frame clips in about 3 minutes.
- LectinDK/DGX-Spark-ComfyUI-Forge - ComfyUI Docker for DGX Spark, xformers CUTLASS off on sm_121, and a --disable-pinned-memory fix that holds DynamicVRAM at a 73-78 GB peak against 88.
- luix93/DGX-Spark-ComfyUI - ComfyUI Docker Compose for DGX Spark with SageAttention 2 built against sm_121, NVFP4 via comfy_kitchen, and a copy=False patch for the unified-memory double-VRAM bug.
- madeye/comfyui-minimax-h3-dgx-spark - MiniMax-H3 video and audio in ComfyUI on GB10 sm_121 with 4-to-8-step turbo LoRAs, where fp8_scaled runs 21% faster than the recommended int8_convrot at 11.50 against 14.26 s/it.
- mmartial/ComfyUI-Nvidia-Docker - Multi-platform ComfyUI Docker (x86_64, Blackwell, DGX Spark) with published aarch64 DGX images and userscripts that build SageAttention 2 and comfy_kitchen from source.
- mvalancy/blender-nvidia-gb10 - Blender 5.0.1 source build for GB10 with Cycles CUDA 13 GPU rendering, via four inherited aarch64 patches and four CUDA-13 patches for OIDN, libglu, Wayland, and libdrm.
- phaserblast/ComfyUI-DGXSparkSafetensorsLoader - Zero-copy model loader for ComfyUI on DGX Spark using the fastsafetensors library.
- AEON-7/qwen3-asr-server - OpenAI /v1/audio/transcriptions server for Qwen3-ASR-0.6B on DGX Spark, vLLM-native with a pinned flashinfer and a soundfile decode path that avoids the PyAV fallback.
- briancaffey/nemotron-asr-server - OpenAI-compatible speech-to-text server for nemotron-3.5-asr-streaming-0.6b on DGX Spark, native NeMo instead of the aarch64-broken Riva path, WebSocket streaming.
- jxlarrea/homeassistant-voice-recipes - Local Wyoming voice pipeline for Home Assistant, aarch64 ONNX Parakeet ASR, ECAPA-TDNN speaker extraction, and Gemma-4-26B-A4B with MTP on llama.cpp built for sm_121.
- kedarpotdar-nv/spark-realtime-chatbot - On-device assistant for voice and video calls on one GB10, ~320 ms and ~850 ms end to end, using Qwen3-VL, Kokoro, and DeepFace.
- Logos-Flux/spark-voice-pipeline - WebSocket voice assistant for DGX Spark, sentence-level streaming across whisper.cpp, Ollama, and VibeVoice-Realtime-0.5B, 766 ms to first audio.
- luka-loehr/qwen3-tts-native - Native Rust and CUDA runtime for Qwen3-TTS-1.7B VoiceDesign on sm_121, 94-96 ms p95 first audio and 5.68 GB peak versus 2.69 s and 108.90 GB for stock SGLang.
- mARTin-B78/dgx-spark-faster-qwen3-tts - Faster-Qwen3-TTS on DGX Spark as an OpenAI-compatible API with CUDA-graph acceleration, four backends including chunk-streaming, and deterministic per-voice seeds.
- Mekopa/whisperx-blackwell - Prebuilt WhisperX image with an sm_121 to sm_90 capability spoof and torchaudio jiterator patch, GPU pyannote diarization, 24 min audio in 62 s.
- ncannings/fastconformer-trt - FP8 TensorRT engine for NeMo FastConformer ASR, parakeet-ultra at 4,903x real time and 1.814% test-clean WER on DGX Spark, against 999x and 1.803% for stock NeMo.
- pipecat-ai/nemotron-voicechat-dgx-spark - Full-duplex speech-to-speech NemotronLabs VoiceChat 11B on one DGX Spark, GPTQ W8 Nano and W8A32 EarTTS at ~66 ms against the 80 ms frame budget.
- Pizzaman213/fish-s2pro-gb10 - GB10-tuned Fish Audio S2-Pro TTS reaching 31.3 tok/s from 1.2 via bit-exact int8 kernels, speculative decode, and quality-gated NVFP4 weights.
- WillIsback/whisperx-gb10 - WhisperX transcription and pyannote diarization REST API for GB10 (aarch64, sm_121), with an async job queue, SRT/VTT/TXT export, and prebuilt Docker Hub/GHCR images on NGC PyTorch 25.05.
Beyond LLMs, GB10's unified memory and aarch64 stack run scientific compute: protein folding, biomolecular prediction, and RAN simulation.
- adrian-greenneuron/openfold3-DGX-Spark - Dockerized OpenFold3 for DGX Spark with evoformer_attn kernels prebuilt, DeepSpeed patched from compute_121 to compute_120, ubiquitin prediction in 55 s.
- chaoticcuriosity-io/g1-humanoid-rl - Unitree G1 humanoid RL on DGX Spark with mjlab, walking policy trained in about 46 min at 2048 envs and an 11 h cartwheel run at 4096 envs.
- chaoticcuriosity-io/regolith - Synthetic-data lunar segmentation on DGX Spark with Isaac Sim 6.0 Replicator and SegFormer, where realistic ground cut real-photo rock over-prediction from 72.6% to 35.7%.
- eetmie/spark-projects - Robotics playbooks for GB10 including openpi pi0.5 vision-language-action inference at 94.7 ms on TensorRT FP8 and NVFP4, 2.12x PyTorch BF16 at cosine 0.997.
- rcbarke/ai-ran-dgx-spark - Bring-up notes for NVIDIA Aerial and Sionna on DGX Spark, with Sionna PHY and SYS working, cuMAC retargeted to run, and Sionna-RT and the 7.2x fronthaul blocked.
- sanjyotshenoy/boltz-gb10-spark - Boltz-2 biomolecular-interaction prediction on DGX Spark with Triton-nightly sm_121 codegen.
- AtomicGaryBusey/DGXSparkGaming - Steam gaming compatibility log for DGX Spark under FEX-Emu, Box64, and Proton, with DLSS 4 frame generation working, DLSS 5 ruled out per frame, and every correction published.
- seanGSISG/dgx-spark-sunshine-setup - One-command Sunshine installer for headless DGX Spark, NVIDIA CustomEDID virtual display with no dummy plug, 4K60 or 1440p120 within the GB10 pixel clock limit.
- agjs/gb10-clock-cap - Clock-cap harness for GB10 inference hosts, where 2200 MHz ran 12 °C cooler at 36% less GPU-rail power for 1.0% decode and 3.9% prefill cost, single-stream at TP=2.
- amer8/pulsebar - Unofficial macOS menu bar monitor that streams GPU and memory telemetry from the DGX Spark dashboard.
- antheas/spark_hwmon - Linux hwmon kernel driver exposing GB10 system power telemetry (per-rail power, energy counters, temperatures) and PL1/PL2 power-cap controls via sysfs.
- ateska/dgx-spark-prometheus - Single-binary Go Prometheus exporter with systemd unit and Grafana dashboard for DGX Spark, GB10 and NIC metrics on port 9835.
- chappa-ai-llc/spark-smi - System-monitor TUI for DGX Spark with unified-memory and Grace P/E-core awareness, a cluster fleet view, an MT2910 fabric bandwidth test, and mixed sm_121 plus sm_86 support.
- christopherowen/dgx-spark-fan-control - Kernel driver and CLI for RPM floors above NVIDIA's fan curve on DGX Spark, 13500 RPM floor measured at 9000 RPM in 6.3 s, with DKMS and Secure Boot signing.
- DanTup/dgx_dashboard - Monitoring dashboard for DGX Spark bound to 0.0.0.0, with GB/GiB-correct memory stats, GPU power draw, and Docker container controls.
- djmad/Spark_Energy_Management - Thermal and power controller for a ThinkStation PGX (GB10) with per-cluster CPU PID loops, a 100 MHz/s GPU clock ramp, and a thermal twin fitted to 1.2 K RMS.
- dorangao/dgx-spark-toolkit - Two-node DGX Spark cluster scripts and manifests: RoCE and NCCL checks on the 200 Gb/s fabric, RDMA pods, MetalLB, pipeline-parallel vLLM Nemotron-3 Nano 30B.
- drowzeys/vllm-gb10-spin-wait-fix - One-command patcher and English write-up for the vLLM spin-wait that heats GB10, 24 °C lower average SoC temperature on TP=2 multi-rank serves and no change at TP=1.
- engineering87/sparkfit - Zero-dependency memory planner for DGX Spark: 128 GB budget split, quantization advisor, and decode roofline on 273 GB/s.
- hectorTSH/dgx-spark-memory-dashboard - Live unified-memory occupancy map for DGX Spark as 1 GiB tiles, hot models against on-disk candidates ranked by free memory, and multi-Spark tabs.
- hoesing/spark-gpu-throttle-check - Throttle test for DGX Spark that loads the GB10 with cuBLAS matmuls and flags clocks staying below a 1400 MHz threshold, a suspected USB-PD power-delivery fault.
- ivanusto/gb10-ops - Host guards for GB10 that terminate the offending GPU process below 3 GiB MemAvailable and abort jobs after 240 s at 88 C, thresholds calibrated on one two-node box.
- jasonacox/dgx-spark - Project hub for GB10 whose nanochat scripts pretrained a chat model from scratch in 9 days for about $8 of power, plus two-Spark InfiniBand training.
- jeffrymahbuubi/dgx-spark-stress-test - Burn-in suite for GB10 unit qualification, 6-24 hour llama.cpp 70B plus SDXL load at ~96% utilization, ~85 GB resident, with 10s temperature and power CSVs.
- jeremyeder/dgx-agentskills - Claude Code integration for DGX Spark: local model serving, GPU monitoring, and VM management.
- joeynyc/spark-doctor - Read-only DGX Spark diagnostic CLI: 14 W power cap, unified-memory pressure, thermal risk, CUDA 13 / sm_121 wheel mismatches, Docker runtime, and recipe checks for tensor-parallel size and memory budget.
- lcasarin-maker/blackbox - Forensic capture for GB10 hangs with Xid, PCIe AER, xHCI and PSI checks, plus a measured gap where 7 GiB held through PyTorch raised memory.current by 15 MiB.
- lynx-lee/lynx-ollama - Ollama manager with a Go web console whose optimize command reads GB10 unified memory to set 131K context, 8-way parallel, and q8_0 KV cache.
- mcampa/sparkrun-ui - Web UI for sparkrun with launch wizard, chat, benchmark charts, and live per-host GPU bars, run via npx or a published aarch64 container.
- mchenetz/sparkd - Localhost dashboard for a DGX Spark fleet, with HF browsing, Claude-generated vLLM recipes, and single-box or Ray-cluster launch.
- MiaAI-Lab/sparkDash - Web dashboard for a DGX Spark fleet with an engine probe for llama.cpp, vLLM, SGLang, ds4, and EXL3, cached against uncached prefill, daily peak tok/s, and Wake-on-LAN.
- parallelArchitect/sparkview - Terminal GPU monitor with GB10-aware unified-memory reporting, memory-pressure (PSI) and power-rail readouts, and an anomaly auto-logger.
- r0b0tlab/hermes-concurrent-agents - Supervised Hermes Agent worker pool for one Linux host, with pre-claim admission, exact process ownership, and optional GB10 memory-pressure presets.
- securitysonar/spark-hashcat - Hashcat REST API service for GB10 aarch64, with an NVRTC CUDA build path that bypasses OpenCL.
- skymaze/Fireworks - Web control plane for a DGX Spark fleet, ARP-probed RoCE rail configuration with rollback, and model direct-pull to workers over the four rails.
- stevibe/SparklingKit - Local-first workspace for OCR, transcription, grounding, and image jobs, routed to six co-resident models on one DGX Spark from Qwen3.6-35B-A3B NVFP4 to LocateAnything-3B.
- TheAwaken1/Spark-Studio - Launch and tuning dashboard for vLLM, SGLang, llama.cpp, and sparkrun recipes on DGX Spark that hands a broken recipe to Claude Code or Codex to patch and relaunch.
- vybe/sparky - Vue 3 web UI for DGX Spark with Ollama chat, a Claude Code agent tab, ComfyUI SDXL and Flux generation, and Docker/systemd control.
- wentbackward/nv-monitor - Terminal monitor and Prometheus exporter for DGX Spark in one zero-dependency C binary, with HugePages-correct unified memory and Grace big.LITTLE core labels.
- Z841973620/dgx-spark-fan-override - Kernel module overriding DGX Spark fan speed through the EC's FF-A eSPI mailbox, with the stock firmware's two temperature curves and 1,260-13,500 RPM ranges documented.
- graham33/nixos-dgx-spark - Nix flake with a NixOS module for NVIDIA's DGX Spark kernel, bootable USB image, and 15 playbook devshells for TRT-LLM, NVFP4, and NCCL over QSFP.
- kindlingai/kindling-spark-os - Read-only OS image for GB10 booted beside DGX OS with auto-revert and a 2 GiB display carveout lent to CUDA, GLM-5.3 TP=4 KV pool 365,440 to 404,608 tokens.
- kyuz0/gb10-toolboxes - Prebuilt aarch64 CUDA 13.x containers for GB10 sm_121 with llama.cpp, antirez/ds4 and vLLM, rebuilt on four-hour upstream polls.
- maxspevack/spark-rocky - Rocky Linux 10.2 Live-USB for DGX Spark on the CIQ 6.18 kernel, shipping 4k pages because 64k faults on every driver branch but the 580, measured 10.4% slower.
- Neural-ICE/ICE-CoreOS - Immutable bootc OS for DGX Spark on CentOS Stream 10 with a 4 KiB-page GB10 kernel, optional TPM2-unlocked LUKS2, and a signed-UKI USB installer.
- RageLtd/arch-dgx-spark-iso - Arch Linux installer ISO builder for DGX Spark, with the linux-dgx-spark kernel and archinstall config.
- AEON-7/AEON-7 - Index of AEON-7's releases, mainly DGX Spark NVFP4 model packs, prebuilt vLLM images, and a voice-AI stack, plus Apple Silicon MLX builds.
- odnodn/dgx-spark - Curated collection of NVIDIA DGX Spark resources and self-hosted AI projects.
Contributions are welcome. Read the contribution guidelines before opening a pull request.