Skip to content

About

A curated list of tools, guides, playbooks, and resources for the NVIDIA DGX Spark (GB10 Grace Blackwell personal AI supercomputer).

Topics

Resources

Contributing

Stars

76 stars

Watchers

0 watching

Forks

Latest commit

 

History

84 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Awesome DGX Spark Awesome

A curated list of awesome tools, guides, playbooks, and resources for the NVIDIA DGX Spark, the GB10 Grace Blackwell personal AI supercomputer.

DGX Spark is a desktop machine built on the GB10 Grace Blackwell Superchip (SM 12.1 / sm_121), with 128 GB of unified CPU+GPU memory. You can link two units over 200 Gb/s networking to run larger models. This list collects community projects for setting it up, serving models, fine-tuning, benchmarking, and day-to-day operation.

Platform essentials: aarch64 · CUDA 13.x · sm_121 · 128 GB unified memory · 200 Gb/s ConnectX-7

Contents

Official

  • NVIDIA/dgx-spark-playbooks - Step-by-step DGX Spark playbooks spanning vLLM, SGLang, llama.cpp, NVFP4 quantization, speculative decoding, cuTile kernels, NCCL, and two- and three-node clustering.

Setup & Configuration

  • a1exus/sparky - Self-hosted DGX Spark LLM stack, vLLM, Ollama, and a llama.cpp router serving every cached GGUF behind Traefik, with mDNS, Cloudflare Tunnel, and Tailscale ingress.
  • Albatross1382/onnxruntime-aarch64-cuda-blackwell - ONNX Runtime 1.24.4 CUDA shared libraries for sm_121 on aarch64, loaded through the Rust ort crate or dlopen.
  • botAGI/AGmind - One-command private RAG stack for DGX Spark, Dify with vLLM, Weaviate, RAGFlow, and Docling across 30+ containers, plus two-node clustering over 200 Gb/s QSFP.
  • christopherowen/dgx-spark-memory-saver - Patch for nvidia-uvm that packs GPU page tables on 64 KiB DGX Spark kernels, 3.03 GiB recovered versus stock 64 KiB and about 1.8 GiB per node over 4 KiB.
  • Chrizz-lab/GB10-Agentig-Coding-Framework - Agentic coding stack for DGX Spark with dual-vLLM Qwen3 and CrewAI orchestration.
  • csabakecskemeti/dgx-spark-community-playbooks - Community playbook collection for DGX Spark, covering dual-Spark RDMA inference, heterogeneous RoCE clustering, and local Claude Code.
  • Entrpi/dgx-spark-serving-mode - Three-state serving-mode script for DGX Spark that pares GNOME and maintenance timers down to multi-user.target for 10-15 GB more unified memory.
  • Fulton-Engineering-Services/dgx-spark-wheels - Wheel index for GB10 covering packages with no upstream aarch64 build: flash-attn, sageattention, flashinfer, and torch 2.13 on CUDA 13.3, each kernel-launch verified.
  • getainode/ainode - Browser-UI AI appliance for GB10 doing inference and LoRA fine-tuning, with UDP-discovered multi-node clustering, TP=2 on the current 0.5.x and TP=4 only on 0.4.x.
  • GuigsEvt/dgx_spark_config - Source-build guide for LLVM, Triton, and PyTorch 2.9.1 against sm_121, with release wheels and ~1.5x on 8192 FP16 GEMM versus the stock cu130 build.
  • HeKun-NVIDIA/dgx-spark-openclaw - Two-script deploy of a local LLM plus OpenClaw frontend, Qwen3.5-35B-A3B and MiniMax-M2.5-REAP-NVFP4 on a GB10 NVFP4-kernel vLLM image, or GLM-4.7-Flash on Ollama.
  • HendrikSchoettle/ragflow-dgx-spark - Build and deploy pipeline for RAGFlow v0.24.0 on DGX Spark aarch64, with a source-built onnxruntime-gpu wheel for sm_121 and multilingual OCR.
  • install-safe-press/gb10-playbooks - Chinese-language walkthrough of NVIDIA's official GB10 playbooks covering the basics and AI-agent sections, with hardware, DAC cabling, and Dell switch-config notes of its own.
  • JetBrains-Hardware/spark-setup - Remote deploy scripts for Qwen, GPT-OSS 120B, Nemotron 3 NVFP4, and Gemma 4 on vLLM, with MTP speculative decoding holding 14.4 tok/s at 200k context.
  • jschmied/dgx-spark-setup-guide - End-to-end DGX Spark guide: llama-server router mode, key-only SSH, DCGM monitoring, SWE-bench evals, and vLLM appendices where MTP-3 takes Qwen3.6-35B-A3B NVFP4 from 73 to 102 tok/s single-stream.
  • m9h/neurocontainers-arm - Prebuilt causal-conv1d wheel plus four published neuroimaging containers built against NGC PyTorch CUDA 13, with locally built FreeSurfer 8.2.0 packages inside two of them.
  • mARTin-B78/dgx-spark_lite-llm_llama-swap_vllm_llama-cpp_ollama - Multi-engine LLM stack for DGX Spark with llama-swap idle eviction behind a LiteLLM gateway, plus a harness pairing llama-benchy with a 69-scenario tool-calling bench.
  • natolambert/dgx-spark-setup - Training setup for CUDA 13.x on aarch64, cu130 vLLM wheels, SDPA over flash-attn, swap-off OOM guards, and profiled SFT, DPO, and LoRA batch ceilings.
  • seitzbg/onnxruntime-gpu-sm121-aarch64 - Prebuilt onnxruntime-gpu 1.27.1 aarch64 wheel, CUDA 13.x execution provider for sm_121, 7.3x over CPU on GB10.
  • Sggin1/DGX-SPARK - Dated GB10 lab notes, TurboQuant 3-bit KV cache at 240K context, dual-Spark 195 Gb/s RDMA, and sm_121a FP4 SASS evidence.
  • sjug/dgx-spark-ethernet-patch - Binary patch for the DGX Spark OOBE ethernet-detection bug, an 8-byte aarch64 HasInternet edit for FastOS 1.120.38.
  • Th0rgal/cuda-blackwell-carry-bug - Repro for a PTXAS bug that drops the carry flag between separate inline-asm blocks on GB10 sm_121, with a Docker harness and a 128-bit reference path to compare against.
  • timothystewart6/ubuntu-gb10 - Ubuntu 24.04 setup for GB10 in place of DGX OS, Ansible roles for NVIDIA driver, CUDA 13.x, DOCA-OFED, dual-node NCCL, and a 33-check read-only verify playbook.
  • tonyd2wild/DGX-Spark-Hard-Poweroff-Fix - Diagnosis of GB10 log-less hard power-offs as an embedded-controller cut, fixed by a 2200 MHz clock cap and page-cache drops at 5% decode cost.

Inference & Serving

vLLM

  • 0xSero/deepseek-v4-flash-0731-spark-sparkinfer - DeepSeek-V4-Flash-0731 on one DGX Spark via EXL3 and SparkInfer sparse MLA, 38.1 tok/s median c1 code decode at a 262K limit, 432-byte NVFP4 KV disabled.
  • AEON-7/vllm-ultimate-dgx-spark - DGX Spark vLLM 0.29.0 image compiled for sm_121a with DSpark quantized Markov heads, DFlash 2 at 3.39x single-stream on Qwen3.8-27B, and Triton NVFP4 KV cache.
  • airawatraj/dgx-spark-nemotron-super-agent - Nemotron-3-Super-120B agentic stack on DGX Spark scoring 93 of 100 on tool-eval-bench, with spark-arena 23.7 tok/s.
  • albond/DenseSpark-Qwen3.8-27B - Qwen3.8-27B INT4 AutoRound on one DGX Spark with vLLM 0.27.1 and sm_121 kernels, MTP drafting, 49.1 tok/s at one request and 260.2 at 16.
  • albond/SingleSpark-Qwen3.8-Flash-Next - Qwen3.8-Flash-Next on one DGX Spark from a 101 GiB four-bit checkpoint with the n-gram table kept resident, 43.34 tok/s median single stream at 65,536 context.
  • Anemll/dspark-vllm-gx10 - Two-node GB10 port of DeepSeek-V4-Flash DSpark to vLLM 0.25.1 with nvfp4_ds_mla KV format and a b12x MXFP4 MoE backend, 48.5 tok/s decode at TP=2.
  • atcuality2021/vllm-gb10-gemma4 - Vendored ManthanQuant 3-bit Lloyd-Max KV-cache compression patched into vLLM's attention backends for Gemma 4 on GB10, 5.12x smaller KV at ~0.978 cosine, quantized in NumPy on the Grace CPU.
  • Avarok-Cybersecurity/dgx-vllm - vLLM image for DGX Spark, tag v22 from February 2026, NVFP4 through a software E2M1 conversion and Marlin backends, ~42 tok/s on Qwen3-Next-80B-A3B-Instruct-NVFP4, 20% over AWQ INT4.
  • bjk110/spark_vllm_docker - vLLM serving from one DGX Spark at TP=1 to two over 200 Gb/s RoCE at TP=2, 42 presets, pinned DeepSeek-V4-Flash-0731 and Solar-Open2-250B production images with rollbacks.
  • blazux/qwen3.8-Flash-DGX - Qwen3.8-Flash-Next NVFP4 on one GB10 with the 48 GiB PLE table mmapped from NVMe, fixing a prefix-cache hit that restored an all-zero Mamba state and a non-deterministic sparse-attention top-k.
  • dolf3131/qwen3.8-flash-next-dgx-spark - NVIDIA's Qwen3.8-Flash-Next NVFP4 on one DGX Spark, 47.7 GiB n-gram table paged to SSD swap, 33.0 tok/s single-stream at 524K context, plus the PLE-offload hang at TP=1.
  • EmilHaase/DGX-Spark-VLLM-Hydra-Manager - vLLM manager for DGX Spark with sm_121a source builds and UMA KV-cache limits for multi-model launch.
  • Entrpi/ds4-spark-vllm - One-command 2-bit DeepSeek-V4-Flash vLLM install on a single DGX Spark, 85 GiB IQ2_XXS plus Q2_K checkpoint validated against antirez/ds4 at 1.75 tok/s under enforce-eager.
  • eugr/spark-vllm-docker - vLLM Docker for one to eight DGX Sparks, native PyTorch distributed by default with Ray opt-in, and 42 run-recipe.sh YAML recipes, Qwen3.8-Flash-Next-NVFP4 solo to GLM-5.2-NVFP4 on eight nodes.
  • gitcommit90/glm-5.3-one-spark - GLM-5.3-Flash EXL3 2.05 bpw on one DGX Spark under vLLM TP1 with DFlash2 K5, 40.1 tok/s on code and 29.9 on prose at 262K context.
  • jordanovski/overdrive - Web console and CLI for launching concurrent vLLM containers on DGX Spark, with preflight GPU-memory admission control and a SWE-bench page comparing resolution rates.
  • mark-ramsey-ri/vllm-dgx-spark - Run vLLM on 1-to-N DGX Spark servers (single Spark, 2 via direct cable, or 3+ via switched fabric) to serve or benchmark LLMs.
  • MiaAI-Lab/Nemotron-Labs-3-Puzzle-75B-DGX-Spark - Nemotron-Labs-3-Puzzle-75B-A9B NVFP4 hybrid Mamba MoE on one DGX Spark, vLLM 0.24 launcher with aarch64 NCCL and FlashInfer cuda_ipc patches, 256K context, MTP k=3.
  • MiaAI-Lab/Ornith-1.5-35B-A3B-DGX-Spark - Ornith-1.5-35B-A3B NVFP4 with in-checkpoint MTP on one DGX Spark, 86.3 to 440 tok/s at 24 streams, plus two b12x patches for CUDA-graph capture.
  • MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark - Qwen3.8-Flash-Next NVFP4 on one DGX Spark under vLLM with the PLE table packed and memory-mapped, 48.7 tok/s prose single stream and 162.9 at eight on 512k YaRN.
  • mouwp2026/qwen3.8-flash-next-gx10-mtp-hashk - Qwen3.8-Flash-Next on one GB10 with the PLE table compressed to a 12.8 GB HashK build, FP8 dense cast, and a shrunk-vocabulary MTP head, 104 tok/s at six streams.
  • mrexodia/Kolibri-1-vLLM-DGX-Spark - Kolibri-1 FP8 on one DGX Spark under vLLM 0.29 with Aleph Alpha's inference plugin, 262,144-token context, 42.0 tok/s generation and 97% prefix-cache hits in a pi coding-agent run.
  • omnia-projetcs/spark-dgx - Interactive vLLM Docker launcher for DGX Spark, 22 preset model configs from single-node NVFP4 to TP=4 Ray clusters, 10 with measured TTFT and concurrency tables.
  • Sapid-Labs/vllm-spark-arena - Crowd-optimization arena for vLLM on sm_121, scoring sitecustomize.py patches over a pinned wheel as paired ratios, gated on byte-identical output and a held-out timed speedup.
  • sayyidfareed/qwen3.8-flash-next-dgx-spark-1m - Qwen3.8-Flash-Next NVFP4 at a validated 989,801-token request on one ASUS GX10, clean PLE pages released by MADV_DONTNEED watermark, 5/5 needles at 26.7 tok/s single stream.
  • spark-arena/sparkrun - One-command launcher for vLLM, SGLang, and llama.cpp on one or more DGX Sparks, where --tp 2 means two hosts over auto-detected RDMA, plus git-based recipe registries.
  • sudoingX/dgx-spark-ling - Official Ling-3.0-flash INT4 on one DGX Spark under vLLM with its MTP layer, 38.7 tok/s against 35.2 for the community GGUF, but 7.9 against 33.6 at 45K context.
  • timothystewart6/vllm-gb10 - Prebuilt GHCR vLLM image for GB10 (sm_121a) with NCCL and FlashInfer built from source, pinned by SHA, digest or version, and latest promoted only after a four-model verification gate.
  • tonyd2wild/Qwen3.8-Flash-Next-NVFP4-DGX-Spark - NVIDIA's own Qwen3.8-Flash-Next NVFP4 checkpoint byte for byte on one DGX Spark, the 47.68 GiB PLE table left on NVMe and gathered 16 rows per token, 43.9 tok/s median.

llama.cpp

  • 0xBakeer/qwen38-flash-next-spark - Qwen3.8-Flash-Next on one DGX Spark with the 51B lookup table on SSD, in two profiles that differ by 2x: 88 tok/s rewriting a file against 32 on prose.
  • cahlen/glm-5.3-flash-GGUF-1bit-dgx-spark - GLM-5.3-Flash as a 1-bit UD-IQ1_S GGUF on one DGX Spark for agentic coding, where leaving --reasoning-budget unset returns nothing on 35% of turns against 0 of 80 with it.
  • croll83/llama.cpp-dgx - Deprecated llama.cpp fork for DGX Spark, kept for its TurboQuant TQ3_0 KV cache and weight kernels, which upstream does not expose.
  • gitcommit90/angelslim-hy3-iq1m-mtp-dgx-spark - AngelSlim Hy3 IQ1_M GGUF with MTP drafting on one DGX Spark through patched llama.cpp, 100K context, 18-19 tok/s on prose and 26-30 on code and JSON.
  • marknx/flash-next-gguf-tools - Qwen3.8-Flash-Next GGUF split across one DGX Spark and an RTX 5090 over 10 GbE, 140 tok/s aggregate at eight streams that OOM-kill the single box, plus converter fixes.
  • phuongncn/qwen3.6-27b-speedhack-gx10-dgx-spark - DFlash block-diffusion spec-decode llama.cpp fork for Qwen3.6-27B on GB10, 7-11 to 38-40 tok/s coding via a p_min drafting threshold, and 60-66 to 113 on a 35B-A3B MoE.
  • Sapid-Labs/llamacpp-spark-arena - Crowd-optimization arena for llama.cpp CUDA kernels on sm_121, with a thermal gate, alternating baseline and candidate runs, and referee-verified held-out speedup.
  • shamily/gemma4-llama-dgx-spark - Dockerized llama.cpp for all four Gemma 4 models on GB10, benchmarked at 69.9 tok/s tg128 for the 26B-A4B MoE against 11.0 for the dense 31B.
  • sxuff/ternary-bonsai-2-27b-gx10 - Ternary-Bonsai-2-27B on one DGX Spark under the PrismML llama.cpp fork, PTQ1_0 (1.75 bpw) at 34.21 tok/s tg128 in 5.53 GiB, build pinned to commit 1a07bfa.
  • Weschera/GLM-5.3-Flash-Unsloth-1x-DGX-Spark - GLM-5.3-Flash UD-IQ3_XXS GGUF on one DGX Spark under Unsloth's llama.cpp fork, MTP n=2 at 20.8 tok/s against 15.5 without, 64K context.

SGLang

  • 0xBakeer/ling3-flash-spark - Ling-3.0-flash MXFP4 on one DGX Spark pairing the Humming MoE backend and an online FP8 LM head with DSpark drafting, with the cookbook and notebook configs selectable for comparison.
  • 0xWhiteMage/qwen3.8-27b-kearuga-sglang-dgx-spark-dflash2 - Qwen3.8-27B Kearuga target with a distilled DFlash 2 drafter on SGLang for one DGX Spark, 35.33 tok/s at C1 and 108.90 at C4, +14.5% and +10.8% over stock.
  • BTankut/dgx-spark-sglang-moe-configs - Tuned Triton MoE configs for GB10's 101,376-byte shared memory limit, where SGLang defaults need 147,456 and EAGLE crashes, GLM-4.7-FP8 at 20-27 tok/s on four nodes.
  • hasso5703/dgx-spark-qwen38 - One-command SGLang service on GB10 for seven switchable Qwen3.8 targets, Qwen3.8-27B NVFP4 at 71.4 tok/s greedy single-stream median via DFlash2 and Qwen3.8-Flash-Next 176B at 262K context.
  • InquiringMinds-AI/longcat-next-multimodal - LongCat-Next 75B-A3B any-to-any multimodal through one SGLang process on a single GB10, image generation and voice-clone TTS on OpenAI endpoints at w8a8_int8 after 4-bit collapsed both.
  • mark-ramsey-ri/sglang-dgx-spark - Run SGLang on 1-to-N DGX Spark servers (single Spark, 2 via direct cable, or 3+ via switched fabric) to serve or benchmark LLMs.
  • MiaAI-Lab/Nemotron3.5-Lightning-DGX-Spark-RTX-5090-6000-PRO - Nemotron 3.5 Lightning 30B-A3B NVFP4 with its DSpark draft model on one DGX Spark via SGLang, 4.93M-token KV pool, 48 max concurrent, up to 1M per request.
  • MiaAI-Lab/Qwen3.8-27B-SGLang-DGX-Spark - Qwen3.8-27B NVFP4 on one DGX Spark with EAGLE/MTP, DSpark and DFlash2 as measured swap-in modes, GDN in bf16 and the scheduler pinned to the X5 cores.
  • pangoleen/qwen3.8-27b-dgx-spark-dflash2 - Qwen3.8-27B NVFP4 on one DGX Spark with a DFlash2 draft budget of 16, 64-78 tok/s single stream on code against ~30 on chat, 387 aggregate at 32 streams.
  • robbiemu/dgx-spark-inference - SGLang services on one DGX Spark under systemd with per-role memory budgets, refusing a launch that will not fit, on a digest-pinned v0.5.14-cu130 runtime.
  • scottgl9/sglang-spark-gb10-optimizations - SGLang fork that routes NVFP4 through Marlin FP4 where CUTLASS returns zeros on sm_121, Qwen3.5-122B-A10B at 43-45 tok/s with ~90% MTP acceptance.
  • ubehera/sglang-spark - SGLang patch stack for DGX Spark with an sm_121a-only sgl-kernel 0.4.4 wheel, model-agnostic TP=2 launch recipes at NEXTN k=3, and a systemd watcher that stops wedged nodes.
  • Weschera/Qwen3.8-27B-NVFP4-DFlash2-DGX-Spark - Qwen3.8-27B NVFP4 with DFlash2 on one DGX Spark, digest-pinned, where a boot-time fp8_gemm autotune race decides between 42 and 33 tok/s for the process lifetime.

Other Engines

  • 0xBakeer/deepseek-v41-flash-spark - DeepSeek-V4.1-Flash on one DGX Spark in a plain-PyTorch engine with routing-ranked expert keep-sets, 24.3-36.6 tok/s by workload at 44% keep, above the reliable 39% default.
  • 0xBakeer/TandemLLM - Inference engine for Qwen3.8-27B on one DGX Spark with StairCut draft-tree sizing and custom NVFP4 kernels, 49.89 tok/s single request with the lookup store off against 13.69 plain greedy.
  • antirez/ds4 - DwarfStar C inference engine for DeepSeek, GLM and Qwen3.8 Flash Next with a make cuda-spark target, DeepSeek V4 Flash Q2 prefill above 820 t/s through 65K, decode 18.1 to 13.8.
  • ashhart/TensorFold - Speculative-decoding LLM server for Apple Silicon and NVIDIA with replies equal to serial decoding, CUDA engines for one or two DGX Spark, GLM-5.3-Flash at 256k tokens on two.
  • Avarok-Cybersecurity/atlas - Pure Rust and CUDA inference engine in one 75 MB binary with GB10 as verified target, ahead of vLLM 0.27.1 at every C=1 to C=128 rung on Qwen3.8-27B NVFP4.
  • Baekpica/ds4-dfm-rs - Rust-host continuation of ds4 with DGX Spark as release reference and recorded gates per model artifact, Qwen3.8 Flash Next Q5 at 1,323 prefill and 28.93 decode tok/s with MTP 2.
  • blake-snc/sm121-kernels - Hand-written PTX kernel library for sm_121 in 259 files, covering flash attention, GEMM, Gated DeltaNet, and MoE, driver-only via cudarc with FP8 attention at ~108 TFLOPS.
  • calico88x/DGX-Model-Manager - Control plane for managing Ollama, SGLang, vLLM, llama.cpp, LocalAI, and ComfyUI on DGX Spark, with roles, API tokens, and Hugging Face cache inventory.
  • HawkBearPig/dgpp - C++/CUDA inference engine for one, two, or four DGX Spark with tensor parallelism over RoCE, 29 benchmarked configurations across GLM, Qwen, DeepSeek, and MiMo models.
  • jdaln/dgx-spark-inference-stack - Docker serving stack for a single DGX Spark with on-demand model loading, automatic idle shutdown, and a unified API gateway.
  • joshhu/meetaclawtaipei - Three concurrent NVFP4 vLLM models on one DGX Spark with a 3-LLM voice-clone roommate demo.
  • kshetrajna12/sparkstation - Headless control plane for DGX Spark model fleets with profile-driven placement over vLLM and SGLang, unified-memory admission control, and LiteLLM gateway sync.
  • lrozewicz/vLLM-Moet-GB10 - vLLM-Moet fork for one GB10 running GLM-5.3-Flash at ~29 tok/s on code and DeepSeek-V4-Flash-0731 at ~21, via 2-bit expert planes plus an FP4 delta tier.
  • mark-ramsey-ri/trt-dgx-spark - TensorRT-LLM serving for 1 to N DGX Spark with the aarch64 nvcr 1.2.1 container and TP set by node count, verified on 1 and 2 nodes.
  • MiaAI-Lab/exllamav3 - ExLlamaV3 fork that builds on aarch64 GB10, with NVFP4 and FP8 paged-attention KV at about 3.5x the context per GB of fp16, plus DFlash2, DSpark, and MTP drafting.
  • MiaAI-Lab/Qwen3.8-Flash-Next-Single-DGX-Spark-TensorFold - Qwen3.8-Flash-Next top-5-expert INT4-AutoRound checkpoint on one DGX Spark under TensorFold's Zig engine, 64.4 tok/s prose at c1 and 200.9 at c8, about 10% lower capability.
  • r0b0tlab/glm53-flash-exl3-exllamav3-gb10 - ExLlamaV3 and TabbyAPI runtime for GLM-5.3-Flash EXL3 2.25 bpw with DFlash2 on one GB10, a fail-closed memory guard, and 259,993-token retrieval at 262,144 context.
  • rdaum/eider - Rust and CUDA inference server for sm_121 NVFP4 with no tensor runtime, Qwen3.8-Flash-Next with its 51B PLE table paged from NVMe at ~300 tok/s uncached prefill, ~13 MTP decode.
  • rdoiron/mimo-mods-for-dgx-spark - Ten vLLM runtime patches for MiMo-V2.5 on sm_121a, with a CUTLASS block-FP8 bypass and a backported tool-call corruption fix (PR #42969).
  • sf-stav/veloGB10 - Rust and CUDA inference engine built only for GB10 with bitwise-lossless MTP, Qwen3.8-Flash-Next EXL3 at 137 tok/s on one machine and 221 on four, pure-code decode.
  • sixteen-miles-labs/sparklab - GB10-only inference runtime for one DGX Spark with NVMe-backed MoE expert banks, GLM-5.3 753B at 1.29 decode tok/s and DeepSeek V4.1 Flash 552B at 1.05, both experimental.
  • Th0rgal/dgx-spark-router - Stdlib-only OpenAI-compatible router that swaps six llama.cpp GGUF and vLLM NVFP4 models in and out of 128 GB unified memory, with Marlin GEMM defaults for GB10.
  • vcruz305/DeepSeek-V4.1-Flash-EXL3-DGX-Spark-recipe - DeepSeek-V4.1-Flash EXL3 1.59 bpw on one DGX Spark in native ExLlamaV3 with the DSpark drafter aliased from mmap over ATS, 17.5 tok/s median on fresh prompts.
  • xangel82/DS4-GB10-GX10-DSpark-CUDA - DS4 fork for DeepSeek-V4-Flash on one GB10, lossless DSpark and HybridLC decode at 24-26 t/s on tool calls against the original 13, 900-953 t/s prefill.

Fine-tuning

Quantization & NVFP4

GB10's Blackwell architecture supports NVFP4 (4-bit floating point) in hardware. It runs faster than INT4 at similar quality.

Models & Benchmarks

  • albond/DGX_Spark_Qwen3.5-122B-A10B-AR-INT4 - Qwen3.5-122B-A10B on DGX Spark, tuned from 28.3 to 52 tok/s (+82%) with a hybrid INT4+FP8 checkpoint and an INT8 LM head.
  • Avarok-Cybersecurity/atlas-recipes - Recipe corpus for Atlas on GB10, run by the atlasctl launcher that replaces sparkrun, with per-model KV dtype and MTP width.
  • Blackwellboy/laguna-s21-lab - Laguna S 2.1 NVFP4 testing lab on one DGX Spark, 20-cell tuning sweep with its losing cells, 12-hour soak of 3,096 turns, and 450-turn thinking-gate study.
  • colonel-otto/3x-dgx-spark-mesh-deepseek-v4-0731-flash - DeepSeek-V4-Flash at TP=3 on three DGX Spark, decode 6.7 to 20.2% faster than two nodes at n=30 on the retired Anemll engine, 84.7 tok/s c1 on eugr's b12x image.
  • DanTup/spark-evals - Harbor-run leaderboard of evo eval, aider polyglot, and swe-rebench scores under Codex and Terminus-2 agents for four model configurations that fit on one DGX Spark.
  • DG1001/local-agentic-coding-128gb - Coding-agent benchmark of 14 local models on one ASUS Ascent GX10 under three harnesses, 86 hidden tests, with Nemotron scoring 47 to 85 across 13 runs.
  • elsung/dgx-spark-deepseek-v4-flash - DeepSeek-V4-Flash official FP8 across two DGX Spark under vLLM TP=2 over a 200 Gb/s QSFP56 RoCE cable, 41 tok/s single stream and ~350 aggregate at c=32.
  • Entrpi/ds4-on-spark - DGX Spark fork of antirez/ds4 at 2.4-3.3x upstream prefill and 1.33-1.47x decode, 59 tok/s over 12 concurrent requests, and 2.26M tokens of active context at shipped defaults.
  • Entrpi/qwen3.5-122B-A10B-on-spark - Qwen3.5-122B-A10B on a single DGX Spark via DFlash block-diffusion spec-decode, 81 tok/s on agent traffic.
  • evanwtf/local-llm - Coding-agent benchmark of model, engine and harness stacks on a two-node DGX Spark cluster, ranked by suite pass rate then wall time, 29 configurations under OpenCode.
  • GaelicThunder/colibri-gb10-attention-cliff - Fixed-size kernel guard at 8192 context tokens in colibri on GB10, silent CPU fallback at 0.45 versus 0.86 tok/s, plus per-branch counters and a sweep harness.
  • GaelicThunder/DeepSeek-V4-Flash-Vision-One-DGX-Spark - DeepSeek-V4-Flash-Vision-Exp in EXL3 on one DGX Spark at 245,760 context, 37.6 tok/s on code with images in the same process, keeping 90% of the original token probability.
  • GaelicThunder/gb10-uma-inference-notes - Five measured GB10 unified-memory properties that hold across models and engines, including host-to-device copies as pure waste worth 63% once removed.
  • GaelicThunder/moe-offload-findings - Nine findings on MoE decode with experts streamed from NVMe, mostly negative: batching does not amortize expert reads and prefetching cannot beat a static pin.
  • gitcommit90/qwen38-27b-dgx-spark - Qwen3.8-27B NVFP4 on one GB10 with Inco DFlash 2 speculative decoding, 44.46 tok/s versus DSpark 31.95 at k=7, plus a BF16 lm_head guard fix.
  • hebo1221/motif3-dgx-spark - Motif-3 315B as mixed IQ2_XXS on one DGX Spark, 83.56 GiB at 16.49 tok/s, published with the author's verdict that the quant misses BF16 quality retention.
  • jeremy-newhouse/dgx-spark-nemotron-super-bench - Single-stream decode benchmark of Nemotron-3-Super-120B-A12B-NVFP4 on one GB10, ~26-27 tok/s realistic with MTP vs ~33.6 microbench.
  • jiayuqi7813/DeepSeek-V4-Flash-0731-CRACK-2x-DGX-Spark - Rank-1 Householder refusal edit of DeepSeek-V4-Flash-0731 in native FP8 UE8M0 blocks on two DGX Spark, compliance 3.53 to 90.59% with HumanEval 148 against 150.
  • jvr0x/dgx-spark-bench - Closed-loop concurrency sweeps on DGX Spark with pinned recipes and a GitHub Pages dashboard, per-session and aggregate tok/s across 43 published runs.
  • k3net/docai-evals - Evidence repository for 19 Hungarian document-AI and GB10 serving experiments, where Marlin beats the publisher-recommended MoE backend by up to 3.3x and NVFP4 roughly doubles counterparty-role errors.
  • Kleybrink/dgx-spark-bench - Ollama benchmarking framework for DGX Spark measuring throughput, latency, memory, and answer quality with an LLM-as-a-judge pipeline over 21 prompts in 9 categories.
  • marksunner/dgx-spark-single-stack - Single-box agent stack on DGX Spark, Hermes runtime and Honcho memory on CPU beside vLLM serving Qwen3.5-122B hybrid INT4+FP8 with MTP at 41-47 tok/s.
  • marksunner/dgx-spark-step37-flash - StepFun's Step 3.7 Flash (198B MoE) on a single DGX Spark with llama.cpp at ~27 tok/s and 96K context on a q8_0 KV cache, reduced from 128K for CUDA-graph stability.
  • martimramos/dgx-spark-ml-guide - Troubleshooting playbook for PyTorch on GB10, 16 numbered failures with an error-to-fix table, from cu128 nightly wheels to aarch64 mmcv builds and NVRTC symlinks.
  • Memoriant/dgx-spark-kv-cache-benchmark - KV cache quantization on GB10: q4_0 costs 37% of generation throughput at 110K context, TurboQuant turbo3 up to 23.6% at 32K, and prompt processing is untouched.
  • mneha05/gb10-attn - Context-split FlashDecoding paged-attention decode kernel at 82-85% of GB10's 273 GB/s on large working sets, fp16 gpt2 geometry, after correcting a 2x roofline and an L2-cached benchmark.
  • msuiche/weightless - Serve-time refusal steering with GGUF layer-projection vectors instead of redistributed weights, validated on Qwen3.8-27B on one DGX Spark and GLM-5.3-Flash on four.
  • nabe2030/dense-27b-31b-dgx-spark - Dense Qwen 3.5/3.6-27B and Gemma 4-31B on llama.cpp b8922, Q4_K_M at 10-12 tok/s versus 3.8-4.5 BF16, JCommonsenseQA cost 0.2-0.6 points.
  • nabe2030/gemma4-vs-qwen35-dgx-spark - Gemma 4 26B-A4B versus Qwen 3.5/3.6-35B-A3B MoE on llama.cpp, F16 thinking-mode bug isolated, 26.5 against 58 tok/s decode and 0.35 point JCommonsenseQA spread.
  • OscarActual/gb10-llm-benchmark - Ollama benchmarks for GB10 across 11 LLMs and 10 embedders: decode tok/s, TTFT, and Czech RAG recall.
  • pendakwahteknologi/gx10-benchmarks - Benchmark roster for the ASUS Ascent GX10 (GB10) with nine published runs across inference, training, efficiency, and generation, each carrying timestamped CSV and log artifacts.
  • r0b0tlab/deepseek-v4-flash-nvfp4-gb10-benchmark - DeepSeek-V4-Flash FP8 on two DGX Spark at TP=2 with MTP over RoCE, 38.4 tok/s c1 and 144.6 aggregate at c16 on driver 580.142, down about 3.5x on 580.159.03.
  • r0b0tlab/diffusiongemma-26b-nvfp4-sm121-vllm - DiffusionGemma 26B-A4B NVFP4 under vLLM on GB10 with the FlashInfer CUTLASS FP4 MoE path, 146.3 tok/s at c1 and 242.9 at c16, no Marlin fallback.
  • r0b0tlab/laguna-s-2.1-nvfp4-sm121-vllm - Laguna S 2.1 NVFP4 on GB10 under vLLM 0.25.1, DFlash K=7 at 22.2 tok/s c1 with thinking off on revision 0761412, plus an 8,620-case scorecard for the earlier 216d1f1.
  • r0b0tlab/nex-n2-mini-nvfp4 - NVFP4 vLLM container for Nex-N2-mini (Qwen3.5-MoE-35B) on GB10, 185 tok/s aggregate at concurrency 8.
  • r0b0tlab/step37-flash-nvfp4-sm121-vllm-docker - vLLM container for StepFun's Step 3.7 Flash NVFP4 (198B MoE VLM) on dual GB10 TP=2, with verified native-CUTLASS sm_121 execution at 16.49 tok/s.
  • ramsred/llm-engine-benchmark - Long-context comparison of vLLM, SGLang, and TensorRT-LLM on GB10 under one client, plus a prefill-budget study cutting P95 TTFT 26.9% by moving 8192 to 2048.
  • styles01/sparkrun-recipes - Source-pinned SparkRun recipes for nine exclusive lanes on one GB10, led by a Qwen3.8-Flash-Next EXL3 native-MTP daily driver at 54.3 tok/s single-stream decode.
  • ThinkCode/glm53-flash-2x-gb10-bench - GLM-5.3-Flash NVFP4 against EXL3 on two GB10, with NVFP4 ahead cold at every concurrency and EXL3 ahead 1.6x at warm C4 on prefix-cache reuse NVFP4 never gets.
  • VincentMarquez/glm52-gb10-colibri - GLM-5.2 744B on one DGX Spark at full top-8 via a colibri engine branch, 11.1 tok/s on a synthetic prompt with PLD, 5-7 on real chat.
  • Weschera/spark-bench - LLM benchmark for DGX Spark across 80 scenarios in 13 domains, with 12 multi-turn agentic workflows, 4 machine-graded long-generation builds, and a TrueScore weighting speed at 5%.

Multi-node

You can connect two DGX Spark units directly over 200 Gb/s QSFP for double the memory and compute.

  • 0xdfi/GLM-5.2-1M-4x-DGX-Spark - Profile index for unpruned GLM-5.2 744B on 4x DGX Spark with NVFP4 DS-MLA KV, O14 Fast at 250K total KV, 819 tok/s prefill, 42.3 peak decode.
  • 0xdfi/GLM-5.2-R9-Adaptive-MTP-FULL-CUDA-4x-DGX-Spark - GLM-5.2 on four DGX Spark nodes with FULL CUDA graphs for every adaptive MTP depth K2/K4/K5, 83.4 tok/s at C4 measured before the 420K retune, 520K balanced profile.
  • 0xSero/glm-5.3-flash-sglang-sm121 - GLM-5.3-Flash NVFP4 under SGLang on two or four DGX Spark, 22.69 tok/s single stream and 82.37 at eight on TP=4, digest-pinned with an acceptance checklist.
  • ajensenwaud/Kolibri-1-2x-DGX-Spark-TensorFold - Kolibri-1 on two DGX Spark under TensorFold with 17 patches over PR #328, 83.7 tok/s single-stream decode at FP8 against 47.1 on one Spark.
  • alexellis/glm-5.3-flash-4x-dgx-spark-switchless - GLM-5.3-Flash NVFP4 at TP4 across four DGX Spark cabled as a switchless RoCE ring, ~45 tok/s on agentic traffic over patched NCCL 2.30.7 that skips the wedging tree connect.
  • ArgentAIOS/dgx-spark-cluster - Two-node DGX Spark guide with DMA-BUF NCCL settings for GB10 where nvidia-peermem GPU Direct fails, 93.5% DDP scaling on a 200 Gb/s RoCEv2 link.
  • bertholomus/deepseek-v4.1-tensorfold-tp2-2xgb10 - DeepSeek-V4.1-Flash EXL3 2.9 bpw on two DGX Spark under a TensorFold TP2 engine with DSpark drafting, 1,048,576-token window at four streams, 100.5 tok/s greedy single-stream code.
  • bird/GLM-spark - GLM-5.2 across four DGX Spark nodes, 1M context at 24.7 tok/s on unpruned Int4-Int8Mix weights, TP4 plus DCP4 KV sharding and lossless speculative decoding.
  • botAGI/dspark-0731-gb10 - DeepSeek-V4-Flash-0731 on two DGX Spark under vLLM 0.25.2, flashinfer_b12x MoE backend worth 9-12% step rate to 256K, 28 tok/s single stream at 1,042,600 tokens.
  • bumasoft/Qwen3.8-Flash-Next-TP3-DGX-Spark-VLLM - Qwen3.8-Flash-Next FP8 across three DGX Spark at tensor-parallel 3, checkpoint padded in place at 60.9 tok/s, where padding key heads alone silently breaks the 1:3 GDN pairing.
  • chadhurley25075-png/pd-bridge - DeepSeek-V4-Flash prefilled on two DGX Spark, decoded on a Mac Studio over plain 10 GbE with no shared cache format, 3.7x vs the Mac alone on a 241K cold prompt.
  • chishiki37/dgx-spark-fabric - Switched MikroTik CRS812 fabric replacing the direct DAC between two DGX Spark nodes, fixing broken RoCE for 26-33% more decode, with 100 and 200 Gb/s breakouts measured identical.
  • christopherowen/dgx-spark-networking - Switchless RoCE collectives library for DGX Spark rings, one-shot BF16 all-reduce at 16.9 µs against NCCL's 90.6 to 92.4 µs for 10 KiB on four nodes.
  • christopherowen/spark-ds41f - DeepSeek-V4.1-Flash on three switchless DGX Spark at TP3 with TileLang sm_121 kernels, 52.2 prose and 63.0 code tok/s single stream at 512K context.
  • ciprianveg/gb10-vllm - Model-agnostic sm_121 vLLM platform, KIMI-K3 v5 at TP16 plus DCP8 with RoCEnante collectives on 16 GB10, 29.81 tok/s at C1 and 87.12 aggregate at C8, GLM-5.3 on GLM-5.2's v19 image.
  • CosmicRaisins/glm-5.2-gb10 - Unpruned GLM-5.2 Int4-Int8Mix on four GB10 nodes with DCP KV sharding, 320K context at 598 t/s prefill or 640K at 430, decode flat near 22 t/s.
  • CosmicRaisins/minimax-m3-awq-gb10 - MiniMax-M3-AWQ-INT4 vLLM serve recipe for 4x GB10, FP8 KV cache, EAGLE3 spec-decode, and indexer-corruption fix.
  • digchick/dgx-spark-200g-link-fix - Troubleshooting playbook for the 200G ConnectX-7 link failing to train between two Sparks (CX7 hotplug power-saving), with the fix and NCCL/RoCE verification.
  • drowzeys/keys-1M-CTX-Inkling-Small-NVFP4-Dspark-NVFP4-KV-Cache-SGlang-SM121-optimized-on-Two-DGX-Sparks - NVFP4 KV cache in the SGLang triton backend for Inkling-Small on two DGX Spark, 3.12x pool for 1M context, 23.9 tok/s on a pooled open-ended mix.
  • drowzeys/keys-MiMo-V2.6-Pro-MOPD-C3-Jarrelscy-ARVQ-4-DGX-Sparks-1M-Context - MiMo-V2.6-Pro-MOPD in Jarrelscy's ARVQ and NVFP4 hybrid on four DGX Spark at TP4 and 1M context, single-stream 32.2 tok/s on code and 24.2 on prose with MTP k=2.
  • drowzeys/Keys-Setup-Autonomous-Self-Improving-Local-Inference-Stack - Mixture-of-Agents stack for four DGX Spark nodes with a DeepSeek-V4-Flash router and an NVFP4 Two-Tower consolidated onto one node at 29 tok/s.
  • drowzeys/keys-TensorFold-GLM-5.3-TP4-4x-DGX-Spark - GLM-5.3 753B at EXL3 2.75 bpw on four DGX Spark under TensorFold TP4 with MTP 2, 40.2 tok/s prose with thinking on, needle found at 905,182 tokens.
  • Enntity/sparkglm - GLM-5.3-Flash NVFP4 on two DGX Spark under the Atlas engine with DFlash2 and prefix caching, 0.67 s median first token on turns 2+ of 30-48K conversations.
  • epappas/gx10-cluster - Ansible provisioning for a GB10 cluster over ConnectX-7 with diagnostic tools, a recipe catalogue, and measured-or-forum-read labels on hardware claims, 22.7 GB/s two-node NCCL busbw.
  • FujitsuPolycom/sparkring - One-command installer and vLLM/SGLang serving for switchless two- or four-Spark GB10 rings, using RoCEnante RDMA collectives for decode and NCCL 2.32.3 for larger ones.
  • HeNryous/mimo-v25-dflash-dgx-spark - MiMo-V2.5 (309B MoE, NVFP4 plus a required MXFP8 o_proj overlay) on two DGX Spark nodes, DFlash drafting ~54 tok/s on structured content, 1.67M-token fp8 KV pool at 500K.
  • idonati/spark-vllm-docker-festr2 - vLLM patches for festr2 MiMo-V2.5-Pro NVFP4/MXFP8 on an 8-node sm_121 cluster, fused-QKV fix for Q mis-slotted as K/V on 7 of 8 ranks, 114 tok/s aggregate at 20 concurrent.
  • jakejharris/jspark3 - GLM-5.3-Flash on three DGX Sparks with a TensorFold fork and DFlash2 drafting, RigMark estimates at reasoning effort low of 91.3 tok/s code, 51.6 prose, 2,124 tok/s cold 64K prefill.
  • jayleaton/glm53-tensorfold-spark - GLM-5.3-Flash abliterated EXL3 4-bit on two DGX Spark under TensorFold plus 77 patches, RigMark 57.5 tok/s at C1 and 95.1 at C4.
  • joesinvestments/DeepSeek-V4-Flash-0731-TP4-4x-DGX-Spark - DeepSeek-V4-Flash-0731 at TP=4 on four DGX Spark under vLLM 0.25.2 with DSpark k=7, 123.13 tok/s single stream at 2K tokens, 57.9 at ~150K agentic prompts.
  • joesinvestments/glm52-spark-kit - GLM-5.2 on four DGX Spark as source overlays over stock vLLM, fused nvfp4_ds_mla KV writer at 51 times the torch reference, sparse indexer unblocked above DCP 1.
  • joeynyc/Hy3-295B-NVFP4-2x-DGX-Spark - Tencent Hunyuan 3 295B MoE in NVFP4 at TP2 on two DGX Spark, 26 tok/s single stream at 262K with TurboQuant k8v4 KV, MTP measured 20% slower.
  • joeynyc/MiniMax-H3-2x-DGX-Spark - Cross-host diffusion executor for one MiniMax H3 video across two DGX Spark nodes, Ulysses sequence parallelism over RoCEv2, 68.8 s against 155.0 s single-box.
  • josephdrose/joe-spark-patches - Out-of-tree vLLM patches for DeepSeek on four DGX Spark, Engram kept on NVMe because cpu_offload frees nothing on GB10, plus a DSpark proposer that survives concurrent requests.
  • josephdrose/nccl-spark-switchless - NCCL 2.30.7 patches for a switchless 4-node GB10 RoCE ring, tree-skip plus 2-hop relay over the uncabled diagonals, MiniMax-M3 229B NVFP4 at ~24 tok/s.
  • karolpalys/glm52-triple-spark-tuning - Tuning log and negative results for GLM-5.2 753B on three DGX Spark nodes, with the measurement methodology and the evidence that killed each rejected change.
  • kindlingai/glm-5.3-flash-gx10 - GLM-5.3-Flash NVFP4 on two, three, four or six GB10 under vLLM with mentat in place of Ray, 117.3 tok/s code decode at TP=4 with DFlash2.
  • kindlingai/glm-5.3-full-exl3-tp6 - Full GLM-5.3 EXL3 3.25 bpw on six DGX Spark at TP6 by re-fragmenting experts without requantizing, 33.8 tok/s prose at MTP k=2, 958 tok/s prefill at 32K.
  • kingjones30/GLM-5.3-Flash-2x-DGX-Spark - GLM-5.3-Flash NVFP4 on two DGX Spark, NoPE latent padded to the rope width sm_121 kernels expect, 24.7 tok/s on code under MTP-5 on the stock image.
  • knapcio/DeepSeek-V4.1-Flash-4x-DGX-Spark-TP4 - DeepSeek-V4.1-Flash TP4 on four DGX Spark under SGLang with RoCEnante all-reduce, 89.7 tok/s prose and 131.9 code at c1 with the SM clock capped at 2200 MHz.
  • knapcio/GLM-5.3-Flash-4x-DGX-Spark-TP4 - GLM-5.3-Flash NVFP4 at TP4 on four DGX Spark with 8-bit dense layers and DFlash2, 90.3 tok/s prose single stream, 3,470 tok/s cold prefill at 32k.
  • Libertai/glm53-flash-vllm-gb10 - GLM-5.3-Flash under vLLM on two DGX Spark at 24.2 tok/s through a hand-written sparse-MLA kernel, with degenerate output traced to an uninitialized MoE activation scale.
  • magicbear/DeepSeek-V4.1-Flash-SGLang-DGX-Spark - DeepSeek-V4.1-Flash under SGLang on four against eight DGX Spark, c1 code decode 57.9 against 67.0 tok/s, 1.63-1.76x prefill at eight, while its own vLLM cross-check reaches 75.5 on four.
  • makiisthenes/dgx-spark-multinode-vllm-ray - Dual-DGX Spark vLLM deployment with NVIDIA vLLM 26.04, Ray, and 200 GbE QSFP.
  • maliubiao/dgx-spark-2-deepseek-flash-0731 - Bilingual manual for DeepSeek-V4-Flash-0731 and Vision-Exp on two DGX Spark, with LMCache v0.5.3 built on aarch64 for NVMe cold KV backup, 1 GiB L1, 500K context.
  • MiaAI-Lab/DeepSeek-v4-Flash-DSpark-2x-DGX-Spark - Two-node vLLM TP=2 recipe for DeepSeek-V4-Flash-Vision-Exp with native image input, DSpark speculative decode and nvfp4_ds_mla KV at a 1M-token ceiling.
  • MiaAI-Lab/DeepSeek-v4.1-Flash-DGX-Sparks - DeepSeek-V4.1-Flash under SGLang on three or four DGX Spark with NVMe-resident Engram tables, TP4 prose c1 87.7 tok/s on a switch, 1M-context needle passed at 1,011,084 tokens.
  • MiaAI-Lab/DeepSeek-v4.1-Flash-EXL3-2x-DGX-Sparks - DeepSeek-V4.1-Flash at 2.9 bpw EXL3 on two DGX Spark with vLLM TP2 and file-backed Engram, 31.6 tok/s single stream with DSpark k=3, 600K context.
  • MiaAI-Lab/GLM-5.2-NVFP4-AQLM-Triple-DGX-Sparks - GLM-5.2 NVFP4 plus AQLM 2-bit hybrid on three DGX Spark at TP3, 21 tok/s structured at 348K vision context with top-4 routing, 25-26 on fp8 KV at 235K.
  • MiaAI-Lab/GLM-5.3-EXL3-3x-DGX-Sparks-TensorFold - GLM-5.3 EXL3 2.75 bpw on three DGX Spark under TensorFold, a 499,712-token window by context parallelism, 41.7 tok/s code, a 94K prompt resumed from NVMe in 3.0 s.
  • MiaAI-Lab/GLM-5.3-Flash-EXL3-2x-DGX-Sparks-TensorFold - GLM-5.3-Flash EXL3 on two DGX Spark under TensorFold v0.6.0 with 96 patches, 1,048,576-token window at 4 concurrent requests, 60.4 tok/s prose and 114.7 structured at c1.
  • MiaAI-Lab/GLM-5.3-Flash-NVFP4-Dual-DGX-Spark - GLM-5.3-Flash NVFP4 on two DGX Spark over Ray TP=2 with image and video input, MTP at four draft tokens, FP8 KV cache and 262K context.
  • MiaAI-Lab/Inkling-Small-NVFP4-Dual-DGX-Sparks - Inkling-Small-NVFP4 across two DGX Spark on SGLang at a 1,142,712-token KV pool, 33.9 tok/s single stream, page size 1 required for the triton DSpark verify path.
  • MiaAI-Lab/MiMo-V2.6-Flash-2x-DGX-Sparks - MiMo-V2.6-Flash MXFP4 under SGLang on two DGX Spark, DFlash at 69.8 tok/s on code and EAGLE MTP at 35.0 on prose, single stream.
  • MiaAI-Lab/Qwen3.8-Flash-Dual-DGX-Sparks-TensorFold - Qwen3.8-Flash-Next NVFP4 on two DGX Spark under TensorFold's Zig engine at TP2 with FP8 KV, 1,048,576-token window, 67.4 tok/s prose at c1 and 390.1 at c16.
  • MiaAI-Lab/Qwen3.8-Flash-Next-Dual-DGX-Sparks - Qwen3.8-Flash-Next NVFP4 on two DGX Spark under vLLM TP2 with expert parallel and MTP-3, 56.8 tok/s prose at one stream and 225 aggregate at eight with 47k draft vocabulary.
  • nabe2030/dgx-spark-2node-rpc - GLM-5.2 GGUF at 228.5 GB split across two DGX Spark nodes over llama.cpp RPC, CX7 measured as two ~100 Gb/s PCIe Gen5 x4 paths, RDMA vs TCP A/B.
  • nacyot/vllm-ds4f-gb10 - vLLM fork for DeepSeek-V4.1-Flash on four GB10 at TP=4 with disk KV offload, a 493K-token session restored from SSD in 7.8 s against a 392 s cold prefill.
  • neko-legends/spark-bench - Four DGX Sparks as one TP4 cluster over six dated model lanes, the live one DeepSeek-V4.1-Flash uncensored on TensorFold at 66.5 tok/s prose and 104.1 code, 1k-160k prompts.
  • ondigo-winder/SWITCHLESS-NCCL-RING-4-GB10 - NCCL 2.30.7 patch for a switchless four-node GB10 ring using both PCIe halves of each QSFP cable, 193.1 Gb/s all-reduce bus bandwidth against 111.6.
  • OsakaTX/qwen3.8-flash-next-vllm-dgx-spark - Qwen3.8-Flash-Next NVFP4 on two DGX Spark under vLLM TP2 with MTP n=3, 41-44 tok/s c1 and 162 aggregate at eight, RDMA passthrough worth 40-45%.
  • pfn/spark-vllm-compose - Head and worker Docker Compose files that run vLLM across DGX Spark nodes with native --nnodes/--node-rank instead of Ray, shipped services for Qwen3.5-397B-A17B-int4 and MiniMax-M2.7.
  • r0b0tlab/DeepSeek-V4-Flash-DSpark-v026-SM121 - Dual-GB10 DeepSeek-V4-Flash-0731 on vLLM 0.26 with a digest-pinned GHCR image and runtime gate, NIAH PASS at 1M, BFCL multi_turn_base 0.755, ~79 tok/s c1 decode.
  • r0b0tlab/qwen38-flash-next-w4a16-sm121-sglang - Qwen3.8-Flash-Next self-quantized to NVFP4 W4A16 on two GB10 under SGLang TP=2 with NEXTN MTP, 63.84 tok/s at c1 and 132.64 at c4 on 1,024-token prompts.
  • rajsinghtechbot/dgx-spark-vllm-k8s - Kubernetes cookbook for DeepSeek-V4-Flash on dual DGX Spark, with Multus/Spiderpool RDMA over RoCEv2, UMA-aware container memory limits, and Prometheus monitoring.
  • raullenchai/twinspark - Self-healing two-node vLLM cluster for DeepSeek-V4-Flash-0731 on DGX Spark at 74.8 tok/s, needle sweep cutting long-context failures from 8/25 to 2/25 by trading nvfp4_ds_mla KV for fp8_ds_mla.
  • Reederey87/dgx-spark-2x-deepseek-v4-flash - DeepSeek-V4-Flash-0731 on two DGX Spark from a full-source vLLM main build, four gated one-file upstream fix layers promoted on Welch non-inferiority, 3,027,217-token KV pool.
  • Reederey87/glm53-flash-exl3-2x-dgx-spark - GLM-5.3-Flash EXL3 on two DGX Spark at 1M context, 110k replay prefix-cache hits lifted from 0 to 97-98% by 3,584-token batches plus a drafter-prune patch.
  • RustRunner/DGX-Llama-Cluster - Three DGX Spark nodes as a switchless ConnectX-7 RDMA star running llama.cpp RPC, 384 GB pooled unified memory for 400B+ models in 4-bit, NFS model share.
  • sfxnz/DeepSeek-V4-Flash-Vision-Exp-vLLM-2x-DGX-Spark - DeepSeek-V4-Flash-Vision-Exp on two DGX Spark at TP=2, ViT loaded as a vLLM plugin that preserves DSpark-6 draft acceptance, 26.2 tok/s prose single stream at 1M context.
  • strusty/GLM-5.3-Flash-NVFP4-3x-DGX-Sparks-kindling-DGXOS - GLM-5.3-Flash NVFP4 at TP=3 on three DGX Spark in a switchless triangle on stock DGX OS, 3,054,135-token KV pool, about 3,500 tok/s cold prefill at 64k.
  • tomsti/guides - GB10 cluster guide for DGX Spark over ConnectX-7 RoCE, covering NCCL rail pinning, the duplicate-MAC workaround, and MikroTik 400G switching.
  • tonyd2wild/DeepSeek-v4-Flash-Vision-Exp-DSpark-1M-NVFP4-KV-2x-DGX-Spark - DeepSeek-V4-Flash-Vision-Exp ported into vLLM on two or four DGX Spark at 1M context, 53 tok/s real-prompt decode single stream and a 2.79M-token NVFP4 KV pool at TP2.
  • tonyd2wild/DeepSeek-V4.1-Flash-vLLM-DGX-Spark - DeepSeek-V4.1-Flash on four DGX Spark at TP4 with Engram rows read from NVMe, 3.5 bpw EXL3 serving lane at 95.3 tok/s idle code and 500K context.
  • tonyd2wild/DS4-H3-Video-Gen-Factory - DeepSeek-V4-Flash at 1M context beside two MiniMax H3 video renders on two DGX Spark, 35% of idle throughput kept, and a load order set by H3's 50 GB footprint.
  • tonyd2wild/GLM-5.2-655K-MTP-4x-DGX-Spark---25-32tok-s - GLM-5.2 744B unpruned (QuantTrio Int4-Int8Mix) at 655,360-token context across four DGX Spark nodes via decode-context-parallelism, 23.0 tok/s single and 47.9 aggregate at 4 concurrent, MTP k=3.
  • tonyd2wild/GLM-5.2-QuantTrio-200K-4x-DGX-Spark--36tok-s - Speed-shape recipe for unpruned GLM-5.2 at 200K context on four DGX Spark nodes, 36 tok/s peak, 75 aggregate at 6 concurrent, 63% of each step overhead.
  • tonyd2wild/GLM-5.3-Flash-NVFP4-DFlash2-2x-DGX-Spark - GLM-5.3-Flash NVFP4 on two DGX Spark with knapcio's TP4 stack ported to TP2 and DFlash2, single-stream code decode 68.9 tok/s against 44.7 for the previous recipe at 262K context.
  • tonyd2wild/GLM-5.3-Int4-Int8Mix-TP4-4x-DGX-Spark - GLM-5.3 743B Int4-Int8Mix at TP4 on four DGX Spark, DFlash2 at 53.32 tok/s structured output (1.98x MTP-4, level on prose), NVFP4 KV pool of 293,447 tokens at 270K.
  • tonyd2wild/MiMo-V2.5-TP2-1M-NVFP4-KV-2xDGX-Spark - MiMo-V2.5 Omni at tensor-parallel 2 on two DGX Spark, NVFP4 4-bit KV cache for a 2.17M-token pool at 1M context, thinking-OFF eval 97.8 against 90.6.
  • tonyd2wild/MiMo-V2.5-TP3-NVFP4-KV-3xDGX-Spark - MiMo V2.5 Omni (310B MoE, text/image/video/audio) at tensor-parallel 3 across three DGX Sparks, with 4-bit NVFP4 KV cache for a ~10.6M-token KV pool at 1M context.
  • tonyd2wild/MiMo-V2.6-Flash-DGX-Spark-Recipe - MiMo-V2.6-Flash-RL on two DGX Spark under vLLM TP2 with DFlash and fp8 KV, 53.3 tok/s single stream under a 300K limit, after four patches to the stock image.
  • tonyd2wild/MiniMax-M3-2x-DGX-Spark-36-tok-s - Reproduction of a3refaat's MiniMax-M3 428B unpruned stack on two DGX Spark, W4A16 GPTQ with NVFP4 KV and EAGLE-3 at 36.6 tok/s single-stream JSON and 31.8 code, 196K context.
  • tonyd2wild/Minimax-M3-NVFP-3x-DGX-Sparks-TP-3 - MiniMax-M3 NVFP4 428B-A23B at tensor-parallel 3 across three DGX Sparks, 10.5 tok/s at 200K context over a switchless 200 Gb/s RoCE mesh.
  • tonyd2wild/nfs-model-weights - NFS recipe for sharing one checkpoint across N DGX Spark nodes, taking a 4-node model library from 6.8 TB to 1.7 TB with no per-node copies.
  • tonyliu312/GLM-5.3-Flash-1M-Context-4x-DGX-Spark - GLM-5.3-Flash at 1,048,576-token context on four DGX Spark, 2.67x decode from dropping --enforce-eager and 1M concurrency from 1.35x to 4.20x by raising --kv-cache-memory, which silently overrides --gpu-memory-utilization.
  • tpurtell/glm-5.2-4x-spark-1x-rtx6k-96gb - Rust attention-FFN disaggregation engine for GLM-5.2 on one RTX PRO 6000 coordinator and four DGX Spark expert nodes, 26.6 NVFP4 and 28.5 EXL3 K3 tok/s weighted decode, 1,807 prefill.
  • tsarihan/qwen3.8-flash-next-nvfp4-2x-dgx-spark-playbook - Qwen3.8-Flash-Next NVFP4 against FP8 on the same two DGX Spark, 3.67x the KV pool and 66% more decode at 245K, needles 5/5 on both lanes.
  • urbanspr1nter/dgx-spark-bare-metal - Four-node DGX Spark Ray and vLLM cluster on a Mikrotik switch, with an 8-node guide on 400 Gb/s breakout cables, plus GLM-5.2 NVFP4, Kimi K2.7 Code and DeepSeek-V4-Flash launchers.
  • ursuciprian/qwen3.8-flash-next-dgx-spark-tp-2 - Qwen3.8-Flash-Next NVFP4 on two DGX Spark under vLLM V2 with b12x kernels, where one replica per node behind a router finishes synthetic agent replays 29-37% sooner than TP=2.
  • vroomfondel/dgxarley - Ansible playbooks for a K3s cluster of four DGX Spark nodes and an x86 control plane, SGLang over SR-IOV RoCE at 9.78 GB/s NCCL bus bandwidth.
  • Weschera/Ling-3.0-flash-VL-2x-DGX-Sparks - Ling-3.0-flash-VL FP8 on two DGX Spark under SGLang TP2 at 131,072 tokens, 26.0 tok/s single-stream on prose and 18.5 on code, with no MTP head.
  • www-ai-rs/gb10-deepseek-v4-flash - Two-node GB10 operator tooling and executed-code test harness for DeepSeek-V4-Flash 304B at 1M context, GPU rail power sampled at 2 Hz.
  • yunwei37/dgx-spark-4-ring-no-switch - Serving profiles and results log for four DGX Spark on a switchless ConnectX ring, DeepSeek-V4.1-Flash TP4 at 1M context 50.14 tok/s C1, gmu 0.88 recorded as unsafe.
  • ZD-AI-Lab/Triple-GB10 - Three-node GB10 QSFP ring with three /30 subnets and forwarding routes, pooling about 300 GB for Ray and vLLM pipeline-parallel across 3 Sparks.
  • zorost/sparkduet - Model-swap layer for two DGX Spark, four checkpoints on disk behind one OpenAI port, DeepSeek-V4-Flash c1 from 72.2 tok/s on math to 33.6 on prose by DSpark draft acceptance.

Image & Media Generation

  • AEON-7/comfyui-aeon-spark - ComfyUI Docker for DGX Spark with SageAttention v3 compiled for sm_121a, CUDA 13, NVFP4, and Flux 2 / LTX 2.3 pre-bundled.
  • alexhegit/h3-spark.c - CUDA port of antirez/h3.c for MiniMax-H3 on GB10, 15.6 s warm for a 512x512 22-frame clip at layers 45 and reuse 2.
  • bjarkebolding/spark-comfyui - Single-script containerized ComfyUI for DGX Spark, sm_121 SageAttention and sha256-verified workflow recipes, with a 2100 MHz clock cap for overcurrent reboots.
  • CoconutMacaroon/blender-arm64 - Blender build for GB10 aarch64 with CUDA, OptiX, and Vulkan, shipping a prebuilt DGX Spark binary release.
  • dr-vij/ComfyUI-DGX-Spark-Docker-opinionated - ComfyUI Docker for DGX Spark with self-built aarch64 wheels for decord NVDEC, flash-attn, onnxruntime-gpu and SageAttention 2.2.0, plus a SAM3 blacklist for WanVAE hangs.
  • dr-vij/Hunyuan3D-2.1-DGX-Spark-Docker - Hunyuan3D-2.1 3D generation on DGX Spark via Docker Compose, building custom_rasterizer and DifferentiableRenderer CUDA components on-box.
  • dr-vij/Trellis2-DGX-Spark-Docker - Dockerized TRELLIS.2 3D generation for DGX Spark, with nvdiffrast, CuMesh, FlexGEMM, and torchvision built from source for sm_121 on CUDA 12.9.
  • drowzeys/keys-SM121-Optimized-MiniMax-H3-Nvidia-Sol-Engine-Kijai-SolAttn_Triton-Single-DGX-Spark - MiniMax H3 in ComfyUI on one DGX Spark with flex_attention forced to Triton where sm_121 miscompiles it, plus Sol-Engine FirstBlockCache and batched VAE ports at 1.54x end to end.
  • Hitheshkaranth/LTX-2.5_Video_DGX_Spark_Setup - LTX-2.5 NVFP4 text and image-to-video with audio on DGX Spark via a three-line sm_121a kernel patch, 768x512 4 s clip in 61 s.
  • jayleaton/localrouter - MCP and OpenAI-style server that loads Qwen-Image 2.1 and MiniMax H3 on demand with Zig engines on DGX Spark, 1024x1024 NVFP4 in 12.7 s warm.
  • joeynyc/cosmos-locateanything-dgx - Two-stage DGX Spark pipeline: Cosmos 3 video generation, then NVIDIA LocateAnything object grounding.
  • joeynyc/MiniMax-H3-DGX-Spark - MiniMax H3 FL2VA video generation on one DGX Spark via vLLM-Omni and online FP8, sm_121 loader and AdaLN patch, 111 s warm request or 80 s with Cache-DiT.
  • kabilankb/cosmos3-nano-gb10 - Cosmos3-Nano 16B video and image generation on a single GB10 instead of NVIDIA's recommended 8x H100, 480p 57-frame clips in about 3 minutes.
  • LectinDK/DGX-Spark-ComfyUI-Forge - ComfyUI Docker for DGX Spark, xformers CUTLASS off on sm_121, and a --disable-pinned-memory fix that holds DynamicVRAM at a 73-78 GB peak against 88.
  • luix93/DGX-Spark-ComfyUI - ComfyUI Docker Compose for DGX Spark with SageAttention 2 built against sm_121, NVFP4 via comfy_kitchen, and a copy=False patch for the unified-memory double-VRAM bug.
  • madeye/comfyui-minimax-h3-dgx-spark - MiniMax-H3 video and audio in ComfyUI on GB10 sm_121 with 4-to-8-step turbo LoRAs, where fp8_scaled runs 21% faster than the recommended int8_convrot at 11.50 against 14.26 s/it.
  • mmartial/ComfyUI-Nvidia-Docker - Multi-platform ComfyUI Docker (x86_64, Blackwell, DGX Spark) with published aarch64 DGX images and userscripts that build SageAttention 2 and comfy_kitchen from source.
  • mvalancy/blender-nvidia-gb10 - Blender 5.0.1 source build for GB10 with Cycles CUDA 13 GPU rendering, via four inherited aarch64 patches and four CUDA-13 patches for OIDN, libglu, Wayland, and libdrm.
  • phaserblast/ComfyUI-DGXSparkSafetensorsLoader - Zero-copy model loader for ComfyUI on DGX Spark using the fastsafetensors library.

Audio & Speech

  • AEON-7/qwen3-asr-server - OpenAI /v1/audio/transcriptions server for Qwen3-ASR-0.6B on DGX Spark, vLLM-native with a pinned flashinfer and a soundfile decode path that avoids the PyAV fallback.
  • briancaffey/nemotron-asr-server - OpenAI-compatible speech-to-text server for nemotron-3.5-asr-streaming-0.6b on DGX Spark, native NeMo instead of the aarch64-broken Riva path, WebSocket streaming.
  • jxlarrea/homeassistant-voice-recipes - Local Wyoming voice pipeline for Home Assistant, aarch64 ONNX Parakeet ASR, ECAPA-TDNN speaker extraction, and Gemma-4-26B-A4B with MTP on llama.cpp built for sm_121.
  • kedarpotdar-nv/spark-realtime-chatbot - On-device assistant for voice and video calls on one GB10, ~320 ms and ~850 ms end to end, using Qwen3-VL, Kokoro, and DeepFace.
  • Logos-Flux/spark-voice-pipeline - WebSocket voice assistant for DGX Spark, sentence-level streaming across whisper.cpp, Ollama, and VibeVoice-Realtime-0.5B, 766 ms to first audio.
  • luka-loehr/qwen3-tts-native - Native Rust and CUDA runtime for Qwen3-TTS-1.7B VoiceDesign on sm_121, 94-96 ms p95 first audio and 5.68 GB peak versus 2.69 s and 108.90 GB for stock SGLang.
  • mARTin-B78/dgx-spark-faster-qwen3-tts - Faster-Qwen3-TTS on DGX Spark as an OpenAI-compatible API with CUDA-graph acceleration, four backends including chunk-streaming, and deterministic per-voice seeds.
  • Mekopa/whisperx-blackwell - Prebuilt WhisperX image with an sm_121 to sm_90 capability spoof and torchaudio jiterator patch, GPU pyannote diarization, 24 min audio in 62 s.
  • ncannings/fastconformer-trt - FP8 TensorRT engine for NeMo FastConformer ASR, parakeet-ultra at 4,903x real time and 1.814% test-clean WER on DGX Spark, against 999x and 1.803% for stock NeMo.
  • pipecat-ai/nemotron-voicechat-dgx-spark - Full-duplex speech-to-speech NemotronLabs VoiceChat 11B on one DGX Spark, GPTQ W8 Nano and W8A32 EarTTS at ~66 ms against the 80 ms frame budget.
  • Pizzaman213/fish-s2pro-gb10 - GB10-tuned Fish Audio S2-Pro TTS reaching 31.3 tok/s from 1.2 via bit-exact int8 kernels, speculative decode, and quality-gated NVFP4 weights.
  • WillIsback/whisperx-gb10 - WhisperX transcription and pyannote diarization REST API for GB10 (aarch64, sm_121), with an async job queue, SRT/VTT/TXT export, and prebuilt Docker Hub/GHCR images on NGC PyTorch 25.05.

Science & HPC

Beyond LLMs, GB10's unified memory and aarch64 stack run scientific compute: protein folding, biomolecular prediction, and RAN simulation.

  • adrian-greenneuron/openfold3-DGX-Spark - Dockerized OpenFold3 for DGX Spark with evoformer_attn kernels prebuilt, DeepSpeed patched from compute_121 to compute_120, ubiquitin prediction in 55 s.
  • chaoticcuriosity-io/g1-humanoid-rl - Unitree G1 humanoid RL on DGX Spark with mjlab, walking policy trained in about 46 min at 2048 envs and an 11 h cartwheel run at 4096 envs.
  • chaoticcuriosity-io/regolith - Synthetic-data lunar segmentation on DGX Spark with Isaac Sim 6.0 Replicator and SegFormer, where realistic ground cut real-photo rock over-prediction from 72.6% to 35.7%.
  • eetmie/spark-projects - Robotics playbooks for GB10 including openpi pi0.5 vision-language-action inference at 94.7 ms on TensorRT FP8 and NVFP4, 2.12x PyTorch BF16 at cosine 0.997.
  • rcbarke/ai-ran-dgx-spark - Bring-up notes for NVIDIA Aerial and Sionna on DGX Spark, with Sionna PHY and SYS working, cuMAC retargeted to run, and Sionna-RT and the 7.2x fronthaul blocked.
  • sanjyotshenoy/boltz-gb10-spark - Boltz-2 biomolecular-interaction prediction on DGX Spark with Triton-nightly sm_121 codegen.

Remote Access & Desktop

  • AtomicGaryBusey/DGXSparkGaming - Steam gaming compatibility log for DGX Spark under FEX-Emu, Box64, and Proton, with DLSS 4 frame generation working, DLSS 5 ruled out per frame, and every correction published.
  • seanGSISG/dgx-spark-sunshine-setup - One-command Sunshine installer for headless DGX Spark, NVIDIA CustomEDID virtual display with no dummy plug, 4K60 or 1440p120 within the GB10 pixel clock limit.

Tools & Monitoring

  • agjs/gb10-clock-cap - Clock-cap harness for GB10 inference hosts, where 2200 MHz ran 12 °C cooler at 36% less GPU-rail power for 1.0% decode and 3.9% prefill cost, single-stream at TP=2.
  • amer8/pulsebar - Unofficial macOS menu bar monitor that streams GPU and memory telemetry from the DGX Spark dashboard.
  • antheas/spark_hwmon - Linux hwmon kernel driver exposing GB10 system power telemetry (per-rail power, energy counters, temperatures) and PL1/PL2 power-cap controls via sysfs.
  • ateska/dgx-spark-prometheus - Single-binary Go Prometheus exporter with systemd unit and Grafana dashboard for DGX Spark, GB10 and NIC metrics on port 9835.
  • chappa-ai-llc/spark-smi - System-monitor TUI for DGX Spark with unified-memory and Grace P/E-core awareness, a cluster fleet view, an MT2910 fabric bandwidth test, and mixed sm_121 plus sm_86 support.
  • christopherowen/dgx-spark-fan-control - Kernel driver and CLI for RPM floors above NVIDIA's fan curve on DGX Spark, 13500 RPM floor measured at 9000 RPM in 6.3 s, with DKMS and Secure Boot signing.
  • DanTup/dgx_dashboard - Monitoring dashboard for DGX Spark bound to 0.0.0.0, with GB/GiB-correct memory stats, GPU power draw, and Docker container controls.
  • djmad/Spark_Energy_Management - Thermal and power controller for a ThinkStation PGX (GB10) with per-cluster CPU PID loops, a 100 MHz/s GPU clock ramp, and a thermal twin fitted to 1.2 K RMS.
  • dorangao/dgx-spark-toolkit - Two-node DGX Spark cluster scripts and manifests: RoCE and NCCL checks on the 200 Gb/s fabric, RDMA pods, MetalLB, pipeline-parallel vLLM Nemotron-3 Nano 30B.
  • drowzeys/vllm-gb10-spin-wait-fix - One-command patcher and English write-up for the vLLM spin-wait that heats GB10, 24 °C lower average SoC temperature on TP=2 multi-rank serves and no change at TP=1.
  • engineering87/sparkfit - Zero-dependency memory planner for DGX Spark: 128 GB budget split, quantization advisor, and decode roofline on 273 GB/s.
  • hectorTSH/dgx-spark-memory-dashboard - Live unified-memory occupancy map for DGX Spark as 1 GiB tiles, hot models against on-disk candidates ranked by free memory, and multi-Spark tabs.
  • hoesing/spark-gpu-throttle-check - Throttle test for DGX Spark that loads the GB10 with cuBLAS matmuls and flags clocks staying below a 1400 MHz threshold, a suspected USB-PD power-delivery fault.
  • ivanusto/gb10-ops - Host guards for GB10 that terminate the offending GPU process below 3 GiB MemAvailable and abort jobs after 240 s at 88 C, thresholds calibrated on one two-node box.
  • jasonacox/dgx-spark - Project hub for GB10 whose nanochat scripts pretrained a chat model from scratch in 9 days for about $8 of power, plus two-Spark InfiniBand training.
  • jeffrymahbuubi/dgx-spark-stress-test - Burn-in suite for GB10 unit qualification, 6-24 hour llama.cpp 70B plus SDXL load at ~96% utilization, ~85 GB resident, with 10s temperature and power CSVs.
  • jeremyeder/dgx-agentskills - Claude Code integration for DGX Spark: local model serving, GPU monitoring, and VM management.
  • joeynyc/spark-doctor - Read-only DGX Spark diagnostic CLI: 14 W power cap, unified-memory pressure, thermal risk, CUDA 13 / sm_121 wheel mismatches, Docker runtime, and recipe checks for tensor-parallel size and memory budget.
  • lcasarin-maker/blackbox - Forensic capture for GB10 hangs with Xid, PCIe AER, xHCI and PSI checks, plus a measured gap where 7 GiB held through PyTorch raised memory.current by 15 MiB.
  • lynx-lee/lynx-ollama - Ollama manager with a Go web console whose optimize command reads GB10 unified memory to set 131K context, 8-way parallel, and q8_0 KV cache.
  • mcampa/sparkrun-ui - Web UI for sparkrun with launch wizard, chat, benchmark charts, and live per-host GPU bars, run via npx or a published aarch64 container.
  • mchenetz/sparkd - Localhost dashboard for a DGX Spark fleet, with HF browsing, Claude-generated vLLM recipes, and single-box or Ray-cluster launch.
  • MiaAI-Lab/sparkDash - Web dashboard for a DGX Spark fleet with an engine probe for llama.cpp, vLLM, SGLang, ds4, and EXL3, cached against uncached prefill, daily peak tok/s, and Wake-on-LAN.
  • parallelArchitect/sparkview - Terminal GPU monitor with GB10-aware unified-memory reporting, memory-pressure (PSI) and power-rail readouts, and an anomaly auto-logger.
  • r0b0tlab/hermes-concurrent-agents - Supervised Hermes Agent worker pool for one Linux host, with pre-claim admission, exact process ownership, and optional GB10 memory-pressure presets.
  • securitysonar/spark-hashcat - Hashcat REST API service for GB10 aarch64, with an NVRTC CUDA build path that bypasses OpenCL.
  • skymaze/Fireworks - Web control plane for a DGX Spark fleet, ARP-probed RoCE rail configuration with rollback, and model direct-pull to workers over the four rails.
  • stevibe/SparklingKit - Local-first workspace for OCR, transcription, grounding, and image jobs, routed to six co-resident models on one DGX Spark from Qwen3.6-35B-A3B NVFP4 to LocateAnything-3B.
  • TheAwaken1/Spark-Studio - Launch and tuning dashboard for vLLM, SGLang, llama.cpp, and sparkrun recipes on DGX Spark that hands a broken recipe to Claude Code or Codex to patch and relaunch.
  • vybe/sparky - Vue 3 web UI for DGX Spark with Ollama chat, a Claude Code agent tab, ComfyUI SDXL and Flux generation, and Docker/systemd control.
  • wentbackward/nv-monitor - Terminal monitor and Prometheus exporter for DGX Spark in one zero-dependency C binary, with HugePages-correct unified memory and Grace big.LITTLE core labels.
  • Z841973620/dgx-spark-fan-override - Kernel module overriding DGX Spark fan speed through the EC's FF-A eSPI mailbox, with the stock firmware's two temperature curves and 1,260-13,500 RPM ranges documented.

Operating Systems & Containers

  • graham33/nixos-dgx-spark - Nix flake with a NixOS module for NVIDIA's DGX Spark kernel, bootable USB image, and 15 playbook devshells for TRT-LLM, NVFP4, and NCCL over QSFP.
  • kindlingai/kindling-spark-os - Read-only OS image for GB10 booted beside DGX OS with auto-revert and a 2 GiB display carveout lent to CUDA, GLM-5.3 TP=4 KV pool 365,440 to 404,608 tokens.
  • kyuz0/gb10-toolboxes - Prebuilt aarch64 CUDA 13.x containers for GB10 sm_121 with llama.cpp, antirez/ds4 and vLLM, rebuilt on four-hour upstream polls.
  • maxspevack/spark-rocky - Rocky Linux 10.2 Live-USB for DGX Spark on the CIQ 6.18 kernel, shipping 4k pages because 64k faults on every driver branch but the 580, measured 10.4% slower.
  • Neural-ICE/ICE-CoreOS - Immutable bootc OS for DGX Spark on CentOS Stream 10 with a 4 KiB-page GB10 kernel, optional TPM2-unlocked LUKS2, and a signed-UKI USB installer.
  • RageLtd/arch-dgx-spark-iso - Arch Linux installer ISO builder for DGX Spark, with the linux-dgx-spark kernel and archinstall config.

Community & Resource Collections

  • AEON-7/AEON-7 - Index of AEON-7's releases, mainly DGX Spark NVFP4 model packs, prebuilt vLLM images, and a voice-AI stack, plus Apple Silicon MLX builds.
  • odnodn/dgx-spark - Curated collection of NVIDIA DGX Spark resources and self-hosted AI projects.

Contributing

Contributions are welcome. Read the contribution guidelines before opening a pull request.

About

A curated list of tools, guides, playbooks, and resources for the NVIDIA DGX Spark (GB10 Grace Blackwell personal AI supercomputer).

Topics

Resources

Contributing

Stars

76 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages