Skip to content

Latest commit

 

History

History
561 lines (458 loc) · 26.4 KB

File metadata and controls

561 lines (458 loc) · 26.4 KB

DeepGEMM

DeepGEMM is a unified, high-performance tensor core kernel library that brings together the key computation primitives of modern large language models — GEMMs (FP8, FP4, BF16), fused MoE with overlapped communication (Mega MoE), MQA scoring for the lightning indexer, HyperConnection (HC), and more — into a single, cohesive CUDA codebase. All kernels are compiled at runtime via a lightweight Just-In-Time (JIT) module, requiring no CUDA compilation during installation.

DeepGEMM leverages some concepts from CUTLASS and CuTe, but avoids heavy reliance on their templates or algebras. The library is designed for simplicity, with only a limited number of core kernel functions, making it a clean and accessible resource for learning NVIDIA GPU kernel optimization techniques.

Despite its lightweight design, DeepGEMM's performance matches or exceeds expert-tuned libraries across various matrix shapes.

News

  • 2026.04.16: Mega MoE, FP8xFP4 GEMM, FP4 Indexer, PDL, faster JIT compilation and more.
    • Please see #304 for more details.
    • For Mega MoE benchmarks, refer to #316.
  • 2025.09.28: DeepGEMM now supports scoring kernels (weighted ReLU MQA logits) for the lightning indexer for DeepSeek v3.2.
    • Please see #200 for more details.
  • 2025.07.20: DeepGEMM now supports both SM90/SM100, and has a full refactor with a low-CPU-overhead JIT CPP module.
    • NVRTC and post-compilation SASS optimization are all disabled.
    • NVRTC will be supported later.
    • As NVCC 12.9 will automatically do the FFMA interleaving, all post optimizations will be no longer supported.
    • Please see #112 for more details.
  • 2025.05.14: DeepGEMM now offers weight gradient kernels for dense and MoE backward! See #95 for details.
  • 2025.05.07: DeepGEMM now supports NVRTC with up to 10x compilation speedup! See #94 for details. Please use DG_JIT_USE_NVRTC=1 to enable it (may have performance loss with some cases).
  • 2025.04.18: DeepGEMM now achieves up to 1550 TFLOPS on H800! See #74, #78, #81, #86 and 340d988 for details.

Quick start

Requirements

  • NVIDIA SM90 or SM100 architecture GPU
  • Python 3.8 or higher
  • Compilers with C++20 support
  • CUDA Toolkit:
    • CUDA 12.3 or higher for SM90
      • We highly recommend 12.9 or higher for the best performance
    • CUDA 12.9 or higher for SM100
  • PyTorch 2.1 or higher
  • CUTLASS 4.0 or higher (could be cloned by Git submodule)
  • {fmt} library (could be cloned by Git submodule)

Development

# Submodule must be cloned
git clone --recursive git@github.com:deepseek-ai/DeepGEMM.git
cd DeepGEMM

# Link some essential includes and build the CPP JIT module
cat develop.sh
./develop.sh

Installation

cat install.sh
./install.sh

Then, import deep_gemm in your Python project, and enjoy!

DCU/HIP W8A8 Mega MoE Quick Start

The DCU path builds a standalone megamoe HIP extension for Hygon gfx938. It is separate from the CUDA deep_gemm JIT flow above and is specialized for the DSV4-Flash W8A8 FP8 channelwise MegaMoE shape:

  • EP size: 8 ranks
  • Experts: 256 total, 32 per rank
  • Top-K: 6
  • Hidden size: 4096
  • Intermediate hidden size: 2048
  • Maximum tokens per rank: set by num_max_tokens_per_rank

Build

Build on a DTK 26.04 environment:

source /opt/dtk-26.04/env.sh
./build_dcu_megamoe.sh

The build script keeps intermediate files under build/ and writes the wheel to build/whl/. It also builds the extension in place, so the local checkout can run the tests directly.

If you do not install the wheel but want to import megamoe from another directory, either set PYTHONPATH to this repository root or install the source tree in editable mode after building:

PYTHONPATH=/workspace/DeepGEMM python your_script.py
# or
pip install -e .

Editable installs point Python back to this checkout, so they use the in-place megamoe/_C*.so, staged k1/k2/k3_fused_ext*.so, and staged .co files created by build_dcu_megamoe.sh. Python-only edits are picked up directly; after changing HIP, asm, or setup.py, rerun build_dcu_megamoe.sh. If you run from an installed wheel instead, reinstall the newly generated wheel after rebuilding.

The optional large-token staged path is built ahead of time as part of the megamoe wheel. Wheel installation places the staged extension modules and asm code objects under the Python package directory, alongside the original MegaMoE fused extension:

  • megamoe/_C*.so
  • megamoe/dcu_megamoe_large_opt/K1_fused/k1_fused_ext*.so
  • megamoe/dcu_megamoe_large_opt/K2_fused/k2_fused_ext*.so
  • megamoe/dcu_megamoe_large_opt/K3_fused/k3_fused_ext*.so
  • megamoe/dcu_megamoe_large_opt/K1_fused/*.co
  • megamoe/dcu_megamoe_large_opt/K3_fused/*.co

If any staged HIP or asm source changes, rebuild and reinstall the wheel. The hygon_tmp directory is only used by test scripts for temporary reports or scratch files; it is not required for installed kernel binaries.

Runtime Routing

The default DCU execution path uses token-threshold auto selection. Without setting MEGAMOE_DCU_USE_LARGE_OPT_3STAGE, the public megamoe.fp8_w8a8_mega_moe API uses the original persistent fused kernel for small token counts and the large-token staged path for larger token counts:

  • K1: dispatch pull + L1 FP8 grouped GEMM
  • K2: SwiGLU + channelwise FP8 quant
  • K3: L2 FP8 grouped GEMM + combine reduce

The default threshold is 128 tokens per rank: num_tokens_per_rank <= 128 uses the original persistent fused kernel, and num_tokens_per_rank > 128 uses the staged K1/K2/K3 path. Override the threshold with MEGAMOE_DCU_LARGE_OPT_3STAGE_TOKEN_THRESHOLD; the comparison is strictly greater than the threshold. Set MEGAMOE_DCU_USE_LARGE_OPT_3STAGE=1 to force the staged path for all token counts, or MEGAMOE_DCU_USE_LARGE_OPT_3STAGE=0 to force the original persistent fused kernel. Set these environment variables before creating SymmBuffer: the mode and threshold are captured on buffer initialization so graph-captured runs see a stable branch choice.

The staged path keeps all K1/K2/K3 implementation files under megamoe.dcu_megamoe_large_opt. Its large temporary activations reuse the original DCU MegaMoE route_scratch allocation; the integration does not allocate a second persistent L1/K2/K3 activation workspace. In auto mode, SymmBuffer prepares the staged path during initialization by creating tensor views into the same route_scratch storage; no extra device kernels, D2H synchronization, or duplicate activation buffers are introduced for the later first large-token call.

By default K3 uses the integrated ASM tail-reduce path for both eager and graph staged execution, so it avoids the separate rank_barrier + reduce tail that can show large latency swings. Set K3_USE_ASM_TAIL_REDUCE=0 only when debugging the older barrier/reduce path. For num_max_tokens_per_rank <= 2048, the tail reducer defaults to 64 reducer workgroups; larger max-token buffers keep the previous 128-workgroup default.

The staged path keeps the tail-reduce signal state in route_scratch; when the large-opt environment is enabled before creating the symmetric buffer, this state is prepared during buffer initialization rather than the timed execution path. The original persistent fused path and the staged path both read the same input slices in the symmetric buffer (x, x_sf, topk_idx, and topk_weights), so switching by token threshold does not require duplicate input copies.

For EP runs with uneven per-rank local token counts, the eager auto threshold decision must be identical on every rank. Frameworks should pass a uniform dispatch_num_tokens to megamoe.fp8_w8a8_mega_moe, typically the EP-group maximum local token count for the current request. If it is omitted, eager auto mode falls back to the local y.size(0) on each rank, which is safe only when every rank has the same token count or when the path is forced by MEGAMOE_DCU_USE_LARGE_OPT_3STAGE. The dispatch_num_tokens value is only a host-side branch selector for big-fused versus staged execution; it does not pad inputs, does not change the valid row count, and does not make a rank compute more than its local y.size(0) rows. Keep it in [0, sym_buffer.num_max_tokens_per_rank]. Without a uniform dispatch value, one rank can enter the persistent fused kernel while another enters staged K1/K2/K3, and their cross-rank barriers will not match.

CUDA Graph Mode

DCU MegaMoE exposes graph-bucket mode through the public megamoe.fp8_w8a8_mega_moe API. The graph bucket size is the symmetric buffer's requested num_max_tokens_per_rank; no separate CUDA Graph max-token environment variable is used. The internal buffer capacity may be aligned up for kernel requirements, but graph replay uses sym_buffer.cuda_graph_max_tokens_per_rank. Pass exactly one graph flag to choose the implementation captured into the graph:

  • big_fused_cuda_graph=True captures the original persistent fused kernel.
  • stages_fused_cuda_graph=True captures the staged K1/K2/K3 large-token path when MEGAMOE_DCU_USE_LARGE_OPT_3STAGE is auto or forced on.

When the current stream is being captured, the eager auto-dispatch path is rejected. Framework integrations should pass one of the graph flags during capture; otherwise a fixed eager launch could be captured with the wrong token count or implementation choice for later replays.

y_graph = torch.empty((sym_buffer.cuda_graph_max_tokens_per_rank, hidden),
                      dtype=torch.bfloat16, device="cuda")
megamoe.fp8_w8a8_mega_moe(
    y_graph,
    l1_weights,
    l2_weights,
    sym_buffer,
    big_fused_cuda_graph=True,
)

For a smaller request, write the actual token count into sym_buffer.cuda_graph_num_tokens, update only the valid input prefix in sym_buffer before replay, replay the graph, and consume only y_graph[:token_count]. The kernel reads the device-side token count during replay, so route building, expert task generation, and local reduce use the valid prefix rather than forcing invalid tail routes through the graph bucket. Each rank owns its local sym_buffer.cuda_graph_num_tokens scalar, so graph replay supports uneven per-rank local token counts. A rank may set this value to 0; kernels publish the count through peer-visible symmetric memory and skip that rank's local output prefix while still serving remote expert work. This matches the usual static-buffer CUDA Graph usage: the graph shape is fixed, while the framework owns the valid-token prefix and chooses which captured graph to replay. A typical framework setup captures one big-fused graph and one staged graph with the same num_max_tokens_per_rank, then replays by token threshold. The replay choice must also be uniform across the EP group, using the same global dispatch token rule as eager mode.

In tests/test_mega_moe_dcu.py, the CUDA Graph test options separate capacity, ordinary correctness input, and replay buckets:

  • --num-max-tokens-per-rank is the symmetric-buffer capacity and graph capture bucket size.
  • --num-tokens is the normal fused-vs-baseline correctness input size. When set to 0, the test enables uneven per-rank local tokens with --num-max-removed-tokens.
  • --num-tokens-per-rank-list overrides --num-tokens with an exact local token count per rank, useful for reproducing framework cases such as 0,133,0,0,0,0,0,0.
  • --dispatch-num-tokens passes the API-level dispatch_num_tokens argument. It is used only by eager auto-dispatch to choose the implementation uniformly across ranks. For uneven-rank tests, set it to the EP-group max local token count, for example 133 for 0,133,0,0,0,0,0,0. The actual per-rank work still comes from each rank's local token count.
  • --cuda-graph-test-tokens is only the list of runtime token counts replayed against the captured graph, for example 32,64,128.
  • --cuda-graph-skip-baseline smoke-tests graph capture/replay without running the DeepEP baseline checker, which is useful when isolating graph compatibility from baseline communication behavior.

Example eager uneven-rank check with a uniform auto-dispatch decision:

source /opt/dtk-26.04/env.sh
python tests/test_mega_moe_dcu.py \
  --num-processes 8 \
  --num-max-tokens-per-rank 256 \
  --num-tokens-per-rank-list 0,133,0,0,0,0,0,0 \
  --dispatch-num-tokens 133 \
  --hidden 4096 \
  --intermediate-hidden 2048 \
  --num-experts 256 \
  --num-topk 6 \
  --correctness-iters 1 \
  --skip-bench

The staged graph bucket supports both K3 combine modes. By default, the captured K3 ASM uses tail-reduce and consumes K1's device-side active-tile count plus the graph runtime token scalar, so replay skips inactive K3 row tiles and reduces only the valid token prefix. Graph mode rejects cumulative_local_expert_recv_stats, because graph replay should not accumulate per-expert statistics across variable-token requests.

Host-side tuning knobs for the staged path do not add device kernels:

  • MEGAMOE_DCU_LARGE_OPT_3STAGE_TOKEN_THRESHOLD controls the auto token cutoff. The default is 128 based on the DSV4-Flash sweep where 32/64/128 favor the persistent fused kernel and 256+ favors the staged path.
  • K2_SKIP_INACTIVE_ROWS_MIN_TOKENS controls when K2 consumes K1's row_combine_ptrs validity metadata. The default is 1536, so larger token counts skip inactive-row activation work while smaller token counts keep the leaner K2 launch path.

Validate

Run the DSV4-Flash correctness and performance check:

source /opt/dtk-26.04/env.sh
python tests/test_mega_moe_dcu.py \
  --num-processes 8 \
  --num-max-tokens-per-rank 2048 \
  --num-tokens 512 \
  --hidden 4096 \
  --intermediate-hidden 2048 \
  --num-experts 256 \
  --num-topk 6 \
  --correctness-iters 1 \
  --warmup 3 \
  --repeat 8 \
  --out hygon_tmp/megamoe_dcu_dsv4_flash_512.json

To exercise CUDA-compatible uneven per-rank local token counts, set --num-tokens 0 --num-max-removed-tokens N. The DCU test follows the CUDA MegaMoE convention local_tokens=max(0, max_tokens-random_remove), so this also covers ranks with zero local tokens:

source /opt/dtk-26.04/env.sh
MEGAMOE_DCU_USE_LARGE_OPT_3STAGE=1 python tests/test_mega_moe_dcu.py \
  --num-processes 8 \
  --num-max-tokens-per-rank 512 \
  --num-tokens 0 \
  --num-max-removed-tokens 768 \
  --hidden 4096 \
  --intermediate-hidden 2048 \
  --num-experts 256 \
  --num-topk 6 \
  --correctness-iters 1 \
  --skip-bench \
  --large-opt-3stage \
  --stages-fused-cuda-graph \
  --cuda-graph-test-tokens 7,32,128,512

Run the requested token-per-rank sweep. The default list includes compact-window representatives around 1025..1441 as well as the main 512/1024/2048 sizes:

source /opt/dtk-26.04/env.sh
bash scripts/run_dcu_megamoe_large_opt.sh

For a correctness-only smoke run, set SKIP_BENCH=1. The staged test keeps the weight FP8 conversion chunked by default; tune MEGAMOE_DCU_WEIGHT_CAST_CHUNK_ROWS if the random-weight setup needs a smaller or larger temporary allocation.

Check one captured persistent-fused graph bucket across several token prefixes:

source /opt/dtk-26.04/env.sh
python tests/test_mega_moe_dcu.py \
  --num-processes 8 \
  --num-max-tokens-per-rank 2048 \
  --num-tokens 2048 \
  --hidden 4096 \
  --intermediate-hidden 2048 \
  --num-experts 256 \
  --num-topk 6 \
  --big-fused-cuda-graph \
  --cuda-graph-test-tokens 32,64,128 \
  --skip-bench

Check one captured staged K1/K2/K3 graph bucket across token prefixes with a 2048-token symmetric buffer:

source /opt/dtk-26.04/env.sh
MEGAMOE_DCU_USE_LARGE_OPT_3STAGE=auto \
python tests/test_mega_moe_dcu.py \
  --num-processes 8 \
  --num-max-tokens-per-rank 2048 \
  --num-tokens 2048 \
  --hidden 4096 \
  --intermediate-hidden 2048 \
  --num-experts 256 \
  --num-topk 6 \
  --stages-fused-cuda-graph \
  --cuda-graph-test-tokens 32,512,1024,2048 \
  --skip-bench

Force one staged-path size directly:

source /opt/dtk-26.04/env.sh
MEGAMOE_DCU_USE_LARGE_OPT_3STAGE=1 python tests/test_mega_moe_dcu.py \
  --num-processes 8 \
  --num-max-tokens-per-rank 2048 \
  --num-tokens 1024 \
  --hidden 4096 \
  --intermediate-hidden 2048 \
  --num-experts 256 \
  --num-topk 6 \
  --correctness-iters 1 \
  --warmup 3 \
  --repeat 8 \
  --out hygon_tmp/large_opt/integrated/dsv4_flash_large_opt_1024.json

Force the original persistent fused path while keeping the symmetric buffer capacity fixed at 2048 tokens per rank:

source /opt/dtk-26.04/env.sh
mkdir -p hygon_tmp/megamoe_dcu_dsv4_flash
for tokens in 512 1024 2048; do
  MEGAMOE_DCU_USE_LARGE_OPT_3STAGE=0 python tests/test_mega_moe_dcu.py \
    --num-processes 8 \
    --num-max-tokens-per-rank 2048 \
    --num-tokens "${tokens}" \
    --hidden 4096 \
    --intermediate-hidden 2048 \
    --num-experts 256 \
    --num-topk 6 \
    --correctness-iters 1 \
    --warmup 3 \
    --repeat 8 \
    --out "hygon_tmp/megamoe_dcu_dsv4_flash/bench_${tokens}.json"
done

Interfaces

Notices

This library provides optimized GEMM kernels for NVIDIA GPUs with a naming convention: D = C + A @ B. The input shape layout is NT (non-transposed A, transposed B). While the SM90 implementation supports only the NT memory layout (row-major, col-major), the SM100 implementation supports all memory layouts (NT, TN, NN, TT). For example, fp8_gemm_nt will do a D = C + A @ B.T

For both architectures, the LHS scaling factor is required to have a TMA-aligned and transposed layout. And the data format for the scaling factor of SM90 and SM100 is different:

  • SM90 requires scaling factors in FP32 format.
  • SM100 requires scaling factors in packed UE8M0 format, which packs 4 UE8M0 into a single torch.int.

Please note that operations like input transposition or FP8 casting must be handled separately by the user, please implement or fuse them into prior kernels independently. While the library provides some simple PyTorch utility functions, these may result in slower performance, but our primary focus is on optimizing the GEMM kernels themselves.

Normal dense GEMMs (non-grouped)

To perform a basic non-grouped FP8 GEMM, call the fp8_gemm_{nt, nn, tn, tt} function. For more details, please refer to the function documentation.

Grouped GEMMs (contiguous layout)

Unlike traditional grouped GEMMs in CUTLASS, DeepGEMM groups only the M-axis, while N and K must remain fixed. This design is tailored for scenarios where experts in an MoE model share the same shape. For training forward passes or inference prefilling, where each expert may process a varying number of tokens, we concatenate these tokens into a single tensor, referred to as the "contiguous" layout. Note that each expert segment must be aligned to the GEMM M block size (get_mk_alignment_for_contiguous_layout()). For more information, please refer to the m_grouped_fp8_gemm_{nt, nn}_contiguous function documentation.

We also provide a K-axis-grouped API for MoE weight backward (with M and N must remain fixed), please refer to k_grouped_fp8_gemm_tn_contiguous for more information.

Grouped GEMMs (masked layout)

During the inference decoding phase, when CUDA graph is enabled and the CPU is unaware of the number of tokens each expert receives, we support masked grouped GEMMs. By providing a mask tensor, the kernel computes only the valid portions.

Use m_grouped_fp8_gemm_nt_masked for this purpose and consult the relevant documentation. An example usage is to use the output of low-latency kernels from DeepEP as input.

V3.2 MQA kernels for the indexer

The kernel family has two versions, non-paged (for prefilling) and paged (for decoding). Take the non-paged version fp8_mqa_logits as an example. It has 6 inputs:

  • q, E4M3 tensor with shape [seq_len, num_heads, head_dim]
  • kv, E4M3 tensor (shaped as [seq_len_kv, head_dim]) with float SF (shaped as [seq_len_kv])
  • weights, float tensor with shape [seq_len, num_heads]
  • cu_seq_len_k_start and cu_seq_len_k_end, int tensor with shape [seq_len]
  • clean_logits, whether to clean the unfilled logits into -inf

The output tensor is shaped as [seq_len, seq_len_kv], indicating token-to-token logits. For each token i in q, it will iterate all tokens j from [cu_seq_len_k_start[i], cu_seq_len_k_end[i]), and calculate the logit out[i, j] as:

kv_j = kv[0][j, :] * kv[1][j].unsqueeze(1)  # [head_dim]
out_ij = q[i, :, :] @ kv_j  # [num_heads]
out_ij = out_ij.relu() * weights[i, :]  # [num_heads]
out_ij = out_ij.sum()  # Scalar

For more details and the paged version fp8_paged_mqa_logits, please refer to tests/test_attention.py.

CUDA Mega MoE

Mega MoE fuses and overlaps EP dispatch, linear 1 (FP8xFP4), SwiGLU, linear 2 (FP8xFP4), and EP combine into a single mega-kernel, overlapping NVLink communication and tensor core computation. It requires multi-process launch with symmetric memory. Usage:

# Allocate symmetric memory buffer
# NOTES: requires PyTorch >= 2.9
buffer = deep_gemm.get_symm_buffer_for_mega_moe(
    group, num_experts, num_max_tokens_per_rank, num_topk, hidden, intermediate_hidden
)

# Transform weights (FP4 with UE8M0 SF) into the required layout
transformed_l1, transformed_l2 = deep_gemm.transform_weights_for_mega_moe(l1_weights, l2_weights)

# Copy inputs into the buffer before each call
# You may fuse these into previous kernels
buffer.x[:num_tokens].copy_(x_fp8)
buffer.x_sf[:num_tokens].copy_(x_sf)
buffer.topk_idx[:num_tokens].copy_(topk_idx)
buffer.topk_weights[:num_tokens].copy_(topk_weights)

# Run the fused mega MoE kernel
y = torch.empty((num_tokens, hidden), dtype=torch.bfloat16, device='cuda')
deep_gemm.fp8_fp4_mega_moe(y, transformed_l1, transformed_l2, buffer)

For the full example with multi-process setup and benchmarking, please refer to tests/test_mega_moe.py.

Utilities

The library provides some utility functions besides the above kernels:

  • deep_gemm.set_num_sms / get_num_sms: set/get the maximum SM count to use
  • deep_gemm.set_tc_util / get_tc_util: set/get an approximated tensor core utilization ratio
  • deep_gemm.set_pdl / get_pdl: enable/disable Programmatic Dependent Launch (PDL)
  • deep_gemm.set_mk_alignment_for_contiguous_layout / get_mk_alignment_for_contiguous_layout: set/get the group-level M/K alignment for contiguous layout
  • deep_gemm.get_theoretical_mk_alignment_for_contiguous_layout: get the theoretical minimum M/K alignment
  • deep_gemm.set_ignore_compile_dims: configure dimensions to ignore during JIT compilation
  • deep_gemm.set_block_size_multiple_of: constrain block sizes to be multiples of a given value
  • deep_gemm.transform_sf_into_required_layout: transform scaling factors into the required layout
  • deep_gemm.get_tma_aligned_size: get the required TMA alignment size
  • deep_gemm.get_mn_major_tma_aligned_tensor: get a MN-major TMA-aligned tensor
  • deep_gemm.get_mn_major_tma_aligned_packed_ue8m0_tensor: get a MN-major TMA-aligned tensor (with packing FP32 into UE8M0)
  • deep_gemm.get_k_grouped_mn_major_tma_aligned_packed_ue8m0_tensor: K-grouped GEMM packing kernel

The library also provides some environment variables, which may be useful:

  • General
    • DG_JIT_DEBUG: 0 or 1, print JIT debugging information, 0 by default
    • DG_PRINT_CONFIGS: 0 or 1, print selected configs for each shape, 0 by default
  • JIT cache
    • DG_JIT_CACHE_DIR: string, cache directory for compiled kernels, $HOME/.deep_gemm by default
  • Compiler selection
    • DG_JIT_USE_NVRTC: 0 or 1, use NVRTC instead of NVCC (faster compilation, may have lower performance for some cases), 0 by default
    • DG_JIT_NVCC_COMPILER: string, NVCC compiler path; defaults to torch.utils.cpp_extension.CUDA_HOME
    • DG_JIT_CPP_STANDARD: integer, C++ standard version, 20 by default
  • Compiler output
    • DG_JIT_PRINT_COMPILER_COMMAND: 0 or 1, print compilation commands, 0 by default
    • DG_JIT_PTXAS_VERBOSE: 0 or 1, show detailed PTXAS output, 0 by default
    • DG_JIT_PTXAS_CHECK: 0 or 1, assert no local memory usage in compiled kernels, 0 by default
    • DG_JIT_PRINT_LOAD_TIME: 0 or 1, print kernel load time, 0 by default
  • Debug and profiling
    • DG_JIT_WITH_LINEINFO: 0 or 1, embed source line info for profiling tools, 0 by default
    • DG_JIT_DUMP_ASM: 0 or 1, dump both PTX and SASS, 0 by default
    • DG_JIT_DUMP_PTX: 0 or 1, dump PTX output, 0 by default
    • DG_JIT_DUMP_SASS: 0 or 1, dump SASS output, 0 by default
    • DG_COMM_KERNEL_DEBUG: 0 or 1, zero symmetric buffer before each Mega MoE call for debugging, 0 by default
    • DG_USE_NVIDIA_TOOLS: 0 or 1, skip internal profiling when running under external NVIDIA tools, 0 by default
  • Build options
    • DG_SKIP_CUDA_BUILD: 0 or 1, skip CUDA extension build during installation, 0 by default
    • DG_FORCE_BUILD: 0 or 1, force local build instead of downloading pre-built wheels, 0 by default
    • DG_JIT_USE_RUNTIME_API: 0 or 1, use CUDA Runtime API for kernel loading (requires CUDA runtime >= 12.8), 0 by default

For additional examples and details, please refer to the test code or review the corresponding Python documentation.

Acknowledgement

DeepGEMM is inspired by the CUTLASS project. Thanks and respect to the developers!

License

This code repository is released under the MIT License.

Citation

@misc{deepgemm2025,
      title={DeepGEMM: clean and efficient BLAS kernel library on GPU}, 
      author={Chenggang Zhao and Zhean Xu and Liang Zhao and Jiashi Li and Chenhao Xu and Anyi Xu and Shengyu Liu and Kexing Zhou and Kuai Yu},
      year={2025},
      publisher = {GitHub},
      howpublished = {\url{https://github.com/deepseek-ai/DeepGEMM}},
}