Popular repositories Loading
-
cuda-optimized-skill
cuda-optimized-skill PublicA CUDA kernel optimization toolkit for validation, benchmarking, Nsight Compute profiling, bottleneck analysis, and iterative tuning. It helps improve custom GPU operators with reproducible workflo…
-
cuda-opt-agent
cuda-opt-agent PublicCUDA Opt Agent is an LLM-driven toolkit for automated CUDA kernel optimization. It generates baselines from text, specs, or existing .cu files, compiles and validates with nvcc, benchmarks on real …
-
cuda-datagen
cuda-datagen PublicMulti-agent LangGraph pipeline that turns CUDA kernel questions into verified SFT data: LLM-generated CUDA/CUTLASS/Triton/TileLang kernels pass nvcc + real GPU numeric checks, critic review and ran…
Python 1
-
cuda-sft-qgen
cuda-sft-qgen Publiccuda-sft-qgen: a LangGraph multi-agent pipeline generating CUDA SFT questions from a versioned knowledge graph. Every question passes 3-layer dedup, independent verification and a separate reviewer…
Python
-
lightly-trian-quant
lightly-trian-quant PublicRTX 5090 optimization of DINOv3 + EoMT segmentation training and inference with fused CUDA kernels, CUDA Graphs, and FP8/FP4 quantization. Achieves 6.3x ViT-L inference and 4.22x training speedups …
Python
-
FlashMoE
FlashMoE PublicForked from osayamenja/FlashMoE
Fork of osayamenja/FlashMoE, the fused distributed MoE kernel, with consumer Blackwell SM120/SM121 (RTX 50 / RTX PRO 6000 / DGX Spark) support and MXFP8, MXFP4 and NVFP4 expert paths: up to 5.6x fa…
Cuda
If the problem persists, check the GitHub status page or contact support.
