MSVC+CUDA llama.cpp fork for Windows: KDA/GDN, long-context KV placement, TurboQuant/TCQ when measured, speculative draft. Lab defaults — TriAttention opt-in only, not recommended.
-
Updated
Jul 28, 2026 - C++
MSVC+CUDA llama.cpp fork for Windows: KDA/GDN, long-context KV placement, TurboQuant/TCQ when measured, speculative draft. Lab defaults — TriAttention opt-in only, not recommended.
Pre-built PyTorch wheels and build scripts for NVIDIA DGX Spark (GB10, sm_121, Blackwell, CUDA 13.0, ARM64)
Measured sm_110 / sm_110a facts for NVIDIA Jetson AGX Thor: tensor-core capability matrix (2:4 sparsity works on plain sm_110; tcgen05 + NVFP4 block-scale is sm_110a-only), CUDA 11/12->13 migration breaks, CPU ISA, and the bandwidth roofline. With probes.
Build the GPU inference stack from source, repeatably: for Blackwell sm_120 on CUDA 13.x and Python 3.14. Machine-readable build state, generated patches, and an abductive-triage skill for build failures.
Windows NVIDIA-only Triton 3.7.0 build pipeline for RTX 5090 / Blackwell sm_120a, with FP8 tl.dot validation and peak benchmark results.
Hunyuan3D-2 fork — image→textured 3D→sliced STL + part segmentation. RTX 50-series (Blackwell/sm_120), CUDA 13.0, Python 3.12, PyTorch 2.11+cu130.
Automated Docker build and serving stack for vLLM with native NVFP4 KV cache quantization on NVIDIA Blackwell (RTX 5090 / SM120)
Accelerate Kimi Delta Attention computations with high-performance CUTLASS kernels designed for NVIDIA SM90 architectures and beyond.
Add a description, image, and links to the cuda-13 topic page so that developers can more easily learn about it.
To associate your repository with the cuda-13 topic, visit your repo's landing page and select "manage topics."