- [2026-08] 🐍 Cache-DiT x FFPA (FP8/FP4) is ready! Feel free to take a try for your Diffusion models. 🎉🎉
- [2026-08] 🚪 FFPA now experimental supports FP4 Attention for headdims [64,1024] (sm_120, forward only), achieving 850-980🎉 TFLOPS (D=128-256) on NVIDIA RTX 5090, 3.8x~4.4x🎉 speedup over PyTorch SDPA (FlashAttention-2 backend), the performance of large headdims is stay tuned for updates. 🎉🎉
- [2026-08] 🦅 FFPA now supports D=512 for NVIDIA B200 via CuTe-DSL tcgen05 2-CTA, 1517 TFLOPS forward and 763 TFLOPS backward, achieving 6x~15x🎉 speedup over standard PyTorch SDPA. 🎉🎉
- [2026-07] 🎯 FFPA now supports FP8 Attention for headdims [64,1024] (sm_120, forward only) and achieving 3x~6x🎉 speedup over PyTorch SDPA for large headdim (D>256). 🎉🎉
- [2026-06] FFPA now supports AMD ROCm/HIP GPUs via the TritonBackend, check #268 for more details. 🎉
- [2026-06] 🦅 NVIDIA-Nemo/AutoModel x FFPA achieving 1.4x~1.5x🎉 End2End training throughput speedup for Gemma4-31B (8xH200, FSDP2 + AC) with FFPA accelerating the 10/60 (D=512) full-attention layers. 🎉🎉
- [2026-06] 🐍 FFPA now supports TritonBackend and CuTeDSLBacked for both forward and backward pass, achieving 1.5x~5x🎉 speedup over standard PyTorch SDPA across many devices. 🎉🎉
- [2026-05] 🚪 FFPA now supports GQA, MQA, cross-attn, causal, attn-mask and dropout with CUDABackend for large headdims (D>256, forward only), achieving 1.3x~2x🎉 speedup over PyTorch SDPA. 🎉🎉
xlite-dev
Pinned Loading
Repositories
- ffpa-attn Public
Fast and Memory-Efficient Exact Attention (BF16/FP16/FP8/FP4) for Large Headdim, 1.5x~15x speedup over PyTorch SDPA.
- .github Public
- LeetCUDA Public
Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.
- sglang Public Forked from sgl-project/sglang
SGLang is a fast serving framework for large language models and vision language models.
- Awesome-LLM-Inference Public
📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉
- diffusers Public Forked from huggingface/diffusers
🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch and FLAX.
- cutlass Public Forked from NVIDIA/cutlass
CUDA Templates and Python DSLs for High-Performance Linear Algebra
- CuTe Public Forked from NVlabs/CuTe
Reference implementation and examples of the CuTe Layout representation and algebra.
Top languages
Loading…
Most used topics
Loading…