Skip to content
@xlite-dev

xlite-dev

Develop ML/AI toolkits and ML/AI/CUDA Learning resources.

Latest News

  • [2026-08] 🐍 Cache-DiT x FFPA (FP8/FP4) is ready! Feel free to take a try for your Diffusion models. 🎉🎉
  • [2026-08] 🚪 FFPA now experimental supports FP4 Attention for headdims [64,1024] (sm_120, forward only), achieving 850-980🎉 TFLOPS (D=128-256) on NVIDIA RTX 5090, 3.8x~4.4x🎉 speedup over PyTorch SDPA (FlashAttention-2 backend), the performance of large headdims is stay tuned for updates. 🎉🎉
  • [2026-08] 🦅 FFPA now supports D=512 for NVIDIA B200 via CuTe-DSL tcgen05 2-CTA, 1517 TFLOPS forward and 763 TFLOPS backward, achieving 6x~15x🎉 speedup over standard PyTorch SDPA. 🎉🎉
  • [2026-07] 🎯 FFPA now supports FP8 Attention for headdims [64,1024] (sm_120, forward only) and achieving 3x~6x🎉 speedup over PyTorch SDPA for large headdim (D>256). 🎉🎉
  • [2026-06] FFPA now supports AMD ROCm/HIP GPUs via the TritonBackend, check #268 for more details. 🎉
  • [2026-06] 🦅 NVIDIA-Nemo/AutoModel x FFPA achieving 1.4x~1.5x🎉 End2End training throughput speedup for Gemma4-31B (8xH200, FSDP2 + AC) with FFPA accelerating the 10/60 (D=512) full-attention layers. 🎉🎉
  • [2026-06] 🐍 FFPA now supports TritonBackend and CuTeDSLBacked for both forward and backward pass, achieving 1.5x~5x🎉 speedup over standard PyTorch SDPA across many devices. 🎉🎉
  • [2026-05] 🚪 FFPA now supports GQA, MQA, cross-attn, causal, attn-mask and dropout with CUDABackend for large headdims (D>256, forward only), achieving 1.3x~2x🎉 speedup over PyTorch SDPA. 🎉🎉

Pinned Loading

  1. LeetCUDA LeetCUDA Public

    Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.

    Cuda 11.8k 1.2k

  2. lite.ai.toolkit lite.ai.toolkit Public

    A lite C++ AI toolkit: 100+ models with MNN, ORT and TRT, including Det, Seg, Stable-Diffusion, Face-Fusion.

    C++ 4.4k 783

  3. Awesome-LLM-Inference Awesome-LLM-Inference Public

    📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

    Python 5.5k 429

  4. Awesome-DiT-Inference Awesome-DiT-Inference Public

    📚A curated list of Awesome Diffusion Inference Papers with Codes: Sampling, Cache, Quantization, Parallelism, etc.🎉

    Python 584 28

  5. torchlm torchlm Public

    💎An easy-to-use PyTorch library for face landmarks detection: training, evaluation, inference, and 100+ data augmentations.🎉

    Python 270 29

  6. ffpa-attn ffpa-attn Public

    Fast and Memory-Efficient Exact Attention (BF16/FP16/FP8/FP4) for Large Headdim, 1.5x~15x speedup over PyTorch SDPA.

    Python 322 25

Repositories

Showing 10 of 73 repositories

Top languages

Loading…

Most used topics

Loading…