Skip to content
@xlite-dev

xlite-dev

Develop ML/AI toolkits and ML/AI/CUDA Learning resources.

ffpa-attn: Fast and Memory-Efficient Exact Attention (BF16/FP16/FP8/FP4) for Large Headdim, 1.5x~15x🔥🔥 speedup over standard PyTorch SDPA.


BF16 Attention for Large Headdim: FFPA vs SDPA (FWD/BWD) across NVIDIA H200 and B200, 6x-15x↑.


FP4 Attention for D=128: FFPA vs SageAttention-3 (FWD) on NVIDIA RTX PRO 6000.

Pinned Loading

  1. LeetCUDA LeetCUDA Public

    Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.

    Cuda 11.9k 1.3k

  2. lite.ai.toolkit lite.ai.toolkit Public

    A lite C++ AI toolkit: 100+ models with MNN, ORT and TRT, including Det, Seg, Stable-Diffusion, Face-Fusion.

    C++ 4.4k 785

  3. Awesome-LLM-Inference Awesome-LLM-Inference Public

    📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

    Python 5.5k 434

  4. Awesome-DiT-Inference Awesome-DiT-Inference Public

    📚A curated list of Awesome Diffusion Inference Papers with Codes: Sampling, Cache, Quantization, Parallelism, etc.🎉

    Python 591 29

  5. torchlm torchlm Public

    💎An easy-to-use PyTorch library for face landmarks detection: training, evaluation, inference, and 100+ data augmentations.🎉

    Python 271 29

  6. ffpa-attn ffpa-attn Public

    Kernel Library for Large Headdim Attention (64~1024, BF16/FP8/FP4), 1.5x~15x↑ vs PyTorch SDPA.

    Python 332 27

Repositories

Showing 10 of 75 repositories
  • ffpa-attn Public

    Kernel Library for Large Headdim Attention (64~1024, BF16/FP8/FP4), 1.5x~15x↑ vs PyTorch SDPA.

    xlite-dev/ffpa-attn's past year of commit activity
    Python 332 Apache-2.0 27 12 0 Updated Sep 11, 2026
  • sglang Public Forked from sgl-project/sglang

    SGLang is a fast serving framework for large language models and vision language models.

    xlite-dev/sglang's past year of commit activity
    Python 3 Apache-2.0 8,823 0 0 Updated Sep 11, 2026
  • LeetCUDA Public

    Modern CUDA Learn Notes with PyTorch for Beginners, 200+ CUDA Kernels, Tensor Cores, HGEMM, FA-2 MMA.

    xlite-dev/LeetCUDA's past year of commit activity
    Cuda 11,932 GPL-3.0 1,256 3 1 Updated Sep 10, 2026
  • lite.ai.toolkit Public

    A lite C++ AI toolkit: 100+ models with MNN, ORT and TRT, including Det, Seg, Stable-Diffusion, Face-Fusion.

    xlite-dev/lite.ai.toolkit's past year of commit activity
    C++ 4,432 GPL-3.0 785 0 1 Updated Sep 5, 2026
  • flashinfer Public Forked from flashinfer-ai/flashinfer

    FlashInfer: Kernel Library for LLM Serving

    xlite-dev/flashinfer's past year of commit activity
    Python 0 Apache-2.0 1,427 0 0 Updated Sep 3, 2026
  • diffusers Public Forked from huggingface/diffusers

    🤗 Diffusers: State-of-the-art diffusion models for image, video, and audio generation in PyTorch and FLAX.

    xlite-dev/diffusers's past year of commit activity
    Python 0 Apache-2.0 7,421 0 0 Updated Sep 1, 2026
  • .github Public
    xlite-dev/.github's past year of commit activity
    1 0 0 0 Updated Aug 27, 2026
  • flash-linear-attention Public Forked from fla-org/flash-linear-attention

    🚀 Efficient implementations for emerging model architectures

    xlite-dev/flash-linear-attention's past year of commit activity
    Python 1 MIT 708 0 0 Updated Aug 27, 2026
  • GCMP Public Forked from VicBilibily/GCMP

    通过集成国内主流原生大模型提供商,为开发者提供更加丰富、更适合本土需求的 AI 编程助手选择。 目前已内置支持 智谱AI、MiniMax、MoonshotAI、DeepSeek、阿里云百炼、快手万擎、火山方舟、腾讯云、Xiaomi MiMo 等原生大模型提供商。 此外,扩展插件已适配支持 OpenAI 与 Anthropic 的 API 接口兼容模型,支持自定义接入任何提供兼容接口的第三方云服务模型。

    xlite-dev/GCMP's past year of commit activity
    TypeScript 0 MIT 58 0 0 Updated Aug 27, 2026
  • Awesome-LLM-Inference Public

    📚A curated list of Awesome LLM/VLM Inference Papers with Codes: Flash-Attention, Paged-Attention, WINT8/4, Parallelism, etc.🎉

    xlite-dev/Awesome-LLM-Inference's past year of commit activity
    Python 5,496 GPL-3.0 434 1 6 Updated Aug 14, 2026