Skip to content

Pinned Loading

  1. PAT PAT Public

    Prefix-Aware Attention for LLM Decoding

    Python 48 3

  2. RAGPulse RAGPulse Public

    An Open-Source RAG Workload Trace to Optimize RAG Serving Systems

    Python 40 4

  3. flash-linear-attention-npu flash-linear-attention-npu Public

    C++ 50 67

Repositories

Showing 5 of 5 repositories
  • OLED-MoE Public

    Accelerating MoE-Based dLLM Inference via Inter-Iteration Locality-Aware Expert Offloading

    flashserve/OLED-MoE's past year of commit activity
    Python 2 0 0 0 Updated Sep 30, 2026
  • flashserve/flash-linear-attention-npu's past year of commit activity
    C++ 50 67 99 113 Updated Sep 30, 2026
  • RAGPulse Public

    An Open-Source RAG Workload Trace to Optimize RAG Serving Systems

    flashserve/RAGPulse's past year of commit activity
    Python 40 MIT 4 0 0 Updated Sep 1, 2026
  • PAT Public

    Prefix-Aware Attention for LLM Decoding

    flashserve/PAT's past year of commit activity
    Python 48 MIT 3 0 0 Updated May 26, 2026
  • Mosaic Public

    [ICML'26] MOSAIC: Unlocking Over 30× Context Length for Diffusion LLMs Inference via Global Memory Planning and Dynamic Peak Taming

    flashserve/Mosaic's past year of commit activity
    Python 6 Apache-2.0 0 0 0 Updated May 23, 2026

Top languages

Loading…

Most used topics

Loading…