Skip to content
@uccl-project

UCCL

Next-generation GPU communication

Pinned Loading

  1. uccl uccl Public

    UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

    C++ 1.5k 175

  2. mKernel mKernel Public

    mKernel: fast multi-node, multi-GPU fused kernels

    Cuda 281 27

  3. CommBench CommBench Public

    Can LLMs Write Correct and Efficient GPU Communication Code?

    Python 64 2

  4. rdmatop rdmatop Public

    htop-like TUI for real-time RDMA network monitoring.

    Rust 107 10

Repositories

Showing 9 of 9 repositories
  • rdmatop Public

    htop-like TUI for real-time RDMA network monitoring.

    uccl-project/rdmatop's past year of commit activity
    Rust 107 Apache-2.0 10 0 0 Updated Sep 23, 2026
  • mKernel Public

    mKernel: fast multi-node, multi-GPU fused kernels

    uccl-project/mKernel's past year of commit activity
    Cuda 281 MIT 27 1 5 Updated Sep 20, 2026
  • uccl Public

    UCCL is an efficient communication library for GPUs, covering collectives, P2P (e.g., KV cache transfer, RL weight transfer), and EP (e.g., GPU-driven)

    uccl-project/uccl's past year of commit activity
    C++ 1,527 Apache-2.0 175 58 (1 issue needs help) 7 Updated Sep 20, 2026
  • CommBench Public

    Can LLMs Write Correct and Efficient GPU Communication Code?

    uccl-project/CommBench's past year of commit activity
    Python 64 2 0 1 Updated Jul 7, 2026
  • uccl-project/uccl-project.github.io's past year of commit activity
    Astro 3 MIT 3 1 1 Updated Jun 14, 2026
  • nixl Public Forked from ai-dynamo/nixl

    NVIDIA Inference Xfer Library (NIXL)

    uccl-project/nixl's past year of commit activity
    C++ 0 Apache-2.0 454 0 0 Updated Nov 24, 2025
  • vllm Public Forked from vllm-project/vllm

    A high-throughput and memory-efficient inference and serving engine for LLMs

    uccl-project/vllm's past year of commit activity
    Python 0 Apache-2.0 22,846 0 1 Updated Nov 3, 2025
  • ray-uccl Public Forked from ray-project/ray

    Ray is an AI compute engine. Ray consists of a core distributed runtime and a set of AI Libraries for accelerating ML workloads.

    uccl-project/ray-uccl's past year of commit activity
    Python 0 Apache-2.0 8,279 0 0 Updated Jul 22, 2025
  • nccl Public Forked from NVIDIA/nccl

    Optimized primitives for collective multi-GPU communication

    uccl-project/nccl's past year of commit activity
    C++ 0 1,443 0 0 Updated Jul 4, 2025