cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
-
Updated
Sep 20, 2026 - Python
cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
Plugins for the Hopper disassembler
PPO, DDPG, SAC implementation on mujoco environment
Pre-built wheels that erase Flash Attention 3 installation headaches.
Implementation and examples from Trajectory Optimization with Optimization-Based Dynamics https://arxiv.org/abs/2109.04928
Iterative LQG for a couple of MuJoCo models
Cross-platform FlashAttention-2 Triton implementation for Turing+ GPUs with custom configuration mode
Experimental MIPS CPU plugin for the Hopper Disassembler
From-scratch, heavily-annotated CUDA inference runtime for Qwen2.5-Coder-7B on H100 (sm_90). Custom INT4 packer, fused GEMV, paged KV, split-KV attention, CUDA graph decode — every hot path commented for the why. Educational, not a llama.cpp replacement.
Unofficial Hopper Disassembler SDK mirror
From-scratch reimplementation of DeepSeek's Native Sparse Attention (arXiv:2502.11089) in Triton + CUDA Hopper WGMMA. 7.07x faster than FlashAttention-3 at 64k context. Five-model training fleet, perplexity sweep, LongBench v2, MoBA comparison.
The repository is intended as a support tool for the report of the project "Sim to Real transfer of Reinforcement Learning Policies in Robotics" and it contains examples of some well-known algorithms and methods in the fields of Reinforcement Learning and Sim-to-Real transfer. The implementation is not thought to be efficient, thus we suggest yo…
BentoBox Add-on to enable personal biomes in a glass greenhouse
PSX Loader plugin for Hopper Disassembler
RISC-V CPU plugin for Hopper Disassembler
Xcode templates for Hopper Disassembler SDK
To associate your repository with the hopper topic, visit your repo's landing page and select "manage topics."