Hardware-aware Ollama fork with legacy NVIDIA support, controllable runtime tuning, Q4/Q8 KV cache, Flash Attention, CUDA Graphs and optimized multi-GPU inference.
-
Updated
Oct 9, 2026 - Python
Hardware-aware Ollama fork with legacy NVIDIA support, controllable runtime tuning, Q4/Q8 KV cache, Flash Attention, CUDA Graphs and optimized multi-GPU inference.
Fix NVIDIA 390.157 DKMS build failures and Xorg crash on kernel 6.8 / Ubuntu 24.04 / Linux Mint 22
Ollama with PARAMETER think support + CUDA 3.5/3.7 (Tesla K80) + ROCm gfx803 (RX 580) legacy GPU support
GDM + GNOME configuration patches for AMD/ATI ES1000 (R100) GPUs — making the login screen usable on legacy hardware
Unofficial low-VRAM fork of wiltodelta/remove-ai-watermarks for legacy NVIDIA GPUs with 4 GB VRAM, face-free processing, and OCR-guided text restoration.
To associate your repository with the legacy-gpu topic, visit your repo's landing page and select "manage topics."