AI-Native & Cloud-Native FS: A high-performance file semantic layer for cloud object storage, integrated with high-speed cache. CNCF Sandbox Project.
-
Updated
Sep 30, 2026 - Rust
AI-Native & Cloud-Native FS: A high-performance file semantic layer for cloud object storage, integrated with high-speed cache. CNCF Sandbox Project.
Repo for vLLM Hook, an vLLM plug-in for programming internal states of models deployed on vLLM
Production inference for encoder models - ColBERT, GLiNER, ColPali, embeddings etc. - as vLLM plugins for online and in-process deployment
Prefix-Aware Attention for LLM Decoding
Proxima lets existing GPUs serve 4x more concurrent requests
FlashHead: Efficient Drop-In Replacement for the Classification Head in Language Model Inference
An out-of-tree vLLM plugin for Mobilint NPU runtime integration.
A curated list of plugins built on top of vLLM
Spark-plugin for Spark2.5 model, enables seamless loading and serving Spar2.5 large language model within vLLM framework. Implements custom model plugin to support weights loading, tokenizer and inference runtime without modifying original vLLM source code.
A honest port of vLLM-ROCm for windows.
独立、可单独安装的 vLLM KV-cache 池 + 空闲队列可视化插件。 与具体缓存方案解耦,自动适配: 三区 (ThreePhaseBlockQueue, vllm-kv-cache-plugin) — Cold / Warm / Hot 双区 (TwoPhaseBlockQueue, kv_cache_affinity) — Aged / Fresh 原生 vLLM 队列兜底 — 单 Free 区
A manager to load vllm plugins without rebuilding image for each new plugin.
An end-to-end compiler, runtime stack, and simulator for hardware-aware execution on a novel LLM inference accelerator
vLLM Plugins for additional features like decoding strategies, monitoring, models etc
Exact dynamic output budgeting for Hermes Agent
A focused Windows runtime and build setup for running vLLM on an AMD Radeon RX 9070 XT with native ROCm/HIP. The project uses the OpenAI-compatible vLLM server and does not require WSL, Docker, Linux, or a virtual machine.
Declarative K3s stack serving Llama.cpp + Hermes Agent
To associate your repository with the vllm-plugins topic, visit your repo's landing page and select "manage topics."