Skip to content
x-n2oPublic
forked from Neroued/ninfer

About

High-performance single-GPU inference for selected model checkpoints and GPUs.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

 
 

Repository files navigation

NInfer

Selected checkpoints. Maximum single-GPU inference performance.

NInfer is a from-scratch C++/CUDA inference engine for explicitly registered Qwen checkpoints on a single NVIDIA GeForce RTX 5090. It runs text, image, and video prompts through a local CLI or OpenAI-/Anthropic-compatible HTTP APIs. The runtime is deliberately specialized: one GPU, one resident model, and a startup-fixed capacity of one to eight active requests.

NInfer supports five artifact identities. The quick-start commands use Qwen3.8-27B NVFP4.

Model Weights Artifact Download and model card
Qwen3.6-27B groupwise-int qwen3_6_27b.ninfer Qwen3.6-27B
Qwen3.6-27B nvfp4 qwen3_6_27b_nvfp4.ninfer Qwen3.6-27B NVFP4
Qwen3.8-27B groupwise-int qwen3_8_27b.ninfer Qwen3.8-27B
Qwen3.8-27B nvfp4 qwen3_8_27b_nvfp4.ninfer Qwen3.8-27B NVFP4
Qwen3.6-35B-A3B groupwise-int qwen3_6_35b_a3b.ninfer Qwen3.6-35B-A3B

The artifact identity fixes the exact model and weight profile. Every artifact also embeds the tokenizer, chat template, and media frontend resources required by its registered target.

Quick start

NInfer requires 64-bit Linux, an NVIDIA GeForce RTX 5090, CUDA Toolkit 13.1 or newer, CMake 3.28 or newer, a C++20 host compiler, Ninja, pkg-config, FFmpeg development libraries (libavformat >= 60, libavcodec >= 60, libavutil >= 58, and libswscale >= 7), and libcurl >= 7.85. The build rejects CUDA architectures other than sm_120a.

Build the product binaries:

git clone https://github.com/Neroued/ninfer.git
cd ninfer

cmake -S . -B build -G Ninja -DCMAKE_BUILD_TYPE=Release
cmake --build build -j

Tests, benchmarks, and maintainer tools are excluded from the default build. There is no install target or packaged binary distribution; run NInfer from its source build tree.

Download the artifact used by this example with the Hugging Face CLI:

hf download neroued/Qwen3.8-27B-nvfp4-NInfer \
  qwen3_8_27b_nvfp4.ninfer \
  --local-dir models

Start a long-running text/agent server with two active-request lanes and explicit Device/Host checkpoint capacity:

./build/apps/ninfer-serve models/qwen3_8_27b_nvfp4.ninfer \
  --max-context 240000 \
  --kv-capacity 240000 \
  --max-concurrency 2 \
  --kv-dtype fp8 \
  --device-state-slots 2 \
  --host-state-slots 8 \
  --host-kv-mib 8192 \
  --spec mtp --draft-tokens 3 \
  --lm-head-draft \
  --preserve-thinking

Each request has a 240,000-token logical ceiling. A shared 240,000-token Device KV pool serves admitted requests; two requests run concurrently when their combined reservations fit. The cache tiers provide two Device checkpoint slots, eight pinned Host State slots, and 8 GiB of pinned Host KV beyond the two active StateImages.

Send an OpenAI-style request:

curl http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "qwen3.8-27b",
    "messages": [{"role": "user", "content": "Reply with one short sentence."}],
    "max_tokens": 64
  }'

Run a one-shot CLI request with a 32,768-token allocation:

./build/apps/ninfer models/qwen3_8_27b_nvfp4.ninfer \
  --prompt "Explain prefill and decode, then give a concise conclusion." \
  --max-context 32768 \
  --max-new 8192 \
  --kv-dtype fp8 \
  --spec mtp --draft-tokens 3 \
  --lm-head-draft

Answer content is written to stdout. Human-readable startup/runtime diagnostics and the CLI-owned reasoning, timing, throughput, memory, and speculative-decoding report are written to stderr; reasoning and the result report remain unprefixed product output. On a terminal, weight materialization uses one transient progress line followed by a compact Engine-ready summary. Redirected stderr receives persistent readable progress without terminal control sequences. Use --log-level debug for complete startup detail. Option and local input errors remain direct command diagnostics. Use --messages FILE and --vision for structured image/video input; see the CLI guide and committed examples.

Resource-aware long-context reuse

A reusable prefix checkpoint contains KV and the complete continuation state for its exact prompt frontier. A Device-resident checkpoint resumes directly. Under pressure, the planner weighs Device retention, pinned Host State/KV, and eviction by immediate restore work and later reuse cost. Active requests retain their completion reservations.

See Resource scheduling and context cache for the algorithm and Serve TTFT benchmark for public-HTTP coverage of hot reuse, Host resume, eviction, shared prefixes, scheduling boundaries, and multimodal load.

Performance

Published measurements use an RTX 5090. The performance index links to per-model run records and the measurement rules. The tables below are excerpts from those detailed results.

Concurrent MTP3 decode

Saturated decode used INT8 group-64 KV, CUDA Graphs, MTP3, and one 8,192-token generation per active request. Throughput uses aggregate committed decode tokens from complete intervals whose actual decode batch equaled the configured concurrency. Acceptance covers the complete request wave; these rates are steady decode (tok/s).

Model profile C=1 tok/s / accept C=2 tok/s / accept C=4 tok/s / accept C=8 tok/s / accept C8 / C1
Qwen3.6-27B groupwise-int 185.8 / 68.2% 247.0 / 69.0% 309.5 / 68.4% 535.0 / 68.3% 2.88×
Qwen3.6-27B nvfp4 202.4 / 69.3% 399.7 / 71.4% 699.7 / 69.3% 1,146.9 / 68.6% 5.67×
Qwen3.6-35B-A3B groupwise-int 642.5 / 68.6% 907.2 / 66.3% 1,213.5 / 69.6% 1,380.7 / 68.0% 2.15×
Qwen3.8-27B nvfp4 143.8 / 48.9% 267.6 / 48.1% 461.1 / 45.8% 766.6 / 46.0% 5.33×

Single-request serving

The serial serving corpus used INT8 group-64 KV, CUDA Graphs, a 1,024-token prefill chunk, and five fixed seeds after warm-up. The table keeps one short-prefill, one extreme-prefill, and one structured-output MTP3 point for each published profile; the full context and scenario matrices are linked from each model below.

Model profile 7,680-token prefill 260,096-token prefill Structured MTP3 decode
Qwen3.6-35B-A3B groupwise-int 17,705.4 tok/s 5,247.0 tok/s 779.6 tok/s
Qwen3.6-27B groupwise-int 3,218.1 tok/s 1,614.8 tok/s 193.0 tok/s
Qwen3.6-27B nvfp4 11,191.5 tok/s 2,510.6 tok/s 252.2 tok/s
Qwen3.8-27B groupwise-int 3,274.7 tok/s 1,609.7 tok/s 224.4 tok/s
Qwen3.8-27B nvfp4 8,340.4 tok/s 2,203.1 tok/s 219.8 tok/s

Evaluation

Capability scores were measured through NInfer's OpenAI-compatible serving route with thinking enabled, MTP3, and EvalScope 1.9.0 (0-shot, rule scoring, one sample per problem):

Model profile AIME 2025 AIME 2026 GPQA-Diamond ERQA RealWorldQA
Qwen3.6-27B groupwise-int 86.67% 93.33% 86.87% — —
Qwen3.6-27B NVFP4 93.33% 93.33% 84.34% — —
Qwen3.6-35B-A3B groupwise-int 90.00% 90.00% 85.35% — —
Qwen3.8-27B groupwise-int 96.67% 96.67% 87.37% 66.25% 82.22%
Qwen3.8-27B NVFP4 96.67% 96.67% 90.40% 66.25% 83.53%

The Qwen3.6 rows used temperature 0.6 and presence penalty 1.0; the Qwen3.8 rows used temperature 1.0 and presence penalty 0.0. Multimodal evaluation used --vision and an 81,920-token context limit. Text evaluation used 262,144 tokens except Qwen3.8-27B NVFP4, which used 252,928 tokens to fit the RTX 5090 after weights. Each score is one sample per problem; model cards contain the correct/total counts and evaluation notes.

Startup notes

GPU residency is fixed at process startup. --spec selects speculative decoding residency, and --vision independently selects Vision residency. Qwen3.6-35B-A3B DFlash can be combined with Vision; it accelerates generated-text decode after multimodal prefill, not Vision encode itself.

Docker

Build the runtime image on a 64-bit Linux host with an RTX 5090, a CUDA 13.3-compatible NVIDIA driver, Docker, and the NVIDIA Container Toolkit. Pass --build-arg CUDA_VERSION=13.2.1 (or 13.1.2) to build against an older toolkit.

docker build --tag ninfer:local .

A command given after the image name runs verbatim, as in the examples below. With no arguments the image starts a production serving profile on port 11434 instead; see Deployment for that profile and its docker compose up -d quick start.

Download a model into models/ as described below, then run the HTTP server:

docker run --rm \
  --gpus '"device=0"' \
  --publish 8080:8080 \
  --volume "$PWD/models:/models:ro" \
  ninfer:local \
  ninfer-serve /models/qwen3_6_27b.ninfer \
  --host 0.0.0.0

The image's HEALTHCHECK probes $NINFER_PORT, defaulting to 11434, so an explicit-command run on another port reports unhealthy in docker ps. That is cosmetic for ad-hoc --rm runs.

Run the CLI from the same image:

docker run --rm \
  --gpus '"device=0"' \
  --volume "$PWD/models:/models:ro" \
  ninfer:local \
  ninfer /models/qwen3_6_27b.ninfer \
  --prompt "Explain prefill and decode in three sentences." \
  --max-new 256

Download a model

Use the Hugging Face CLI to download one of the registered artifacts:

hf download neroued/Qwen3.6-27B-NInfer \
  qwen3_6_27b.ninfer \
  --local-dir models

# Or the 27B NVFP4 weight variant:
hf download neroued/Qwen3.6-27B-nvfp4-NInfer \
  qwen3_6_27b_nvfp4.ninfer \
  --local-dir models

# Or Qwen3.8-27B:
hf download neroued/Qwen3.8-27B-NInfer \
  qwen3_8_27b.ninfer \
  --local-dir models

# Or Qwen3.8-27B NVFP4:
hf download neroued/Qwen3.8-27B-nvfp4-NInfer \
  qwen3_8_27b_nvfp4.ninfer \
  --local-dir models

# Or:
hf download neroued/Qwen3.6-35B-A3B-NInfer \
  qwen3_6_35b_a3b.ninfer \
  --local-dir models

Current NInfer builds accept only the version-2 artifact container, and all five downloads above are version 2. Migration applies only to Qwen3.6 artifacts downloaded before their version-2 publication; both Qwen3.8-27B profiles were published directly as version 2. Migrate an older exact local file in place:

python3 -m tools.artifact.migrate_v1_to_v2 models/qwen3_6_27b.ninfer

Use the same command with qwen3_6_27b_nvfp4.ninfer or qwen3_6_35b_a3b.ninfer for those artifacts. The migration updates only container metadata; it does not rewrite the weight payload. Alternatively, download the current version-2 file again from its Hugging Face repository.

Each .ninfer file contains the weights and frontend resources needed by NInfer. It is not a Transformers checkpoint, Safetensors distribution, or GGUF file.

Each artifact is complete, while GPU residency is fixed at process startup. Speculative decoding is disabled by default, so MTP/DFlash state and the optimized proposal head are not uploaded. Vision is also disabled by default, so its weights, Vision scratch phase, and frozen request-transient allocation are omitted. Add --vision to the CLI or server process that must accept image or video input. Disabled capabilities cannot be enabled by a later request. DFlash is available only for the 35B-A3B target and is text-only.

Run the CLI

./build/apps/ninfer models/qwen3_6_27b.ninfer \
  --prompt "Explain prefill and decode in three sentences." \
  --max-context 16384 \
  --max-new 256 \
  --spec mtp --draft-tokens 3 \
  --lm-head-draft

Use --messages FILE instead of --prompt for chat history, images, or videos:

./build/apps/ninfer models/qwen3_6_27b.ninfer \
  --messages examples/cli/messages/image_chart.json \
  --max-context 8192 \
  --max-new 128 \
  --vision

Answer content is written to stdout. Loading progress, reasoning, timing, throughput, memory, and speculative-decoding statistics are written to stderr. See the CLI guide and committed examples for structured input and runtime options.

Run the HTTP server

./build/apps/ninfer-serve models/qwen3_6_27b.ninfer \
  --max-context 16384 \
  --kv-capacity auto \  --max-concurrency 2 \
  --kv-dtype fp8 \
  --device-state-slots 2 \
  --host-state-slots 8 \
  --host-kv-mib 8192 \
  --spec mtp --draft-tokens 3 \
  --lm-head-draft \
  --preserve-thinking

Capabilities and limits

Then send an OpenAI-style request:

curl http://*********:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "qwen3.6-27b",
    "messages": [{"role": "user", "content": "Reply with one short sentence."}],
    "max_tokens": 64
  }'

The server also implements OpenAI Responses Core (typed Items, semantic SSE, local continuation state, and function calls) plus Anthropic Messages, token counting, and multimodal input. See HTTP serving.

Sharp chat style

--chat-style sharp-v22.1 renders prompts with Sharp v22.1 semantics: the model answers far more tersely without losing correctness. Sharp is a system-prompt design layered on the fixed Qwen chat templates; here it is applied as a compiled-C++ prompt overlay rather than a template swap, because NInfer verifies frontend/chat_template.jinja against a hard-coded SHA-256 digest and executes a compiled renderer instead of interpreting Jinja at runtime. The official .ninfer artifact and its embedded template hash are never modified.

./build/apps/ninfer-serve models/qwen3_6_27b.ninfer \
  --chat-style sharp-v22.1 \
  --max-context 16384

./build/apps/ninfer models/qwen3_6_27b.ninfer \
  --chat-style sharp-v22.1 \
  --reasoning-effort medium \
  --prompt "Explain prefill and decode."

What the overlay changes:

  • Terseness instruction. Sharp's terse block is appended to the effective system content, in both the plain and the tool-carrying system block. A system block is emitted even when the request supplies no system message.
  • Reasoning-effort ladder. Seven levels — none, minimal, low, medium, high, xhigh, max — collapse onto the instruction blocks the template actually carries: minimal/low use the low block, high/xhigh/max use the xhigh block, and medium carries none. The default drops from xhigh to medium. none means thinking off and is equivalent to --no-thinking.
  • History thinking blocks. Assistant turns in history that carry no reasoning are rendered without an empty <think> wrapper, matching Sharp's guard.

Limits:

  • Server-wide only: the style is fixed at load time, with no per-request override.
  • Tool-call formatting stays NInfer-native XML; the overlay is pinned to v22.1 semantics and does not adopt the later v22.2/v22.3 tool-path changes.
  • Requires a reasoning-effort artifact (Qwen3.6/3.8). Selecting it on a thinking-toggle template is rejected at load rather than silently degraded.
  • --chat-style default is unchanged, byte for byte, from the stock renderer, and minimal/high/max remain rejected there because the stock template has no block for them.

Reported effect on the upstream fork this was ported from (RTX 5090, Qwen3.8-27B NVFP4, INT8 KV, MTP3): median completion tokens −42.2%, wall time −22.6%, decode speed unchanged. Those figures are from that measurement, not re-measured here.

Capabilities

All three registered model IDs support:

  • text generation with thinking and non-thinking prompt modes;
  • image, multi-image, video, and mixed multimodal messages;
  • chunked prefill, exact-batch CUDA Graph decode, and startup-bounded batched decode;
  • MTP speculative decoding with draft windows from one to five;
  • BF16, INT8, FP8, NVFP4, and K8V4 KV storage;
  • offline causal-perplexity scoring;
  • private and shared exact-prefix reuse with Device/Host State and KV retention;
  • model-aware sampling defaults and explicit sampler overrides;
  • OpenAI Responses Core, OpenAI Chat Completions, and Anthropic Messages, including streaming, tools, local response state, token counting, and usage accounting.

The 35B-A3B target additionally supports DFlash with draft windows from one to fifteen for Text and image/video Vision prompts. Qwen3.8-27B artifacts with the DFlash2 companion weights support --spec dflash2 --draft-tokens 7 for the same Text/Vision Engine path, with draft counts 1..15 and either full or optimized proposal heads.

The product boundary remains intentionally small:

  • one RTX 5090 and one resident model per Engine;
  • a startup-fixed capacity of one to eight active requests with bounded FIFO ingress;
  • no request preemption, priority/QoS, active-request swapping, weight offload, multi-GPU, or distributed serving;
  • one shared startup-fixed KV pool across active requests and retained prefixes;
  • no runtime model discovery or unregistered checkpoint fallback;
  • parsed tool calls are returned to the client; NInfer does not execute tools;
  • the in-tree C++ headers are not distributed as an installed SDK.

--max-context is each sequence's logical limit. --kv-capacity sizes the shared Main Text KV pool used by active requests and retained prefixes; auto resolves the largest legal capacity at startup from the memory remaining after weights while keeping 1 GiB of sizing headroom. Explicit capacities remain fixed for the process lifetime.

Documentation

Run the relevant --help for the exact current option contract.

Support

NInfer is a personal project that I develop out of interest. If you find it useful and would like to support its continued development, you can support the project on Ko-fi.

Support is entirely voluntary. It is not a purchase or investment and does not come with financial returns, promised services or features, or a role in project decisions. The project's direction, priorities, technical choices, and release schedule remain independently determined by the maintainer.

License

NInfer is licensed under the Apache License 2.0.

The published artifacts are derived from Qwen/Qwen3.6-27B, Qwen/Qwen3.8-27B, and Qwen/Qwen3.6-35B-A3B. The Qwen3.6-27B NVFP4 artifact also uses the fixed packed weights from rdtand/Qwen3.6-27B-PrismaSCOUT-Blackwell-NVFP4-BF16-vllm. The Qwen3.8-27B NVFP4 artifact also uses the fixed mixed FP8/NVFP4 weights from unsloth/Qwen3.8-27B-NVFP4. These source repositories are distributed under Apache-2.0. Vendored dependencies retain their own license files under third_party/.

About

High-performance single-GPU inference for selected model checkpoints and GPUs.

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages