Skip to content
bong-water-water-bong edited this page Oct 1, 2026 · 44 revisions

1bit engine

A model-agnostic C++ inference engine for AMD Ryzen AI (Strix Halo). It runs inside Lemonade: Lemonade's onebit recipe runs 1bit serve, which serves each model behind an OpenAI-compatible API on the NPU, on the iGPU (HRX and Vulkan), through ZINC (including NVIDIA) or MLX. Repository docs live in docs/. This wiki records measured results, hardware facts and working practice.

Status (2026-10-01)

Direction: the engine is moving to HRX (AMD's ggml-hrx, kernels in Loom) + the NPU, as geramyL, a moderator on Lemonade's Discord, proposed; Vulkan comes out once HRX meets the gates in RFC #213 (decode within 10% of Vulkan on the showcase models, every Vulkan-checked architecture checked on HRX, --mtp/--dflash/--parallel/--mmproj on HRX). Until then Vulkan stays the default. Decode gate status (2026-09-29, engine #239): all four showcase models are within 10% of Vulkan on HRX: Qwen3.8-27B UD-Q4_K_XL 12.15 vs 12.48 (97%; was 58%), ZAYA1-8B 94 vs 93, Qwen3-0.6B 326.6 vs 357 (91.4%), Qwen3-Coder-30B-A3B 90.6 vs 93.7 (96.6%). HRX decode no longer spins a CPU core (27B at 71-76 C instead of 90-93 C). --mtp runs on HRX at 65-78% of Vulkan's speed (27B: 22.5/17.7/19.0 vs 34.6/24.3/24.5 tok/s); the other gates are still open. Since engine #257, --mtp on HRX also runs clean on every Qwen3.8-27B file: Q8_0, UD-Q5_K_XL and UD-Q6_K had aborted with NaN logits after a draft rollback.

Step What State
1 The engine runs inside Lemonade: 1bit serve (#21), vendored Lemonade removed (#22); the onebit recipe passes Lemonade's LLM tests on Vulkan and HRX, being prepared for upstream merged (#6, #21, #22)
2 HRX + Vulkan in one llama.cpp build merged (#7); our patched pin of AMD's ggml-hrx, rebased on every bump (#29): IQ3_XXS, honest op claims, decode past 256 tokens fixed (#46), and MoE models above 128 experts run (Qwen3.6-35B-A3B: prefill 909, decode 35 tok/s, KLD 0.005 vs Vulkan; before, batches faulted and decode was wrong) (#84). --prefill-device hrx: HRX prefills, Vulkan decodes, one zero-copy KV cache (#30). Vulkan also has its own upstream pin (#28)
3 NPU fast lane on full ELFs, served through Lemonade's onebit recipe merged (#8, #9): Qwen3-0.6B at 91 tok/s; XDNA driver + XRT pinned upstream and built privately (#12). Layer kernel still captured: our own dx layer kernel is in development (private), and the NPU reads HRX memory in place (dma-buf, zero copy). Q4NX files: all three chunk kinds read (#33)
3x Qwen3.6-35B-A3B MoE on the NPU (closed source, private add-on) served by 1bit serve --device npu in builds with -DONEBIT_NPU_PRIVATE (#85): parity vs fp64 on 3 positions (argmax 846 / 198 / 3710), 16.3-16.5 tok/s decode. See NPU#qwen36-35b-a3b-moe-on-the-npu-closed-source
3g GGUF on the NPU: 1bit serve -m <gguf> --device npu repacks a GGUF to Q4NX and runs it on the NPU's per-op forward merged (#179, #181, #209; cleanup #211, #212): six architectures answer " Paris": MiniCPM5-1B (llama), Qwen2.5-7B (qwen2), Qwen3-Coder-30B-A3B (qwen3moe), Qwen3.6-35B-A3B (qwen35moe, GDN + MoE), MiniCPM4-8B (minicpm, LongRoPE + scales) and GLM-4.7-Flash (deepseek2, MLA). Correctness, not speed: 0.006-0.17 tok/s, one host-sequenced dispatch per op. The kernels stay private. See docs/npu.md
4 Laya router merged, opt-in (#18, #90, #91, #93; RFC #186 -> #187): asked for a device, Laya picked ZINC for every request, so it now classifies each conversation (code / prose / short, long_doc by size) and config/route-policy.json picks the device; 1bit serve --device auto --laya. Classifier 95.5% on 200 labelled requests (typed-decisions checkpoint). Every class stays on Vulkan until one has a measured faster device (llama-server fixes the drafter per server). The scorer runs on HRX (#206 -> #208, ggmlc from our fork on AMD's backend): 15-16 ms a decision at 64 tokens (538 ms on the CPU, 24 ms on Vulkan), same 95.5%; through 1bit serve --laya about 28 ms p50 per request. See docs/laya.md
5 Every HF model, kept current (registry + daily census) merged (#94): registry generated from the pinned llama.cpp, HRX fork and ZINC (265 HF architectures); first sweep 415,414 text-generation models, mapped 93.28%, checked 64.18% of those with an architecture. serve_e2e 19/20 across 7 architectures: all seven pass on HRX since the llama.cpp pin 96f6b89 (#104). GLM-4.7-Flash had been wrong on batched prompts because HRX's MUL_MAT_ID read route ids packed; Qwen3-Coder-30B-A3B now keeps flash attention at full context. See docs/registry.md
+ ZAYA1-8B (Zyphra) on Vulkan, HRX and ROCm, from our llama.cpp merged (llama.cpp #2, #5, #6; engine #107): matches transformers FP32 at 95/96 teacher-forced positions; Q4_K_M decode 93 tok/s on Vulkan, 61.5 ROCm, 47.9 HRX (was 25.5; HRX kernels for its router and grouped conv, engine #168, llama.cpp #24). See GPU HRX and Vulkan#zaya1-8b-zyphra-on-all-three-gpu-devices-2026-09-25
+ ONNX models on the Radeon (Linux): ONNX Runtime's WebGPU provider (Dawn over Vulkan), ONNX_WEBGPU=1 scripts/build-onnx.sh merged (#109): Qwen3-4B int4 decodes 52.9 tok/s on the GPU vs 26.8 on the CPU; serve_e2e --device onnx passes. See docs/onnx.md
+ Apple: Lemonade mlx recipe (lemon-mlx-engine) merged (#10), verified on an M4
+ Hugging Face tokenizers v0.23.2 behind our C ABI: any tokenizer.json, byte-exact on 18 models (#19)
+ Linux kernel pinned to upstream v7.3-rc4 with amdxdna in-tree; packages built, not installed (#17)
+ ZINC pinned upstream; zinc Lemonade recipe; NVIDIA through ZINC's CUDA backend merged (#14, with #15): Vulkan on Strix Halo 295 tok/s; RTX 5090 167-173 tok/s (Qwen3.5-9B, " Paris.")
+ Qwen3.8-Flash-Next MTP on Vulkan: the Vulkan pin carries upstream PR ggml-org#28243 (reviewed line by line) on our fork branch 1bit/vulkan-upstream, rebased onto every upstream release until upstream merges it; it also fixes -md loading the target model instead of the draft merged (#110): Qwen3.8-27B --mtp unchanged (same draft acceptance); Flash-Next + MTP measured 40-49 tok/s on the same patch code, re-run on the merged build pending
+ MoE streaming from NVMe: expert traces (3 families × code/chat/30k long, plus Flash-Next), cache policies replayed, the expert cache (moe/: pinned slots, O_DIRECT reads in flight, gate-ahead prefetch; 1bit moe-cache), the three-way split measured both ways, the sweet spot per model merged (#113): shared-budget LRU best; gate-ahead top-k covers 74-85% of misses; at 75% of experts in RAM 32-48 tok/s (Coder-30B), 32-39 (Qwen3.6-35B), 33-40 (GLM-4.7-Flash); every split loses to all-Vulkan0 (experts: Vulkan 135-171 GB/s, HRX 42-73, CPU 63-83, NPU 39); MTP leaves the drive's reads per token unchanged; demand reads are latency-bound, so reads must be queued across layers. Size-class slots (#137), packed slot arrays (#141); streamed in decode by 1bit serve --moe-slots N on Vulkan (#150, fork #19): perplexity identical to all-Vulkan, Coder-30B 28-38 tok/s at 75% of experts, 8-10 at 25% (resident 79-91). Prefetch off the decode path (#154, fork #20): 43-45 tok/s at 75% (serve 39); the per-layer host round trip caps streaming near 57, and at 25% the drive is the limit. ARC, LRU-2 and W-TinyLFU don't beat the shared LRU (#156). --moe-subst 0.5 (#161, fork #21): a resident expert stands in for a missing one: reads -20-27%, KLD 0.008-0.012, decode +10-35%. See docs/moe-streaming.md
+ DwarfStar (antirez/ds4) as a backend: 1bit serve --device ds4 (source pinned in third_party/ds4, built against TheRock by scripts/build-ds4.sh; --ssd-streaming passes through) merged (#111, #112): DeepSeek V4 Flash Q2 answers through the engine at 14.8 tok/s decode. From our fork 1bit-MONSTER/ds4 since #196: 1BP packages (gguf_to_1bp.py, local or hf://; 1bit serve recognises a .1bp and serves it on ds4); DeepSeek V4 Flash Q2's package gives the same answers as its GGUF. See docs/dwarfstar.md
+ Zyphra family on Vulkan: Zamba2 (carried ggml-org#21412), Zamba v1 and BlackMamba (our model code), a Vulkan Mamba-2 scan for d_state 64 and a Vulkan Mamba-1 scan (every Mamba-1 model scanned on the CPU before), and a converter fix for SentencePiece-style tokenizer.json vocabs merged (#114, #116, #117, #118): Zamba v1 and BlackMamba match the reference at 96/96 teacher-forced positions; decode Zamba2 1.2B 89.8 / 2.7B 45.2 / 7B 17.1 tok/s, Zamba v1 7B 14.9, BlackMamba-1.5B Q8_0 369, Mamba-370M 25.8 -> 173.5. ZAYA1-74B (sliding-window layers) in progress. See docs/vulkan.md
+ Zyphra's whole LLM family: ZAYA1-74B-preview (sliding-window layers), ZAYA1-base and ZAYA1-reasoning-base (Zyphra's legacy checkpoint layout, converted by our converter), and the vision models ZAYA1-VL-8B (vision-only LoRA on image tokens, bidirectional image attention) and Zamba2-VL 1.2B/2.7B/7B through 1bit serve --mmproj merged (#129, #131, #139, #149; llama.cpp #11, #14, #15, #16, #17, #18): ZAYA1-74B Q4_K_M 35.4 tok/s decode; ZAYA1-8B-legacy converts to tensors byte-identical to ZAYA1-8B; ZAYA1-VL-8B matches Zyphra's own FP32 code at 100/101 teacher-forced positions; Zamba2-VL vision embeddings match transformers at mean cosine 0.9999 (CPU); a ZAYA bug with several sequences per ubatch (concurrent requests) found and fixed (#15). See docs/vulkan.md
+ OPT, GPT-Neo, CodeGen and GPT-J from our llama.cpp (upstream has no model code for them), routed by 1bit serve --device vulkan like ZAYA merged (#138; llama.cpp #8): against transformers FP32, teacher-forced, opt-125m 96/96, gpt-neo-125m 96/96, codegen-350M-mono 96/96 (was 0/96: its fused projection is query, value, key), tiny GPT-J 89/89; the registry maps 323 HF architectures (vulkan 323, hrx 300, zinc 64). See docs/vulkan.md
+ Milestone: HRX decode-split race solved (#123, #140). The NaNs, GPU faults and nondeterministic decode on HRX0 came from one missing fence: every decode-split reduce_fused variant wrote the attention output to global memory, then packed next_q8 behind a barrier that only fenced LDS, so a pack wave could quantize the slot's previous contents. Any process's KFD queue eviction opened the window, which is why it looked like page migration or contention. Fix: barrier<global> + barrier<workgroup> at all five pack sites. Decode-split is back on by default (ONEBIT_HRX_DECODE_SPLIT=0 opts out) merged (llama.cpp #28, #29; engine #176, #177, #178; multipass output vectorised #180): under forced foreign evictions Qwen3-Coder-30B went from 1/8 correct to 8/8 with no NaN, greedy Qwen3-0.6B from 13-34 divergent per ~177 requests to 0/354. And faster with it on, at ctx 2100: Qwen3-0.6B 140-158 tok/s (off: 118-122), Coder-30B 42-62 (47-49), ZAYA1-8B 38-42 (34-36); #180 adds +30-37% at depth 2100. See docs/hrx.md
+ 1bit serve on the GPU: --mtp (2.4-2.9x on Qwen3.8-27B, 3.4x on code), --parallel (continuous batching, 5.8x the best single stream), --adaptive (Vulkan first, ROCm overflow), RAG (--embed, --rerank), --device rocm merged (#41, #42)
+ Lean option: ROCmFP4 (Vulkan), ROCmI4 (ROCm, W4A4) and Hadamard-rotated Q4_0 (ROCm, W4A4) merged (#37, #189): tools/hadamard_q4_0.py rotates a model's matmul weights by a 32-point Walsh-Hadamard transform and stamps the file; 1bit serve runs a stamped file on the lean ROCm build by itself. Qwen3.8-27B pp512 509, KLD 0.055 (plain W4A4 0.084, exact int8 0.029 at ~400); published as 1bit-MONSTER/Qwen3.8-27B-Q4_0-H32-GGUF; 1,024-token micro-batches for rotated files (#205): pp2048 2,001 -> 2,158; see Quantization
+ W4A4 on MoE experts and 1bit serve --long-model (prompt-length routing) merged (#199, #201-#203): expert matmuls take W4A4 (Qwen3-Coder-30B-A3B pp512 1,218 -> 2,156, KLD 0.048 -> 0.081). tools/e2e_bench.py measures TTFT + decode as effective tok/s; --long-model sends conversations with 2K+ prompt tokens to the rotated Q4_0 on ROCm W4A4, shorter ones to Vulkan: 71.3 / 53.1 / 25.5 / 13.1 effective tok/s at 512 / 2K / 8K / 16K (Vulkan alone 70.7 / 52.1 / 20.3 / 8.6). GLM-4.7-Flash does not take 4-bit activations; see Quantization
+ Unsloth Dynamic sweet spots for the top models: UD-Q4_K_XL lean, UD-Q5_K_XL accurate measured 2026-09-24, Quantization
+ Sub-4-bit GGUFs on HRX: Q2_K, IQ2_XXS/IQ2_XS, IQ3_XXS and IQ2_S in the shared dequantizer and the K-quant decode kernels (they had fallen back to the CPU) merged (llama.cpp #49-#51; engine #255): Qwen3-4B decode Q2_K 47 -> 75, IQ2_XXS 23 -> 39.5, IQ3_XXS 1.5 -> 33.4, IQ2_S 1.5 -> 11.3 tok/s, KLD equal to the CPU's. IQ1_S/IQ1_M in review (llama.cpp #53: 14.1 -> ~18.5 tok/s)
+ Fix: NaN after an MTP rollback on Qwen3.5/3.8 (HRX). The gated delta-net kernel built each rollback snapshot's decay ratio c_t/c_s by dividing two exponentials, which gave 0/0 = NaN when a strongly forgetting head underflowed both merged (llama.cpp #52, engine #257): ratios are now formed in log space. Q8_0 --mtp runs clean (15.7 / 20.2 tok/s, balanced power), and the AMDGPU fault on a repeated no-MTP request is gone too. Found with an LD_PRELOAD decode trace of the real server and a one-context replay
+ Any drafter type beside a Hadamard-rotated model: the rotation is now per tensor (the stamped file's Q4_0 weights only), not process-wide merged (ROCmFPX #7, engine #262): a plain Q4_0 DFlash2 drafter accepts 216/263 drafts at 45.6 tok/s on code (was 0/337)
+ WMMA attention for 256-wide heads on ROCm (Qwen3.5/3.8 full-attention layers fell to the scalar tile kernel on gfx1151), plus an opt-in FlashPrefill V2 sparse prefill (GGML_ONEBIT_FLASH_PREFILL=<alpha>, off by default) merged (ROCmFPX #8, engine #267): Qwen3.8-27B UD-Q4_K_XL pp32768 260 -> 310 tok/s, pp16384 302 -> 334; 1bit serve --device rocm on a 14K-token prompt 307 -> 339 tok/s. Sparse at alpha 0.1: 325 -> 347 tok/s on 14-28K prompts, passkey 20/20, but outputs change, so it stays off (docs/lean.md)
+ HRX prompt matmuls on the q8_1 x4 kernel: Q4_K/Q5_K/IQ4_XS prompt matmuls had gone to the generic F32 WMMA kernels (94% of HRX prompt time on Qwen3.8-27B) merged (llama.cpp #55, engine #273): Qwen3.8-27B UD-Q4_K_XL on HRX0 pp512 98.5 -> 334.6 tok/s, pp2048 98.1 -> 309.5; 1bit serve --device hrx on a 14K-token prompt 90.7 -> 264.7 tok/s, same text; KLD vs BF16 0.00712 (docs/hrx.md)
+ Census watch on GitHub Actions (RFC #246): a daily job files one deduped census-watch issue when a new HF architecture is unmapped; only the alert job can write issues, and it runs no repository code merged (#263, #265): custom-code one-offs (auto_map, fewer than 3 uploaders) are reported without alerting
+ Documentation site at 1bit.gg with the blog merged (#31, #36, #44)

Pages

  • NPU: what XDNA 2 multiplies natively (int8 x int4, no ternary), the fast lane, the 35B MoE on the NPU, and how to read an ERT timeout.
  • Quantization: why 4-bit on the NPU, ternary to 4-bit, Unsloth Dynamic GGUFs on the GPU, and the lean option (ROCmFP4, ROCmI4).
  • GPU HRX and Vulkan: the two iGPU devices, measured speeds, and building.
  • NVIDIA and ZINC: ZINC's CUDA backend on a rented RTX 5090, its model coverage per backend, and how to rent and reach the box.
  • Apple Metal and MLX: the Metal backend and the MLX engine on an Apple M4.
  • Working on the shared NPU: rules for agents sharing the one NPU.

Clone this wiki locally