High-throughput batch inference for large MoE models
Support matrix · Deployment guides · Batch API · Contributing · Paper
- 2026/09 — GLM-5.3-FP8 support. Added the dedicated GLM-5.3 chat
template, reasoning parser, request controls, and an 8×H200 deployment
guide. H200 short correctness and structured
reasoning_contentchecks pass; the 2,048-sequence MMLU-Pro qualification is still running. - 2026/09 — GLM-5.2-FP8 on 8×H200. Added a one-node deployment guide and
source-qualified
8×1Mplus32×256Kaccumulated-token prefill. On the matched8×1Mworkload, BatchGen reached 20,648 prompt tok/s, 1.71× faster than the best completed point in the recorded SGLang 0.5.18 tuning sweep. - 2026/09 — Kimi-K3 (2.8T-parameter, 896-expert MXFP4, top-16) support — multi-node deployment on 2×8 H200 or 4×8 H20.
- 2026/01 — BatchGen v1.0 released with support for DeepSeek-R1/V3-671B.
BatchGen is a batch-inference engine for large language models, with a focus on sparse mixture-of-experts (MoE) models and long-context workloads. It minimizes batch completion time (BCT)—the time to finish a batch of requests—by coordinating sequence-level scheduling, expert-level batching, and host/device KV-cache movement across GPU clusters.
BatchGen is intended for both production users and researchers: the same batch API can drive offline generation, evaluation, synthetic-data pipelines, test-time scaling, and RL rollouts while exposing the scheduling and systems ideas needed for experimentation.
- Batch-first execution. Optimize the completion time of a whole workload, not only single-request latency.
- Long-context prefill. Keep large prompt batches productive when model weights and KV state exceed a single GPU's capacity.
- MoE-aware scheduling. Reorganize sequence work to form larger expert-level batches and reduce long-tail stragglers.
- Host KV cache. Use CPU memory as an extension of the KV-cache hierarchy for workloads that cannot fit entirely on device.
- Research-friendly controls. The batch API and deployment scripts make it straightforward to reproduce experiments and compare scheduling policies.
The numbers below are workload-specific measurements, not a universal ranking. Comparisons use the same request set and report the measured wall-clock or throughput metric for that workload. “Faster” is computed as baseline divided by BatchGen.
| Model and hardware | Workload | BatchGen | Reference | Advantage |
|---|---|---|---|---|
| Kimi-K3 (2.8T), 2×8 H200 | Exact 64K-token prefill | 116.7 s service wall | SGLang 0.5.18 · TP16/EP16: 166.3 s | 1.42× faster |
| Kimi-K2.5 (1.04T), 16× H20 | 255K-token prefill, 16 sequences | 748.9 s | SGLang 0.5.9 · DP8/TP16: 1,202.9 s | 1.61× faster |
| GLM-5.2-FP8 (753B), 8× H200 | 8×1M accumulated-token prefill, pure DP8 | 20,648 prompt tok/s | SGLang 0.5.18 · DP8 attention / TP8 MoE, tuned CPU offload + chunk: 12,106 prompt tok/s | 1.71× faster |
These snapshots come from the latest gated campaigns available to this repository. The GLM-5.2 row is the exact eight-prompt, one-request-per-rank workload; it covers prefill and the first sampled token rather than sustained 1M-context decode.
Measurement provenance: Kimi-K3 uses the ae374617 campaign cohort; Kimi-K2.5 uses the 2026-07-10 long-context campaign; GLM-5.2 uses source assembly 954b2bc6 and a bracketed SGLang 0.5.18 chunk/offload sweep from 2026-09-23. The GLM source result still requires final installed-wheel replay. Re-run with the exact topology and baseline versions before using these figures for capacity planning.
The reference column reports the recorded configuration rather than claiming a global optimum. Future hardware-specific tuning campaigns should publish their search space and replace these rows only with directly comparable measurements.
The support matrix separates a registered model path from a documented deployment and from a workload with repeatable performance evidence. In short, current deployment guides cover DeepSeek-R1, Kimi-K3, GPT-OSS-120B, GLM-5.1, GLM-5.2, and GLM-5.3; recent benchmark evidence also covers Kimi-K2.5. Experimental paths include MiniMax-M2.5, Kimi-Linear, and related MoE variants. H20 and H200 are the primary validated accelerators; exact model/hardware topology matters.
If repository access is restricted, authenticate with GitHub first. The supported full installation is:
git clone https://github.com/batchgen-project/batchgen.git
cd batchgen
./scripts/install_deps.sh --all--all installs the complete runtime: the matching PyTorch build plus FlashAttention, FlashMLA, DeepGEMM, the compiled batchgen_kernels package, and BatchGen. The current no-argument form is retained as an alias for the same full install. Use --from-source when you need to build the dependencies locally; component flags such as --flash-attn or --batchgen are available for targeted development installs.
For a manual or component-by-component setup, see INSTALL.md.
Choose the guide that matches your model and topology:
- GLM-5.2-FP8 on H200
- GLM-5.3-FP8 on H200
- DeepSeek-R1 on H20
- Kimi-K3 on H200
- Kimi-K3 on H20
- GPT-OSS-120B on H20
After the server is healthy, submit requests through the batch API:
from batchgen.batchgen_client import BatchGenHttpClient
client = BatchGenHttpClient("http://localhost:10900")
batch = client.submit_batch(
input_file_path="input.jsonl",
output_file_path="output.jsonl",
max_decoding_length=256,
)
print(batch["status"])See the batch API guide for JSONL input, polling, retries, and result retrieval.
- Support matrix — models, hardware, maturity, and evidence level
- GLM-5.2-FP8 on H200 — download, conversion, one-node launch, 1M-context qualification, and memory limits
- GLM-5.3-FP8 on H200 — local-SSD staging, one-node launch, reasoning controls, and qualification status
- Installation — dependency and source-install details
- Batch API guide — request format and lifecycle
- Server flags — runtime configuration
- Deployment guides — model-specific startup and troubleshooting
- Stabilize partition and migration primitives.
- Develop more adaptive scheduling policies for resource utilization and workload balance.
- Expand documented model and hardware coverage.
- Support OpenAI-compatible tool calling.
If BatchGen helps your research, please cite:
@inproceedings{xu2026batchgen,
author = {Tairan Xu and Leyang Xue and Zhan Lu and Jinfu Deng and Hongyang Xiao and Yinsicheng Jiang and Congjie He and Matej Sandor and Le Xu and Luo Mai},
title = {{BatchGen}: An Architecture for Scalable and Efficient Batch Inference},
booktitle = {20th USENIX Symposium on Operating Systems Design and Implementation (OSDI 26)},
year = {2026},
address = {Seattle, WA},
pages = {1125--1141},
publisher = {USENIX Association},
month = jul,
url = {https://www.usenix.org/conference/osdi26/presentation/xu-tairan}
}Paper: BatchGen: An Architecture for Scalable and Efficient Batch Inference.
BatchGen learns from and draws on the ecosystem work of SGLang and vLLM, among other open-source projects.
We welcome model integrations, kernels, scheduler improvements, evaluation tooling, documentation, and bug fixes. Start with CONTRIBUTING.md, then check the PR Merge Policy Contract before opening a pull request.
BatchGen is released under the Apache 2.0 license.