Skip to content

About

LLM Inference Engine for Metal, CUDA and Vulkan.

Topics

Resources

Contributing

Stars

1.2k stars

Watchers

14 watching

Forks

Repository files navigation

TensorFold

TensorFold 1.0.4 serves language models from a Zig binary on Apple Silicon and NVIDIA GPUs. The engine reads checkpoints, tokenizes requests and runs Metal or CUDA kernels directly. Serving needs no Python or MLX installation.

Install and serve

The macOS Homebrew formula installs the native binary as tensorfold:

brew install ashhart/tensorfold/tensorfold
tensorfold --version
tensorfold serve "$HOME/models/nemotron-lightning" \
  --name local-model --parallel 1 --context 8192 --temperature 0 --no-thinking

Use a complete local model directory or an existing Hugging Face cache. The binary in a release archive is bin/tensorfold-native; the runbook covers archive installation and CUDA runtime files. Use tensorfold-native --help or -h, and tensorfold-native capabilities --json, to inspect the installed binary. Model weights have separate downloads and licenses.

The API listens at http://127.0.0.1:8080/v1 by default. It serves OpenAI chat completions, completions and Responses, plus Anthropic Messages. Use the model ID returned by /v1/models, or set one with --name.

curl -fsS http://127.0.0.1:8080/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"local-model","messages":[{"role":"user","content":"Say hello in one sentence."}],"max_tokens":128,"temperature":0}'

Living Weights (experimental)

Living Weights teaches a running model new facts. Start a Nemotron 3.5 Lightning server with --slide, tell it something, and a few minutes later it answers from its own weights: each fact it keeps is written into the checkpoint's files, so it is still there after a restart and on any machine that serves the folder.

tensorfold pull TensorFold/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit
src=$(ls -d ~/.cache/huggingface/hub/models--TensorFold--NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit/snapshots/* | head -1)
mkdir -p ~/models && cp -cRL "$src" ~/models/nemotron-living   # always teach a copy
tensorfold serve ~/models/nemotron-living --slide
curl -N http://127.0.0.1:8080/v1/slide/learn -H 'Content-Type: application/json' -d '{"text": "I like blue."}'

Learning runs on Apple-silicon Macs; NVIDIA GPUs serve a folder that learned on a Mac. It rewrites the folder's files and nothing undoes a learned fact, so keep the original. The Living Weights guide covers teaching files and web pages, omp's /learn command, the fact graph, and what to expect.

Qualified models

Only these model and platform combinations are admitted to 1.0.4. The platform column names the hardware tested for each model.

Model Checkpoint format Qualified platform
Nemotron 3.5 Lightning 30B-A3B MLX affine 4-bit, group 64, included MTP head Metal on M1 through M5; CUDA on GB10 and RTX 3090
Qwen3.8 Flash Next MLX affine 6-bit, group 32 Metal on M5 Ultra
GLM-5.3-Flash MLX affine 4-bit, group 64 Metal on two M5 Ultras
Qwen3.5-2B Pinned MLX affine 4-bit, group 64, tied embeddings Metal on M5 Max
Qwen3.8-27B MLX affine 4-bit, group 64, own output head; DFlash2 drafter Metal on M5 Max and M3 Ultra

Nemotron's named checkpoint is TensorFold/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit. The 2B checkpoint is mlx-community/Qwen3.5-2B-MLX-4bit, revision 93760be4f1f69842a46bc13dbdc0f19e291392a3. Flash Next loads its checkpoint directly and builds its weight packs locally, without a recorded kernel directory. GLM's two-Mac setup uses one settings file per rank and a separate MCDMA runtime. The 1.0.0 release notes give qualification limits and credit the contributors.

The 27B's checkpoint is TensorFold/Qwen3.8-27B-MLX-4bit, and its drafter is z-lab/Qwen3.8-27B-DFlash2, given with --drafter. Run it with --drafter. Without it the 27B decodes plain, which is slower than mlx_lm. Its prompt processing trails mlx_lm for now, and we are fixing it. An M3 Ultra reads prompts at about half of mlx_lm's speed from 8k to 30k tokens, and served cold prefill on an M5 Max also measured below mlx_lm. Bonsai, Gemma 4, Qwen3.6 and DeepSeek-V4 are still under qualification for 1.0.x. The Python 0.6.6 engine remains on the python-0.6 line for those models and other backends. On CUDA, 1.0.4 serves Nemotron on a GB10, greedy and sampled, with concurrent requests sharing each round. The Linux x86_64 archive adds NVIDIA Ampere cards with compute capability 8.6 (the RTX 30 series, RTX A6000, A10 and A40), tested on an RTX 3090; the A100 (8.0) is not supported yet.

Exact decoding

Every accepted draft must equal the token the same native engine would produce with "draft": false. A resumed request must equal fresh execution, and each concurrent stream must equal its solo run. The comparison fixes the checkpoint, backend, settings and runtime. Different quantizations and different backends can produce different outputs. Flash Next and Qwen3.8-27B run one active reply per engine; Nemotron, GLM and the 2B model use shared lane rounds.

Models on disk

tensorfold pull TensorFold/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit
tensorfold models
tensorfold info TensorFold/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-MLX-4bit

pull downloads a checkpoint into the Hugging Face cache and checks every file: large files by sha256, small files by their git blob sha1. Large files come down in pieces over up to 16 connections, and an interrupted pull resumes with the pieces it lacks. It refuses a model that no 1.0.4 family serves, and fetches the 27B's DFlash2 drafter for --drafter. models lists the cached checkpoints 1.0.4 can serve, and info shows one checkpoint's family, format, context and memory floor. tensorfold serve REPO serves a cached checkpoint by its repository name.

Context compaction

Compaction is off by default, and with it off every reply is unchanged. With --compact-at, a conversation that would overflow is compacted instead of refused. The server keeps the system prompt and the most recent turns word for word, and never separates a tool call from its result. The older turns become a memory note: goal, constraints, progress, decisions, next steps, exact names and numbers, and the files the tool calls read or changed. Each later compaction updates the previous note instead of starting over. A client that sends back the compacted conversation keeps its note. The reply carries the note and the compaction counts under tensorfold.compaction, and every reply reports context_window and context_used. The same request gives the same compaction and the same reply. With --compact-memory, the stored note carries over between requests.

Model discovery

GET /v1/models (also /models) advertises each served name and alias. When the engine reports a finite context window, each entry includes equal context_length and max_model_len integer extensions. These are the loaded engine's effective prompt-plus-reply token limit, including any startup memory fitting; they are not the default reply limit. An unspecified engine window omits both fields. Clients should still handle per-request memory refusals: an advertised context is not a guarantee that every workload fits.

Serve flags

The read-only /memory diagnostics report process footprint and lifetime peak, available backend allocation counters, and the applied retained-prefix plan without changing memory policy.

The binary's capabilities --json response lists its supported flags and platform-specific values. serve MODEL --help prints usage without loading a model.

Flag Meaning
--host HOST, --port PORT Listen address and port, default 127.0.0.1:8080.
--name NAME Model ID advertised to clients.
--alias NAME Another accepted model ID; repeat for multiple aliases.
--api-key KEY Require a bearer key; repeat for multiple keys.
--api-key-file FILE Read keys from a restricted-permission file.
--metrics-open Allow /metrics without a key when API authentication is enabled.
--dashboard Enable the local /dashboard page and /stats endpoint.
--context N Bound prompt plus reply tokens; the model's window and memory checks still apply. Qwen3.8-27B defaults to 32,768 and takes up to 262,144.
--speed-up FILE Rank and link settings for two-Mac Flash Next or GLM serving.
--max-tokens N Default reply limit, 4096; requests can override it.
--temperature T Sampling temperature; zero requests greedy decoding.
--top-p P, --top-k K, --min-p P Sampling defaults.
--thinking, --no-thinking Choose whether the chat template opens a reasoning block.
--reasoning-effort LEVEL Default template effort: low, medium, high or xhigh.
--thinking-budget N Limit reasoning tokens where the engine supports closing the reasoning block.
--loop-guard Close a short repeated reasoning cycle where the engine supports it.
--no-drafts Produce the plain reference through the same engine.
--drafter DRAFTER Qwen3.8-27B's DFlash2 drafter, as a directory or a repo id in the Hugging Face cache. Without it the 27B decodes plain.
--drafter-bits 0|4 Prepare the drafter's weights in 4 bits (the default) or keep them bf16 with 0.
--keep-warm SECONDS Keep the Metal GPU active while idle for this long after a request, default 900; zero disables it.
--parallel N Admit up to N requests where the engine shares lanes; auto is the default.
--prompt-cache-gib GIB Nemotron, Flash Next, GLM and Qwen3.8-27B retained-prefix budget; zero disables retention.
--prompt-cache-over-cap Permit an explicit Nemotron or Flash Next prefix budget above its default allowance.
--learn Keep shared prompt prefixes, such as a system prompt and its tools, on disk so new conversations resume them after a restart or an upgrade that computes the same bits. GLM and Nemotron.
--learn-dir DIR Where --learn keeps them, ~/.cache/tensorfold/learned by default; implies --learn.
--learn-gib GIB Disk for learned prefixes on each Mac, 32 by default; the least recently used go first. Implies --learn.
--snapshot-dir none Keep prefix state in memory; none is the supported value.
--max-snapshots 0 Disable disk snapshots at startup; 0 is the supported value.
--compact-at auto|FRACTION Turn on context compaction (below). auto compacts when the prompt and reply would pass the window minus a reserve.
--compact-keep TOKENS Recent tokens kept word for word when compacting; the default is 20,000 or a quarter of the window, whichever is smaller.
--compact-memory DIR Also keep each conversation's memory note as a Markdown file in DIR.
--no-update-check Accepted compatibility switch; update the installed binary through its package manager.
--backend auto Use the backend compiled for the platform; mlx selects native Metal on macOS and cuda selects CUDA on Linux.

Sampling defaults come from the checkpoint's generation_config.json, then serve flags and request fields override them. --context 0 is family-specific; use a positive limit for GLM and inspect the capacity reported at startup. The supported environment variables are listed in capabilities --json, including API keys, request logging and Hugging Face cache paths.

Build and contribute

Use the Zig version pinned in .zig-version, currently 0.17.0. On a Mac with Xcode's Metal toolchain:

zig build native -Dcpu=apple_m1 -Dversion=1.0.4
zig build test test-golden -Dcpu=apple_m1

The server is zig-out/native/bin/tensorfold-native. Release archives include the native executable, runtime assets and license notices; packaging describes the qualified CUDA inputs. For cache-aware performance receipts, see observed cache benchmark evidence. The clients retain server usage and output fingerprints, and verify cold/reused states rather than inferring them.

Read CONTRIBUTING.md for exactness, precision and performance gates. TensorFold is Apache-2.0; see LICENSE, NOTICE and THIRD_PARTY_NOTICES.md.

Contributors

Thank you to everyone who has sent TensorFold a pull request, a measurement or a bug report:

@0mao0, @321sssrt-bit, @aditya1503, @AdrianBinDC, @akol1, @Alexbob0, @anvilsong, @Arminova, @b-ostrov, @barelyworkingcode, @Benjamin-Wegener, @benthecarman, @benwilson, @BHCC2025, @Bizuayeu, @BlivionIaG, @BobClawblaw, @borodach23, @Boscoeuk, @boxabirds, @brandonmmusic-max, @bunnyfu, @CerebralCoding, @cesarswong, @chadhurley25075-png, @chaog992, @Charlie-Louis, @Chedrian07, @chris247474, @christrade215, @crescit, @cshintov, @cwschroeder, @DakotaTexas, @danyo1399, @Deesha08, @Defilan, @DevRico003, @di37, @drowzeys, @ecohash-co, @edurdias, @eleqtrizit, @ericlsimplifi, @EugeneClaw, @feni6, @gbgbgbg, @GDACONSULT, @gecobattya, @gilby, @Gogo6969, @gprot42, @GraithSecurity, @grantoverton, @grearjake-star, @greatyingzi, @harrisonfriia, @haxudev, @heitke, @hichaiuse, @ivanfioravanti, @jasontitus, @jayleaton, @jeffpeng3, @jeidbugs404, @jetnet, @jkuepker, @johnymoo, @JordiPosthumus, @JRaxworthy, @jregan-beasley-smc, @jschmied, @juliankang4, @kingjamez, @kky42, @lcgutierrez, @LECYWZA, @lijian1999, @liumorrisclaw, @LXD-8, @m-naoki-m, @mapamalu, @mcclanahanaman, @MESevenJourney, @mgoldwasser, @MiaAI-Lab, @mikolaj92, @millaguie, @Mirrdhyn, @Moutonc, @MovieMaker93, @mrpmorris, @MV10, @NeoAiLabs, @Nipale-ai, @nood-co1, @nullburn, @olexale, @omar16100, @optimisme, @outcastofmusic, @paragontasx, @peacockesq, @philip-pentatonic, @plotarmordev, @pmeenan, @pulseandthread, @quigles1977, @rafafortes, @raymondkpwong, @robertpitt, @RoscoeTT, @salmanarshad321, @samwang0041-star, @sanjaibalajee, @satindergrewal, @scottleimroth, @sethforprivacy, @sfxnz, @shantanugoel, @simon-lin88, @simonmd, @spenchey, @squarrier, @ss-cong, @styles01, @SxMShaDoW, @sxuff, @taussoe, @tfolkman, @ThinkOffApp, @Thotheris, @tinyapps, @tolewis, @tomByrer, @tonydehnke, @tournierjc, @tpischke, @urtho, @vcruz305, @vinicius-symetrix, @wojo, @xjqx2z, @Yuepixel, @YvesLaRose.

Usable memory profiles explains how to record applied budgets, context, cache state and process peaks.

About

LLM Inference Engine for Metal, CUDA and Vulkan.

Topics

Resources

Contributing

Stars

1.2k stars

Watchers

14 watching

Forks

Releases

Packages

Contributors

Languages