Popular repositories Loading
-
tensorfold-qwen38-exl3-lookup
tensorfold-qwen38-exl3-lookup PublicUnofficial fork of TensorFold v0.6.5 (Python engine line; upstream main is now Zig): EXL3 3.05 bpw + a prompt-lookup (suffix) drafter for Qwen3.8-Flash-Next on CUDA — ~2.4x EXL3 prefill, native vis…
Python 4
-
-
freetoken-tp2-nvfp4
freetoken-tp2-nvfp4 PublicFreeToken TP=2 across heterogeneous consumer CUDA GPUs (Ampere+Ada, no native FP4): NVFP4 quantized-MoE expert sharding, bit-exact vs single-GPU greedy. Port of issue #478 methodology.
Python
-
TensorFold
TensorFold PublicForked from ashhart/TensorFold
LLM Inference Engine for Metal, CUDA and Vulkan.
Python
Something went wrong, please refresh the page to try again.
If the problem persists, check the GitHub status page or contact support.
If the problem persists, check the GitHub status page or contact support.