Switchless RoCE collectives, NCCL profiles and fabric tooling for DGX Spark inference recipes
-
Updated
Oct 7, 2026 - Python
Switchless RoCE collectives, NCCL profiles and fabric tooling for DGX Spark inference recipes
Practical AI homelab setup guides for GB10, Mac Studio Ultra, RoCE/RDMA, MikroTik switching, NCCL, and heterogeneous workload experiments.
Clustering NVIDIA GB10 workstations (DGX Spark and OEM boxes): two nodes, a three-node ring, and a switched 200G RoCE fabric for four or more. By Petronella Technology Group, Inc.
NCCL over a switchless 4-node DGX Spark / GB10 ring using both PCIe halves of every QSFP cable: ~193 Gb/s per cable instead of ~112, up to +34% vLLM prefill. One patch on top of switchless-nccl.
Deploy and manage OpenShift clusters on NVIDIA Air with Red Hat Assisted Installer.
Field notes, benchmarks, and turnkey scripts for running 284B LLM inference and multi-modal workflows across 2x NVIDIA DGX Spark (GB10) with 200GbE RoCE. 双机 NVIDIA DGX Spark (GB10) 284B 大模型分布式推理与多模态部署实战笔记、真机实测数据与避坑指南。
practical guide to multi-node NCCL over switched RoCE fabric on NVIDIA GB10 (DGX Spark class) — documenting the gaps in NVIDIA's official playbooks
GLM-5.3-Flash on 1x RTX 5090 + 2x DGX Spark: attention on the 5090, routed experts on the Sparks over MCDMA RoCE, on TensorFold v0.6.5 (+ 0.6.6's commits) with Mia's GLM work. 2.0 at tag v2.0, v1.0 (glm53f-afd) at tag v1.0. Deployment recipe, as-is.
Read-only health check for the ConnectX-7 cluster fabric on NVIDIA DGX Spark and GB10 workstations: link, cable, MTU, addressing, RoCE/NCCL, peers.
To associate your repository with the connectx-7 topic, visit your repo's landing page and select "manage topics."