Publish and pull precompiled tt-metal kernel caches over Hugging Face Hub for Tenstorrent accelerators
-
Updated
Oct 9, 2026 - Python
Publish and pull precompiled tt-metal kernel caches over Hugging Face Hub for Tenstorrent accelerators
A streamlined CLI tool for profiling Tenstorrent's TT-Metal tests and extracting device kernel performance metrics
Demo of doing division in Tenstorrent
ULP accuracy and kernel timing measurements for TT-Metal eltwise ops, benchmarked against fp64 torch goldens across architectures (Wormhole, Blackhole) and dtypes (bf16, fp32).
Compatibility guardrail: tt-metal / tt-installer build status across community Linux distributions
Lifting Wavelet Transform (LWT) library optimized for Tenstorrent accelerators using tt-metal.
Day-by-day performance tracking dashboard for Tenstorrent TT-Metal TTNN eltwise operations.
A visual, plain-English guide to Tenstorrent's Tensix processor architecture — from chip-level mesh to per-core compute pipeline.
A curated collection of academic research papers, peer-reviewed publications, preprints, and system studies on Tenstorrent hardware, the Tensix spatial architecture, and the TT-Metalium software stack.
Fused and block-float ML kernels for Tenstorrent Tensix cores in TT-Lang and TT-Metalium C++, measured on Tenstorrent's simulators, with CI on simulated chips, containers, a Kubernetes JobSet sweep and Ansible provisioning.
To associate your repository with the tt-metal topic, visit your repo's landing page and select "manage topics."