Skip to content

Implement v5e-8 XLA port of pinned ANVIL2 nanoGPT leaderboard model - #168

Merged
charlesmartin14 merged 6 commits into
mainfrom
codex/nanogpt-leaderboard-tpu-port
Oct 6, 2026
Merged

charlesmartin14 merged 6 commits into
mainfrom
codex/nanogpt-leaderboard-tpu-port

Conversation

@charlesmartin14

@charlesmartin14 charlesmartin14 commented Oct 6, 2026 •

Copy link
Copy Markdown
Member

Implements an isolated full-size ANVIL2 TPU port for the existing v5litepod-8. The 58 byte-verified CUDA reference files are unchanged.

Latest review found and fixed an actual launch blocker: CPU-to-XLA migration replaces Parameters and loses custom bank metadata. Optimizer layouts now derive from tensor shapes; the frozen layer-7 MLP is explicit. Full dense model migration plus optimizer and averaging-buffer allocation passed on XLA CPU.

Incremental WeightWatcher is now required and automatic: a separate CPU worker analyzes 50 active attention/MLP projections at initialization, every 100 updates and after final averaging. Immutable snapshots are paired with validation NLL, perplexity and token accuracy/error; raw alpha is never replaced by clipped alpha. Partial 131,072-token diagnostics cannot qualify the benchmark. Final acceptance still requires 1,194 updates, 328,663,040 train tokens and NLL <= 3.28 on all 10,485,760 validation tokens. Completion waits for spectral analysis; failures are visible and fail the run.

Also adds per-rank HBM telemetry, host row/cache/sparse-update timing, and an early disk-space gate. Full dense and n-gram weights remain exported for final analysis.

Validation: 41 local tests pass, one Gloo test skipped because the local sandbox blocks sockets; dedicated CI exercises Gloo. Real WeightWatcher analysis/queue/CSV tests pass. Full dense CPU forward/backward and evaluation pass, as do XLA optimizer/untie/tail and migration tests. The same six ANVIL polynomial maps use FP32/highest matmul precision because BF16 recurrence was unstable in XLA testing; all numerical differences are documented.

Still unqualified on real TPU: full-size graph compilation, HBM peak, throughput and convergence require preflight and a complete run on the VM. No guarantee of 3.28, bitwise equivalence or H100 timing is made.

@charlesmartin14 charlesmartin14 changed the title Draft: full leaderboard TPU capacity audit and port requirements Draft: target existing v5e-8 for full leaderboard TPU port Oct 6, 2026
@charlesmartin14 charlesmartin14 changed the title Draft: target existing v5e-8 for full leaderboard TPU port Implement v5e-8 XLA port of pinned ANVIL2 nanoGPT leaderboard model Oct 6, 2026
@charlesmartin14
charlesmartin14 marked this pull request as ready for review October 6, 2026 21:31
@charlesmartin14
charlesmartin14 merged commit e9d2569 into main Oct 6, 2026
4 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant