Repository navigation
Implement v5e-8 XLA port of pinned ANVIL2 nanoGPT leaderboard model - #168
Merged
Merged
Conversation
…flight and weight export
… optimizer cadence
…imizer regression
charlesmartin14
marked this pull request as ready for review
October 6, 2026 21:31
…r with paired validation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Implements an isolated full-size ANVIL2 TPU port for the existing v5litepod-8. The 58 byte-verified CUDA reference files are unchanged.
Latest review found and fixed an actual launch blocker: CPU-to-XLA migration replaces Parameters and loses custom bank metadata. Optimizer layouts now derive from tensor shapes; the frozen layer-7 MLP is explicit. Full dense model migration plus optimizer and averaging-buffer allocation passed on XLA CPU.
Incremental WeightWatcher is now required and automatic: a separate CPU worker analyzes 50 active attention/MLP projections at initialization, every 100 updates and after final averaging. Immutable snapshots are paired with validation NLL, perplexity and token accuracy/error; raw alpha is never replaced by clipped alpha. Partial 131,072-token diagnostics cannot qualify the benchmark. Final acceptance still requires 1,194 updates, 328,663,040 train tokens and NLL <= 3.28 on all 10,485,760 validation tokens. Completion waits for spectral analysis; failures are visible and fail the run.
Also adds per-rank HBM telemetry, host row/cache/sparse-update timing, and an early disk-space gate. Full dense and n-gram weights remain exported for final analysis.
Validation: 41 local tests pass, one Gloo test skipped because the local sandbox blocks sockets; dedicated CI exercises Gloo. Real WeightWatcher analysis/queue/CSV tests pass. Full dense CPU forward/backward and evaluation pass, as do XLA optimizer/untie/tail and migration tests. The same six ANVIL polynomial maps use FP32/highest matmul precision because BF16 recurrence was unstable in XLA testing; all numerical differences are documented.
Still unqualified on real TPU: full-size graph compilation, HBM peak, throughput and convergence require preflight and a complete run on the VM. No guarantee of 3.28, bitwise equivalence or H100 timing is made.