Skip to content

PLAN: Vulkan through the apps on three GPUs, and the 16k MediaTek chunk rule - #7

Merged
alpharomercoma merged 1 commit into
mainfrom
docs-vulkan-on-device
Oct 4, 2026
Merged

alpharomercoma merged 1 commit into
mainfrom
docs-vulkan-on-device

Conversation

@alpharomercoma

@alpharomercoma alpharomercoma commented Oct 4, 2026 •

Copy link
Copy Markdown
Member

Documentation only: no pipeline code changes.

  • Out-of-date statements fixed: docs/PLAN.md said the apps can't load Vulkan files and pinned ExecuTorch 1.4.0. Both changed on 2026-10-04; the Vulkan row in the phase table now reads done.
  • Finding 38 (Vulkan on the phone): Qwen3-0.6B 2k, run through openweights and ExecuServe on a Mali-G925, an Adreno 750 and the SM8850's Adreno.
    • Correct on every GPU.
    • On Mali the CPU file decodes about 3× faster.
    • On the SM8850 the GPU reads long prompts 1.4–1.9× as fast, but decodes short replies at 0.7× the CPU's speed.
    • The study and its raw logs are in openweights docs/research/vulkan-on-device.md and tools/eval/results/vulkan-2026-10-04/.
  • Finding 39 (MediaTek 16k): LFM2.5 at 16k built only once each chunk held at most one attention layer: 8 chunks for the 1.2B, 10 for the 2.6B. The 32k runs were cancelled for budget, so it's not known whether they build.
    • The max_chunks input's description in export-mtk.yml now says this.
  • CLAUDE.md: the local test pins follow the 1.5.1 bump.

ruff check and pytest pass (205 passed, 4 skipped).

🤖 Generated with Claude Code


View with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is enabled.

…nk rule, as findings 38 and 39

PLAN still said the apps could not load Vulkan files and pinned ExecuTorch 1.4.0; both changed
on 2026-10-04. Finding 38 records what the apps measured on a Mali-G925, an Adreno 750 and the
SM8850's Adreno (correct everywhere; the CPU file decodes three times faster on Mali, the GPU
reads long prompts faster on the SM8850), with the study and raw logs in openweights. Finding
39 records why LFM2.5 at 16k built only once each chunk held one attention layer (8 chunks
for the 1.2B, 10 for the 2.6B), which the max_chunks input's description now says, and that
32k was cancelled rather than shown to fail. CLAUDE.md's local test pins follow the 1.5.1 bump.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@alpharomercoma
alpharomercoma merged commit 7070ac5 into main Oct 4, 2026
2 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant