Skip to content
XtrComSuPublic

About

OpenCL compatibility and acceleration research for Exynos 9611 / Mali-G72, Android, Termux, and ggml.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

1 Commit

Folders and files

Repository files navigation

XT-AX9

XT-Artificial-Xeno-9611 — an OpenCL compatibility and acceleration layer for Exynos 9611 / Mali-G72 devices running Android and Termux, targeting the llama.cpp/ggml OpenCL backend.

XT-AX9 is experimental. It does not guarantee acceleration. It is layered so each stage can be verified separately:

Stage Status
Loader compatibility (vendor library reachable across the linker namespace) working, verified
OpenCL device discovery (platform and GPU enumeration) working, verified
Kernel execution (a kernel is compiled, run and read back correctly) working, verified
ggml backend support (llama.cpp accepting this GPU) not done, blocked upstream
llama.cpp inference performance not achieved; measured ceiling documented

1. Project overview

llama.cpp's OpenCL backend cannot use this GPU, for two independent reasons, and XT-AX9 addresses the first and documents the second.

The first is an Android loader problem. The vendor OpenCL implementation is loadable only from a vendor/SP-HAL namespace that an ordinary application cannot enter, so the backend fails to load and llama.cpp silently reports no devices. XT-AX9 solves this with a small shared library that presents an ordinary libOpenCL.so to the consumer and loads the real driver through Android's namespace-aware loader. That layer is complete and validated.

The second is a ggml limitation. ggml's OpenCL backend supports only Adreno and Intel GPUs by name, and its fast matmul kernels depend on an OpenCL subgroup API that this Mali driver does not implement. Supporting Mali means porting ggml; XT-AX9 records what that requires and has not done it.

Where the CPU backend reaches 1.63 tok/s on the reference model, XT-AX9's measurements show that this device cannot reach 20–30 tok/s with any kernel work, because DRAM bandwidth caps it near 10.8 tok/s. That analysis is in docs/benchmarks.md.

2. Architecture

llama-cli
  |
  v
libggml-opencl.so                 application linker namespace
  |  DT_NEEDED: libOpenCL.so
  v
libXTAX9OpenCL.so                 SONAME: libOpenCL.so, a forwarder only
  |
  |  android_dlopen_ext(), namespace "sphal", flag 0x200
  v
Android vendor / SP-HAL namespace
  |
  v
/vendor/lib64/egl/libGLES_mali.so  Samsung/ARM, proprietary, never modified
  |
  v
Mali-G72 r0p1, OpenCL 3.0

3. Supported hardware

Component Value
SoC Samsung Exynos 9611
GPU ARM Mali-G72 r0p1 (MP2)
Driver v1.r38p1-01bet0-mbs2v41_0, reports OpenCL 3.0
ABI arm64-v8a
OS Android 16 (SDK 36)
Host Termux
Vendor library /vendor/lib64/egl/libGLES_mali.so

The loader work is generic to any device where a vendor OpenCL implementation lives behind a namespace boundary. The measurements in docs/benchmarks.md are specific to this GPU and driver.

4. The Android linker namespace problem

/vendor/lib64 is not in the application namespace's permitted paths. The vendor OpenCL library needs libion_exynos.so, which exists only there:

libion_exynos.so    /system/lib64: no    /vendor/lib64: yes

So the obvious workaround — copy libGLES_mali.so next to the consumer — cannot work. The copy cannot resolve its own dependencies, dlopen("libggml-opencl.so") fails inside ggml, and the failure happens before device enumeration, which is why llama.cpp reports an empty device list with no explanation.

LD_LIBRARY_PATH cannot fix this: it adds search paths, it cannot extend a namespace's permitted-path set. Neither can LD_PRELOAD of the vendor library, nor su.

5. XT-AX9 solution

A shared library with SONAME = libOpenCL.so that:

  1. resolves android_dlopen_ext and android_get_exported_namespace from wherever they actually live;
  2. obtains the sphal namespace handle;
  3. loads /vendor/lib64/egl/libGLES_mali.so inside that namespace with android_dlopen_ext;
  4. resolves the OpenCL entry points with dlsym;
  5. forwards every call unchanged.

Nothing under /vendor, /system, /system_ext, /product or /odm is modified, and no linker configuration or SELinux policy is touched.

6. OpenCL loading mechanism

Two details are not documented in any public header and were determined by measuring the loader on this device. Full derivation in docs/linker-namespace.md.

android_get_exported_namespace is not exported by libdl.so, which is why linking against it fails with undefined symbol. It is found in libdl_android.so, which is not a DT_NEEDED of libc and must be dlopened explicitly, or in ld-android.so as __loader_android_get_exported_namespace.

ANDROID_DLEXT_USE_NAMESPACE (0x1) does not select a namespace on Android 16. It is accepted and silently ignored, and the loader then reports that the library is inaccessible to the caller's default namespace. The working combination is flag 0x200 with the namespace pointer at offset 40 of struct android_dlextinfo, not offset 8:

+0   uint32_t flags
+8   void*                ext_namespace        legacy union slot
+16  const char*          ext_relpath
+24  uint64_t             ext_flags
+32  void*                reserved_addr
+40  android_namespace_t* library_namespace    the field that works

XT-AX9 tries 0x200/offset 40 first and keeps 0x1/offset 8 as a fallback for older Android releases. examples/xt_ax9_dlext_probe.c re-derives all of this on any device.

7. ggml integration

The shim's exported symbols are derived from the consumer's actual imports rather than from the OpenCL specification:

llvm-readelf -W --dyn-syms $PREFIX/lib/libggml-opencl.so \
    | awk 'NR>3 && $7=="UND" && $8 ~ /^cl/'

For llama.cpp b9590 that is 29 symbols, with version requirements OPENCL_1.0 (23), OPENCL_1.1 (1), OPENCL_1.2 (4) and OPENCL_3.0 (1). All 29 are exported at exactly those versions. Five further symbols are exported as a documented superset: the two clGetExtensionFunctionAddress* entry points, clGetEventProfilingInfo (used by the benchmark to separate kernel time from dispatch cost), and clReleaseContext / clReleaseCommandQueue / clReleaseKernel.

With the shim in place, ggml loads the platform, enumerates the device and receives its name correctly:

ggml_opencl: selected platform: 'ARM Platform'
ggml_opencl: device: 'Mali-G72 r0p1 (OpenCL 3.0 v1.r38p1-01bet0-mbs2v41_0...)'
ggml_opencl: unsupported GPU 'Mali-G72 r0p1'.
ggml_opencl: drop unsupported device 'Mali-G72 r0p1'.

The last two lines are ggml's own GPU-family allowlist, not an XT-AX9 failure.

8. Mali-G72 limitations

The driver reports OpenCL 3.0 but implements no subgroup API. It does not advertise cl_khr_subgroups, and the OpenCL 2.0/3.0 core subgroup builtins are undeclared by its compiler under every -cl-std value tested, including -cl-std=CL3.0. This is not a build-flag problem.

Consequences for ggml: 84 of its 138 kernels cannot compile on this device, including every _1d_* and _8x_flat matmul variant, and ggml's own gate rejects the device for lacking cl_khr_subgroups or cl_intel_subgroups.

XT-AX9 therefore requires capability-driven, subgroup-free execution. It does not fake a subgroup size, a wave size, or a device identity. xt-ax9-probe performs the check at run time.

Two further driver constraints, both invisible in CL_DEVICE_MAX_WORK_GROUP_SIZE and both affecting every kernel launched:

  • Only power-of-two work-group sizes are accepted. 96, 192 and 384 fail with CL_INVALID_WORK_GROUP_SIZE (-54), although 384 is reported as the maximum.
  • A kernel using local memory and barriers requires global_work_size >= local_work_size; fewer work items than the local size is rejected with the same error.

9. Current performance

Reference model: MiniCPM5 1B Q4_K_M, 651.29 MiB, 1.08 B parameters, on Exynos 9611 / Mali-G72. Decoding streams every weight once per token, about 544 MB, of which the 1536x130560 output projection is 107.6 MiB.

Measurement Value
CPU prompt processing, pp512 16.38 ± 2.21 tok/s
CPU decode, tg128 1.63 ± 0.19 tok/s (614 ms/token)
Streaming read bandwidth 5.9 GB/s (flat 64–512 MiB)
Decode-shaped mat-vec, best measured 0.45 GB/s (8 % of peak)
Same traffic with dequantisation removed 3.4 GB/s (58 % of peak)
llama.cpp inference on the GPU not measured; backend not supported

The last row is not a rounding of the others: the ggml backend does not run.

The dequantisation control kernel isolates the bottleneck. The access pattern sustains 58 % of peak, so mat-vec is bound by dequantisation instruction issue — roughly six instructions per weight — not by memory.

The 20–30 tok/s objective is not reachable on this hardware. Reaching 20 tok/s needs 10.9 GB/s of weight streaming per token and 30 tok/s needs 16.3 GB/s; the device measures 5.9 GB/s. Even a hypothetical zero-cost dequantisation path caps at 10.8 tok/s, and 30 tok/s would exceed the memory bus's theoretical bandwidth. Methodology and full tables: docs/benchmarks.md.

10. Building

git clone https://github.com/XtrComSu/XT-AX9.git
cd XT-AX9
scripts/build.sh

Produces in build/:

Artifact Purpose
libXTAX9OpenCL.so the shim, SONAME libOpenCL.so
xt-ax9-test loader validation
xt-ax9-probe device capability probe
xt-ax9-bench kernel microbenchmarks
xt-ax9-bench-wg work-group launch constraints
xt-ax9-dlext-probe android_dlextinfo layout discovery

scripts/build.sh clean removes build/.

Requirements: Termux, clang, python3, and the OpenCL headers from $PREFIX/include/CL. No CMake, Ninja or NDK is needed.

Two Termux-specific notes, both handled by the build script:

  • clang runs its frontend through su. With TERMUX_EXEC__PROC_SELF_EXE set, which termux-exec installs, the driver misfires with Unrecognized option: 't' and clang frontend command failed. The script unsets it for every invocation. The build consequently runs elevated, and the script restores ownership of build/.
  • Termux readelf produces empty output for large binaries. Use llvm-readelf for ELF inspection.

11. Installation

Two options. Neither modifies the vendor partition.

Preload, writing nothing:

LD_PRELOAD=$PWD/build/libXTAX9OpenCL.so llama-cli --list-devices

The shim's SONAME is libOpenCL.so, so it satisfies the consumer's DT_NEEDED entry from the global scope.

Install as a dependency:

scripts/install.sh install     # copies to $PREFIX/lib/libOpenCL.so
scripts/install.sh remove      # restores the previous file

install backs up any existing $PREFIX/lib/libOpenCL.so to libOpenCL.so.xt-ax9-backup before replacing it.

12. Testing

XT_AX9_LOG=quiet ./build/xt-ax9-test

Prints which libOpenCL.so the process bound to, the shim's own status, the platform and device inventory, and then compiles and runs a kernel and reads the result back. Exit status is 0 only if a platform, a GPU device and a correct round trip were all obtained, so it works as a build gate. The round trip exercises context, queue, buffer, program, build, kernel, argument marshalling, NDRange submission, blocking readback and release.

With the shim installed:

llama-cli --list-devices

Expected on the current state: llama.cpp enumerates the Mali-G72 through XT-AX9 and then drops it at its own GPU-family check.

13. Benchmark methodology

Full detail in docs/benchmarks.md. In summary:

  • Program build time is printed separately and never counted as throughput.
  • Kernel-only time comes from OpenCL profiling events, with a wall-clock fallback when event timestamps are not orderable.
  • Every configuration is warmed up once and that run discarded.
  • Every kernel is validated against a host reference over the same input; the reported figure is the maximum relative error. Weight scales are generated as valid fp16 across normal, subnormal, exact-power, zero and negative values.
  • Bandwidth is reported at several buffer sizes so a flat curve can be distinguished from a cache artefact.
  • CPU and GPU figures use the same model, context size and thread count.

14. Known limitations

  • ggml does not support this GPU. Its family allowlist is Adreno and Intel only. Porting it is described, not implemented, in docs/ggml-integration.md.
  • No subgroup support in the driver, so ggml's subgroup-dependent kernels cannot compile here at all.
  • Bandwidth caps decode near 10.8 tok/s even with a perfect dequantisation path, against a 1.63 tok/s CPU baseline.
  • The shim covers one consumer. It exports the 29 OpenCL symbols ggml b9590 imports plus 5 more. A different ggml build may import a different set; the exported list is derived from the binary rather than assumed, so it must be re-derived for another build.
  • android_dlopen_ext usage is version-specific. The flag and struct offset are measured, not documented. Re-derive them on a different Android release with xt-ax9-dlext-probe before trusting the shim there.
  • Prompt-processing GEMM performance is not yet benchmarked, so the comparison between the generic GEMM path and prompt processing is open.

15. Safety and rollback

Nothing outside the XT-AX9 checkout and one file in the Termux prefix is modified. Specifically untouched: /vendor, /system, /system_ext, /product, /odm, /linkerconfig/ld.config.txt, the SELinux policy and the GPU driver.

scripts/install.sh remove          # restore the previous libOpenCL.so
rm -rf ~/XT-AX9                     # remove the checkout
LD_PRELOAD= llama-cli --list-devices # confirm the previous behaviour

install.sh remove restores $PREFIX/lib/libOpenCL.so from libOpenCL.so.xt-ax9-backup if one exists. On this device the file that was there before any XT-AX9 work was a byte copy of libGLES_mali.so, which does not work for the reason given in section 4.

No privilege escalation is required to use the shim through LD_PRELOAD. The installation path writes only inside the user's own Termux prefix.

16. Development

src/       the shim and its symbol version map
include/   the loader ABI declarations, with the measured layout
kernels/   OpenCL C used by the benchmark and probes, embedded at build time
tests/     loader validation and device capability probe
bench/     kernel microbenchmarks and driver constraint probes
scripts/   build and install
docs/      loader derivation, benchmarks, ggml integration analysis
examples/  the android_dlextinfo layout discovery probe

Build flags live in scripts/build.sh. kernels/embed_kernel.py turns each .cl file into a C string constant so kernels stay reviewable as OpenCL C rather than being buried in C; add a kernel there and to the list in the build script.

Renaming note: this project was developed under the working name XT-Amelia. The shim was previously built as libXTAmeliaOpenCL.so and, before this release, a copy of it may be installed at $PREFIX/lib/libOpenCL.so. Remove it with scripts/install.sh remove or by deleting that file; the current artifact is libXTAX9OpenCL.so.

17. Future work

In the order the measurements suggest:

  1. Vectorise the q4_0 dequantisation through Mali's SIMD unit using OpenCL vector types (uchar16, short8, half8). This needs no subgroups, so it is available on this driver, and it targets the measured bottleneck.
  2. Add GPU_FAMILY::GENERIC to ggml and route it to the existing _l4_lm kernels, skipping the 84 programs that cannot compile.
  3. Specialise the 1536x130560 output projection, which is 20 % of per-token traffic and is a pure bandwidth problem at N = 1.
  4. Measure prompt-processing throughput, which is still unmeasured.
  5. Reduce CPU-GPU synchronisation in the decode graph once kernels exist.

18. License

XT-AX9 is licensed under the GNU Affero General Public License, version 3 or later (AGPL-3.0-or-later). The complete licence text is in LICENSE and COPYING.

Every original XT-AX9 source file carries an SPDX identifier:

SPDX-License-Identifier: AGPL-3.0-or-later

If you distribute a modified version of XT-AX9, you must comply with the AGPL-3.0-or-later requirements. In particular, section 13 requires that users interacting with a modified version over a network be offered the Corresponding Source of your modified version, and sections 4 and 5 require that the Corresponding Source be conveyed with the work under the same or a compatible licence. Conveying a modified binary without its Corresponding Source, or offering network interaction without offering the source, does not satisfy the licence.

No exceptions weaken these terms. There is no proprietary-licence carve-out and no "source available" alternative.

Third-party components are not relicensed. llama.cpp/ggml is MIT-licensed and is an external dependency, not vendored here; the Khronos OpenCL headers are MIT and used from the toolchain; the Android bionic loader interface is Apache-2.0 and only an ABI is declared; and Samsung's proprietary Mali driver is never redistributed. See THIRD-PARTY-LICENSES.md for the full inventory with versions and locations.

About

OpenCL compatibility and acceleration research for Exynos 9611 / Mali-G72, Android, Termux, and ggml.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages