XT-Artificial-Xeno-9611 — an OpenCL compatibility and acceleration layer for Exynos 9611 / Mali-G72 devices running Android and Termux, targeting the llama.cpp/ggml OpenCL backend.
XT-AX9 is experimental. It does not guarantee acceleration. It is layered so each stage can be verified separately:
| Stage | Status |
|---|---|
| Loader compatibility (vendor library reachable across the linker namespace) | working, verified |
| OpenCL device discovery (platform and GPU enumeration) | working, verified |
| Kernel execution (a kernel is compiled, run and read back correctly) | working, verified |
| ggml backend support (llama.cpp accepting this GPU) | not done, blocked upstream |
| llama.cpp inference performance | not achieved; measured ceiling documented |
llama.cpp's OpenCL backend cannot use this GPU, for two independent reasons, and XT-AX9 addresses the first and documents the second.
The first is an Android loader problem. The vendor OpenCL implementation is
loadable only from a vendor/SP-HAL namespace that an ordinary application cannot
enter, so the backend fails to load and llama.cpp silently reports no devices.
XT-AX9 solves this with a small shared library that presents an ordinary
libOpenCL.so to the consumer and loads the real driver through Android's
namespace-aware loader. That layer is complete and validated.
The second is a ggml limitation. ggml's OpenCL backend supports only Adreno and Intel GPUs by name, and its fast matmul kernels depend on an OpenCL subgroup API that this Mali driver does not implement. Supporting Mali means porting ggml; XT-AX9 records what that requires and has not done it.
Where the CPU backend reaches 1.63 tok/s on the reference model, XT-AX9's measurements show that this device cannot reach 20–30 tok/s with any kernel work, because DRAM bandwidth caps it near 10.8 tok/s. That analysis is in docs/benchmarks.md.
llama-cli
|
v
libggml-opencl.so application linker namespace
| DT_NEEDED: libOpenCL.so
v
libXTAX9OpenCL.so SONAME: libOpenCL.so, a forwarder only
|
| android_dlopen_ext(), namespace "sphal", flag 0x200
v
Android vendor / SP-HAL namespace
|
v
/vendor/lib64/egl/libGLES_mali.so Samsung/ARM, proprietary, never modified
|
v
Mali-G72 r0p1, OpenCL 3.0
| Component | Value |
|---|---|
| SoC | Samsung Exynos 9611 |
| GPU | ARM Mali-G72 r0p1 (MP2) |
| Driver | v1.r38p1-01bet0-mbs2v41_0, reports OpenCL 3.0 |
| ABI | arm64-v8a |
| OS | Android 16 (SDK 36) |
| Host | Termux |
| Vendor library | /vendor/lib64/egl/libGLES_mali.so |
The loader work is generic to any device where a vendor OpenCL implementation lives behind a namespace boundary. The measurements in docs/benchmarks.md are specific to this GPU and driver.
/vendor/lib64 is not in the application namespace's permitted paths. The
vendor OpenCL library needs libion_exynos.so, which exists only there:
libion_exynos.so /system/lib64: no /vendor/lib64: yes
So the obvious workaround — copy libGLES_mali.so next to the consumer — cannot
work. The copy cannot resolve its own dependencies, dlopen("libggml-opencl.so")
fails inside ggml, and the failure happens before device enumeration, which is
why llama.cpp reports an empty device list with no explanation.
LD_LIBRARY_PATH cannot fix this: it adds search paths, it cannot extend a
namespace's permitted-path set. Neither can LD_PRELOAD of the vendor library,
nor su.
A shared library with SONAME = libOpenCL.so that:
- resolves
android_dlopen_extandandroid_get_exported_namespacefrom wherever they actually live; - obtains the
sphalnamespace handle; - loads
/vendor/lib64/egl/libGLES_mali.soinside that namespace withandroid_dlopen_ext; - resolves the OpenCL entry points with
dlsym; - forwards every call unchanged.
Nothing under /vendor, /system, /system_ext, /product or /odm is
modified, and no linker configuration or SELinux policy is touched.
Two details are not documented in any public header and were determined by measuring the loader on this device. Full derivation in docs/linker-namespace.md.
android_get_exported_namespace is not exported by libdl.so, which is why
linking against it fails with undefined symbol. It is found in
libdl_android.so, which is not a DT_NEEDED of libc and must be dlopened
explicitly, or in ld-android.so as __loader_android_get_exported_namespace.
ANDROID_DLEXT_USE_NAMESPACE (0x1) does not select a namespace on Android
16. It is accepted and silently ignored, and the loader then reports that the
library is inaccessible to the caller's default namespace. The working
combination is flag 0x200 with the namespace pointer at offset 40 of
struct android_dlextinfo, not offset 8:
+0 uint32_t flags
+8 void* ext_namespace legacy union slot
+16 const char* ext_relpath
+24 uint64_t ext_flags
+32 void* reserved_addr
+40 android_namespace_t* library_namespace the field that works
XT-AX9 tries 0x200/offset 40 first and keeps 0x1/offset 8 as a fallback for
older Android releases. examples/xt_ax9_dlext_probe.c re-derives all of this
on any device.
The shim's exported symbols are derived from the consumer's actual imports rather than from the OpenCL specification:
llvm-readelf -W --dyn-syms $PREFIX/lib/libggml-opencl.so \
| awk 'NR>3 && $7=="UND" && $8 ~ /^cl/'For llama.cpp b9590 that is 29 symbols, with version requirements
OPENCL_1.0 (23), OPENCL_1.1 (1), OPENCL_1.2 (4) and OPENCL_3.0 (1). All
29 are exported at exactly those versions. Five further symbols are exported as
a documented superset: the two clGetExtensionFunctionAddress* entry points,
clGetEventProfilingInfo (used by the benchmark to separate kernel time from
dispatch cost), and clReleaseContext / clReleaseCommandQueue /
clReleaseKernel.
With the shim in place, ggml loads the platform, enumerates the device and receives its name correctly:
ggml_opencl: selected platform: 'ARM Platform'
ggml_opencl: device: 'Mali-G72 r0p1 (OpenCL 3.0 v1.r38p1-01bet0-mbs2v41_0...)'
ggml_opencl: unsupported GPU 'Mali-G72 r0p1'.
ggml_opencl: drop unsupported device 'Mali-G72 r0p1'.
The last two lines are ggml's own GPU-family allowlist, not an XT-AX9 failure.
The driver reports OpenCL 3.0 but implements no subgroup API. It does not
advertise cl_khr_subgroups, and the OpenCL 2.0/3.0 core subgroup builtins are
undeclared by its compiler under every -cl-std value tested, including
-cl-std=CL3.0. This is not a build-flag problem.
Consequences for ggml: 84 of its 138 kernels cannot compile on this device,
including every _1d_* and _8x_flat matmul variant, and ggml's own gate
rejects the device for lacking cl_khr_subgroups or cl_intel_subgroups.
XT-AX9 therefore requires capability-driven, subgroup-free execution. It does
not fake a subgroup size, a wave size, or a device identity. xt-ax9-probe
performs the check at run time.
Two further driver constraints, both invisible in
CL_DEVICE_MAX_WORK_GROUP_SIZE and both affecting every kernel launched:
- Only power-of-two work-group sizes are accepted. 96, 192 and 384 fail with
CL_INVALID_WORK_GROUP_SIZE(-54), although 384 is reported as the maximum. - A kernel using local memory and barriers requires
global_work_size >= local_work_size; fewer work items than the local size is rejected with the same error.
Reference model: MiniCPM5 1B Q4_K_M, 651.29 MiB, 1.08 B parameters, on Exynos 9611 / Mali-G72. Decoding streams every weight once per token, about 544 MB, of which the 1536x130560 output projection is 107.6 MiB.
| Measurement | Value |
|---|---|
| CPU prompt processing, pp512 | 16.38 ± 2.21 tok/s |
| CPU decode, tg128 | 1.63 ± 0.19 tok/s (614 ms/token) |
| Streaming read bandwidth | 5.9 GB/s (flat 64–512 MiB) |
| Decode-shaped mat-vec, best measured | 0.45 GB/s (8 % of peak) |
| Same traffic with dequantisation removed | 3.4 GB/s (58 % of peak) |
| llama.cpp inference on the GPU | not measured; backend not supported |
The last row is not a rounding of the others: the ggml backend does not run.
The dequantisation control kernel isolates the bottleneck. The access pattern sustains 58 % of peak, so mat-vec is bound by dequantisation instruction issue — roughly six instructions per weight — not by memory.
The 20–30 tok/s objective is not reachable on this hardware. Reaching 20 tok/s needs 10.9 GB/s of weight streaming per token and 30 tok/s needs 16.3 GB/s; the device measures 5.9 GB/s. Even a hypothetical zero-cost dequantisation path caps at 10.8 tok/s, and 30 tok/s would exceed the memory bus's theoretical bandwidth. Methodology and full tables: docs/benchmarks.md.
git clone https://github.com/XtrComSu/XT-AX9.git
cd XT-AX9
scripts/build.shProduces in build/:
| Artifact | Purpose |
|---|---|
libXTAX9OpenCL.so |
the shim, SONAME libOpenCL.so |
xt-ax9-test |
loader validation |
xt-ax9-probe |
device capability probe |
xt-ax9-bench |
kernel microbenchmarks |
xt-ax9-bench-wg |
work-group launch constraints |
xt-ax9-dlext-probe |
android_dlextinfo layout discovery |
scripts/build.sh clean removes build/.
Requirements: Termux, clang, python3, and the OpenCL headers from
$PREFIX/include/CL. No CMake, Ninja or NDK is needed.
Two Termux-specific notes, both handled by the build script:
clangruns its frontend throughsu. WithTERMUX_EXEC__PROC_SELF_EXEset, whichtermux-execinstalls, the driver misfires withUnrecognized option: 't'andclang frontend command failed. The script unsets it for every invocation. The build consequently runs elevated, and the script restores ownership ofbuild/.- Termux
readelfproduces empty output for large binaries. Usellvm-readelffor ELF inspection.
Two options. Neither modifies the vendor partition.
Preload, writing nothing:
LD_PRELOAD=$PWD/build/libXTAX9OpenCL.so llama-cli --list-devicesThe shim's SONAME is libOpenCL.so, so it satisfies the consumer's
DT_NEEDED entry from the global scope.
Install as a dependency:
scripts/install.sh install # copies to $PREFIX/lib/libOpenCL.so
scripts/install.sh remove # restores the previous fileinstall backs up any existing $PREFIX/lib/libOpenCL.so to
libOpenCL.so.xt-ax9-backup before replacing it.
XT_AX9_LOG=quiet ./build/xt-ax9-testPrints which libOpenCL.so the process bound to, the shim's own status, the
platform and device inventory, and then compiles and runs a kernel and reads
the result back. Exit status is 0 only if a platform, a GPU device and a correct
round trip were all obtained, so it works as a build gate. The round trip
exercises context, queue, buffer, program, build, kernel, argument marshalling,
NDRange submission, blocking readback and release.
With the shim installed:
llama-cli --list-devicesExpected on the current state: llama.cpp enumerates the Mali-G72 through XT-AX9 and then drops it at its own GPU-family check.
Full detail in docs/benchmarks.md. In summary:
- Program build time is printed separately and never counted as throughput.
- Kernel-only time comes from OpenCL profiling events, with a wall-clock fallback when event timestamps are not orderable.
- Every configuration is warmed up once and that run discarded.
- Every kernel is validated against a host reference over the same input; the reported figure is the maximum relative error. Weight scales are generated as valid fp16 across normal, subnormal, exact-power, zero and negative values.
- Bandwidth is reported at several buffer sizes so a flat curve can be distinguished from a cache artefact.
- CPU and GPU figures use the same model, context size and thread count.
- ggml does not support this GPU. Its family allowlist is Adreno and Intel only. Porting it is described, not implemented, in docs/ggml-integration.md.
- No subgroup support in the driver, so ggml's subgroup-dependent kernels cannot compile here at all.
- Bandwidth caps decode near 10.8 tok/s even with a perfect dequantisation path, against a 1.63 tok/s CPU baseline.
- The shim covers one consumer. It exports the 29 OpenCL symbols ggml
b9590imports plus 5 more. A different ggml build may import a different set; the exported list is derived from the binary rather than assumed, so it must be re-derived for another build. android_dlopen_extusage is version-specific. The flag and struct offset are measured, not documented. Re-derive them on a different Android release withxt-ax9-dlext-probebefore trusting the shim there.- Prompt-processing GEMM performance is not yet benchmarked, so the comparison between the generic GEMM path and prompt processing is open.
Nothing outside the XT-AX9 checkout and one file in the Termux prefix is
modified. Specifically untouched: /vendor, /system, /system_ext,
/product, /odm, /linkerconfig/ld.config.txt, the SELinux policy and the
GPU driver.
scripts/install.sh remove # restore the previous libOpenCL.so
rm -rf ~/XT-AX9 # remove the checkout
LD_PRELOAD= llama-cli --list-devices # confirm the previous behaviourinstall.sh remove restores $PREFIX/lib/libOpenCL.so from
libOpenCL.so.xt-ax9-backup if one exists. On this device the file that was
there before any XT-AX9 work was a byte copy of libGLES_mali.so, which does
not work for the reason given in section 4.
No privilege escalation is required to use the shim through LD_PRELOAD. The
installation path writes only inside the user's own Termux prefix.
src/ the shim and its symbol version map
include/ the loader ABI declarations, with the measured layout
kernels/ OpenCL C used by the benchmark and probes, embedded at build time
tests/ loader validation and device capability probe
bench/ kernel microbenchmarks and driver constraint probes
scripts/ build and install
docs/ loader derivation, benchmarks, ggml integration analysis
examples/ the android_dlextinfo layout discovery probe
Build flags live in scripts/build.sh. kernels/embed_kernel.py turns each
.cl file into a C string constant so kernels stay reviewable as OpenCL C rather
than being buried in C; add a kernel there and to the list in the build script.
Renaming note: this project was developed under the working name XT-Amelia.
The shim was previously built as libXTAmeliaOpenCL.so and, before this
release, a copy of it may be installed at $PREFIX/lib/libOpenCL.so. Remove it
with scripts/install.sh remove or by deleting that file; the current artifact
is libXTAX9OpenCL.so.
In the order the measurements suggest:
- Vectorise the q4_0 dequantisation through Mali's SIMD unit using OpenCL
vector types (
uchar16,short8,half8). This needs no subgroups, so it is available on this driver, and it targets the measured bottleneck. - Add
GPU_FAMILY::GENERICto ggml and route it to the existing_l4_lmkernels, skipping the 84 programs that cannot compile. - Specialise the 1536x130560 output projection, which is 20 % of per-token traffic and is a pure bandwidth problem at N = 1.
- Measure prompt-processing throughput, which is still unmeasured.
- Reduce CPU-GPU synchronisation in the decode graph once kernels exist.
XT-AX9 is licensed under the GNU Affero General Public License, version 3 or
later (AGPL-3.0-or-later). The complete licence text is in
LICENSE and COPYING.
Every original XT-AX9 source file carries an SPDX identifier:
SPDX-License-Identifier: AGPL-3.0-or-later
If you distribute a modified version of XT-AX9, you must comply with the AGPL-3.0-or-later requirements. In particular, section 13 requires that users interacting with a modified version over a network be offered the Corresponding Source of your modified version, and sections 4 and 5 require that the Corresponding Source be conveyed with the work under the same or a compatible licence. Conveying a modified binary without its Corresponding Source, or offering network interaction without offering the source, does not satisfy the licence.
No exceptions weaken these terms. There is no proprietary-licence carve-out and no "source available" alternative.
Third-party components are not relicensed. llama.cpp/ggml is MIT-licensed and is an external dependency, not vendored here; the Khronos OpenCL headers are MIT and used from the toolchain; the Android bionic loader interface is Apache-2.0 and only an ABI is declared; and Samsung's proprietary Mali driver is never redistributed. See THIRD-PARTY-LICENSES.md for the full inventory with versions and locations.