Termux llama.cpp source installers for Samsung S24 Ultra: CPU + Adreno OpenCL GPU, or CPU + Hexagon v75 NPU. Complete NPU setup: install-npu.sh and docs/NPU_INSTALL.md. Tested inference; performance and compatibility limits documented.
Python
0
3 commits
updated Oct 4, 2026
Source-built llama.cpp inference on the S24 Ultra: CPU, Adreno OpenCL GPU, and a separate Hexagon v75 NPU installer.
Builds a pinned llama.cpp revision with Android sphal vendor-runtime loading. install.sh builds CPU + GPU; install-npu.sh separately builds CPU + NPU. Tested on Samsung Galaxy S24 Ultra, Snapdragon 8 Gen 3 / Adreno 750, Android 16, F-Droid Termux 0.118.3, aarch64.
The current generic OpenCL configuration is slower than CPU on the tested phone. This project proves working GPU computation; it is not optimized production acceleration.
| Original matched test, 4 threads | CPU | OpenCL |
|---|---|---|
| Prompt processing, 128 tokens | 122.74 tok/s | 82.45 tok/s |
| Generation, 64 tokens | 28.72 tok/s | 9.50 tok/s |
OpenCL was 32.8% slower for prompt processing and 66.9% slower for generation. Measurements and reproduction details: Benchmarks.
The full CPU/OpenCL/NPU/hybrid report and machine-readable evidence are now available. The investigation is paused at the owner's request; its final characterization is incomplete.
Experimental native Hexagon NPU tests show repeated prompt-processing gains of 7.14x for the tested 1.5B model and 7.83x for Qwen3-4B against matched common-setting CPU tests. NPU generation averaged 32.97 and 12.94 tok/s respectively; the best practical CPU comparison remains unfinished. Long-prompt GPU/hybrid results are candidates, and no sustained thermal winner is established. The GPU installer preserves the proven generic OpenCL baseline. The separate NPU installer packages the working native Hexagon implementation; experimental specialized GPU and handoff patches remain research snapshots. The historical report predates NPU packaging; see the current NPU guide.
For the separate CPU + Hexagon v75 NPU build, use:
pkg update
pkg install git
git clone https://github.com/Ishabdullah/OpenCL-S24-Ultra.git
cd OpenCL-S24-Ultra
./install-npu.sh --download-test-model
This installs missing Termux dependencies, prepares checksum-pinned SDK/compiler tools in project storage, patches pinned llama.cpp, source-builds native ARM64/Hexagon binaries, executes numerical DSP tests, and generates 32 tokens from the pinned official Qwen 1.5B GGUF. QEMU emulates compiler tools only; inference is native. No root or system/vendor changes are required. Allow at least 8 GiB free plus model space and tens of minutes for the build.
Use an existing model instead with ./install-npu.sh --model /path/to/model.gguf, then launch it through ./scripts/run-npu.sh /path/to/model.gguf. A model-free install verifies DSP matrices but does not test LLM generation. The optional model download is explicit; no model/SDK/compiler/vendor binary is hosted in this repository.
Start here: NPU installation, operation, licensing and troubleshooting. A fresh source/build reproduction passed eight numerical DSP tests and generated text with all 29 Qwen 1.5B layers offloaded: NPU validation. NPU inference is experimental: no universal performance/thermal benefit or arbitrary model/device compatibility is claimed. The GPU and NPU installers coexist in separate directories and do not yet provide one validated CPU/GPU/NPU executable.
libOpenCL.so, libOpenCL_adreno.so, libCB.so, libgsl.so in /vendor/lib64.Confirmed: the S24 Ultra configuration above. Detected but unverified: other phones whose vendor stack initializes and exposes an Adreno device. Potentially compatible: other Snapdragon/Adreno devices with the required namespace APIs and stack; file presence alone is not proof of compatibility.
In a fresh F-Droid Termux terminal:
pkg update
pkg install git
git clone https://github.com/Ishabdullah/OpenCL-S24-Ultra.git
cd OpenCL-S24-Ultra
./install.sh
The installer installs missing git clang cmake ninja python opencl-headers openssl packages. Clang pulls in LLVM (llvm-readelf); Python/Termux provide supporting runtime dependencies. Make, pkg-config and an OpenCL ICD loader are not required by this Ninja/source build.
It downloads upstream llama.cpp at e358d59178377be4c58ba567925e05faadbccb57, checks/applies the maintained patch, builds llama-completion, llama-bench, and test-backend-ops, verifies ELF dependencies, and detects the GPU. It does not modify an existing ~/llama.cpp, install llama libraries into $PREFIX, use root, or modify Android system/vendor files.
Generated source, binaries, cache and logs live under .work/ inside this repository. First kernel compilation and the source build take time; stages print the relevant log path. A successful installer run without a model proves device initialization, not model inference.
Rerunning ./install.sh checks an existing installation. Options:
./install.sh --help
./install.sh --verify
./install.sh --rebuild # configure and incrementally rebuild
./install.sh --clean-build # replace only the generated build directory
./install.sh --jobs 2 # lower build concurrency
Local source modifications are rejected rather than silently reset. Upstream is pinned; arbitrary HEAD updates are unsupported.
Put models outside .work/, for example under ~/models/. Obtain them from a source you trust and comply with their licenses. The validated model is Qwen2.5-Coder-1.5B-Instruct Q4_K_M; other models have not been validated here.
./scripts/run-model.sh /path/to/model.gguf
The launcher uses llama-completion; models with a chat template normally enter conversation mode. For a deterministic single raw prompt:
./scripts/run-model.sh /path/to/model.gguf \
-no-cnv -p "def add(a, b):" -n 32 --temp 0 --seed 1234
Defaults are -ngl 99 -c 512 -b 128 -ub 128 -t 4 -fa off -n 128. Additional llama arguments pass through as separate arguments and override defaults, for example -ngl 10 for less offload, -c 1024, or -st -p "Explain this code" for a single chat turn. Not every model fits GPU-accessible memory. See Troubleshooting.
The script derives paths from its own location and replaces LD_LIBRARY_PATH with only its own build/bin directory. Do not append $PREFIX/lib, /system/lib64, or /vendor/lib64. This avoids the vendor binder dependency trap and mixed llama libraries.
./scripts/check-device.sh
./scripts/verify-opencl.sh
./scripts/verify-opencl.sh --model /path/to/model.gguf
./scripts/verify-opencl.sh --matrix-tests
The model-free verifier checks the architecture, vendor files, exact patched source, build configuration, ELF dependencies, successful InitOpenCLDriver() = 0, and an Adreno OpenCL device. With a model it also requires positive layer offload, OpenCL model storage, completed generation evaluations, and exit status 0.
On the tested model the logs showed QUALCOMM Adreno(TM) 750, 5542 MiB, offloaded 29/29 layers, and about 739 MiB OpenCL weights. This does not mean every operation executes on GPU: embedding and large output-projection paths have CPU fallbacks. Separate profiling measured 20,539 actual GPU kernel executions.
Logs and machine-readable summaries are saved in .work/logs/.
./scripts/benchmark.sh /path/to/model.gguf
./scripts/benchmark.sh /path/to/model.gguf \
--threads 4 --prompt-tokens 128 --generation-tokens 64 --repetitions 3
CPU and OpenCL run sequentially against the same model, with matching thread counts, prompt/generation lengths, batch sizes, f16 KV and attention settings. Profiling must be OFF. Raw JSON, commands, standard deviations and a comparison summary are retained. A negative percentage means OpenCL is slower. Avoid concurrent builds/inference; thermal state and phone scheduling affect results.
./uninstall.sh # dry run
./uninstall.sh --yes # remove only this repository's .work
Shared Termux packages and models outside .work/ are retained. Uninstall protects user GGUFs inside .work/. Only unchanged vocabulary fixtures tracked at the pinned upstream commit are treated as generated source data.
Built on ggml-org/llama.cpp and Qualcomm's device-provided OpenCL stack. The upstream MIT license is preserved in docs/LLAMA_CPP_LICENSE.txt. Original scripts/documentation and patch contributions use MIT; model and proprietary vendor-library licenses remain separate. No upstream endorsement or universal Qualcomm compatibility is claimed.
Python
52.3%
Shell
34.0%
C++
12.2%
C
1.6%
Termux llama.cpp source installers for Samsung S24 Ultra: CPU + Adreno OpenCL GPU, or CPU + Hexagon v75 NPU. Complete NPU setup: install-npu.sh and docs/NPU_INSTALL.md. Tested inference; performance and compatibility limits documented.
Python
0
3 commits
updated Oct 4, 2026
Source-built llama.cpp inference on the S24 Ultra: CPU, Adreno OpenCL GPU, and a separate Hexagon v75 NPU installer.
Builds a pinned llama.cpp revision with Android sphal vendor-runtime loading. install.sh builds CPU + GPU; install-npu.sh separately builds CPU + NPU. Tested on Samsung Galaxy S24 Ultra, Snapdragon 8 Gen 3 / Adreno 750, Android 16, F-Droid Termux 0.118.3, aarch64.
The current generic OpenCL configuration is slower than CPU on the tested phone. This project proves working GPU computation; it is not optimized production acceleration.
| Original matched test, 4 threads | CPU | OpenCL |
|---|---|---|
| Prompt processing, 128 tokens | 122.74 tok/s | 82.45 tok/s |
| Generation, 64 tokens | 28.72 tok/s | 9.50 tok/s |
OpenCL was 32.8% slower for prompt processing and 66.9% slower for generation. Measurements and reproduction details: Benchmarks.
The full CPU/OpenCL/NPU/hybrid report and machine-readable evidence are now available. The investigation is paused at the owner's request; its final characterization is incomplete.
Experimental native Hexagon NPU tests show repeated prompt-processing gains of 7.14x for the tested 1.5B model and 7.83x for Qwen3-4B against matched common-setting CPU tests. NPU generation averaged 32.97 and 12.94 tok/s respectively; the best practical CPU comparison remains unfinished. Long-prompt GPU/hybrid results are candidates, and no sustained thermal winner is established. The GPU installer preserves the proven generic OpenCL baseline. The separate NPU installer packages the working native Hexagon implementation; experimental specialized GPU and handoff patches remain research snapshots. The historical report predates NPU packaging; see the current NPU guide.
For the separate CPU + Hexagon v75 NPU build, use:
pkg update
pkg install git
git clone https://github.com/Ishabdullah/OpenCL-S24-Ultra.git
cd OpenCL-S24-Ultra
./install-npu.sh --download-test-model
This installs missing Termux dependencies, prepares checksum-pinned SDK/compiler tools in project storage, patches pinned llama.cpp, source-builds native ARM64/Hexagon binaries, executes numerical DSP tests, and generates 32 tokens from the pinned official Qwen 1.5B GGUF. QEMU emulates compiler tools only; inference is native. No root or system/vendor changes are required. Allow at least 8 GiB free plus model space and tens of minutes for the build.
Use an existing model instead with ./install-npu.sh --model /path/to/model.gguf, then launch it through ./scripts/run-npu.sh /path/to/model.gguf. A model-free install verifies DSP matrices but does not test LLM generation. The optional model download is explicit; no model/SDK/compiler/vendor binary is hosted in this repository.
Start here: NPU installation, operation, licensing and troubleshooting. A fresh source/build reproduction passed eight numerical DSP tests and generated text with all 29 Qwen 1.5B layers offloaded: NPU validation. NPU inference is experimental: no universal performance/thermal benefit or arbitrary model/device compatibility is claimed. The GPU and NPU installers coexist in separate directories and do not yet provide one validated CPU/GPU/NPU executable.
libOpenCL.so, libOpenCL_adreno.so, libCB.so, libgsl.so in /vendor/lib64.Confirmed: the S24 Ultra configuration above. Detected but unverified: other phones whose vendor stack initializes and exposes an Adreno device. Potentially compatible: other Snapdragon/Adreno devices with the required namespace APIs and stack; file presence alone is not proof of compatibility.
In a fresh F-Droid Termux terminal:
pkg update
pkg install git
git clone https://github.com/Ishabdullah/OpenCL-S24-Ultra.git
cd OpenCL-S24-Ultra
./install.sh
The installer installs missing git clang cmake ninja python opencl-headers openssl packages. Clang pulls in LLVM (llvm-readelf); Python/Termux provide supporting runtime dependencies. Make, pkg-config and an OpenCL ICD loader are not required by this Ninja/source build.
It downloads upstream llama.cpp at e358d59178377be4c58ba567925e05faadbccb57, checks/applies the maintained patch, builds llama-completion, llama-bench, and test-backend-ops, verifies ELF dependencies, and detects the GPU. It does not modify an existing ~/llama.cpp, install llama libraries into $PREFIX, use root, or modify Android system/vendor files.
Generated source, binaries, cache and logs live under .work/ inside this repository. First kernel compilation and the source build take time; stages print the relevant log path. A successful installer run without a model proves device initialization, not model inference.
Rerunning ./install.sh checks an existing installation. Options:
./install.sh --help
./install.sh --verify
./install.sh --rebuild # configure and incrementally rebuild
./install.sh --clean-build # replace only the generated build directory
./install.sh --jobs 2 # lower build concurrency
Local source modifications are rejected rather than silently reset. Upstream is pinned; arbitrary HEAD updates are unsupported.
Put models outside .work/, for example under ~/models/. Obtain them from a source you trust and comply with their licenses. The validated model is Qwen2.5-Coder-1.5B-Instruct Q4_K_M; other models have not been validated here.
./scripts/run-model.sh /path/to/model.gguf
The launcher uses llama-completion; models with a chat template normally enter conversation mode. For a deterministic single raw prompt:
./scripts/run-model.sh /path/to/model.gguf \
-no-cnv -p "def add(a, b):" -n 32 --temp 0 --seed 1234
Defaults are -ngl 99 -c 512 -b 128 -ub 128 -t 4 -fa off -n 128. Additional llama arguments pass through as separate arguments and override defaults, for example -ngl 10 for less offload, -c 1024, or -st -p "Explain this code" for a single chat turn. Not every model fits GPU-accessible memory. See Troubleshooting.
The script derives paths from its own location and replaces LD_LIBRARY_PATH with only its own build/bin directory. Do not append $PREFIX/lib, /system/lib64, or /vendor/lib64. This avoids the vendor binder dependency trap and mixed llama libraries.
./scripts/check-device.sh
./scripts/verify-opencl.sh
./scripts/verify-opencl.sh --model /path/to/model.gguf
./scripts/verify-opencl.sh --matrix-tests
The model-free verifier checks the architecture, vendor files, exact patched source, build configuration, ELF dependencies, successful InitOpenCLDriver() = 0, and an Adreno OpenCL device. With a model it also requires positive layer offload, OpenCL model storage, completed generation evaluations, and exit status 0.
On the tested model the logs showed QUALCOMM Adreno(TM) 750, 5542 MiB, offloaded 29/29 layers, and about 739 MiB OpenCL weights. This does not mean every operation executes on GPU: embedding and large output-projection paths have CPU fallbacks. Separate profiling measured 20,539 actual GPU kernel executions.
Logs and machine-readable summaries are saved in .work/logs/.
./scripts/benchmark.sh /path/to/model.gguf
./scripts/benchmark.sh /path/to/model.gguf \
--threads 4 --prompt-tokens 128 --generation-tokens 64 --repetitions 3
CPU and OpenCL run sequentially against the same model, with matching thread counts, prompt/generation lengths, batch sizes, f16 KV and attention settings. Profiling must be OFF. Raw JSON, commands, standard deviations and a comparison summary are retained. A negative percentage means OpenCL is slower. Avoid concurrent builds/inference; thermal state and phone scheduling affect results.
./uninstall.sh # dry run
./uninstall.sh --yes # remove only this repository's .work
Shared Termux packages and models outside .work/ are retained. Uninstall protects user GGUFs inside .work/. Only unchanged vocabulary fixtures tracked at the pinned upstream commit are treated as generated source data.
Built on ggml-org/llama.cpp and Qualcomm's device-provided OpenCL stack. The upstream MIT license is preserved in docs/LLAMA_CPP_LICENSE.txt. Original scripts/documentation and patch contributions use MIT; model and proprietary vendor-library licenses remain separate. No upstream endorsement or universal Qualcomm compatibility is claimed.
Python
52.3%
Shell
34.0%
C++
12.2%
C
1.6%