The unmodified upstream Linux amdgpu + amdkfd driver running on macOS: AMD Radeon GPUs over Thunderbolt on Apple Silicon, inside a DriverKit extension.
C
2
327 commits
updated Oct 7, 2026
The upstream Linux amdgpu driver, unmodified, running on macOS.
AMD Radeon GPUs over Thunderbolt on Apple Silicon Macs, driven by the same
amdgpu + amdkfd code Linux uses, inside a DriverKit system extension.
mac_linuxgpu · amdgpu_mtopg · LemonSeed Engine
Apple Silicon Macs have Thunderbolt 5 but no AMD GPU driver. mac_linuxgpu
runs the real Linux kernel driver instead of reimplementing one. It compiles
the upstream amdgpu, amdkfd, TTM, DRM core and GPU scheduler sources
(468 files from a pinned Linux commit, unmodified) against linuxu, a
userspace implementation of the Linux kernel APIs those sources need, and
hosts the result in a PCIDriverKit extension.
The GPU initializes the way it does on Linux: IP discovery, PSP, SMU, GMC, GFX12, SDMA and MES. Compute works the way it does on Linux too. Every client is a KFD process with its own GPU address space, and its HSA queues are MES-scheduled user queues. The goal is that GPU software written for Linux/ROCm runs on a Mac with nothing Mac-specific.
LemonSeed Engine v0.5.4, Qwen3.8-27B Q4 with the Q8 DFlash2 draft, on an AMD Radeon AI PRO R9700 over Thunderbolt 5 (MacBook Pro, Apple M5 Max), driver build 262:
| Measured | |
|---|---|
| Prefill, 137-token prompt (time to first token, warm) | 0.17 s |
| Prefill, 1060 / 2118 / 4230-token prompt (warm) | 1,421 / 1,490 / 1,496 tok/s |
| First request after server start (137 tokens) | 0.34 s |
| Decode, DFlash2 speculative (640 tokens) | 50.5 tok/s |
| Decode, MTP=3 speculative (640 tokens) | 46.7–50.8 tok/s |
| Decode, plain | 27.7 tok/s |
| Model load (Q4 + DFlash2) | 3.4 s |
| Operations that fell back to the CPU | 0 |
llama.cpp (Vulkan through RADV, Qwen3.6-27B Q4_K_XL): pp512 1,100 tok/s, tg128 24.8 tok/s.
Against stock Linux on the same GPU model (amdgpu + ROCm 7.13, PCIe 5.0 x16, same LSE v0.5.4 build and model files), this driver over Thunderbolt is 12–43% faster on LSE decode, at parity on prefill from 532 to 2118 tokens, and 8% faster at 4K. Linux remains faster on per-dispatch host cost (doorbell and signal writes are direct memory writes there) and on bulk copies (an 8x wider link).
amdgpu_pci_probe completes
on the R9700, and every firmware image is loaded on demand by name.ALLOC_MEMORY_OF_GPU / MAP_MEMORY_TO_GPU, and
queues go through CREATE_QUEUE → MES ADD_QUEUE, exactly as on Linux.libhsa-runtime64.dylib is installed with the driver, in
/usr/local/lib and /Library/MacAMDGPU/runtime. HRX/Loom-based software
such as LemonSeed Engine finds it without configuration.scripts/read-driver-log.py). While compute runs,
it also reads upstream's own telemetry: sysfs attributes such as
gpu_busy_percent, mem_busy_percent, gpu_metrics and hwmon, and
AMDGPU_INFO. That's the same data amdgpu_top reads on Linux, and what
amdgpu_mtopg displays.
Events that matter after a hang or a reboot (driver start, probe result,
GPU ring timeouts and device dumps, errors from upstream, removal,
quarantine, session close) also go to the macOS unified log, prefixed
mac.linuxgpu: EVENT; routine lines stay in the retained log only. Read
them, from any boot the log still holds, with
log show --last 1d --predicate 'eventMessage BEGINSWITH "mac.linuxgpu: EVENT"'
(DriverKit logs through the kernel, so they show as kernel).amdgpu_mtopg monitoring the
R9700 through this driver while LemonSeed Engine generates text: GPU load from
GRBM_STATUS samples, SMU clocks and power, hwmon temperatures and fan, and
VRAM in use by the model.

your app (LemonSeed, HRX/Loom, other HSA clients)
│ HSA API
libhsa-runtime64.dylib installed with the driver
│ IOKit user client
┌──────┴──────────────────── DriverKit extension ────────────────────┐
│ per-client KFD process: /dev/kfd + DRM render node, in-process │
│ upstream amdkfd + amdgpu + TTM + DRM + drm_sched (unmodified) │
│ linuxu: the Linux kernel API in userspace │
│ workqueues · timers · RCU · fences · mm/VMA · mmu notifiers · │
│ sysfs · firmware loader · DMA through DART · MMIO · MSI-X │
└──────┬─────────────────────────────────────────────────────────────┘
│ PCIDriverKit: BARs, DMA, interrupts
AMD GPU over Thunderbolt
xcode-select pointing at Xcode)python3, git and curl (from the Xcode command line tools)cmake for the HSA runtime (brew install cmake)ripgrep for some test scripts (brew install ripgrep)There are two ways. Both end in the same tree.
Option 1: the setup script (recommended). It downloads only what the build needs.
git clone https://github.com/lemonade-sdk/mac_linuxgpu.git
cd mac_linuxgpu
scripts/bootstrap.sh
scripts/bootstrap.sh does all one-time setup and is safe to re-run:
third_party/linux as a shallow,
partial, sparse checkout of only the paths the build reads (about 70 MB
downloaded, 600 MB on disk);patches/linux/;build/firmware/ and
verifies their hashes;build/setup/.Option 2: plain git.
git clone --recurse-submodules https://github.com/lemonade-sdk/mac_linuxgpu.git
# or, in an existing clone:
git submodule update --init
This works, but git checks out the entire kernel. Measured: about 3.3 GB
downloaded and 1.7 GB on disk, against about 70 MB with the setup script. The
first make (or scripts/bootstrap.sh) then converts the checkout to the
sparse set the build uses, applies the patches and fetches the firmware.
Mesa (third_party/mesa) and llama.cpp (third_party/llama.cpp) are
optional submodules for the Vulkan build only. They are marked
update = none, so neither option fetches them; make radv and
make llama-vulkan do (scripts/bootstrap.sh --mesa-only, --llama-only).
To keep the kernel at the right commit when you pull updates, use
git pull --recurse-submodules, or set it once with:
git config submodule.recurse true
Either way, you don't have to run setup by hand. Every make target
(except clean, distclean, hsa and hsa-test) checks the setup stamp. If
the kernel checkout is missing, at the wrong commit, or unpatched, or the
firmware is missing, make runs scripts/bootstrap.sh first.
scripts/activate.sh and the Xcode project's "Build Linux KMD" phase both go
through make, so a fresh clone builds from any entry point.
make lib-dext # the DriverKit build of the driver library
make test # the offline test suite (no GPU needed)
scripts/activate.sh # build, sign, install and activate the driver
| Command | Result |
|---|---|
make / make lib | build/libmacamgdu.a, host platform, used by the tests |
make lib-dext | build-dk/libmacamgdu-dk.a, DriverKit platform |
make dext | links the dext with make (unsigned unless DIDENTITY is set) |
make test | the offline test suite |
make hsa, make hsa-test | the HSA runtime (build/hsa) and its unit tests |
make libdrm-mlg | libdrm and libdrm_amdgpu over the Linux-file transport (build/libdrm-mlg) |
make radv | Mesa's RADV Vulkan driver and its ICD (build/radv/install) |
make llama-vulkan | llama.cpp with its Vulkan backend (build/llama.cpp/bin) |
make verify-source | checks the upstream tree against the pin and patch set |
make clean | removes build outputs, keeps setup and firmware |
make distclean | removes build/ and build-dk/ |
The Xcode project mac_linuxgpu.xcodeproj builds the signed dext
(MacLinuxGPU) and the host app (MacLinuxGPUHost). Its first build phase
runs make lib-dext, and the app bundles build/firmware/amdgpu.
scripts/build-installer-dmg.sh produces a signed disk image whose app
installs the driver, the HSA runtime and the firmware.
scripts/activate.sh builds everything, signs with Xcode automatic signing or
locally approved provisioning profiles, installs the app to /Applications,
installs the firmware and HSA runtime, and waits for the driver to attach. Set
XCODE_TEAM_ID to sign with your own team. See scripts/activate.sh --help.
third_party/linux is a git submodule of
torvalds/linux pinned to
1f63dd8ca0dc05a8272bb8155f643c691d29bb11. Its sparse checkout covers the
amdgpu tree, the DRM core, TTM, the scheduler, a few lib/ helpers and
include/. Paths that differ only in letter case are excluded, so the
checkout stays clean on case-insensitive APFS.mk/upstream_sources.mk lists every upstream .c file the build
compiles. Nothing is globbed. The include paths in mk/host_clang.mk point
straight into the submodule.patches/linux/*.patch are applied to the submodule working tree by the
setup script, never committed inside it. Re-applying is detected and
skipped. To change a patch, edit it, update its sha256 in
patches/manifest.json, and run scripts/bootstrap.sh --reset-linux.scripts/verify-upstream.py runs as make verify-source and before every
lib-dext. It checks that:
linuxu/headers still match;linux.pin in
patches/manifest.json and mk/upstream_sources.mk, then run
scripts/bootstrap.sh.GPU firmware is not stored in this repository. firmware/firmware.lock pins
a linux-firmware release and lists each needed file with its SHA-256 and
size. scripts/fetch-firmware.sh downloads them from
gitlab.com/kernel-firmware/linux-firmware, falling back to the
git.kernel.org mirror, and verifies every hash.
scripts/fetch-firmware.sh --add amdgpu/<file>.bin fetches more files at
the pinned release and adds them to the lock.scripts/fetch-firmware.sh --all fetches the whole amdgpu directory, for
GPUs beyond the ones in the lock. Bundle it with
LINUX_FIRMWARE_DIR=build/firmware/full scripts/activate.sh.At run time the driver asks for firmware by name. The installed host
component serves it from
/Library/Application Support/MacLinuxGPU/firmware/amdgpu.
dext/ DriverKit extension sources, Info.plist, entitlements
host/ host app: installer, activation, firmware servicer, CLI
hsa/ userspace HSA runtime (CMake) and its tests
libmlg_drm/ the Linux-file client library, and libdrm over it (libdrm/)
vulkan/ the Vulkan path: RADV notes, offline test
linuxu/headers/ Linux kernel API headers for userspace
linuxu/src/ their implementation: memory, locking, PCI, DMA, ...
linuxu/tests/ unit and integration tests
mk/ build configuration: CONFIG table, flags, source list
patches/ patches/linux/*.patch and the upstream manifest
firmware/ firmware.lock (the files themselves are fetched)
scripts/ setup, firmware fetch, verifier, installers, tests
tools/ CONFIG table consistency checks
third_party/linux pinned upstream Linux (submodule)
third_party/mesa pinned Mesa for RADV (optional submodule)
third_party/llama.cpp pinned llama.cpp (optional submodule)
gfx1201), over
Thunderbolt 5. Other AMD GPUs supported by upstream amdgpu should match and
probe, but are untested.IOPCIFamily). If scripts/read-driver-log.py reports a
quarantined session, restart the Mac instead.systemextensionsctl list shows it terminating for upgrade via delegate) and the new one attaches only once every old instance is gone.
The installer therefore asks the running driver to close its session the
normal way before it requests the replacement, and to terminate itself once
macOS has accepted it (the Retire selector, dext/sources/session_state.h).
It reports apps that still hold a session and a quarantine that needs a
restart. To finish a stuck upgrade, quit the apps it lists and run
MacLinuxGPUHost retire-previous (--force closes their sessions);
MacLinuxGPUHost instances shows what is attached. Drivers up to 0.1.129
(build 233) have no Retire: from 0.1.128 (232) they leave cleanly when the
GPU enclosure is switched off; older ones need a restart.MAC_HSA_STATUS_DEVICE_LOST). Upstream's
only suspend path for a discrete GPU evicts all of VRAM to system memory,
which doesn't fit a large model on the iPad. A low-power period without
host sleep keeps VRAM: compute is quiesced through upstream KFD suspend
and resumes where it left off (mac_hsa_agent_prepare_low_power /
mac_hsa_agent_resume, or MacLinuxGPUHostApp power-watch on a Mac).
See dext/sources/power_state.h.Original code in this repository is available under either the
MIT license or the GNU GPL v2, at your
option (SPDX-License-Identifier: MIT OR GPL-2.0-only); see
LICENSE. Upstream Linux code in third_party/linux, and files in
linuxu/ that state a Linux-derived license, keep their own licenses.
Firmware fetched at build time is distributed under its own terms (see its
WHENCE file) and is not part of this repository.
This driver runs the stack behind these results. They are LemonSeed Engine's HumanEval+ runs of Qwen3.8-27B Q4 on the AMD R9700 with HRX/Loom: 1,368 completed generations at standard, 16K and 32K context. They were measured through the earlier mac_amdgpu driver path, and they are the target for this driver.

The unmodified upstream Linux amdgpu + amdkfd driver running on macOS: AMD Radeon GPUs over Thunderbolt on Apple Silicon, inside a DriverKit extension.
C
2
327 commits
updated Oct 7, 2026
The upstream Linux amdgpu driver, unmodified, running on macOS.
AMD Radeon GPUs over Thunderbolt on Apple Silicon Macs, driven by the same
amdgpu + amdkfd code Linux uses, inside a DriverKit system extension.
mac_linuxgpu · amdgpu_mtopg · LemonSeed Engine
Apple Silicon Macs have Thunderbolt 5 but no AMD GPU driver. mac_linuxgpu
runs the real Linux kernel driver instead of reimplementing one. It compiles
the upstream amdgpu, amdkfd, TTM, DRM core and GPU scheduler sources
(468 files from a pinned Linux commit, unmodified) against linuxu, a
userspace implementation of the Linux kernel APIs those sources need, and
hosts the result in a PCIDriverKit extension.
The GPU initializes the way it does on Linux: IP discovery, PSP, SMU, GMC, GFX12, SDMA and MES. Compute works the way it does on Linux too. Every client is a KFD process with its own GPU address space, and its HSA queues are MES-scheduled user queues. The goal is that GPU software written for Linux/ROCm runs on a Mac with nothing Mac-specific.
LemonSeed Engine v0.5.4, Qwen3.8-27B Q4 with the Q8 DFlash2 draft, on an AMD Radeon AI PRO R9700 over Thunderbolt 5 (MacBook Pro, Apple M5 Max), driver build 262:
| Measured | |
|---|---|
| Prefill, 137-token prompt (time to first token, warm) | 0.17 s |
| Prefill, 1060 / 2118 / 4230-token prompt (warm) | 1,421 / 1,490 / 1,496 tok/s |
| First request after server start (137 tokens) | 0.34 s |
| Decode, DFlash2 speculative (640 tokens) | 50.5 tok/s |
| Decode, MTP=3 speculative (640 tokens) | 46.7–50.8 tok/s |
| Decode, plain | 27.7 tok/s |
| Model load (Q4 + DFlash2) | 3.4 s |
| Operations that fell back to the CPU | 0 |
llama.cpp (Vulkan through RADV, Qwen3.6-27B Q4_K_XL): pp512 1,100 tok/s, tg128 24.8 tok/s.
Against stock Linux on the same GPU model (amdgpu + ROCm 7.13, PCIe 5.0 x16, same LSE v0.5.4 build and model files), this driver over Thunderbolt is 12–43% faster on LSE decode, at parity on prefill from 532 to 2118 tokens, and 8% faster at 4K. Linux remains faster on per-dispatch host cost (doorbell and signal writes are direct memory writes there) and on bulk copies (an 8x wider link).
amdgpu_pci_probe completes
on the R9700, and every firmware image is loaded on demand by name.ALLOC_MEMORY_OF_GPU / MAP_MEMORY_TO_GPU, and
queues go through CREATE_QUEUE → MES ADD_QUEUE, exactly as on Linux.libhsa-runtime64.dylib is installed with the driver, in
/usr/local/lib and /Library/MacAMDGPU/runtime. HRX/Loom-based software
such as LemonSeed Engine finds it without configuration.scripts/read-driver-log.py). While compute runs,
it also reads upstream's own telemetry: sysfs attributes such as
gpu_busy_percent, mem_busy_percent, gpu_metrics and hwmon, and
AMDGPU_INFO. That's the same data amdgpu_top reads on Linux, and what
amdgpu_mtopg displays.
Events that matter after a hang or a reboot (driver start, probe result,
GPU ring timeouts and device dumps, errors from upstream, removal,
quarantine, session close) also go to the macOS unified log, prefixed
mac.linuxgpu: EVENT; routine lines stay in the retained log only. Read
them, from any boot the log still holds, with
log show --last 1d --predicate 'eventMessage BEGINSWITH "mac.linuxgpu: EVENT"'
(DriverKit logs through the kernel, so they show as kernel).amdgpu_mtopg monitoring the
R9700 through this driver while LemonSeed Engine generates text: GPU load from
GRBM_STATUS samples, SMU clocks and power, hwmon temperatures and fan, and
VRAM in use by the model.

your app (LemonSeed, HRX/Loom, other HSA clients)
│ HSA API
libhsa-runtime64.dylib installed with the driver
│ IOKit user client
┌──────┴──────────────────── DriverKit extension ────────────────────┐
│ per-client KFD process: /dev/kfd + DRM render node, in-process │
│ upstream amdkfd + amdgpu + TTM + DRM + drm_sched (unmodified) │
│ linuxu: the Linux kernel API in userspace │
│ workqueues · timers · RCU · fences · mm/VMA · mmu notifiers · │
│ sysfs · firmware loader · DMA through DART · MMIO · MSI-X │
└──────┬─────────────────────────────────────────────────────────────┘
│ PCIDriverKit: BARs, DMA, interrupts
AMD GPU over Thunderbolt
xcode-select pointing at Xcode)python3, git and curl (from the Xcode command line tools)cmake for the HSA runtime (brew install cmake)ripgrep for some test scripts (brew install ripgrep)There are two ways. Both end in the same tree.
Option 1: the setup script (recommended). It downloads only what the build needs.
git clone https://github.com/lemonade-sdk/mac_linuxgpu.git
cd mac_linuxgpu
scripts/bootstrap.sh
scripts/bootstrap.sh does all one-time setup and is safe to re-run:
third_party/linux as a shallow,
partial, sparse checkout of only the paths the build reads (about 70 MB
downloaded, 600 MB on disk);patches/linux/;build/firmware/ and
verifies their hashes;build/setup/.Option 2: plain git.
git clone --recurse-submodules https://github.com/lemonade-sdk/mac_linuxgpu.git
# or, in an existing clone:
git submodule update --init
This works, but git checks out the entire kernel. Measured: about 3.3 GB
downloaded and 1.7 GB on disk, against about 70 MB with the setup script. The
first make (or scripts/bootstrap.sh) then converts the checkout to the
sparse set the build uses, applies the patches and fetches the firmware.
Mesa (third_party/mesa) and llama.cpp (third_party/llama.cpp) are
optional submodules for the Vulkan build only. They are marked
update = none, so neither option fetches them; make radv and
make llama-vulkan do (scripts/bootstrap.sh --mesa-only, --llama-only).
To keep the kernel at the right commit when you pull updates, use
git pull --recurse-submodules, or set it once with:
git config submodule.recurse true
Either way, you don't have to run setup by hand. Every make target
(except clean, distclean, hsa and hsa-test) checks the setup stamp. If
the kernel checkout is missing, at the wrong commit, or unpatched, or the
firmware is missing, make runs scripts/bootstrap.sh first.
scripts/activate.sh and the Xcode project's "Build Linux KMD" phase both go
through make, so a fresh clone builds from any entry point.
make lib-dext # the DriverKit build of the driver library
make test # the offline test suite (no GPU needed)
scripts/activate.sh # build, sign, install and activate the driver
| Command | Result |
|---|---|
make / make lib | build/libmacamgdu.a, host platform, used by the tests |
make lib-dext | build-dk/libmacamgdu-dk.a, DriverKit platform |
make dext | links the dext with make (unsigned unless DIDENTITY is set) |
make test | the offline test suite |
make hsa, make hsa-test | the HSA runtime (build/hsa) and its unit tests |
make libdrm-mlg | libdrm and libdrm_amdgpu over the Linux-file transport (build/libdrm-mlg) |
make radv | Mesa's RADV Vulkan driver and its ICD (build/radv/install) |
make llama-vulkan | llama.cpp with its Vulkan backend (build/llama.cpp/bin) |
make verify-source | checks the upstream tree against the pin and patch set |
make clean | removes build outputs, keeps setup and firmware |
make distclean | removes build/ and build-dk/ |
The Xcode project mac_linuxgpu.xcodeproj builds the signed dext
(MacLinuxGPU) and the host app (MacLinuxGPUHost). Its first build phase
runs make lib-dext, and the app bundles build/firmware/amdgpu.
scripts/build-installer-dmg.sh produces a signed disk image whose app
installs the driver, the HSA runtime and the firmware.
scripts/activate.sh builds everything, signs with Xcode automatic signing or
locally approved provisioning profiles, installs the app to /Applications,
installs the firmware and HSA runtime, and waits for the driver to attach. Set
XCODE_TEAM_ID to sign with your own team. See scripts/activate.sh --help.
third_party/linux is a git submodule of
torvalds/linux pinned to
1f63dd8ca0dc05a8272bb8155f643c691d29bb11. Its sparse checkout covers the
amdgpu tree, the DRM core, TTM, the scheduler, a few lib/ helpers and
include/. Paths that differ only in letter case are excluded, so the
checkout stays clean on case-insensitive APFS.mk/upstream_sources.mk lists every upstream .c file the build
compiles. Nothing is globbed. The include paths in mk/host_clang.mk point
straight into the submodule.patches/linux/*.patch are applied to the submodule working tree by the
setup script, never committed inside it. Re-applying is detected and
skipped. To change a patch, edit it, update its sha256 in
patches/manifest.json, and run scripts/bootstrap.sh --reset-linux.scripts/verify-upstream.py runs as make verify-source and before every
lib-dext. It checks that:
linuxu/headers still match;linux.pin in
patches/manifest.json and mk/upstream_sources.mk, then run
scripts/bootstrap.sh.GPU firmware is not stored in this repository. firmware/firmware.lock pins
a linux-firmware release and lists each needed file with its SHA-256 and
size. scripts/fetch-firmware.sh downloads them from
gitlab.com/kernel-firmware/linux-firmware, falling back to the
git.kernel.org mirror, and verifies every hash.
scripts/fetch-firmware.sh --add amdgpu/<file>.bin fetches more files at
the pinned release and adds them to the lock.scripts/fetch-firmware.sh --all fetches the whole amdgpu directory, for
GPUs beyond the ones in the lock. Bundle it with
LINUX_FIRMWARE_DIR=build/firmware/full scripts/activate.sh.At run time the driver asks for firmware by name. The installed host
component serves it from
/Library/Application Support/MacLinuxGPU/firmware/amdgpu.
dext/ DriverKit extension sources, Info.plist, entitlements
host/ host app: installer, activation, firmware servicer, CLI
hsa/ userspace HSA runtime (CMake) and its tests
libmlg_drm/ the Linux-file client library, and libdrm over it (libdrm/)
vulkan/ the Vulkan path: RADV notes, offline test
linuxu/headers/ Linux kernel API headers for userspace
linuxu/src/ their implementation: memory, locking, PCI, DMA, ...
linuxu/tests/ unit and integration tests
mk/ build configuration: CONFIG table, flags, source list
patches/ patches/linux/*.patch and the upstream manifest
firmware/ firmware.lock (the files themselves are fetched)
scripts/ setup, firmware fetch, verifier, installers, tests
tools/ CONFIG table consistency checks
third_party/linux pinned upstream Linux (submodule)
third_party/mesa pinned Mesa for RADV (optional submodule)
third_party/llama.cpp pinned llama.cpp (optional submodule)
gfx1201), over
Thunderbolt 5. Other AMD GPUs supported by upstream amdgpu should match and
probe, but are untested.IOPCIFamily). If scripts/read-driver-log.py reports a
quarantined session, restart the Mac instead.systemextensionsctl list shows it terminating for upgrade via delegate) and the new one attaches only once every old instance is gone.
The installer therefore asks the running driver to close its session the
normal way before it requests the replacement, and to terminate itself once
macOS has accepted it (the Retire selector, dext/sources/session_state.h).
It reports apps that still hold a session and a quarantine that needs a
restart. To finish a stuck upgrade, quit the apps it lists and run
MacLinuxGPUHost retire-previous (--force closes their sessions);
MacLinuxGPUHost instances shows what is attached. Drivers up to 0.1.129
(build 233) have no Retire: from 0.1.128 (232) they leave cleanly when the
GPU enclosure is switched off; older ones need a restart.MAC_HSA_STATUS_DEVICE_LOST). Upstream's
only suspend path for a discrete GPU evicts all of VRAM to system memory,
which doesn't fit a large model on the iPad. A low-power period without
host sleep keeps VRAM: compute is quiesced through upstream KFD suspend
and resumes where it left off (mac_hsa_agent_prepare_low_power /
mac_hsa_agent_resume, or MacLinuxGPUHostApp power-watch on a Mac).
See dext/sources/power_state.h.Original code in this repository is available under either the
MIT license or the GNU GPL v2, at your
option (SPDX-License-Identifier: MIT OR GPL-2.0-only); see
LICENSE. Upstream Linux code in third_party/linux, and files in
linuxu/ that state a Linux-derived license, keep their own licenses.
Firmware fetched at build time is distributed under its own terms (see its
WHENCE file) and is not part of this repository.
This driver runs the stack behind these results. They are LemonSeed Engine's HumanEval+ runs of Qwen3.8-27B Q4 on the AMD R9700 with HRX/Loom: 1,368 completed generations at standard, 16K and 32K context. They were measured through the earlier mac_amdgpu driver path, and they are the target for this driver.
