Local, reproducible llama.cpp benchmarking and role-prompt assets for
GLaDOS.
This repository contains the benchmark wrapper, documentation, role prompts, and compact normalized benchmark CSV results. It intentionally excludes GGUF model weights, llama.cpp build trees, runtime binaries, credentials, and verbose benchmark sidecar logs.
| Component | Configuration |
|---|---|
| Host | GLaDOS |
| CPU | AMD Ryzen 7 8745H — 8 cores / 16 threads |
| GPU | AMD Radeon 780M integrated GPU |
| GPU backend | Mesa RADV / Vulkan |
| llama.cpp device | Vulkan0 |
| Usable memory | About 29 GiB shared system memory |
| Active model storage | PVE CephFS at /mnt/pve/llama_models |
| Expanded model library | /mnt/pve/AcornStash-LLMModels |
The Radeon 780M shares system memory with the host. Model selection must leave room for the operating system, context/KV cache, model runtime overhead, and other services.
.
├── README.md
├── scripts/
│ └── run-llama-bench.sh
└── docs/
├── prompts/
│ ├── README.md
│ ├── agentic-orchestrator.md
│ ├── coding-engineer.md
│ ├── evaluator-red-team.md
│ ├── financial-planning.md
│ └── systems-administration.md
└── stats/
Benchmarking is local and offline:
llama-bench directly; do not benchmark through a network service.CPU round 1 → Vulkan round 1 → CPU round 2 → Vulkan round 2 → CPU round 3 → Vulkan round 3
llama-bench:
/mnt/pve/llama_models/llama.cpp/build/bin/llama-bench
working directory:
/mnt/pve/llama_models/bin/llama
default operational output directory:
/mnt/pve/llama_models/stats
The repository can store reviewable compact benchmark CSVs under docs/stats/.
Raw and stderr sidecar logs remain ignored.
The documents in docs/prompts/ are application-level system prompts. They
are not model-specific GGUF chat templates and must not contain model control
tokens.
| Prompt | Purpose | Suggested model class |
|---|---|---|
agentic-orchestrator.md | Bounded planning and coordination | Qwen 3.8 4B or tool-use candidate |
coding-engineer.md | Code, tests, refactors, and review | Qwen2.5-Coder 7B or 14B |
systems-administration.md | Read-first infrastructure operations | Phi-4-mini-Instruct or Qwen 8B |
financial-planning.md | Budgeting and scenario analysis | Qwen2.5 14B Instruct |
evaluator-red-team.md | Critical independent review | Phi-4-mini or reasoning reviewer |
Use each model’s embedded chat template through llama.cpp, then provide the selected role prompt as the system message. Prompts improve consistency but do not make model output authoritative.
These are candidates for the Ryzen 7 8745H / Radeon 780M shared-memory Vulkan system. They must be tested against real tasks; throughput alone is not enough.
| Priority | Model | Quantization | Approx. size | Primary role |
|---|---|---|---|---|
| 1 | Qwen2.5-Coder-7B-Instruct | Q4_K_M | 4.4 GiB | Coding default |
| 2 | Qwen3.8-4B | Q6_K | 3.4 GiB | Agentic/orchestrator default |
| 3 | Phi-4-mini-Instruct | Q4_K_M | 2.4 GiB | Systems administration |
| 4 | Phi-4-mini-Reasoning | Q4_K_M | 2.4 GiB | Reasoning reviewer |
| 5 | NVIDIA-Nemotron-Nano-9B-v2 | Q4_K_M | 6.1 GiB | Agentic/tool-use candidate |
| 6 | Qwen3-8B | Q4_K_M | 4.7 GiB | General technical assistant |
| 7 | Qwen2.5-Coder-14B-Instruct | Q4_K_M | 8.4 GiB | Deep coding |
| 8 | Nous Hermes 4 14B | Q4_K_M | 8.4 GiB | Agentic escalation |
| 9 | NousCoder 14B | Q4_K_M | 8.4 GiB | Deep coding alternative |
| 10 | DeepSeek-R1-Distill-Qwen-14B | Q4_K_M | 8.4 GiB | Hard reasoning |
| 11 | Qwen2.5-14B-Instruct | Q4_K_M | 8.4 GiB | Financial planning |
Benchmark these first:
Qwen2.5-Coder-7B-Instruct-Q4_K_M
Qwen3.8-4B-Q6_K
Phi-4-mini-Instruct-Q4_K_M
Phi-4-mini-Reasoning-Q4_K_M
NVIDIA-Nemotron-Nano-9B-v2-Q4_K_M
Qwen3-8B-Q4_K_M
Treat 14B models as on-demand escalation candidates. They may fit in memory, but leave less headroom for large contexts, concurrent requests, and host services.
Do not benchmark standalone projector/draft artifacts as chat models:
mmproj-*.ggufmtp-*.gguf*-mmproj*.ggufRank candidates using real work as well as throughput:
Representative tasks include systemd diagnosis from logs, idempotent Ansible repair, shell-script safety review, Proxmox/Ceph troubleshooting, GitLab CI drafting, code review, and financial scenario comparison with explicit assumptions.
This repository stores source configuration and benchmark tooling for llama.cpp AI engines. It does not store GGUF weights, credentials, runtime binaries, caches, logs, or generated service state.
presets/ contains reviewed model-preset source files.docs/ai-engines/ documents each inference host.deploy/ contains benchmark deployment payloads.A Git preset change does not alter a running llama-server. Host changes require separate validation and an explicit deployment procedure.
Local, reproducible llama.cpp benchmarking and role-prompt assets for
GLaDOS.
This repository contains the benchmark wrapper, documentation, role prompts, and compact normalized benchmark CSV results. It intentionally excludes GGUF model weights, llama.cpp build trees, runtime binaries, credentials, and verbose benchmark sidecar logs.
| Component | Configuration |
|---|---|
| Host | GLaDOS |
| CPU | AMD Ryzen 7 8745H — 8 cores / 16 threads |
| GPU | AMD Radeon 780M integrated GPU |
| GPU backend | Mesa RADV / Vulkan |
| llama.cpp device | Vulkan0 |
| Usable memory | About 29 GiB shared system memory |
| Active model storage | PVE CephFS at /mnt/pve/llama_models |
| Expanded model library | /mnt/pve/AcornStash-LLMModels |
The Radeon 780M shares system memory with the host. Model selection must leave room for the operating system, context/KV cache, model runtime overhead, and other services.
.
├── README.md
├── scripts/
│ └── run-llama-bench.sh
└── docs/
├── prompts/
│ ├── README.md
│ ├── agentic-orchestrator.md
│ ├── coding-engineer.md
│ ├── evaluator-red-team.md
│ ├── financial-planning.md
│ └── systems-administration.md
└── stats/
Benchmarking is local and offline:
llama-bench directly; do not benchmark through a network service.CPU round 1 → Vulkan round 1 → CPU round 2 → Vulkan round 2 → CPU round 3 → Vulkan round 3
llama-bench:
/mnt/pve/llama_models/llama.cpp/build/bin/llama-bench
working directory:
/mnt/pve/llama_models/bin/llama
default operational output directory:
/mnt/pve/llama_models/stats
The repository can store reviewable compact benchmark CSVs under docs/stats/.
Raw and stderr sidecar logs remain ignored.
The documents in docs/prompts/ are application-level system prompts. They
are not model-specific GGUF chat templates and must not contain model control
tokens.
| Prompt | Purpose | Suggested model class |
|---|---|---|
agentic-orchestrator.md | Bounded planning and coordination | Qwen 3.8 4B or tool-use candidate |
coding-engineer.md | Code, tests, refactors, and review | Qwen2.5-Coder 7B or 14B |
systems-administration.md | Read-first infrastructure operations | Phi-4-mini-Instruct or Qwen 8B |
financial-planning.md | Budgeting and scenario analysis | Qwen2.5 14B Instruct |
evaluator-red-team.md | Critical independent review | Phi-4-mini or reasoning reviewer |
Use each model’s embedded chat template through llama.cpp, then provide the selected role prompt as the system message. Prompts improve consistency but do not make model output authoritative.
These are candidates for the Ryzen 7 8745H / Radeon 780M shared-memory Vulkan system. They must be tested against real tasks; throughput alone is not enough.
| Priority | Model | Quantization | Approx. size | Primary role |
|---|---|---|---|---|
| 1 | Qwen2.5-Coder-7B-Instruct | Q4_K_M | 4.4 GiB | Coding default |
| 2 | Qwen3.8-4B | Q6_K | 3.4 GiB | Agentic/orchestrator default |
| 3 | Phi-4-mini-Instruct | Q4_K_M | 2.4 GiB | Systems administration |
| 4 | Phi-4-mini-Reasoning | Q4_K_M | 2.4 GiB | Reasoning reviewer |
| 5 | NVIDIA-Nemotron-Nano-9B-v2 | Q4_K_M | 6.1 GiB | Agentic/tool-use candidate |
| 6 | Qwen3-8B | Q4_K_M | 4.7 GiB | General technical assistant |
| 7 | Qwen2.5-Coder-14B-Instruct | Q4_K_M | 8.4 GiB | Deep coding |
| 8 | Nous Hermes 4 14B | Q4_K_M | 8.4 GiB | Agentic escalation |
| 9 | NousCoder 14B | Q4_K_M | 8.4 GiB | Deep coding alternative |
| 10 | DeepSeek-R1-Distill-Qwen-14B | Q4_K_M | 8.4 GiB | Hard reasoning |
| 11 | Qwen2.5-14B-Instruct | Q4_K_M | 8.4 GiB | Financial planning |
Benchmark these first:
Qwen2.5-Coder-7B-Instruct-Q4_K_M
Qwen3.8-4B-Q6_K
Phi-4-mini-Instruct-Q4_K_M
Phi-4-mini-Reasoning-Q4_K_M
NVIDIA-Nemotron-Nano-9B-v2-Q4_K_M
Qwen3-8B-Q4_K_M
Treat 14B models as on-demand escalation candidates. They may fit in memory, but leave less headroom for large contexts, concurrent requests, and host services.
Do not benchmark standalone projector/draft artifacts as chat models:
mmproj-*.ggufmtp-*.gguf*-mmproj*.ggufRank candidates using real work as well as throughput:
Representative tasks include systemd diagnosis from logs, idempotent Ansible repair, shell-script safety review, Proxmox/Ceph troubleshooting, GitLab CI drafting, code review, and financial scenario comparison with explicit assumptions.
This repository stores source configuration and benchmark tooling for llama.cpp AI engines. It does not store GGUF weights, credentials, runtime binaries, caches, logs, or generated service state.
presets/ contains reviewed model-preset source files.docs/ai-engines/ documents each inference host.deploy/ contains benchmark deployment payloads.A Git preset change does not alter a running llama-server. Host changes require separate validation and an explicit deployment procedure.