phillf/llama-bench

Shell

0

9 commits

updated Oct 5, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built a reproducible llama.cpp benchmark harness for AMD Vulkan and NVIDIA CUDA — GitLab CI automation is next (r/LocalLLM)

I have been trying to make local llama.cpp benchmark results easier to reproduce across two pretty different setups: \- AMDGPU + Vulkan \- NVIDIA + CUDA The recurring issue for me was that a CSV with tokens/sec did not provide enough context to compare someone else’s result—or even to reproduce my…

1

Oct 5, 2026

README

llama-bench

Local, reproducible llama.cpp benchmarking and role-prompt assets for GLaDOS.

This repository contains the benchmark wrapper, documentation, role prompts, and compact normalized benchmark CSV results. It intentionally excludes GGUF model weights, llama.cpp build trees, runtime binaries, credentials, and verbose benchmark sidecar logs.

GLaDOS profile

ComponentConfiguration
HostGLaDOS
CPUAMD Ryzen 7 8745H — 8 cores / 16 threads
GPUAMD Radeon 780M integrated GPU
GPU backendMesa RADV / Vulkan
llama.cpp deviceVulkan0
Usable memoryAbout 29 GiB shared system memory
Active model storagePVE CephFS at /mnt/pve/llama_models
Expanded model library/mnt/pve/AcornStash-LLMModels

The Radeon 780M shares system memory with the host. Model selection must leave room for the operating system, context/KV cache, model runtime overhead, and other services.

Repository layout

.
├── README.md
├── scripts/
│   └── run-llama-bench.sh
└── docs/
    ├── prompts/
    │   ├── README.md
    │   ├── agentic-orchestrator.md
    │   ├── coding-engineer.md
    │   ├── evaluator-red-team.md
    │   ├── financial-planning.md
    │   └── systems-administration.md
    └── stats/

Benchmark policy

Benchmarking is local and offline:

  • Use llama-bench directly; do not benchmark through a network service.
  • Do not bind, expose, or connect to a network port.
  • Supply an explicit GGUF path.
  • Keep CPU and GPU runs serial.
  • For paired comparisons, run:
CPU round 1 → Vulkan round 1 → CPU round 2 → Vulkan round 2 → CPU round 3 → Vulkan round 3
  • Keep CPU, Vulkan, CUDA, different quantizations, different contexts, and different llama.cpp builds in separate comparison groups.
  • Record model checksum, runtime commit/build, backend/device, thread count, GPU layers, prompt tokens, generation tokens, and repetitions.

Wrapper defaults

llama-bench:
/mnt/pve/llama_models/llama.cpp/build/bin/llama-bench

working directory:
/mnt/pve/llama_models/bin/llama

default operational output directory:
/mnt/pve/llama_models/stats

The repository can store reviewable compact benchmark CSVs under docs/stats/. Raw and stderr sidecar logs remain ignored.

Role prompts

The documents in docs/prompts/ are application-level system prompts. They are not model-specific GGUF chat templates and must not contain model control tokens.

PromptPurposeSuggested model class
agentic-orchestrator.mdBounded planning and coordinationQwen 3.8 4B or tool-use candidate
coding-engineer.mdCode, tests, refactors, and reviewQwen2.5-Coder 7B or 14B
systems-administration.mdRead-first infrastructure operationsPhi-4-mini-Instruct or Qwen 8B
financial-planning.mdBudgeting and scenario analysisQwen2.5 14B Instruct
evaluator-red-team.mdCritical independent reviewPhi-4-mini or reasoning reviewer

Use each model’s embedded chat template through llama.cpp, then provide the selected role prompt as the system message. Prompts improve consistency but do not make model output authoritative.

Suggested models

These are candidates for the Ryzen 7 8745H / Radeon 780M shared-memory Vulkan system. They must be tested against real tasks; throughput alone is not enough.

PriorityModelQuantizationApprox. sizePrimary role
1Qwen2.5-Coder-7B-InstructQ4_K_M4.4 GiBCoding default
2Qwen3.8-4BQ6_K3.4 GiBAgentic/orchestrator default
3Phi-4-mini-InstructQ4_K_M2.4 GiBSystems administration
4Phi-4-mini-ReasoningQ4_K_M2.4 GiBReasoning reviewer
5NVIDIA-Nemotron-Nano-9B-v2Q4_K_M6.1 GiBAgentic/tool-use candidate
6Qwen3-8BQ4_K_M4.7 GiBGeneral technical assistant
7Qwen2.5-Coder-14B-InstructQ4_K_M8.4 GiBDeep coding
8Nous Hermes 4 14BQ4_K_M8.4 GiBAgentic escalation
9NousCoder 14BQ4_K_M8.4 GiBDeep coding alternative
10DeepSeek-R1-Distill-Qwen-14BQ4_K_M8.4 GiBHard reasoning
11Qwen2.5-14B-InstructQ4_K_M8.4 GiBFinancial planning

Initial paired benchmark shortlist

Benchmark these first:

Qwen2.5-Coder-7B-Instruct-Q4_K_M
Qwen3.8-4B-Q6_K
Phi-4-mini-Instruct-Q4_K_M
Phi-4-mini-Reasoning-Q4_K_M
NVIDIA-Nemotron-Nano-9B-v2-Q4_K_M
Qwen3-8B-Q4_K_M

Treat 14B models as on-demand escalation candidates. They may fit in memory, but leave less headroom for large contexts, concurrent requests, and host services.

Do not benchmark standalone projector/draft artifacts as chat models:

  • mmproj-*.gguf
  • mtp-*.gguf
  • *-mmproj*.gguf

Evaluation criteria

Rank candidates using real work as well as throughput:

  • Correctness and explicit uncertainty.
  • Safe handling of destructive infrastructure actions.
  • Valid Bash, Python, YAML, Ansible, Docker, and GitLab CI.
  • Structured output and tool-use reliability where applicable.
  • Quality at the intended context size and concurrency.
  • Stable prompt-processing and text-generation throughput.

Representative tasks include systemd diagnosis from logs, idempotent Ansible repair, shell-script safety review, Proxmox/Ceph troubleshooting, GitLab CI drafting, code review, and financial scenario comparison with explicit assumptions.

Safety

  • Prefer observation, a bounded plan, validation, and explicit approval before infrastructure mutation.
  • Use Ansible and GitLab CI/CD as the routine mutation and audit path.
  • Do not commit credentials, private keys, tokens, GGUF files, or raw logs containing sensitive content.
  • Financial-planning prompts support education and scenario analysis only; they are not fiduciary, tax, legal, accounting, lending, or investment advice.
  • Verify time-sensitive technical, legal, tax, market, and product facts using current authoritative sources before acting.

AI-engine configuration

This repository stores source configuration and benchmark tooling for llama.cpp AI engines. It does not store GGUF weights, credentials, runtime binaries, caches, logs, or generated service state.

A Git preset change does not alter a running llama-server. Host changes require separate validation and an explicit deployment procedure.

Documentation

  • Architecture — repository and benchmark design.
  • Operations — normal operational procedures.
  • Troubleshooting — diagnostics and recovery guidance.
  • OpenClaw integration — consuming benchmark and model-selection data from OpenClaw.
  • Validated GLaDOS example — known-good GLaDOS benchmark example.
  • AI engines — host-specific inference-engine runtime contracts and controlled configuration procedures.
  • GLaDOS AI engine — GLaDOS llama.cpp router, host-local preset contract, and linked virtualization inventory.
  • Model presets — version-controlled llama.cpp preset source policy and profile conventions.

phillf/llama-bench

Shell

0

9 commits

updated Oct 5, 2026

See the code

See what people are saying

SourceMessageScoreDate

I built a reproducible llama.cpp benchmark harness for AMD Vulkan and NVIDIA CUDA — GitLab CI automation is next (r/LocalLLM)

I have been trying to make local llama.cpp benchmark results easier to reproduce across two pretty different setups: \- AMDGPU + Vulkan \- NVIDIA + CUDA The recurring issue for me was that a CSV with tokens/sec did not provide enough context to compare someone else’s result—or even to reproduce my…

1

Oct 5, 2026

README

llama-bench

Local, reproducible llama.cpp benchmarking and role-prompt assets for GLaDOS.

This repository contains the benchmark wrapper, documentation, role prompts, and compact normalized benchmark CSV results. It intentionally excludes GGUF model weights, llama.cpp build trees, runtime binaries, credentials, and verbose benchmark sidecar logs.

GLaDOS profile

ComponentConfiguration
HostGLaDOS
CPUAMD Ryzen 7 8745H — 8 cores / 16 threads
GPUAMD Radeon 780M integrated GPU
GPU backendMesa RADV / Vulkan
llama.cpp deviceVulkan0
Usable memoryAbout 29 GiB shared system memory
Active model storagePVE CephFS at /mnt/pve/llama_models
Expanded model library/mnt/pve/AcornStash-LLMModels

The Radeon 780M shares system memory with the host. Model selection must leave room for the operating system, context/KV cache, model runtime overhead, and other services.

Repository layout

.
├── README.md
├── scripts/
│   └── run-llama-bench.sh
└── docs/
    ├── prompts/
    │   ├── README.md
    │   ├── agentic-orchestrator.md
    │   ├── coding-engineer.md
    │   ├── evaluator-red-team.md
    │   ├── financial-planning.md
    │   └── systems-administration.md
    └── stats/

Benchmark policy

Benchmarking is local and offline:

  • Use llama-bench directly; do not benchmark through a network service.
  • Do not bind, expose, or connect to a network port.
  • Supply an explicit GGUF path.
  • Keep CPU and GPU runs serial.
  • For paired comparisons, run:
CPU round 1 → Vulkan round 1 → CPU round 2 → Vulkan round 2 → CPU round 3 → Vulkan round 3
  • Keep CPU, Vulkan, CUDA, different quantizations, different contexts, and different llama.cpp builds in separate comparison groups.
  • Record model checksum, runtime commit/build, backend/device, thread count, GPU layers, prompt tokens, generation tokens, and repetitions.

Wrapper defaults

llama-bench:
/mnt/pve/llama_models/llama.cpp/build/bin/llama-bench

working directory:
/mnt/pve/llama_models/bin/llama

default operational output directory:
/mnt/pve/llama_models/stats

The repository can store reviewable compact benchmark CSVs under docs/stats/. Raw and stderr sidecar logs remain ignored.

Role prompts

The documents in docs/prompts/ are application-level system prompts. They are not model-specific GGUF chat templates and must not contain model control tokens.

PromptPurposeSuggested model class
agentic-orchestrator.mdBounded planning and coordinationQwen 3.8 4B or tool-use candidate
coding-engineer.mdCode, tests, refactors, and reviewQwen2.5-Coder 7B or 14B
systems-administration.mdRead-first infrastructure operationsPhi-4-mini-Instruct or Qwen 8B
financial-planning.mdBudgeting and scenario analysisQwen2.5 14B Instruct
evaluator-red-team.mdCritical independent reviewPhi-4-mini or reasoning reviewer

Use each model’s embedded chat template through llama.cpp, then provide the selected role prompt as the system message. Prompts improve consistency but do not make model output authoritative.

Suggested models

These are candidates for the Ryzen 7 8745H / Radeon 780M shared-memory Vulkan system. They must be tested against real tasks; throughput alone is not enough.

PriorityModelQuantizationApprox. sizePrimary role
1Qwen2.5-Coder-7B-InstructQ4_K_M4.4 GiBCoding default
2Qwen3.8-4BQ6_K3.4 GiBAgentic/orchestrator default
3Phi-4-mini-InstructQ4_K_M2.4 GiBSystems administration
4Phi-4-mini-ReasoningQ4_K_M2.4 GiBReasoning reviewer
5NVIDIA-Nemotron-Nano-9B-v2Q4_K_M6.1 GiBAgentic/tool-use candidate
6Qwen3-8BQ4_K_M4.7 GiBGeneral technical assistant
7Qwen2.5-Coder-14B-InstructQ4_K_M8.4 GiBDeep coding
8Nous Hermes 4 14BQ4_K_M8.4 GiBAgentic escalation
9NousCoder 14BQ4_K_M8.4 GiBDeep coding alternative
10DeepSeek-R1-Distill-Qwen-14BQ4_K_M8.4 GiBHard reasoning
11Qwen2.5-14B-InstructQ4_K_M8.4 GiBFinancial planning

Initial paired benchmark shortlist

Benchmark these first:

Qwen2.5-Coder-7B-Instruct-Q4_K_M
Qwen3.8-4B-Q6_K
Phi-4-mini-Instruct-Q4_K_M
Phi-4-mini-Reasoning-Q4_K_M
NVIDIA-Nemotron-Nano-9B-v2-Q4_K_M
Qwen3-8B-Q4_K_M

Treat 14B models as on-demand escalation candidates. They may fit in memory, but leave less headroom for large contexts, concurrent requests, and host services.

Do not benchmark standalone projector/draft artifacts as chat models:

  • mmproj-*.gguf
  • mtp-*.gguf
  • *-mmproj*.gguf

Evaluation criteria

Rank candidates using real work as well as throughput:

  • Correctness and explicit uncertainty.
  • Safe handling of destructive infrastructure actions.
  • Valid Bash, Python, YAML, Ansible, Docker, and GitLab CI.
  • Structured output and tool-use reliability where applicable.
  • Quality at the intended context size and concurrency.
  • Stable prompt-processing and text-generation throughput.

Representative tasks include systemd diagnosis from logs, idempotent Ansible repair, shell-script safety review, Proxmox/Ceph troubleshooting, GitLab CI drafting, code review, and financial scenario comparison with explicit assumptions.

Safety

  • Prefer observation, a bounded plan, validation, and explicit approval before infrastructure mutation.
  • Use Ansible and GitLab CI/CD as the routine mutation and audit path.
  • Do not commit credentials, private keys, tokens, GGUF files, or raw logs containing sensitive content.
  • Financial-planning prompts support education and scenario analysis only; they are not fiduciary, tax, legal, accounting, lending, or investment advice.
  • Verify time-sensitive technical, legal, tax, market, and product facts using current authoritative sources before acting.

AI-engine configuration

This repository stores source configuration and benchmark tooling for llama.cpp AI engines. It does not store GGUF weights, credentials, runtime binaries, caches, logs, or generated service state.

A Git preset change does not alter a running llama-server. Host changes require separate validation and an explicit deployment procedure.

Documentation

  • Architecture — repository and benchmark design.
  • Operations — normal operational procedures.
  • Troubleshooting — diagnostics and recovery guidance.
  • OpenClaw integration — consuming benchmark and model-selection data from OpenClaw.
  • Validated GLaDOS example — known-good GLaDOS benchmark example.
  • AI engines — host-specific inference-engine runtime contracts and controlled configuration procedures.
  • GLaDOS AI engine — GLaDOS llama.cpp router, host-local preset contract, and linked virtualization inventory.
  • Model presets — version-controlled llama.cpp preset source policy and profile conventions.