AvilaCarlosDev/polaris-local-ai

Local AI stack for the AMD RX 580 8 GB: OpenAI-compatible router with 6 text/vision models, SD 3.5 + SD 1.5 image generation, Whisper ASR and an agent engine. Vulkan on Polaris, one install script, measured benchmarks. MIT. Inspired by Strata.

Python

1

15 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

I run 6 LLMs + SD 3.5 + Whisper behind an OpenAI-compatible API on an RX 580 8 GB - the card ROCm forgot. Benchmarks inside. (r/LocalLLaMA)

ROCm dropped Polaris years ago, so every RX 580 owner gets told to buy a new card. I went the Vulkan route instead: the whole stack runs on Mesa's RADV. What's running (one Debian LXC, one setup script): \- OpenAI-compatible router, one API key, swaps models on demand \- qwen2.5-coder 1.5B/7B,…

1

Oct 6, 2026

README

polaris-local-ai

English · Español

CI License: MIT Release Runs on AMD RX 580 8 GB

A full local AI stack — text, vision and image generation — on an AMD Radeon RX 580 2048SP (8 GB VRAM) with 32 GB of RAM. No cloud, no API bills. It also serves as the inference engine for an AI coding/assistant agent (Hermes Agent).

The stack listing its models by category, answering a chat request in 6.1 s and starting an image generation with SD 3.5
full video (29 s) — one real run: 57 s wall clock, played at 2×

FLOATING ISLAND MIRAGE: a voxel diorama with a floating island, a giant straw hat, palms and a sailboat, orbited by the camera
FLOATING ISLAND MIRAGE (29 s) — 870 frames rendered and encoded on the local machine, title generated by the local qwen2.5-7b

Inspired by Strata. Strata proved that a serious local inference stack can be packaged so a normal person can run it on a normal PC. This repository follows that same idea — one install script, an OpenAI-compatible API on localhost, docs with real numbers — but for older hardware and with a different engine. No Strata code is included. Credits to its author for the concept and the bar it set. See NOTICE for the full attribution notice.


Why this exists

Strata targets 12 GB+ cards from the RX 6800 series upward, and its AMD path requires ROCm 7, which dropped the Polaris architecture years ago. An RX 580 — still a very common card — is out of scope for it.

This stack does not care. It runs over Vulkan through Mesa's RADV driver, which supports Polaris perfectly, and leans on 32 GB of system RAM for the models that do not fit in 8 GB of VRAM. The result: six text/vision models and two image models, all reachable through a single OpenAI-compatible endpoint.

What you get

Text6 models, from a 1.5 B coder up to a 30 B MoE, swapped on demand
VisionQwen2.5-VL 3 B with its mmproj projector
ImagesSD 3.5 Medium and SD 1.5 through stable-diffusion.cpp
APIOne OpenAI-compatible endpoint, one API key, 6 text/vision ids + image generation
Selectoria-models lists every id the router serves, grouped by category
AgentWorks as the backend engine for Hermes Agent, with MCP tools
FootprintRuns in a Debian LXC on a Proxmox host, or on bare metal

Measured numbers, the install steps, the model guide and everything that broke along the way are in docs/:

What it actually feels like

Measured on this machine, not estimated. Same prompt, same model, two states:

SituationWall time
Warm API call, model already loaded3.1 – 14.1 s
Cold API call, right after a model swap11.0 – 55.0 s
First message to an agent, new session~161 s
Every later message in that same session~21 – 25 s

Row 1 vs row 2 is the swap — 7.9–38.5 s to restart llama-server with new weights. Generation speed is identical in both states (21–99 tok/s); a cold model is not slower, it is slow to arrive. Keep one model loaded for the whole session and you live in row 1.

Row 2 vs row 3 is the agent, not the GPU. Hermes sends a ~20,100-token prompt (55 KB of identity plus 27 tool schemas), so the model spends ~124 s reading your sentence before answering it in ~2.5 s. Prompt caching collapses that to 25–71 tokens on turns 2+, and ~11.7 s of every turn is the agent CLI restarting per invocation.

Full breakdown: BENCHMARKS.md · What to expect and how to use it.

Quick start

./setup.sh

The script checks your GPU, Vulkan driver, RAM and disk, builds what is missing, installs the systemd units and starts the router. It refuses to proceed if your hardware cannot run the stack — see HARDWARE.md for the exact requirements.

Then:

export IA_API_KEY=your-key
curl http://127.0.0.1:8090/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $IA_API_KEY" \
  -d '{"model":"qwen2.5-7b-instruct-q4_k_m",
       "messages":[{"role":"user","content":"Hello"}]}'

Picking a model

Ask the router what it serves instead of guessing from the docs:

ia-models               # every id, grouped by category
ia-models -c vision     # one category only
ia-models --json        # raw list for scripts
CategoryIdsReach for it when
texto1.5 B coder → 7 B, plus ornith 9 Bchat, summaries, code
visionQwen2.5-VL 3 Bthe job involves an image
multitareaQwen3-30B-A3B (3 B active)long reasoning, agent work
imagenSD 3.5 Medium, SD 1.5you want a picture
audiowhisper mediumtranscription — its own service, listed only when installed

The categories come from the endpoint itself (GET /v1/models), so the list stays truthful: an * marks the model loaded right now, and nothing in the table can drift away from what the router actually runs. Model-by-model trade-offs are in MODELS.md.

Repository layout

router.py              OpenAI-compatible router: swaps models on demand
image-mcp.py           Image generation bridge
clients/ia-imagen      CLI to generate an image and save it locally
clients/ia-models      CLI listing every model, grouped by category
scripts/hero-demo.sh   The script behind the video above — run it yourself
systemd/               The units that run the stack
docs/                  Hardware notes, benchmarks, setup, gotchas
setup.sh               One-shot installer

Credits and licence

  • Strata — MIT, by its author. The inspiration for this project's scope and packaging. Not a fork; no Strata source code is present here.
  • llama.cpp — the text and vision inference engine.
  • stable-diffusion.cpp — the image engine.
  • Qwen, Ornith, ISTA-DASLab — model families used here; each carries its own licence.

This repository is MIT-licensed. Model weights are not included and keep their own licences.

ai-agent
amd-gpu
image-generation
llama-cpp
llm
local-ai
local-llm
openai-api
openai-compatible
self-hosted
stable-diffusion
vulkan
whisper

AvilaCarlosDev/polaris-local-ai

Local AI stack for the AMD RX 580 8 GB: OpenAI-compatible router with 6 text/vision models, SD 3.5 + SD 1.5 image generation, Whisper ASR and an agent engine. Vulkan on Polaris, one install script, measured benchmarks. MIT. Inspired by Strata.

Python

1

15 commits

updated Oct 6, 2026

See the code

See what people are saying

SourceMessageScoreDate

I run 6 LLMs + SD 3.5 + Whisper behind an OpenAI-compatible API on an RX 580 8 GB - the card ROCm forgot. Benchmarks inside. (r/LocalLLaMA)

ROCm dropped Polaris years ago, so every RX 580 owner gets told to buy a new card. I went the Vulkan route instead: the whole stack runs on Mesa's RADV. What's running (one Debian LXC, one setup script): \- OpenAI-compatible router, one API key, swaps models on demand \- qwen2.5-coder 1.5B/7B,…

1

Oct 6, 2026

README

polaris-local-ai

English · Español

CI License: MIT Release Runs on AMD RX 580 8 GB

A full local AI stack — text, vision and image generation — on an AMD Radeon RX 580 2048SP (8 GB VRAM) with 32 GB of RAM. No cloud, no API bills. It also serves as the inference engine for an AI coding/assistant agent (Hermes Agent).

The stack listing its models by category, answering a chat request in 6.1 s and starting an image generation with SD 3.5
full video (29 s) — one real run: 57 s wall clock, played at 2×

FLOATING ISLAND MIRAGE: a voxel diorama with a floating island, a giant straw hat, palms and a sailboat, orbited by the camera
FLOATING ISLAND MIRAGE (29 s) — 870 frames rendered and encoded on the local machine, title generated by the local qwen2.5-7b

Inspired by Strata. Strata proved that a serious local inference stack can be packaged so a normal person can run it on a normal PC. This repository follows that same idea — one install script, an OpenAI-compatible API on localhost, docs with real numbers — but for older hardware and with a different engine. No Strata code is included. Credits to its author for the concept and the bar it set. See NOTICE for the full attribution notice.


Why this exists

Strata targets 12 GB+ cards from the RX 6800 series upward, and its AMD path requires ROCm 7, which dropped the Polaris architecture years ago. An RX 580 — still a very common card — is out of scope for it.

This stack does not care. It runs over Vulkan through Mesa's RADV driver, which supports Polaris perfectly, and leans on 32 GB of system RAM for the models that do not fit in 8 GB of VRAM. The result: six text/vision models and two image models, all reachable through a single OpenAI-compatible endpoint.

What you get

Text6 models, from a 1.5 B coder up to a 30 B MoE, swapped on demand
VisionQwen2.5-VL 3 B with its mmproj projector
ImagesSD 3.5 Medium and SD 1.5 through stable-diffusion.cpp
APIOne OpenAI-compatible endpoint, one API key, 6 text/vision ids + image generation
Selectoria-models lists every id the router serves, grouped by category
AgentWorks as the backend engine for Hermes Agent, with MCP tools
FootprintRuns in a Debian LXC on a Proxmox host, or on bare metal

Measured numbers, the install steps, the model guide and everything that broke along the way are in docs/:

What it actually feels like

Measured on this machine, not estimated. Same prompt, same model, two states:

SituationWall time
Warm API call, model already loaded3.1 – 14.1 s
Cold API call, right after a model swap11.0 – 55.0 s
First message to an agent, new session~161 s
Every later message in that same session~21 – 25 s

Row 1 vs row 2 is the swap — 7.9–38.5 s to restart llama-server with new weights. Generation speed is identical in both states (21–99 tok/s); a cold model is not slower, it is slow to arrive. Keep one model loaded for the whole session and you live in row 1.

Row 2 vs row 3 is the agent, not the GPU. Hermes sends a ~20,100-token prompt (55 KB of identity plus 27 tool schemas), so the model spends ~124 s reading your sentence before answering it in ~2.5 s. Prompt caching collapses that to 25–71 tokens on turns 2+, and ~11.7 s of every turn is the agent CLI restarting per invocation.

Full breakdown: BENCHMARKS.md · What to expect and how to use it.

Quick start

./setup.sh

The script checks your GPU, Vulkan driver, RAM and disk, builds what is missing, installs the systemd units and starts the router. It refuses to proceed if your hardware cannot run the stack — see HARDWARE.md for the exact requirements.

Then:

export IA_API_KEY=your-key
curl http://127.0.0.1:8090/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $IA_API_KEY" \
  -d '{"model":"qwen2.5-7b-instruct-q4_k_m",
       "messages":[{"role":"user","content":"Hello"}]}'

Picking a model

Ask the router what it serves instead of guessing from the docs:

ia-models               # every id, grouped by category
ia-models -c vision     # one category only
ia-models --json        # raw list for scripts
CategoryIdsReach for it when
texto1.5 B coder → 7 B, plus ornith 9 Bchat, summaries, code
visionQwen2.5-VL 3 Bthe job involves an image
multitareaQwen3-30B-A3B (3 B active)long reasoning, agent work
imagenSD 3.5 Medium, SD 1.5you want a picture
audiowhisper mediumtranscription — its own service, listed only when installed

The categories come from the endpoint itself (GET /v1/models), so the list stays truthful: an * marks the model loaded right now, and nothing in the table can drift away from what the router actually runs. Model-by-model trade-offs are in MODELS.md.

Repository layout

router.py              OpenAI-compatible router: swaps models on demand
image-mcp.py           Image generation bridge
clients/ia-imagen      CLI to generate an image and save it locally
clients/ia-models      CLI listing every model, grouped by category
scripts/hero-demo.sh   The script behind the video above — run it yourself
systemd/               The units that run the stack
docs/                  Hardware notes, benchmarks, setup, gotchas
setup.sh               One-shot installer

Credits and licence

  • Strata — MIT, by its author. The inspiration for this project's scope and packaging. Not a fork; no Strata source code is present here.
  • llama.cpp — the text and vision inference engine.
  • stable-diffusion.cpp — the image engine.
  • Qwen, Ornith, ISTA-DASLab — model families used here; each carries its own licence.

This repository is MIT-licensed. Model weights are not included and keep their own licences.

ai-agent
amd-gpu
image-generation
llama-cpp
llm
local-ai
local-llm
openai-api
openai-compatible
self-hosted
stable-diffusion
vulkan
whisper