Hitheshkaranth/Ornith-1.5_A3B_Model_DGX_Spark_Setup

Ornith-1.5-35B-A3B (NVFP4 + FP8 + MTP) on NVIDIA DGX Spark (GB10)

Python

3

5 commits

updated Sep 26, 2026

See the code

See what people are saying

SourceMessageScoreDate

Self-hosting a 35B MoE for a 20-person team on one desktop box: what it took and what it does (all open source) (r/LocalLLM)

We've been running our own LLM for the team for about a week, and I wanted to share how it turned out, since "can one machine actually serve a whole team?" comes up here a lot. Short answer: yes, if you pick the right model. **Hardware and model** * One NVIDIA DGX Spark (128GB unified memory) *…

12

Oct 6, 2026

README

NVIDIA DGX Spark

Ornith-1.5-35B-A3B (NVFP4 + FP8 + MTP) on NVIDIA DGX Spark (GB10)

Serve an official 4-bit MoE with multi-token prediction and vision: 20 concurrent users, 262K context, 490 tok/s peak aggregate

Model Quant vLLM Hardware License Context


Overview

This repository contains the pipeline used to run ornith-ai/Ornith-1.5-35B-A3B-NVFP4 on a single NVIDIA DGX Spark (GB10, Blackwell, 128GB unified memory), replacing an earlier Qwen3.8-35B-A3B deployment on the same box. Ornith-1.5 is a Qwen3_5Moe fine-tune (35B total / 3B active MoE) whose self-reported scores beat Qwen3.6-35B-A3B by roughly 15 points on Terminal-Bench 2.1 and 10 points on SWE-bench Pro.

Unlike Qwen3.8, this checkpoint is officially pre-quantized by ornith-ai (ModelOpt MIXED_PRECISION: W4A16 NVFP4 routed + shared experts, FP8 attention, FP8 KV) — there is no local quantization step, just download and serve. It also ships two capabilities the previous Qwen3.8 deployment didn't have:

  • Multi-token prediction (MTP) speculative decoding — the checkpoint carries its own mtp.* draft-model weights, so vLLM can self-speculate with no separate draft model.
  • A real vision tower (333 model.visual.* weights) — this is a genuine vision-language checkpoint, not a text-only distill, so it accepts image input.

Every number in this README was measured on the deployment itself. See Benchmarks for methodology and raw data in benchmarks/.

Table of Contents


Architecture

Build pipeline

flowchart LR
    HF(["Hugging Face\nornith-ai/Ornith-1.5-35B-A3B-NVFP4\nNVFP4 + FP8 · 21.8 GiB"])
    DL["download.sh\nhf CLI · plain HTTP"]
    CK[("Pre-quantized checkpoint\nNVFP4 experts · FP8 attn/KV\nmtp.* draft weights included")]
    RUN["run.sh\nvLLM 0.24.0 · :8004\nMTP + vision enabled"]

    HF ==> DL ==> CK ==> RUN

    classDef source fill:#5794F2,stroke:#2D5FA3,color:#ffffff,stroke-width:2px
    classDef step fill:#FF9830,stroke:#C46F1F,color:#1a1a1a,stroke-width:2px
    classDef artifact fill:#73BF69,stroke:#3F7A39,color:#0a1f08,stroke-width:2px

    class HF source
    class DL,RUN step
    class CK artifact

No quantization step and no accuracy-verification gate are needed here — unlike the Qwen3.8 deployment, ornith-ai ships the NVFP4/FP8 checkpoint directly.

Request flow

flowchart LR
    Client(["Client\ncurl · OpenAI SDK · Open WebUI"])
    API["API Server :8004\nOpenAI-compatible"]
    Sched["Scheduler\nchunked prefill · prefix cache\nmax-num-seqs = 20"]
    Draft["MTP Drafter\n1 layer · shares embed/lm_head"]
    Engine["Engine Core\nNVFP4 MoE (auto) · FP8 attention · vision tower"]
    GPU(["NVIDIA GB10 GPU\n21.8 GiB weights · FP8 KV cache"])
    Prom[("Prometheus")]
    Graf["Grafana"]

    Client == "POST /v1/chat/completions" ==> API
    API ==> Sched ==> Engine
    Engine <-. "draft tokens" .-> Draft
    Engine ==> GPU
    GPU -. tokens .-> Engine -. "SSE stream" .-> API
    API == "streamed response" ==> Client
    API -. "/metrics" .-> Prom ==> Graf

    classDef client fill:#5794F2,stroke:#2D5FA3,color:#ffffff,stroke-width:2px
    classDef serving fill:#FF9830,stroke:#C46F1F,color:#1a1a1a,stroke-width:2px
    classDef hardware fill:#73BF69,stroke:#3F7A39,color:#0a1f08,stroke-width:2px
    classDef obs fill:#8E8E93,stroke:#5A5A5E,color:#ffffff,stroke-width:2px

    class Client client
    class API,Sched,Engine serving
    class Draft serving
    class GPU hardware
    class Prom,Graf obs

Hardware & Software Requirements

ComponentRequirementNotes
GPU1x NVIDIA Blackwell GPU (SM ≥ 12.0)Tested on DGX Spark (GB10), compute capability 12.1
Memory≥ 32GB GPU-addressable memoryWeights take 21.8 GiB; GB10's 128GB is unified CPU+GPU memory
DriverNVIDIA driver with CUDA 13.x support580.173.02 used here
OSLinux (aarch64 or x86_64)Ubuntu 24.04.4, kernel 6.17.0-1031-nvidia
DockerDocker Engine + NVIDIA Container Toolkit--gpus all must work
NGCPull access to nvcr.io/nvidia/vllm:26.07-py3docker login nvcr.io if needed
Disk~25GB free21.8 GiB NVFP4 checkpoint, no BF16 source needed
NetworkOutbound HTTPS to huggingface.co~22 GB download on first run

Quick Start

One line, on any single-Blackwell-GPU box with Docker and the NVIDIA Container Toolkit already set up:

curl -fsSL https://raw.githubusercontent.com/Hitheshkaranth/Ornith-1.5_A3B_Model_DGX_Spark_Setup/main/install.sh | bash

Prefer to clone it yourself first? Same result:

git clone https://github.com/Hitheshkaranth/Ornith-1.5_A3B_Model_DGX_Spark_Setup.git && cd Ornith-1.5_A3B_Model_DGX_Spark_Setup
./download.sh && ./run.sh

[!WARNING] On GB10, run one 0.70-utilization vLLM server at a time. If another model is already serving, stop it first (docker stop <container>).

Configuration Reference

FlagValueWhy
--gpu-memory-utilization0.70Hard ceiling on this unified-memory machine; higher values made concurrency collapse under load.
--max-model-len262144The model's native maximum context.
--max-num-seqs20 (was 16, matching the previous Qwen3.8 deployment)Raised after the KV cache logged Maximum concurrency for 262,144 tokens per request: 19.75x, showing headroom for 20 concurrent sequences even at full context.
--max-num-batched-tokens8192Keeps prefill batching efficient under concurrent load.
--enable-prefix-cachingonReuses KV for shared system prompts and multi-turn chats.
--kv-cache-dtypefp8Halves KV memory versus BF16.
--speculative-config{"model": "<same checkpoint>", "num_speculative_tokens": 1}MTP speculative decoding. The checkpoint ships mtp.* weights (mtp_num_hidden_layers: 1, 785 tensors); vLLM auto-detects the qwen3_5_mtp draft architecture from model_type=qwen3_5_moe and shares the target model's embedding/lm_head with the draft. Observed ~86-88% draft acceptance rate in production.
--moe-backend(unset → auto)Deliberately not forced to marlin. The MTP drafter's expert layers are unquantized BF16, and Marlin only supports quantized MoE — forcing marlin crashed the drafter on load (moe_backend='marlin' is not supported for unquantized MoE). auto still picks marlin for the main model's NVFP4 experts on this GPU (no native FP4 support).
--language-model-only(unset)Deliberately not set. This checkpoint has a real vision tower (333 model.visual.* weights + vision_config); setting this flag (copied from the text-only Qwen3.8 config) silently disabled all image input (At most 0 image(s) may be provided).
--served-model-nameornith-1.5-35b-a3bThe name clients put in "model".
--reasoning-parserqwen3Moves the <think> block into message.reasoning instead of the answer.
--enable-auto-tool-choice / --tool-call-parserqwen3_coderOpenAI-style function calling.

Recommended sampling: temperature=0.6, top_p=0.95, top_k=20. Allow a generous max_tokens (16,384+) since every answer starts with a reasoning block.

Hosting on This System

This is how the model is hosted on the DGX Spark, behind the same metering gateway used by other deployments on this box. Full setup, Tailscale keys and gateway operations are in docs/GATEWAY.md.

Services on the box

PortContainerWhat it is
8000vllmQwen3.6-35B-A3B-NVFP4 (stopped; kept for reference/rollback)
8002vllm-qwen38-a3bQwen3.8-35B-A3B (stopped 2026-09-26, replaced by this deployment)
8004vllm-ornith-a3bThis server: Ornith-1.5-35B-A3B, model name ornith-1.5-35b-a3b (localhost + docker bridge only)
8080llm-gateway (systemd)Metering gateway: the only client entry point; per-user/per-client token metrics
3000open-webui (systemd)Open WebUI, default model set to ornith-1.5-35b-a3b
3001vllm-grafanaGrafana dashboards
9090vllm-prometheusPrometheus, scraping :8004 (and other vLLM ports) in the vllm job
9400 / 9100vllm-dcgm-exporter / vllm-node-exporterGPU / host metrics

Clients don't call vLLM directly. It's published only on 127.0.0.1 and the docker bridge (BIND_ADDRS="127.0.0.1 172.17.0.1" ./run.sh, the default), and every client goes through the metering gateway at :8080. The gateway polls /v1/models on each backend and routes by model name, so swapping models needs no gateway config change beyond adding the new port once — already done here (GATEWAY_BACKENDS includes :8004; Prometheus's vllm job includes it too).

Switching models

run.sh defaults to --restart unless-stopped, since this is the active production model on this box — it needs to survive a reboot or crash without manual intervention. Only one 0.70-utilization container should carry that policy at a time, so swap it explicitly:

# Ornith -> something else
docker update --restart no vllm-ornith-a3b && docker stop vllm-ornith-a3b
docker update --restart unless-stopped <other-container> && docker start <other-container>

# something else -> Ornith
docker update --restart no <other-container> && docker stop <other-container>
cd Ornith-1.5_A3B_Model_DGX_Spark_Setup && ./run.sh   # ready in ~5 min (weights + MTP drafter + vision tower + compile)

For a one-off test run that shouldn't affect the production restart policy, override it: RESTART=no ./run.sh.

Open WebUI default model

Open WebUI persists its default/pinned model in its own sqlite config table (ui.default_models, ui.default_pinned_models, ui.model_order_list), which overrides the DEFAULT_MODELS environment variable once set. Swapping the default model requires updating both the systemd unit's env var (for a fresh install) and the existing database (for an already-deployed instance) — see Troubleshooting.

Metering Gateway & Tailscale

Every client, whether Open WebUI, opencode or scripts on laptops across the tailnet, reaches the model through a small metering gateway on :8080. It attributes each request to a user, device or app, and counts tokens live. vLLM itself is locked so nothing can go around it.

flowchart LR
    D(["Tailnet devices
one key each"]) == ":8080" ==> GW["llm-gateway
auth · routing · live metering"]
    O(["Open WebUI
forwards user email"]) ==> GW
    GW == "upstream key" ==> V["vLLM :8004
localhost + docker bridge
key required"]
    GW -.-> P[("Prometheus → Grafana
per-user · per-device · bypass alert")]
    D -. "direct → refused" .-x V
    classDef c fill:#5794F2,stroke:#2D5FA3,color:#fff,stroke-width:2px
    classDef g fill:#FF9830,stroke:#C46F1F,color:#1a1a1a,stroke-width:2px
    classDef m fill:#73BF69,stroke:#3F7A39,color:#0a1f08,stroke-width:2px
    classDef o fill:#8E8E93,stroke:#5A5A5E,color:#fff,stroke-width:2px
    class D,O c
    class GW g
    class V m
    class P o
./gateway/setup.sh                                                        # 1. install the gateway (systemd user service, :8080)
VLLM_API_KEY=$(cat gateway/upstream.key) BIND_ADDRS="127.0.0.1 172.17.0.1" ./run.sh   # 2. lock vLLM behind it
cd gateway && venv/bin/python tailscale_keys.py && systemctl --user restart llm-gateway  # 3. one key per tailnet device

Clients then use http://<server-tailscale-ip-or-name>:8080/v1 with their own key.

WhatWhere
Full setup, client configs (opencode, Open WebUI, SDK), Tailscale (per-device keys, tailnet-only binding, HTTPS via tailscale serve, ACLs), operations, troubleshootingdocs/GATEWAY.md
Gateway code and helpersgateway/
Prometheus scrape snippet, Grafana dashboard, bypass alertmonitoring/

Input tokens are counted when a request starts (exact count via vLLM /tokenize). Output tokens are counted as they stream, or on completion for non-streaming calls. Keys, the upstream secret and the usage database are generated locally and never committed. This gateway is shared with the Qwen3.8 deployment on the same box — see What's in gateway/ for why the code is published in both repos rather than referenced from one.

Vision & PDF Support

Unlike the text-only Qwen3.8 deployment, this model accepts image input directly:

curl http://localhost:8080/v1/chat/completions \
  -H "Authorization: Bearer $GATEWAY_KEY" -H "Content-Type: application/json" \
  -d '{"model":"ornith-1.5-35b-a3b","messages":[{"role":"user","content":[
        {"type":"text","text":"What is in this image?"},
        {"type":"image_url","image_url":{"url":"data:image/png;base64,<...>"}}
      ]}]}'

Confirmed working end to end with a real test image through the gateway.

PDFs work too, through Open WebUI's existing pipeline — not through the model's API directly (vLLM's OpenAI-compatible endpoint has no native PDF/file content type, only text and image_url). Open WebUI extracts PDF text locally with pypdf, embeds it with a local sentence-transformers model, and retrieves the top-k relevant chunks into the prompt. Verified end to end by running a real PDF through Open WebUI's own extraction code path. Embedded images inside PDFs are not currently extracted (rag.pdf_extract_images is off — enabling it needs an additional rapidocr-onnxruntime dependency that isn't installed, and would do OCR on those images rather than route them through this model's vision tower).

Benchmarks

Both models ran on the same box, one at a time, with identical prompts. Scripts and raw results are in benchmarks/ (copied from the Qwen3.8 repo's benchmark suite, with a GATEWAY_KEY auth header added to reasoning_bench.py and throughput_bench.py since vLLM now requires a bearer key on /v1/*).

Throughput & MTP Speculative Decoding

Throughput comparison

MetricQwen3.8-35B-A3B (no MTP, max-num-seqs 16)Ornith-1.5-35B-A3B (this repo, MTP on, max-num-seqs 20)
Decode, 1 user56.1 tok/s79.0 tok/s (+41%)
Peak aggregate decode372.7 tok/s (at its cap of 16 users)490.3 tok/s (at its cap of 20 users, +32%)
Per-user decode, at peak23.3 tok/s24.9 tok/s
Checkpoint size20.23 GiB21.81 GiB
KV cache (fp8)6.05M tokens5.29M tokens
Max concurrency at 262K context23.1x20.2x
Cold start to ready258 s290 s

Peak token output was found by sweeping concurrency past each server's own --max-num-seqs (1 → 16 → 20 → 24 for Ornith). Aggregate throughput peaks exactly at the configured cap — 490.3 tok/s at 20 concurrent requests — then drops to 405.7 tok/s at 24, because the extra requests queue (Waiting > 0) instead of decoding concurrently, adding scheduling overhead for no throughput gain. This confirms --max-num-seqs 20 (raised from Qwen3.8's 16 after the KV cache showed headroom — see Configuration Reference) is set at the right point, not left on the table.

MTP gives a large single-stream win (+41% at 1 user) since there's a full model's worth of spare decode bandwidth to spend on speculation, and the higher --max-num-seqs gives Ornith a 32% higher peak aggregate ceiling than Qwen3.8's, despite MTP's extra per-step drafter overhead. In live production traffic, the SpecDecoding metrics show a consistent ~86-88% per-position draft-acceptance rate (vllm:spec_decode / engine logs), meaning the 1-token draft is right the large majority of the time.

Method: every request generates exactly 512 tokens (ignore_eos), after one warm-up request (throughput_bench.py · throughput_result.json, which has the full 1/16/20/24-user sweep).

Reasoning & Output Length

Measured with the model card's thinking-mode sampling (temperature=0.6, top_p=0.95, top_k=20) and 16 requests in flight:

Reasoning comparison

TestOrnith-1.5-35B-A3B (this repo)Qwen3.8-35B-A3B
GSM8K, first 100 test problems98% (98/100)99% (99/100)
avg / max output tokens384 / 3,630303 / 1,251
hit the 16,384-token limit00
MATH-500, first 100 with integer answers97% (97/100)94% (94/100)
avg / max output tokens1,937 / 16,3842,263 / 16,384
hit the 16,384-token limit36
Long-form, 4 prompts, 32,768-token capavg 17,804 tokensavg 20,723 tokens

Ornith trades a single point on GSM8K for three points on the harder MATH-500 benchmark, with fewer truncated answers. Raw results: reasoning_result_ornith.json · script: reasoning_bench.py.

Knowledge: MMLU-Pro

MMLU-Pro (TIGER-Lab, NeurIPS 2024): 10 options per question, so random guessing scores 10%. Stratified sample of 700 questions (50 per category, fixed seed), zero-shot, thinking on, max_tokens 16,384, 8 requests in flight through the metering gateway.

Knowledge comparison

Ornith-1.5-35B-A3B (this repo)Qwen3.8-35B-A3B
MMLU-Pro accuracy81.1% (568/700, 95% CI ±2.9)80.9% (566/700, 95% CI ±2.9)
unparsed / truncated answers37 (5.3%)11 (1.6%)
avg output tokens2,2731,155
CategoryOrnithQwen3.8CategoryOrnithQwen3.8
Math94%94%Psychology88%82%
Physics94%92%Business84%82%
Biology92%94%Philosophy80%76%
Computer science92%86%Health76%74%
Economics92%90%History68%68%
Chemistry90%86%Other66%68%
Law62%70%
Engineering58%70%

Overall accuracy is essentially tied (81.1% vs 80.9%, well within the ±2.9-point sampling error of either single run), but the two models trade strength by category: Ornith is stronger on computer science and psychology, Qwen3.8 is stronger on law and engineering. Ornith also produced roughly 2x the average output tokens per answer and 3x the unparsed-answer rate, suggesting it reasons longer and occasionally fails to land on the requested ANSWER: <letter> format — a prompt-format issue more than a knowledge gap.

Raw results: knowledge_result_ornith.json (summary) · knowledge_result_ornith.jsonl (per question) · script: knowledge_bench.py. This run took several hours of wall time interleaved with three server restarts (for MTP, vision, and the max-num-seqs change) and live user traffic sharing the same concurrency slots, so the wall-clock figure isn't a speed measure — use the throughput benchmark above for that.

Live Monitoring

The server exposes Prometheus metrics at /metrics, scraped in the same vllm job as the other vLLM ports on this box; Grafana's panels key off the model/model_name labels, so this model's traffic appears automatically with no dashboard changes needed:

scrape_configs:
  - job_name: vllm
    static_configs:
      - targets: ["host.docker.internal:8000", "host.docker.internal:8002", "host.docker.internal:8004"]

If prometheus.yml is bind-mounted as a single file, run docker restart <prometheus> after editing it — a SIGHUP won't pick up a file replaced by sed -i.

Client Usage

from openai import OpenAI

client = OpenAI(api_key="<gateway key>", base_url="http://localhost:8080/v1")

resp = client.chat.completions.create(
    model="ornith-1.5-35b-a3b",
    messages=[{"role": "user", "content": "A snail climbs 3 m a day and slips 2 m a night in a 10 m well. How many days?"}],
    max_tokens=16384,
    temperature=0.6, top_p=0.95,
    extra_body={"top_k": 20},
)
print(resp.choices[0].message.reasoning)   # the <think> trace
print(resp.choices[0].message.content)     # the answer: 8

Repository Layout

Ornith-1.5_A3B_Model_DGX_Spark_Setup/
├── README.md                        # this file
├── install.sh                       # one-line installer (curl | bash)
├── Dockerfile                       # nvcr vLLM 26.07 base + xgrammar patch
├── download.sh                      # fetches the official pre-quantized NVFP4 checkpoint
├── run.sh                           # builds the image and runs the server on :8004
├── gateway/                         # metering gateway (see docs/GATEWAY.md)
│   ├── gateway.py                   #   auth · routing · live token metering · /metrics
│   ├── setup.sh                     #   venv + upstream key + systemd user service
│   ├── add_key.py                   #   issue a key for a user/app
│   ├── tailscale_keys.py            #   one key per tailnet device (+ CSV to hand out)
│   └── keys.example.json            #   key file format (real keys.json is git-ignored)
├── monitoring/
│   ├── prometheus-scrape.yml        # vllm + llm-gateway scrape jobs
│   ├── grafana-dashboard.json       # vLLM Command Center (per-user/per-device panels)
│   └── gateway-bypass-alert.yml     # Grafana alert: traffic bypassing the gateway
├── benchmarks/
│   ├── throughput_bench.py          # decode tok/s at N concurrent users
│   ├── throughput_result.json       # raw results cited above (Ornith vs Qwen3.8, combined)
│   ├── throughput_comparison.png    # chart
│   ├── reasoning_bench.py           # GSM8K / MATH-500 accuracy + long-form length (+ gateway auth)
│   ├── reasoning_result_ornith.json # raw reasoning results cited above
│   ├── reasoning_result_qwen38.json # baseline, copied from the Qwen3.8 repo
│   ├── reasoning_comparison.png     # chart
│   ├── knowledge_bench.py           # MMLU-Pro accuracy (700-question stratified sample)
│   ├── knowledge_result_ornith.json # MMLU-Pro summary + per-category scores
│   ├── knowledge_result_ornith.jsonl # MMLU-Pro per-question records (resume log)
│   ├── knowledge_result_qwen38.json # baseline, copied from the Qwen3.8 repo
│   ├── knowledge_comparison.png     # chart
│   └── make_charts.py               # renders all three PNGs from the JSON results above
├── assets/
│   ├── ornith_logo.png              # from the model card
│   ├── ornith_35b_eval.png          # official eval chart, from the model card
│   └── dgx-spark-banner-new.png     # hardware banner
└── docs/
    └── GATEWAY.md                   # metering gateway + Tailscale: setup, clients, ops

models/ (the downloaded checkpoint, ~22 GiB) is git-ignored and Docker-ignored. gateway/ and monitoring/ are a published copy of the same shared infrastructure documented in the Qwen3.8 repo — see What's in gateway/ for why it's duplicated rather than referenced.

Troubleshooting

SymptomFix
moe_backend='marlin' is not supported for unquantized MoE (on startup, after "Loading drafter model...")The MTP drafter's expert layers are unquantized. Don't pass --moe-backend marlin; leave it unset (auto).
At most 0 image(s) may be provided in one prompt.The server was started with --language-model-only, which strips the vision tower. Remove that flag — this checkpoint has real vision weights.
knowledge_bench.py exits with ModuleNotFoundError: No module named 'pyarrow'Run it via uv run --no-project --with pyarrow python3 benchmarks/knowledge_bench.py ... (its own docstring documents this), not plain python3.
Benchmark scripts get 401 UnauthorizedvLLM requires the gateway's upstream key (VLLM_API_KEY). Pass GATEWAY_KEY=<key> in the environment when running reasoning_bench.py / knowledge_bench.py / throughput_bench.py.
Open WebUI still shows the old default model after changing DEFAULT_MODELSOpen WebUI persists ui.default_models in its sqlite config table, which overrides the env var once set. Stop the service, update ui.default_models / ui.default_pinned_models / ui.model_order_list directly in webui.db, then restart.
CUDA out of memory at startupAnother 0.70-utilization vLLM server is running. Stop it first.
Answers are cut off (finish_reason: length)Raise max_tokens; the reasoning block comes first and can use several thousand tokens.

Credits & License


Ornith logo

Hitheshkaranth/Ornith-1.5_A3B_Model_DGX_Spark_Setup

Ornith-1.5-35B-A3B (NVFP4 + FP8 + MTP) on NVIDIA DGX Spark (GB10)

Python

3

5 commits

updated Sep 26, 2026

See the code

See what people are saying

SourceMessageScoreDate

Self-hosting a 35B MoE for a 20-person team on one desktop box: what it took and what it does (all open source) (r/LocalLLM)

We've been running our own LLM for the team for about a week, and I wanted to share how it turned out, since "can one machine actually serve a whole team?" comes up here a lot. Short answer: yes, if you pick the right model. **Hardware and model** * One NVIDIA DGX Spark (128GB unified memory) *…

12

Oct 6, 2026

README

NVIDIA DGX Spark

Ornith-1.5-35B-A3B (NVFP4 + FP8 + MTP) on NVIDIA DGX Spark (GB10)

Serve an official 4-bit MoE with multi-token prediction and vision: 20 concurrent users, 262K context, 490 tok/s peak aggregate

Model Quant vLLM Hardware License Context


Overview

This repository contains the pipeline used to run ornith-ai/Ornith-1.5-35B-A3B-NVFP4 on a single NVIDIA DGX Spark (GB10, Blackwell, 128GB unified memory), replacing an earlier Qwen3.8-35B-A3B deployment on the same box. Ornith-1.5 is a Qwen3_5Moe fine-tune (35B total / 3B active MoE) whose self-reported scores beat Qwen3.6-35B-A3B by roughly 15 points on Terminal-Bench 2.1 and 10 points on SWE-bench Pro.

Unlike Qwen3.8, this checkpoint is officially pre-quantized by ornith-ai (ModelOpt MIXED_PRECISION: W4A16 NVFP4 routed + shared experts, FP8 attention, FP8 KV) — there is no local quantization step, just download and serve. It also ships two capabilities the previous Qwen3.8 deployment didn't have:

  • Multi-token prediction (MTP) speculative decoding — the checkpoint carries its own mtp.* draft-model weights, so vLLM can self-speculate with no separate draft model.
  • A real vision tower (333 model.visual.* weights) — this is a genuine vision-language checkpoint, not a text-only distill, so it accepts image input.

Every number in this README was measured on the deployment itself. See Benchmarks for methodology and raw data in benchmarks/.

Table of Contents


Architecture

Build pipeline

flowchart LR
    HF(["Hugging Face\nornith-ai/Ornith-1.5-35B-A3B-NVFP4\nNVFP4 + FP8 · 21.8 GiB"])
    DL["download.sh\nhf CLI · plain HTTP"]
    CK[("Pre-quantized checkpoint\nNVFP4 experts · FP8 attn/KV\nmtp.* draft weights included")]
    RUN["run.sh\nvLLM 0.24.0 · :8004\nMTP + vision enabled"]

    HF ==> DL ==> CK ==> RUN

    classDef source fill:#5794F2,stroke:#2D5FA3,color:#ffffff,stroke-width:2px
    classDef step fill:#FF9830,stroke:#C46F1F,color:#1a1a1a,stroke-width:2px
    classDef artifact fill:#73BF69,stroke:#3F7A39,color:#0a1f08,stroke-width:2px

    class HF source
    class DL,RUN step
    class CK artifact

No quantization step and no accuracy-verification gate are needed here — unlike the Qwen3.8 deployment, ornith-ai ships the NVFP4/FP8 checkpoint directly.

Request flow

flowchart LR
    Client(["Client\ncurl · OpenAI SDK · Open WebUI"])
    API["API Server :8004\nOpenAI-compatible"]
    Sched["Scheduler\nchunked prefill · prefix cache\nmax-num-seqs = 20"]
    Draft["MTP Drafter\n1 layer · shares embed/lm_head"]
    Engine["Engine Core\nNVFP4 MoE (auto) · FP8 attention · vision tower"]
    GPU(["NVIDIA GB10 GPU\n21.8 GiB weights · FP8 KV cache"])
    Prom[("Prometheus")]
    Graf["Grafana"]

    Client == "POST /v1/chat/completions" ==> API
    API ==> Sched ==> Engine
    Engine <-. "draft tokens" .-> Draft
    Engine ==> GPU
    GPU -. tokens .-> Engine -. "SSE stream" .-> API
    API == "streamed response" ==> Client
    API -. "/metrics" .-> Prom ==> Graf

    classDef client fill:#5794F2,stroke:#2D5FA3,color:#ffffff,stroke-width:2px
    classDef serving fill:#FF9830,stroke:#C46F1F,color:#1a1a1a,stroke-width:2px
    classDef hardware fill:#73BF69,stroke:#3F7A39,color:#0a1f08,stroke-width:2px
    classDef obs fill:#8E8E93,stroke:#5A5A5E,color:#ffffff,stroke-width:2px

    class Client client
    class API,Sched,Engine serving
    class Draft serving
    class GPU hardware
    class Prom,Graf obs

Hardware & Software Requirements

ComponentRequirementNotes
GPU1x NVIDIA Blackwell GPU (SM ≥ 12.0)Tested on DGX Spark (GB10), compute capability 12.1
Memory≥ 32GB GPU-addressable memoryWeights take 21.8 GiB; GB10's 128GB is unified CPU+GPU memory
DriverNVIDIA driver with CUDA 13.x support580.173.02 used here
OSLinux (aarch64 or x86_64)Ubuntu 24.04.4, kernel 6.17.0-1031-nvidia
DockerDocker Engine + NVIDIA Container Toolkit--gpus all must work
NGCPull access to nvcr.io/nvidia/vllm:26.07-py3docker login nvcr.io if needed
Disk~25GB free21.8 GiB NVFP4 checkpoint, no BF16 source needed
NetworkOutbound HTTPS to huggingface.co~22 GB download on first run

Quick Start

One line, on any single-Blackwell-GPU box with Docker and the NVIDIA Container Toolkit already set up:

curl -fsSL https://raw.githubusercontent.com/Hitheshkaranth/Ornith-1.5_A3B_Model_DGX_Spark_Setup/main/install.sh | bash

Prefer to clone it yourself first? Same result:

git clone https://github.com/Hitheshkaranth/Ornith-1.5_A3B_Model_DGX_Spark_Setup.git && cd Ornith-1.5_A3B_Model_DGX_Spark_Setup
./download.sh && ./run.sh

[!WARNING] On GB10, run one 0.70-utilization vLLM server at a time. If another model is already serving, stop it first (docker stop <container>).

Configuration Reference

FlagValueWhy
--gpu-memory-utilization0.70Hard ceiling on this unified-memory machine; higher values made concurrency collapse under load.
--max-model-len262144The model's native maximum context.
--max-num-seqs20 (was 16, matching the previous Qwen3.8 deployment)Raised after the KV cache logged Maximum concurrency for 262,144 tokens per request: 19.75x, showing headroom for 20 concurrent sequences even at full context.
--max-num-batched-tokens8192Keeps prefill batching efficient under concurrent load.
--enable-prefix-cachingonReuses KV for shared system prompts and multi-turn chats.
--kv-cache-dtypefp8Halves KV memory versus BF16.
--speculative-config{"model": "<same checkpoint>", "num_speculative_tokens": 1}MTP speculative decoding. The checkpoint ships mtp.* weights (mtp_num_hidden_layers: 1, 785 tensors); vLLM auto-detects the qwen3_5_mtp draft architecture from model_type=qwen3_5_moe and shares the target model's embedding/lm_head with the draft. Observed ~86-88% draft acceptance rate in production.
--moe-backend(unset → auto)Deliberately not forced to marlin. The MTP drafter's expert layers are unquantized BF16, and Marlin only supports quantized MoE — forcing marlin crashed the drafter on load (moe_backend='marlin' is not supported for unquantized MoE). auto still picks marlin for the main model's NVFP4 experts on this GPU (no native FP4 support).
--language-model-only(unset)Deliberately not set. This checkpoint has a real vision tower (333 model.visual.* weights + vision_config); setting this flag (copied from the text-only Qwen3.8 config) silently disabled all image input (At most 0 image(s) may be provided).
--served-model-nameornith-1.5-35b-a3bThe name clients put in "model".
--reasoning-parserqwen3Moves the <think> block into message.reasoning instead of the answer.
--enable-auto-tool-choice / --tool-call-parserqwen3_coderOpenAI-style function calling.

Recommended sampling: temperature=0.6, top_p=0.95, top_k=20. Allow a generous max_tokens (16,384+) since every answer starts with a reasoning block.

Hosting on This System

This is how the model is hosted on the DGX Spark, behind the same metering gateway used by other deployments on this box. Full setup, Tailscale keys and gateway operations are in docs/GATEWAY.md.

Services on the box

PortContainerWhat it is
8000vllmQwen3.6-35B-A3B-NVFP4 (stopped; kept for reference/rollback)
8002vllm-qwen38-a3bQwen3.8-35B-A3B (stopped 2026-09-26, replaced by this deployment)
8004vllm-ornith-a3bThis server: Ornith-1.5-35B-A3B, model name ornith-1.5-35b-a3b (localhost + docker bridge only)
8080llm-gateway (systemd)Metering gateway: the only client entry point; per-user/per-client token metrics
3000open-webui (systemd)Open WebUI, default model set to ornith-1.5-35b-a3b
3001vllm-grafanaGrafana dashboards
9090vllm-prometheusPrometheus, scraping :8004 (and other vLLM ports) in the vllm job
9400 / 9100vllm-dcgm-exporter / vllm-node-exporterGPU / host metrics

Clients don't call vLLM directly. It's published only on 127.0.0.1 and the docker bridge (BIND_ADDRS="127.0.0.1 172.17.0.1" ./run.sh, the default), and every client goes through the metering gateway at :8080. The gateway polls /v1/models on each backend and routes by model name, so swapping models needs no gateway config change beyond adding the new port once — already done here (GATEWAY_BACKENDS includes :8004; Prometheus's vllm job includes it too).

Switching models

run.sh defaults to --restart unless-stopped, since this is the active production model on this box — it needs to survive a reboot or crash without manual intervention. Only one 0.70-utilization container should carry that policy at a time, so swap it explicitly:

# Ornith -> something else
docker update --restart no vllm-ornith-a3b && docker stop vllm-ornith-a3b
docker update --restart unless-stopped <other-container> && docker start <other-container>

# something else -> Ornith
docker update --restart no <other-container> && docker stop <other-container>
cd Ornith-1.5_A3B_Model_DGX_Spark_Setup && ./run.sh   # ready in ~5 min (weights + MTP drafter + vision tower + compile)

For a one-off test run that shouldn't affect the production restart policy, override it: RESTART=no ./run.sh.

Open WebUI default model

Open WebUI persists its default/pinned model in its own sqlite config table (ui.default_models, ui.default_pinned_models, ui.model_order_list), which overrides the DEFAULT_MODELS environment variable once set. Swapping the default model requires updating both the systemd unit's env var (for a fresh install) and the existing database (for an already-deployed instance) — see Troubleshooting.

Metering Gateway & Tailscale

Every client, whether Open WebUI, opencode or scripts on laptops across the tailnet, reaches the model through a small metering gateway on :8080. It attributes each request to a user, device or app, and counts tokens live. vLLM itself is locked so nothing can go around it.

flowchart LR
    D(["Tailnet devices
one key each"]) == ":8080" ==> GW["llm-gateway
auth · routing · live metering"]
    O(["Open WebUI
forwards user email"]) ==> GW
    GW == "upstream key" ==> V["vLLM :8004
localhost + docker bridge
key required"]
    GW -.-> P[("Prometheus → Grafana
per-user · per-device · bypass alert")]
    D -. "direct → refused" .-x V
    classDef c fill:#5794F2,stroke:#2D5FA3,color:#fff,stroke-width:2px
    classDef g fill:#FF9830,stroke:#C46F1F,color:#1a1a1a,stroke-width:2px
    classDef m fill:#73BF69,stroke:#3F7A39,color:#0a1f08,stroke-width:2px
    classDef o fill:#8E8E93,stroke:#5A5A5E,color:#fff,stroke-width:2px
    class D,O c
    class GW g
    class V m
    class P o
./gateway/setup.sh                                                        # 1. install the gateway (systemd user service, :8080)
VLLM_API_KEY=$(cat gateway/upstream.key) BIND_ADDRS="127.0.0.1 172.17.0.1" ./run.sh   # 2. lock vLLM behind it
cd gateway && venv/bin/python tailscale_keys.py && systemctl --user restart llm-gateway  # 3. one key per tailnet device

Clients then use http://<server-tailscale-ip-or-name>:8080/v1 with their own key.

WhatWhere
Full setup, client configs (opencode, Open WebUI, SDK), Tailscale (per-device keys, tailnet-only binding, HTTPS via tailscale serve, ACLs), operations, troubleshootingdocs/GATEWAY.md
Gateway code and helpersgateway/
Prometheus scrape snippet, Grafana dashboard, bypass alertmonitoring/

Input tokens are counted when a request starts (exact count via vLLM /tokenize). Output tokens are counted as they stream, or on completion for non-streaming calls. Keys, the upstream secret and the usage database are generated locally and never committed. This gateway is shared with the Qwen3.8 deployment on the same box — see What's in gateway/ for why the code is published in both repos rather than referenced from one.

Vision & PDF Support

Unlike the text-only Qwen3.8 deployment, this model accepts image input directly:

curl http://localhost:8080/v1/chat/completions \
  -H "Authorization: Bearer $GATEWAY_KEY" -H "Content-Type: application/json" \
  -d '{"model":"ornith-1.5-35b-a3b","messages":[{"role":"user","content":[
        {"type":"text","text":"What is in this image?"},
        {"type":"image_url","image_url":{"url":"data:image/png;base64,<...>"}}
      ]}]}'

Confirmed working end to end with a real test image through the gateway.

PDFs work too, through Open WebUI's existing pipeline — not through the model's API directly (vLLM's OpenAI-compatible endpoint has no native PDF/file content type, only text and image_url). Open WebUI extracts PDF text locally with pypdf, embeds it with a local sentence-transformers model, and retrieves the top-k relevant chunks into the prompt. Verified end to end by running a real PDF through Open WebUI's own extraction code path. Embedded images inside PDFs are not currently extracted (rag.pdf_extract_images is off — enabling it needs an additional rapidocr-onnxruntime dependency that isn't installed, and would do OCR on those images rather than route them through this model's vision tower).

Benchmarks

Both models ran on the same box, one at a time, with identical prompts. Scripts and raw results are in benchmarks/ (copied from the Qwen3.8 repo's benchmark suite, with a GATEWAY_KEY auth header added to reasoning_bench.py and throughput_bench.py since vLLM now requires a bearer key on /v1/*).

Throughput & MTP Speculative Decoding

Throughput comparison

MetricQwen3.8-35B-A3B (no MTP, max-num-seqs 16)Ornith-1.5-35B-A3B (this repo, MTP on, max-num-seqs 20)
Decode, 1 user56.1 tok/s79.0 tok/s (+41%)
Peak aggregate decode372.7 tok/s (at its cap of 16 users)490.3 tok/s (at its cap of 20 users, +32%)
Per-user decode, at peak23.3 tok/s24.9 tok/s
Checkpoint size20.23 GiB21.81 GiB
KV cache (fp8)6.05M tokens5.29M tokens
Max concurrency at 262K context23.1x20.2x
Cold start to ready258 s290 s

Peak token output was found by sweeping concurrency past each server's own --max-num-seqs (1 → 16 → 20 → 24 for Ornith). Aggregate throughput peaks exactly at the configured cap — 490.3 tok/s at 20 concurrent requests — then drops to 405.7 tok/s at 24, because the extra requests queue (Waiting > 0) instead of decoding concurrently, adding scheduling overhead for no throughput gain. This confirms --max-num-seqs 20 (raised from Qwen3.8's 16 after the KV cache showed headroom — see Configuration Reference) is set at the right point, not left on the table.

MTP gives a large single-stream win (+41% at 1 user) since there's a full model's worth of spare decode bandwidth to spend on speculation, and the higher --max-num-seqs gives Ornith a 32% higher peak aggregate ceiling than Qwen3.8's, despite MTP's extra per-step drafter overhead. In live production traffic, the SpecDecoding metrics show a consistent ~86-88% per-position draft-acceptance rate (vllm:spec_decode / engine logs), meaning the 1-token draft is right the large majority of the time.

Method: every request generates exactly 512 tokens (ignore_eos), after one warm-up request (throughput_bench.py · throughput_result.json, which has the full 1/16/20/24-user sweep).

Reasoning & Output Length

Measured with the model card's thinking-mode sampling (temperature=0.6, top_p=0.95, top_k=20) and 16 requests in flight:

Reasoning comparison

TestOrnith-1.5-35B-A3B (this repo)Qwen3.8-35B-A3B
GSM8K, first 100 test problems98% (98/100)99% (99/100)
avg / max output tokens384 / 3,630303 / 1,251
hit the 16,384-token limit00
MATH-500, first 100 with integer answers97% (97/100)94% (94/100)
avg / max output tokens1,937 / 16,3842,263 / 16,384
hit the 16,384-token limit36
Long-form, 4 prompts, 32,768-token capavg 17,804 tokensavg 20,723 tokens

Ornith trades a single point on GSM8K for three points on the harder MATH-500 benchmark, with fewer truncated answers. Raw results: reasoning_result_ornith.json · script: reasoning_bench.py.

Knowledge: MMLU-Pro

MMLU-Pro (TIGER-Lab, NeurIPS 2024): 10 options per question, so random guessing scores 10%. Stratified sample of 700 questions (50 per category, fixed seed), zero-shot, thinking on, max_tokens 16,384, 8 requests in flight through the metering gateway.

Knowledge comparison

Ornith-1.5-35B-A3B (this repo)Qwen3.8-35B-A3B
MMLU-Pro accuracy81.1% (568/700, 95% CI ±2.9)80.9% (566/700, 95% CI ±2.9)
unparsed / truncated answers37 (5.3%)11 (1.6%)
avg output tokens2,2731,155
CategoryOrnithQwen3.8CategoryOrnithQwen3.8
Math94%94%Psychology88%82%
Physics94%92%Business84%82%
Biology92%94%Philosophy80%76%
Computer science92%86%Health76%74%
Economics92%90%History68%68%
Chemistry90%86%Other66%68%
Law62%70%
Engineering58%70%

Overall accuracy is essentially tied (81.1% vs 80.9%, well within the ±2.9-point sampling error of either single run), but the two models trade strength by category: Ornith is stronger on computer science and psychology, Qwen3.8 is stronger on law and engineering. Ornith also produced roughly 2x the average output tokens per answer and 3x the unparsed-answer rate, suggesting it reasons longer and occasionally fails to land on the requested ANSWER: <letter> format — a prompt-format issue more than a knowledge gap.

Raw results: knowledge_result_ornith.json (summary) · knowledge_result_ornith.jsonl (per question) · script: knowledge_bench.py. This run took several hours of wall time interleaved with three server restarts (for MTP, vision, and the max-num-seqs change) and live user traffic sharing the same concurrency slots, so the wall-clock figure isn't a speed measure — use the throughput benchmark above for that.

Live Monitoring

The server exposes Prometheus metrics at /metrics, scraped in the same vllm job as the other vLLM ports on this box; Grafana's panels key off the model/model_name labels, so this model's traffic appears automatically with no dashboard changes needed:

scrape_configs:
  - job_name: vllm
    static_configs:
      - targets: ["host.docker.internal:8000", "host.docker.internal:8002", "host.docker.internal:8004"]

If prometheus.yml is bind-mounted as a single file, run docker restart <prometheus> after editing it — a SIGHUP won't pick up a file replaced by sed -i.

Client Usage

from openai import OpenAI

client = OpenAI(api_key="<gateway key>", base_url="http://localhost:8080/v1")

resp = client.chat.completions.create(
    model="ornith-1.5-35b-a3b",
    messages=[{"role": "user", "content": "A snail climbs 3 m a day and slips 2 m a night in a 10 m well. How many days?"}],
    max_tokens=16384,
    temperature=0.6, top_p=0.95,
    extra_body={"top_k": 20},
)
print(resp.choices[0].message.reasoning)   # the <think> trace
print(resp.choices[0].message.content)     # the answer: 8

Repository Layout

Ornith-1.5_A3B_Model_DGX_Spark_Setup/
├── README.md                        # this file
├── install.sh                       # one-line installer (curl | bash)
├── Dockerfile                       # nvcr vLLM 26.07 base + xgrammar patch
├── download.sh                      # fetches the official pre-quantized NVFP4 checkpoint
├── run.sh                           # builds the image and runs the server on :8004
├── gateway/                         # metering gateway (see docs/GATEWAY.md)
│   ├── gateway.py                   #   auth · routing · live token metering · /metrics
│   ├── setup.sh                     #   venv + upstream key + systemd user service
│   ├── add_key.py                   #   issue a key for a user/app
│   ├── tailscale_keys.py            #   one key per tailnet device (+ CSV to hand out)
│   └── keys.example.json            #   key file format (real keys.json is git-ignored)
├── monitoring/
│   ├── prometheus-scrape.yml        # vllm + llm-gateway scrape jobs
│   ├── grafana-dashboard.json       # vLLM Command Center (per-user/per-device panels)
│   └── gateway-bypass-alert.yml     # Grafana alert: traffic bypassing the gateway
├── benchmarks/
│   ├── throughput_bench.py          # decode tok/s at N concurrent users
│   ├── throughput_result.json       # raw results cited above (Ornith vs Qwen3.8, combined)
│   ├── throughput_comparison.png    # chart
│   ├── reasoning_bench.py           # GSM8K / MATH-500 accuracy + long-form length (+ gateway auth)
│   ├── reasoning_result_ornith.json # raw reasoning results cited above
│   ├── reasoning_result_qwen38.json # baseline, copied from the Qwen3.8 repo
│   ├── reasoning_comparison.png     # chart
│   ├── knowledge_bench.py           # MMLU-Pro accuracy (700-question stratified sample)
│   ├── knowledge_result_ornith.json # MMLU-Pro summary + per-category scores
│   ├── knowledge_result_ornith.jsonl # MMLU-Pro per-question records (resume log)
│   ├── knowledge_result_qwen38.json # baseline, copied from the Qwen3.8 repo
│   ├── knowledge_comparison.png     # chart
│   └── make_charts.py               # renders all three PNGs from the JSON results above
├── assets/
│   ├── ornith_logo.png              # from the model card
│   ├── ornith_35b_eval.png          # official eval chart, from the model card
│   └── dgx-spark-banner-new.png     # hardware banner
└── docs/
    └── GATEWAY.md                   # metering gateway + Tailscale: setup, clients, ops

models/ (the downloaded checkpoint, ~22 GiB) is git-ignored and Docker-ignored. gateway/ and monitoring/ are a published copy of the same shared infrastructure documented in the Qwen3.8 repo — see What's in gateway/ for why it's duplicated rather than referenced.

Troubleshooting

SymptomFix
moe_backend='marlin' is not supported for unquantized MoE (on startup, after "Loading drafter model...")The MTP drafter's expert layers are unquantized. Don't pass --moe-backend marlin; leave it unset (auto).
At most 0 image(s) may be provided in one prompt.The server was started with --language-model-only, which strips the vision tower. Remove that flag — this checkpoint has real vision weights.
knowledge_bench.py exits with ModuleNotFoundError: No module named 'pyarrow'Run it via uv run --no-project --with pyarrow python3 benchmarks/knowledge_bench.py ... (its own docstring documents this), not plain python3.
Benchmark scripts get 401 UnauthorizedvLLM requires the gateway's upstream key (VLLM_API_KEY). Pass GATEWAY_KEY=<key> in the environment when running reasoning_bench.py / knowledge_bench.py / throughput_bench.py.
Open WebUI still shows the old default model after changing DEFAULT_MODELSOpen WebUI persists ui.default_models in its sqlite config table, which overrides the env var once set. Stop the service, update ui.default_models / ui.default_pinned_models / ui.model_order_list directly in webui.db, then restart.
CUDA out of memory at startupAnother 0.70-utilization vLLM server is running. Stop it first.
Answers are cut off (finish_reason: length)Raise max_tokens; the reasoning block comes first and can use several thousand tokens.

Credits & License


Ornith logo