Chida82/StarForge

Shell

2

14 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

I took antirez's ds4, stripped it down to Qwen3.8 Flash Next on Metal, ported a bunch of improvements, and it's now ~10% faster with bit-exact output (r/LocalLLaMA)

I've had one pull request merged into ds4 (DwarfStar), a tiny one. There are a few more still waiting in the queue. I’m not complaining. Antirez says it clearly in the README: with coding agents everyone can tune the engine for their hardware and model and he can’t review everything. That made me…

0

Oct 4, 2026

README

StarForge

One inference engine per model, Metal only, forked from antirez/ds4 (DwarfStar).

ds4 runs several large open models on several backends from one codebase. It is excellent and it is also a lot to hold in your head. StarForge produces children: full forks of ds4 reduced to one model on Apple Metal, each a normal GitHub repo you can build, read and modify without the other models and backends in the way. Children keep merging upstream through plain git merge, so fixes and speedups from ds4 keep flowing in.

ChildModelBinaryServer portSpeculative decoding
sf-ds4flashDeepSeek V4 Flashsf-ds4flash8001DSpark only; separate 0731 support GGUF; no legacy MTP
sf-ds4-1flashDeepSeek V4.1 Flashsf-ds4-1flash8002none
sf-glm5-3flashGLM 5.3 Flashsf-glm5-3flash8003built-in MTP
sf-q3-8flashQwen3.8 Flash Nextsf-q3-8flash8004built-in MTP

Each child is a separate repository. This repository is the orchestrator: rules, procedures, tools, and a read-only mirror of upstream. No inference code lives here.

Why

ds4 is built around a few models rather than as a general GGUF runner, and it is meant to be read and changed with a coding agent: a working template to adapt to your model and hardware, not a product that covers every setup. StarForge pushes both ideas to the end and makes the result repeatable, one child per model:

  • Smaller code, clearer to an LLM. SPEC.md decides what every child deletes and what it must keep, so each child shrinks the same way. What is left is a tree a person or a coding agent can load whole and reason about, which makes it cheap to try an idea on one model, measure it and keep or drop it.
  • Metal only. The children are developed and tested on one machine, an Apple M5 Max, so Metal is the only GPU backend any child keeps.
  • Small, but not frozen. The orchestrator's tools keep every child merging upstream/main with plain git merge, rerere and the parity oracle, so ds4's fixes and speedups keep flowing into a tree that stays small.
  • Room to investigate. Because each child holds one model, open ds4 pull requests can be analysed against that model and ported where they hold up, and further improvements can be investigated for Apple Silicon. That work happens and is recorded inside each child (its docs/upstream-prs.md), not here.

If you just want to run a model

Go to the child's repo. Its README tells you make, ./download.sh <quant>, ./sf-<model>. download.sh uses the Hugging Face CLI and creates a gguf/ symlink to the shared Hugging Face cache; it does not duplicate the model in the repo. You do not need this repo.

If you maintain the children

git clone https://github.com/Chida82/StarForge && cd StarForge
tools/clone-all.sh          # upstream/ds4 (full) + children/* (those that exist on GitHub)
tools/status.sh             # where each child stands vs upstream/main
I want to…ReadRun
understand the rulesAGENTS.md (short), then SPEC.md (the law)
create a childBOOTSTRAP.mdtools/new-child.sh <child> <org> then the checklist
see what upstream changed for a childSYNC.md §1tools/sync-preview.sh [child] (no arg = all)
pull upstream into a childSYNC.mdtools/sync-start.sh <child> → resolve → trace → parity → explicit commit/push request → PR → tools/sync-finish.sh <child> --push
prove a child still matches upstreamSPEC.md §F.4tools/parity-check.sh <child> [model.gguf]
work with a coding agentopen children/<child> (it has its own AGENTS.md) or this folder

Tools

All bash + git + coreutils, in tools/:

ScriptDoes
clone-all.shclone/fetch upstream/ds4 and every child in the registry; sets upstream remote and rerere
status.shper child: upstream merge-base, last sync-* tag, commits behind, dirty tree
new-child.sh <child> [org]fresh child clone with remotes, rerere, rendered AGENTS.md, parity prompts. No commit/tag/push
sync-preview.sh [child]classify pending upstream commits: only-removed / docs-only / touches-live / HOT. No argument: every cloned child
sync-start.sh <child>branch sync/<sha7>, apply upstream/main as an uncommitted merge, list conflicts
rm-deleted-conflicts.sh <child>resolve modify/delete conflicts by keeping the child's deletions
parity-check.sh <child> [gguf]build upstream at the child's merge-base and the child; same cache-backed GGUF/symlink, same prompts, greedy; token-identical + speed ±2%
speed-compare.sh <child> <gguf>same bench sweep on the child, a 3-minute rest, then upstream at the child's base (SF_SPEED_FLAGS, default --ssd-streaming); CSVs and per-frontier deltas in tools/speed/out/; informational, gates nothing
sync-finish.sh <child> --pushafter the PR is merged and push explicitly authorized: tag main as sync-<sha7>, push
parity/<child>.txtprompt sets for the oracle (TAB → extra flags: steering plus DSpark or MTP per child)
lib.shshared helpers and the registry (keep in sync with AGENTS.md)

Layout

AGENTS.md  SPEC.md  BOOTSTRAP.md  SYNC.md  README.md
templates/child-AGENTS.md        copied into a child at bootstrap
tools/                           scripts + parity prompt sets
children/<name>/                 plain clones (gitignored)
upstream/ds4/                    plain full clone of ds4 (gitignored, read-only)

Children and upstream are plain clones, not submodules: each child already records its own upstream position (tags, merge-base). There is deliberately no central SHA/status file; tools/status.sh derives live state from Git. See SPEC.md §H.

Design in five lines

  1. A child is a full git fork; git merge upstream/main is how it stays current.
  2. ds4_* file names and identifiers never change; only binary names, paths and help do.
  3. Removal is deletion, not #ifdef. Cuts inside live files carry /* sf-ablate(...) */.
  4. git rerere replays conflict resolutions, so ablated regions stop hurting after the first sync.
  5. The parity oracle (same GGUF, same prompts, greedy, token-identical, speed ±2%) is the definition of correct.

Always kept in every child: steering, TP/RDMA, CPU reference path, server, CLI, bench, eval. Speculative decoding follows the table: notably, sf-ds4flash keeps DSpark but deletes legacy MTP. Never in a child: ds4-agent, CUDA, ROCm, other models.

Agents may prepare and test changes, but commit and push only after an explicit user request.

Status

sf-q3-8flash and sf-ds4-1flash are bootstrapped, published and synced, and each now carries its own performance work. sf-ds4flash and sf-glm5-3flash have not been started.

Which upstream commit each child sits on is deliberately not written here -- tools/status.sh reads it from Git (SPEC.md §A, §H).

Licence

Tools and documents in this repository: MIT. Children inherit ds4's MIT licence and keep its LICENSE file, including the GGML notice, unchanged. Everything works because of ds4, llama.cpp and GGML.

Chida82/StarForge

Shell

2

14 commits

updated Oct 3, 2026

See the code

See what people are saying

SourceMessageScoreDate

I took antirez's ds4, stripped it down to Qwen3.8 Flash Next on Metal, ported a bunch of improvements, and it's now ~10% faster with bit-exact output (r/LocalLLaMA)

I've had one pull request merged into ds4 (DwarfStar), a tiny one. There are a few more still waiting in the queue. I’m not complaining. Antirez says it clearly in the README: with coding agents everyone can tune the engine for their hardware and model and he can’t review everything. That made me…

0

Oct 4, 2026

README

StarForge

One inference engine per model, Metal only, forked from antirez/ds4 (DwarfStar).

ds4 runs several large open models on several backends from one codebase. It is excellent and it is also a lot to hold in your head. StarForge produces children: full forks of ds4 reduced to one model on Apple Metal, each a normal GitHub repo you can build, read and modify without the other models and backends in the way. Children keep merging upstream through plain git merge, so fixes and speedups from ds4 keep flowing in.

ChildModelBinaryServer portSpeculative decoding
sf-ds4flashDeepSeek V4 Flashsf-ds4flash8001DSpark only; separate 0731 support GGUF; no legacy MTP
sf-ds4-1flashDeepSeek V4.1 Flashsf-ds4-1flash8002none
sf-glm5-3flashGLM 5.3 Flashsf-glm5-3flash8003built-in MTP
sf-q3-8flashQwen3.8 Flash Nextsf-q3-8flash8004built-in MTP

Each child is a separate repository. This repository is the orchestrator: rules, procedures, tools, and a read-only mirror of upstream. No inference code lives here.

Why

ds4 is built around a few models rather than as a general GGUF runner, and it is meant to be read and changed with a coding agent: a working template to adapt to your model and hardware, not a product that covers every setup. StarForge pushes both ideas to the end and makes the result repeatable, one child per model:

  • Smaller code, clearer to an LLM. SPEC.md decides what every child deletes and what it must keep, so each child shrinks the same way. What is left is a tree a person or a coding agent can load whole and reason about, which makes it cheap to try an idea on one model, measure it and keep or drop it.
  • Metal only. The children are developed and tested on one machine, an Apple M5 Max, so Metal is the only GPU backend any child keeps.
  • Small, but not frozen. The orchestrator's tools keep every child merging upstream/main with plain git merge, rerere and the parity oracle, so ds4's fixes and speedups keep flowing into a tree that stays small.
  • Room to investigate. Because each child holds one model, open ds4 pull requests can be analysed against that model and ported where they hold up, and further improvements can be investigated for Apple Silicon. That work happens and is recorded inside each child (its docs/upstream-prs.md), not here.

If you just want to run a model

Go to the child's repo. Its README tells you make, ./download.sh <quant>, ./sf-<model>. download.sh uses the Hugging Face CLI and creates a gguf/ symlink to the shared Hugging Face cache; it does not duplicate the model in the repo. You do not need this repo.

If you maintain the children

git clone https://github.com/Chida82/StarForge && cd StarForge
tools/clone-all.sh          # upstream/ds4 (full) + children/* (those that exist on GitHub)
tools/status.sh             # where each child stands vs upstream/main
I want to…ReadRun
understand the rulesAGENTS.md (short), then SPEC.md (the law)
create a childBOOTSTRAP.mdtools/new-child.sh <child> <org> then the checklist
see what upstream changed for a childSYNC.md §1tools/sync-preview.sh [child] (no arg = all)
pull upstream into a childSYNC.mdtools/sync-start.sh <child> → resolve → trace → parity → explicit commit/push request → PR → tools/sync-finish.sh <child> --push
prove a child still matches upstreamSPEC.md §F.4tools/parity-check.sh <child> [model.gguf]
work with a coding agentopen children/<child> (it has its own AGENTS.md) or this folder

Tools

All bash + git + coreutils, in tools/:

ScriptDoes
clone-all.shclone/fetch upstream/ds4 and every child in the registry; sets upstream remote and rerere
status.shper child: upstream merge-base, last sync-* tag, commits behind, dirty tree
new-child.sh <child> [org]fresh child clone with remotes, rerere, rendered AGENTS.md, parity prompts. No commit/tag/push
sync-preview.sh [child]classify pending upstream commits: only-removed / docs-only / touches-live / HOT. No argument: every cloned child
sync-start.sh <child>branch sync/<sha7>, apply upstream/main as an uncommitted merge, list conflicts
rm-deleted-conflicts.sh <child>resolve modify/delete conflicts by keeping the child's deletions
parity-check.sh <child> [gguf]build upstream at the child's merge-base and the child; same cache-backed GGUF/symlink, same prompts, greedy; token-identical + speed ±2%
speed-compare.sh <child> <gguf>same bench sweep on the child, a 3-minute rest, then upstream at the child's base (SF_SPEED_FLAGS, default --ssd-streaming); CSVs and per-frontier deltas in tools/speed/out/; informational, gates nothing
sync-finish.sh <child> --pushafter the PR is merged and push explicitly authorized: tag main as sync-<sha7>, push
parity/<child>.txtprompt sets for the oracle (TAB → extra flags: steering plus DSpark or MTP per child)
lib.shshared helpers and the registry (keep in sync with AGENTS.md)

Layout

AGENTS.md  SPEC.md  BOOTSTRAP.md  SYNC.md  README.md
templates/child-AGENTS.md        copied into a child at bootstrap
tools/                           scripts + parity prompt sets
children/<name>/                 plain clones (gitignored)
upstream/ds4/                    plain full clone of ds4 (gitignored, read-only)

Children and upstream are plain clones, not submodules: each child already records its own upstream position (tags, merge-base). There is deliberately no central SHA/status file; tools/status.sh derives live state from Git. See SPEC.md §H.

Design in five lines

  1. A child is a full git fork; git merge upstream/main is how it stays current.
  2. ds4_* file names and identifiers never change; only binary names, paths and help do.
  3. Removal is deletion, not #ifdef. Cuts inside live files carry /* sf-ablate(...) */.
  4. git rerere replays conflict resolutions, so ablated regions stop hurting after the first sync.
  5. The parity oracle (same GGUF, same prompts, greedy, token-identical, speed ±2%) is the definition of correct.

Always kept in every child: steering, TP/RDMA, CPU reference path, server, CLI, bench, eval. Speculative decoding follows the table: notably, sf-ds4flash keeps DSpark but deletes legacy MTP. Never in a child: ds4-agent, CUDA, ROCm, other models.

Agents may prepare and test changes, but commit and push only after an explicit user request.

Status

sf-q3-8flash and sf-ds4-1flash are bootstrapped, published and synced, and each now carries its own performance work. sf-ds4flash and sf-glm5-3flash have not been started.

Which upstream commit each child sits on is deliberately not written here -- tools/status.sh reads it from Git (SPEC.md §A, §H).

Licence

Tools and documents in this repository: MIT. Children inherit ds4's MIT licence and keep its LICENSE file, including the GGML notice, unchanged. Everything works because of ds4, llama.cpp and GGML.

Languages

Shell

100.0%