ds4go is a zero-CGO Go wrapper for the ds4 inference engine. Applications using ds4go loads a pre-built libds4 shared library at runtime with github.com/ebitengine/purego. The shared library owns hardware acceleration. Use a Metal, CUDA, or CPU build of ds4 that matches your machine and model.
ds4 itself is an inference engine focused on
large mixture-of-experts models, including
DeepSeek V4 Flash,
GLM 5.3 Flash,
GLM 5.2, and
Qwen3.8 Flash Next.
These models target machines with substantial GPU-accessible memory.
We try to maintain parity with the upstream ds4 library, wrapping its C API. We build slightly-opinionated tools to facilitate using ds4.
C is a wonderful language for low-level, high-performance, portable code; a clean C API can be wrapped and used by other laguages. Golang is a wonderful language for systems and tools development, and generally more friendly for developers, esepecially when creating networked applications. LLMs are great at programming both. We take the high-performance C engine of ds4 and allow Golang to directly utilize it, simplifying local LLM application development.
Install the ds4go CLI with the quick-install script, Homebrew, or the Go toolchain:
# Quick install script (Linux/macOS)
curl -fsSL https://nimblemarkets.github.io/ds4go/install.sh | sh
# Homebrew (macOS/Linux)
brew install --cask nimblemarkets/tap/ds4go
# or with the Go toolchain
go install github.com/NimbleMarkets/ds4go/cmd/ds4go@latest
To use ds4go as a library:
go get github.com/NimbleMarkets/ds4go
Once the CLI is installed, fetch a prebuilt native libds4 from GitHub Releases:
ds4go install --backend auto
The installer downloads from github.com/NimbleMarkets/ds4 by default. Use
--repo, --version, --backend, or --url to select a fork, release, build,
or direct archive. It installs into $DS4_DIR/lib, defaulting to ~/.ds4/lib.
--backend auto selects metal on macOS arm64, cuda or rocm on Linux when detected, and cpu elsewhere.
On a DGX Spark (GB10) the installer picks the linux-arm64-gb10-cuda asset, whose sm_121a kernels
the generic arm64 build lacks; --variant gb10|sbsa overrides that detection.
Use ds4go install catalog to list release assets before choosing a build;
add --json for machine-readable output.
If the library is already installed and up-to-date, the installer exits successfully
without re-downloading. If a different version is present, it will prompt to replace it
(or require --force in non-interactive environments).
DS4_DIR is the ds4 home directory used by ds4go tooling:
$DS4_DIR/lib/ native shared libraries
$DS4_DIR/models/ GGUF model files
Manage curated DeepSeek V4 Flash, GLM 5.3, and GLM 5.2 models with:
ds4go model list # --installed, --available, or --json
ds4go model download q2-imatrix
ds4go model set q2-imatrix
# GLM 5.3 Flash fits one 128 GB machine; glm53-q4 and glm53-full-q2 need more
ds4go model download glm53-q2
ds4go model set glm53-q2
# GLM 5.2 catalog aliases include glm-iq2xxs, glm-q2, and glm-q4
ds4go model download glm-iq2xxs
ds4go model set glm-iq2xxs
The default model path for commands and examples is
$DS4_DIR/models/ds4flash.gguf.
ListModels() exposes the curated catalog with local installation/default state,
paths, descriptions, size/RAM guidance, model-family and vision metadata, and
companion aliases. It reads local metadata without loading libds4 or downloading
anything. Use Installed && IsChatModel() for a selector of available standalone
chat models; retain Path as each option's value and Default for its initial
selection. Encoder and speculative-draft files, and distributed pieces, are
excluded by IsChatModel(). Group options by Family to collect quantization
variants of the same checkpoint (for example, q2-imatrix and q2-q4-imatrix
share deepseek-v4-flash). Companion entries have an empty Family.
catalog, err := ds4.ListModels()
if err != nil {
return err
}
for _, model := range catalog {
if model.Installed && model.IsChatModel() {
fmt.Printf("%s — %s (%s)\n", model.Alias, model.Notes, model.Path)
}
}
Vision describes the checkpoint's capability. Look up its Encoder alias in
that same catalog and check Installed to tell whether its companion is present;
use ApplyVisionDefaults(&opts) after assigning the selected Path to
opts.ModelPath. Hardware/runtime compatibility still depends on the engine.
Keep unavailable entries to offer download choices, using Partial and
PartialBytes to display unfinished downloads. Custom GGUFs are not catalog
entries, so offer a file picker separately if your app supports them.
For a flag or text field, ResolveModelPath("vision-q2") resolves an installed
alias to its GGUF path. Paths, unknown aliases, and uninstalled aliases pass
through unchanged for the app to validate. ResolveModelInfo(path) provides the
same metadata for an installed file, including the default hard link. Selecting
an entry in your application does not change ds4go's global default.
DeepSeek Flash Vision-Exp, GLM 5.3 Flash, DeepSeek V4.1 Flash, and Qwen3.8 Flash Next accept images in conversations. Download a vision model together with its encoder, then attach an image to a prompt:
ds4go model download glm53-q2 # also fetches glm53-vision, its encoder
ds4go prompt -m glm53-q2 --image photo.png -p "What is this?"
model status shows which processes have an engine loaded and the models,
MTP/DSpark companions, and vision encoders they hold open, by catalog alias
and role, then every partial download with its state (downloading, stalled,
interrupted), the downloader's PID, progress, a sampled rate, and an ETA;
--watch 10s refreshes and --json reports both sections with
seconds-based fields.
model download adds a vision model's encoder when it is not installed
(--no-encoder skips that), accepts several aliases, and fetches them one
after another, checking the combined size against free space before the
first byte moves.
DeepSeek V4.1 Flash (v41-q2, v41-q4, encoder v41-vision) is a separate
model family: Metal upstream, with text inference on CUDA; no DSpark or
external MTP; and thinking as a numeric reasoning effort. --think-level 25 (or /think 25 in chat) sets
1 to 100, 0 disables thinking, --think is 75 and --think-max 100. Q2 runs
on one 128 GB Mac with --ssd-streaming; Q4 is published in two parts that
model download fetches, joins, and verifies (allow 37 GiB extra while
joining). GLM 5.3 Flash FP8 is glm53-fp8. V4.1 speaks a spaced DSML
dialect for tool calls; ToolSyntax picks it automatically, so tool loops
need no changes.
Qwen3.8 Flash Next (qwen38-q2, qwen38-q4k, encoder qwen38-vision) runs
on Metal and CUDA. Each GGUF holds the main weights, the built-in MTP, and a
95 GiB BF16 n-gram table that stays on disk, so keep it on a fast local SSD;
qwen38-q2 fits a 64 GB Mac started with --ctx 8192 --prefill-chunk 1024.
Its reasoning effort is an instruction at the head of the system turn rather
than a think prefix, and ThinkLow / ThinkMedium (the OpenAI low,
minimal, and medium efforts) map onto it; BuildChatPrompt handles that
through Engine.IsQwen4 and Engine.Qwen4ReasoningEffortText. Speculation
is built in: --mtp-timing enables it and prints acceptance/timing counters,
as for GLM. In Go, set EngineOptions.GLMMTP to enable it without counters.
Qwen's XML tool-call
dialect (<tool_call><function=...><parameter=...>) is dsml.SyntaxQwen;
ToolSyntax selects it for a Qwen engine, so tool loops need no changes.
Vision here means image understanding: the encoder turns a PNG or JPEG into embedding rows that are spliced into the prompt, and the model answers in text. Nothing generates or edits images.
In chat mode, /read photo.png sends an image as the next turn (the bytes are
kept, so the file may change or disappear afterwards); /read notes.txt still
reads a text prompt. Vision-Exp is a different DeepSeek checkpoint from Flash
0731 and pins its own DSpark drafter, vision-dspark-support. Image prompts
think before answering, so give them a few hundred tokens of budget.
From Go, an image is a content part holding encoded PNG or JPEG bytes (or a
path). ImageInputPNG and ImageInputJPEG encode an image.Image for you.
The prompt is built with the multimodal builder and run through
GeneratePrompt, which syncs the image spans; free the Prompt afterwards:
images := ds4.NewImageEncoder(engine) // caches embeddings by image bytes
in, err := ds4.ImageInputPNG(img) // or ds4.ImageInputJPEG(img, 85) for photos
parts := []ds4.ContentPart{
{Text: "What is this?"},
{Image: &in},
}
history := []ds4.ChatMessage{
{Role: "user", Parts: parts},
}
prompt, err := ds4.BuildChatPromptMultimodal(engine, images, "You are a helpful assistant", nil, history, ds4.ThinkHigh)
defer prompt.Free()
_, err = (ds4.Generator{Engine: engine, Session: session}).GeneratePrompt(prompt, ds4.GenerateOptions{
MaxTokens: 1024, StopOnEOS: true, ThinkMode: ds4.ThinkHigh,
})
Only user and tool messages may carry images. A ToolLoop handles them the
same way once its Images field is set to the encoder; workspacetool's
view_image returns an image observation the loop feeds back to the model.
Over HTTP, examples/openai-compatible accepts images the way upstream
ds4-server does: image_url parts carrying inline data:image/png;base64,...
or data:image/jpeg;base64,... URIs in user and tool messages, at most 16
images per request and a 64 MiB body. Remote URLs and file paths are rejected.
ds4.ParseOpenAIContent is the reusable parser behind it.
DeepSeek speculative decoding uses a separate MTP support-model GGUF. GLM's
optional next-token predictor is embedded in the base model instead, and
libds4 rejects an external MTP path for GLM. When a curated GLM model is
selected directly or through the active-model link, ds4go therefore suppresses
--mtp / EngineOptions.MTPPath; DeepSeek models keep the configured external
MTP model.
Place the shared library in ~/.ds4/lib/, $DS4_DIR/lib/, next to your
executable, or in a lib/ directory next to your executable. You can also point
at it explicitly. The current working directory and the repository root are not
searched, to avoid loading a planted library:
export DS4_LIB=/absolute/path/to/libds4.dylib
# or
export DS4_DIR=/opt/ds4
Platform defaults are:
| Platform | Library |
|---|---|
| macOS | libds4.dylib |
| Linux | libds4.so |
| Windows | libds4.dll (not a current target) |
Windows is not a supported target at present: the loader's optional symbol
lookups use purego.Dlsym, which purego does not provide on Windows, so the
tree does not build there, and stderr redirection has no Windows path either.
The libds4.dll name and the Windows-specific files are placeholders for a
future port, not a working configuration.
CUDA and ROCm are separate libds4 build flavors, but upstream ds4.h uses the
same backend enum value for both: DS4_BACKEND_CUDA. A ROCm-built library
interprets that value as ROCm and reports rocm through ds4_backend_name().
For that reason, Go code should load distinct shared-library files explicitly
and create engines from those Library handles:
cudaLib, err := ds4.Load("/opt/ds4/cuda/libds4.so")
if err != nil {
panic(err)
}
rocmLib, err := ds4.Load("/opt/ds4/rocm/libds4.so")
if err != nil {
panic(err)
}
cudaEngine, err := cudaLib.NewEngine(ds4.EngineOptions{
ModelPath: "/models/ds4flash.gguf",
Backend: ds4.BackendCUDA,
})
if err != nil {
panic(err)
}
defer cudaEngine.Close()
rocmEngine, err := rocmLib.NewEngine(ds4.EngineOptions{
ModelPath: "/models/ds4flash.gguf",
Backend: ds4.BackendCUDA, // ROCm uses the CUDA ABI backend slot.
})
if err != nil {
panic(err)
}
defer rocmEngine.Close()
Use explicit Library handles for side-by-side builds; package-level helpers
such as ds4.NewEngine use the singleton default library. Keep engines,
sessions, and token vectors with the library that created them. ds4go currently
serializes calls into libds4 across the process, so multiple engines can be
loaded independently but inference is not run concurrently by the Go binding.
import ds4 "github.com/NimbleMarkets/ds4go"
engine, err := ds4.NewEngine(ds4.EngineOptions{
ModelPath: "/models/ds4flash.gguf",
Backend: ds4.BackendMetal,
})
if err != nil {
panic(err)
}
defer engine.Close()
session, err := engine.NewSession(32768)
if err != nil {
panic(err)
}
defer session.Close()
prompt, err := engine.EncodeChatPrompt("", "Explain Redis streams briefly.", ds4.ThinkHigh)
if err != nil {
panic(err)
}
defer prompt.Free()
_, err = ds4.Generator{Engine: engine, Session: session}.GenerateTokens(prompt, ds4.GenerateOptions{
MaxTokens: 128,
StopOnEOS: true,
OnToken: func(token int) {
text, _ := engine.TokenText(token)
fmt.Print(text)
},
})
go run ./cmd/ds4go prompt --model ./ds4flash.gguf -p "Explain Redis streams in one paragraph."
go run ./cmd/ds4go prompt --model ./ds4flash.gguf
cmd/ds4go prompt and the examples accept the same arguments as the upstream ds4 C programs, parsed with pflag so options take the --option form. cmd/ds4go prompt, examples/simple, and examples/chat mirror the ds4 CLI (ds4_cli.c); examples/openai-compatible mirrors ds4-server (ds4_server.c). Run any of them with --help for the full list.
The only addition with no C equivalent is --lib, which points at the libds4 shared library the pure-Go wrapper loads at runtime. When empty, ds4go searches DS4_LIB, $DS4_DIR/lib (or ~/.ds4/lib), and executable-local paths; it never hands a bare name to the OS loader, whose search includes the working directory on macOS and Windows.
$ ds4go help cheat
ds4go — command cheat sheet
├── completion Generate the autocompletion script for the specified shell
│ ├── bash Generate the autocompletion script for bash
│ ├── fish Generate the autocompletion script for fish
│ ├── powershell Generate the autocompletion script for powershell
│ └── zsh Generate the autocompletion script for zsh
│
├── install Download a prebuilt libds4 shared library
│ └── catalog List available libds4 release assets
│
├── model Browse, download, and manage curated ds4 models
│ ├── status Show loaded engines and in-progress, stalled, or interrupted downloads
│ ├── delete Delete a downloaded model from disk
│ ├── download Download a curated model from Hugging Face
│ ├── info Show details for a curated model
│ ├── list List installed and available models
│ └── set Set the default chat model
│
├── prompt Run prompt or interactive chat inference
│
├── status Find processes holding or using the libds4 shared library
│
├── uninstall Uninstall the installed libds4 shared library
│
├── validate Validate the installed libds4 shared library
│
└── web Test browser-backed web tools
├── search Execute Google search and print Markdown links
└── visit Visit a web page and print extracted Markdown
Run 'ds4go help <command>' for detailed usage.
ds4go includes optional tool packages that can be registered on a
ToolRegistry for model-driven workflows:
workspacetool exposes local workspace tools:
read, more, list, search, view_image (with a vision encoder),
opt-in write / edit, and opt-in shell jobs.scratchtool exposes a session-scoped key/value
scratchpad (scratch_list/get/set/append/delete): durable private working
memory for plans and intermediate notes, persisted under
$DS4_DIR/scratch/<session>/.webtool exposes browser-backed google_search and
visit_page tools, and fetch_image for pulling a web image into a vision
model's context.lsp/lsptool adapts a language-server client as
diagnostics, hover, symbols, and completion tools.Workspace tools are conservative by default: paths are confined to the
configured root, symlink traversal is rejected, writes require AllowWrite, and
shell commands require AllowShell.
go run ./examples/simple --model ./ds4flash.gguf
go run ./examples/chat --model ./ds4flash.gguf
go run ./examples/toolloop --mock
go run ./examples/toolloop --mock --scratch # add the scratchpad tools
go run ./examples/toolloop --model ./ds4flash.gguf --nothink --tokens 512
go run ./examples/openai-compatible --model ./ds4flash.gguf --host 127.0.0.1 --port 8000
The toolloop example registers a Go add tool and exercises model-native
tool-call parsing (DeepSeek DSML or GLM <tool_call> markup), tool dispatch,
tool-result rendering, and replay. Use --mock for a no-model smoke test. The
OpenAI-compatible example exposes POST /v1/chat/completions for a minimal
local test server.
Most users should import the root package ds4 from github.com/NimbleMarkets/ds4go. It provides Go-native runtime policy and convenience helpers on top of the raw API. This includes DetectDefaultBackend(libPath), which queries backend preferences from installation metadata (ds4go-install.json) or falls back to system checks (probes for /dev/nvidia0 or nvidia-smi on Linux to select CUDA; defaults to Metal on macOS arm64, and CPU reference otherwise).
The strict binding layer lives in package ds4api, imported as github.com/NimbleMarkets/ds4go/ds4api. It mirrors the public ds4.h API: engines, sessions, token vectors, chat prompt rendering, tokenization, logprob helpers, MTP metadata, GLM model-family helpers, engine context and placement hints, directional steering options, snapshot/payload save-load, and DS4 context-memory helpers. APIs that take FILE * use the package's opaque ds4api.File wrapper around a C FILE*.
EngineOptions.ContextSize communicates the largest planned session context to
the engine. PlacementCtxHint, PlacementSessionCountHint, and
ShareSessionPrefillWorkspace provide the corresponding GPU-placement planning
hints. The CLI maps --ctx to both ContextSize and PlacementCtxHint.
ds4_log is exposed as LogString, which safely calls it with a fixed "%s" format. Arbitrary C varargs are intentionally not surfaced as a Go variadic API. SetStderr/SetStderrFd redirect libds4's diagnostic stream to a file or descriptor (see below). SetAbortFunc exposes libds4's fatal-invariant hook, which fires immediately before libds4 aborts the process.
libds4 writes its diagnostics — including Metal/CUDA backend messages — to its
own stderr stream. ds4go redirects that stream to a file or descriptor with
SetStderr, SetStderrFd, and DiscardLogs:
f, _ := os.Create("ds4.log")
err := ds4.SetStderr(f) // redirect libds4 diagnostics to f
err = ds4.DiscardLogs() // or send them to the null device
err = ds4.SetStderr(nil) // restore the native stderr
libds4 dups the descriptor internally and writes unbuffered, so you may close
your file once it is no longer the active target. The redirect target is
process-global inside libds4, not per engine, so install it once during startup,
before NewEngine. It is targeted at libds4's own output — not a process-wide
dup2 — so anything other libraries write directly to file descriptor 2 is
unaffected. Diagnostics are redirected as plain text; libds4 uses log levels only
to colorize TTY output, so no per-message level is surfaced to Go.
To capture diagnostics into an in-process io.Writer — a TUI log overlay, a ring
buffer, or an slog adapter — use CaptureStderr, which bridges the descriptor
redirect to a writer with an internal pipe and pump goroutine:
cap, err := ds4.CaptureStderr(myWriter) // libds4 diagnostics stream into myWriter
defer cap.Close() // restore native stderr and drain on exit
Redirection is not supported on Windows: os.File.Fd returns a Win32
HANDLE, which libds4's CRT-based ds4_set_stderr_fd cannot accept, so these
calls return ErrStderrUnsupportedOnWindows there.
For CLI use, you can also redirect stderr with your shell:
ds4go prompt ... 2>ds4.log
ds4go prompt ... 2>/dev/null
Recent libds4 builds expose ds4_abort_set, and ds4go wraps it as
SetAbortFunc. This is a last-chance fatal-invariant hook: libds4 calls it
after logging the fatal message at LogError and immediately before native
abort().
err := ds4.SetAbortFunc(func(msg string) {
crashReporter.Record("libds4 fatal invariant", msg)
})
Returning from the callback does not recover the engine. The native library
still calls abort() because the invariant is already broken. Use the hook for
crash telemetry, flushing logs, or deliberate process termination. Do not call
back into ds4go/libds4 from the callback; it can run from native worker threads
while an FFI call is active.
Do not use signal.NotifyContext around C FFI calls. SIGINT (Ctrl+C) can be delivered to any OS thread, including C worker threads inside libds4 (Metal, CUDA, or CPU). When that happens the C runtime aborts and the process segfaults.
Safe cancellation is programmatic only — pass a context.Context to GenerateOptions.Context and cancel it from Go code. The generator checks ctx.Done() between tokens, so cancellation never interrupts an active FFI call:
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
_, err = ds4.Generator{Engine: engine, Session: session}.GenerateTokens(prompt, ds4.GenerateOptions{
MaxTokens: 128,
Context: ctx,
OnToken: func(token int) {
text, _ := engine.TokenText(token)
fmt.Print(text)
},
})
This is exactly how examples/openai-compatible handles client disconnects — it wires r.Context() into generation so the engine stops cleanly when the HTTP connection drops.
Bindings are generated by hand against the public ds4 header at https://github.com/antirez/ds4/blob/main/ds4.h.
Inference runs in-process. The Golang wrapper adds FFI calls but does not proxy tokens through a server or copy model weights. Prefill, generation, Metal/CUDA/CPU execution, MTP, KV reuse, and disk KV payload serialization are all handled by the loaded ds4 shared library.
We welcome contributions and feedback. Please adhere to our Code of Conduct when engaging our community.
Thanks to @antirez for his work on ds4 and for his local-LLM advocacy. Thanks to DeepSeek for their public contributions.
Released under the MIT License, see LICENSE.txt.
Copyright (c) 2026 Neomantra Corp.
Made with :heart: and :fire: by the team behind Nimble.Markets.
Go
99.0%
ds4go is a zero-CGO Go wrapper for the ds4 inference engine. Applications using ds4go loads a pre-built libds4 shared library at runtime with github.com/ebitengine/purego. The shared library owns hardware acceleration. Use a Metal, CUDA, or CPU build of ds4 that matches your machine and model.
ds4 itself is an inference engine focused on
large mixture-of-experts models, including
DeepSeek V4 Flash,
GLM 5.3 Flash,
GLM 5.2, and
Qwen3.8 Flash Next.
These models target machines with substantial GPU-accessible memory.
We try to maintain parity with the upstream ds4 library, wrapping its C API. We build slightly-opinionated tools to facilitate using ds4.
C is a wonderful language for low-level, high-performance, portable code; a clean C API can be wrapped and used by other laguages. Golang is a wonderful language for systems and tools development, and generally more friendly for developers, esepecially when creating networked applications. LLMs are great at programming both. We take the high-performance C engine of ds4 and allow Golang to directly utilize it, simplifying local LLM application development.
Install the ds4go CLI with the quick-install script, Homebrew, or the Go toolchain:
# Quick install script (Linux/macOS)
curl -fsSL https://nimblemarkets.github.io/ds4go/install.sh | sh
# Homebrew (macOS/Linux)
brew install --cask nimblemarkets/tap/ds4go
# or with the Go toolchain
go install github.com/NimbleMarkets/ds4go/cmd/ds4go@latest
To use ds4go as a library:
go get github.com/NimbleMarkets/ds4go
Once the CLI is installed, fetch a prebuilt native libds4 from GitHub Releases:
ds4go install --backend auto
The installer downloads from github.com/NimbleMarkets/ds4 by default. Use
--repo, --version, --backend, or --url to select a fork, release, build,
or direct archive. It installs into $DS4_DIR/lib, defaulting to ~/.ds4/lib.
--backend auto selects metal on macOS arm64, cuda or rocm on Linux when detected, and cpu elsewhere.
On a DGX Spark (GB10) the installer picks the linux-arm64-gb10-cuda asset, whose sm_121a kernels
the generic arm64 build lacks; --variant gb10|sbsa overrides that detection.
Use ds4go install catalog to list release assets before choosing a build;
add --json for machine-readable output.
If the library is already installed and up-to-date, the installer exits successfully
without re-downloading. If a different version is present, it will prompt to replace it
(or require --force in non-interactive environments).
DS4_DIR is the ds4 home directory used by ds4go tooling:
$DS4_DIR/lib/ native shared libraries
$DS4_DIR/models/ GGUF model files
Manage curated DeepSeek V4 Flash, GLM 5.3, and GLM 5.2 models with:
ds4go model list # --installed, --available, or --json
ds4go model download q2-imatrix
ds4go model set q2-imatrix
# GLM 5.3 Flash fits one 128 GB machine; glm53-q4 and glm53-full-q2 need more
ds4go model download glm53-q2
ds4go model set glm53-q2
# GLM 5.2 catalog aliases include glm-iq2xxs, glm-q2, and glm-q4
ds4go model download glm-iq2xxs
ds4go model set glm-iq2xxs
The default model path for commands and examples is
$DS4_DIR/models/ds4flash.gguf.
ListModels() exposes the curated catalog with local installation/default state,
paths, descriptions, size/RAM guidance, model-family and vision metadata, and
companion aliases. It reads local metadata without loading libds4 or downloading
anything. Use Installed && IsChatModel() for a selector of available standalone
chat models; retain Path as each option's value and Default for its initial
selection. Encoder and speculative-draft files, and distributed pieces, are
excluded by IsChatModel(). Group options by Family to collect quantization
variants of the same checkpoint (for example, q2-imatrix and q2-q4-imatrix
share deepseek-v4-flash). Companion entries have an empty Family.
catalog, err := ds4.ListModels()
if err != nil {
return err
}
for _, model := range catalog {
if model.Installed && model.IsChatModel() {
fmt.Printf("%s — %s (%s)\n", model.Alias, model.Notes, model.Path)
}
}
Vision describes the checkpoint's capability. Look up its Encoder alias in
that same catalog and check Installed to tell whether its companion is present;
use ApplyVisionDefaults(&opts) after assigning the selected Path to
opts.ModelPath. Hardware/runtime compatibility still depends on the engine.
Keep unavailable entries to offer download choices, using Partial and
PartialBytes to display unfinished downloads. Custom GGUFs are not catalog
entries, so offer a file picker separately if your app supports them.
For a flag or text field, ResolveModelPath("vision-q2") resolves an installed
alias to its GGUF path. Paths, unknown aliases, and uninstalled aliases pass
through unchanged for the app to validate. ResolveModelInfo(path) provides the
same metadata for an installed file, including the default hard link. Selecting
an entry in your application does not change ds4go's global default.
DeepSeek Flash Vision-Exp, GLM 5.3 Flash, DeepSeek V4.1 Flash, and Qwen3.8 Flash Next accept images in conversations. Download a vision model together with its encoder, then attach an image to a prompt:
ds4go model download glm53-q2 # also fetches glm53-vision, its encoder
ds4go prompt -m glm53-q2 --image photo.png -p "What is this?"
model status shows which processes have an engine loaded and the models,
MTP/DSpark companions, and vision encoders they hold open, by catalog alias
and role, then every partial download with its state (downloading, stalled,
interrupted), the downloader's PID, progress, a sampled rate, and an ETA;
--watch 10s refreshes and --json reports both sections with
seconds-based fields.
model download adds a vision model's encoder when it is not installed
(--no-encoder skips that), accepts several aliases, and fetches them one
after another, checking the combined size against free space before the
first byte moves.
DeepSeek V4.1 Flash (v41-q2, v41-q4, encoder v41-vision) is a separate
model family: Metal upstream, with text inference on CUDA; no DSpark or
external MTP; and thinking as a numeric reasoning effort. --think-level 25 (or /think 25 in chat) sets
1 to 100, 0 disables thinking, --think is 75 and --think-max 100. Q2 runs
on one 128 GB Mac with --ssd-streaming; Q4 is published in two parts that
model download fetches, joins, and verifies (allow 37 GiB extra while
joining). GLM 5.3 Flash FP8 is glm53-fp8. V4.1 speaks a spaced DSML
dialect for tool calls; ToolSyntax picks it automatically, so tool loops
need no changes.
Qwen3.8 Flash Next (qwen38-q2, qwen38-q4k, encoder qwen38-vision) runs
on Metal and CUDA. Each GGUF holds the main weights, the built-in MTP, and a
95 GiB BF16 n-gram table that stays on disk, so keep it on a fast local SSD;
qwen38-q2 fits a 64 GB Mac started with --ctx 8192 --prefill-chunk 1024.
Its reasoning effort is an instruction at the head of the system turn rather
than a think prefix, and ThinkLow / ThinkMedium (the OpenAI low,
minimal, and medium efforts) map onto it; BuildChatPrompt handles that
through Engine.IsQwen4 and Engine.Qwen4ReasoningEffortText. Speculation
is built in: --mtp-timing enables it and prints acceptance/timing counters,
as for GLM. In Go, set EngineOptions.GLMMTP to enable it without counters.
Qwen's XML tool-call
dialect (<tool_call><function=...><parameter=...>) is dsml.SyntaxQwen;
ToolSyntax selects it for a Qwen engine, so tool loops need no changes.
Vision here means image understanding: the encoder turns a PNG or JPEG into embedding rows that are spliced into the prompt, and the model answers in text. Nothing generates or edits images.
In chat mode, /read photo.png sends an image as the next turn (the bytes are
kept, so the file may change or disappear afterwards); /read notes.txt still
reads a text prompt. Vision-Exp is a different DeepSeek checkpoint from Flash
0731 and pins its own DSpark drafter, vision-dspark-support. Image prompts
think before answering, so give them a few hundred tokens of budget.
From Go, an image is a content part holding encoded PNG or JPEG bytes (or a
path). ImageInputPNG and ImageInputJPEG encode an image.Image for you.
The prompt is built with the multimodal builder and run through
GeneratePrompt, which syncs the image spans; free the Prompt afterwards:
images := ds4.NewImageEncoder(engine) // caches embeddings by image bytes
in, err := ds4.ImageInputPNG(img) // or ds4.ImageInputJPEG(img, 85) for photos
parts := []ds4.ContentPart{
{Text: "What is this?"},
{Image: &in},
}
history := []ds4.ChatMessage{
{Role: "user", Parts: parts},
}
prompt, err := ds4.BuildChatPromptMultimodal(engine, images, "You are a helpful assistant", nil, history, ds4.ThinkHigh)
defer prompt.Free()
_, err = (ds4.Generator{Engine: engine, Session: session}).GeneratePrompt(prompt, ds4.GenerateOptions{
MaxTokens: 1024, StopOnEOS: true, ThinkMode: ds4.ThinkHigh,
})
Only user and tool messages may carry images. A ToolLoop handles them the
same way once its Images field is set to the encoder; workspacetool's
view_image returns an image observation the loop feeds back to the model.
Over HTTP, examples/openai-compatible accepts images the way upstream
ds4-server does: image_url parts carrying inline data:image/png;base64,...
or data:image/jpeg;base64,... URIs in user and tool messages, at most 16
images per request and a 64 MiB body. Remote URLs and file paths are rejected.
ds4.ParseOpenAIContent is the reusable parser behind it.
DeepSeek speculative decoding uses a separate MTP support-model GGUF. GLM's
optional next-token predictor is embedded in the base model instead, and
libds4 rejects an external MTP path for GLM. When a curated GLM model is
selected directly or through the active-model link, ds4go therefore suppresses
--mtp / EngineOptions.MTPPath; DeepSeek models keep the configured external
MTP model.
Place the shared library in ~/.ds4/lib/, $DS4_DIR/lib/, next to your
executable, or in a lib/ directory next to your executable. You can also point
at it explicitly. The current working directory and the repository root are not
searched, to avoid loading a planted library:
export DS4_LIB=/absolute/path/to/libds4.dylib
# or
export DS4_DIR=/opt/ds4
Platform defaults are:
| Platform | Library |
|---|---|
| macOS | libds4.dylib |
| Linux | libds4.so |
| Windows | libds4.dll (not a current target) |
Windows is not a supported target at present: the loader's optional symbol
lookups use purego.Dlsym, which purego does not provide on Windows, so the
tree does not build there, and stderr redirection has no Windows path either.
The libds4.dll name and the Windows-specific files are placeholders for a
future port, not a working configuration.
CUDA and ROCm are separate libds4 build flavors, but upstream ds4.h uses the
same backend enum value for both: DS4_BACKEND_CUDA. A ROCm-built library
interprets that value as ROCm and reports rocm through ds4_backend_name().
For that reason, Go code should load distinct shared-library files explicitly
and create engines from those Library handles:
cudaLib, err := ds4.Load("/opt/ds4/cuda/libds4.so")
if err != nil {
panic(err)
}
rocmLib, err := ds4.Load("/opt/ds4/rocm/libds4.so")
if err != nil {
panic(err)
}
cudaEngine, err := cudaLib.NewEngine(ds4.EngineOptions{
ModelPath: "/models/ds4flash.gguf",
Backend: ds4.BackendCUDA,
})
if err != nil {
panic(err)
}
defer cudaEngine.Close()
rocmEngine, err := rocmLib.NewEngine(ds4.EngineOptions{
ModelPath: "/models/ds4flash.gguf",
Backend: ds4.BackendCUDA, // ROCm uses the CUDA ABI backend slot.
})
if err != nil {
panic(err)
}
defer rocmEngine.Close()
Use explicit Library handles for side-by-side builds; package-level helpers
such as ds4.NewEngine use the singleton default library. Keep engines,
sessions, and token vectors with the library that created them. ds4go currently
serializes calls into libds4 across the process, so multiple engines can be
loaded independently but inference is not run concurrently by the Go binding.
import ds4 "github.com/NimbleMarkets/ds4go"
engine, err := ds4.NewEngine(ds4.EngineOptions{
ModelPath: "/models/ds4flash.gguf",
Backend: ds4.BackendMetal,
})
if err != nil {
panic(err)
}
defer engine.Close()
session, err := engine.NewSession(32768)
if err != nil {
panic(err)
}
defer session.Close()
prompt, err := engine.EncodeChatPrompt("", "Explain Redis streams briefly.", ds4.ThinkHigh)
if err != nil {
panic(err)
}
defer prompt.Free()
_, err = ds4.Generator{Engine: engine, Session: session}.GenerateTokens(prompt, ds4.GenerateOptions{
MaxTokens: 128,
StopOnEOS: true,
OnToken: func(token int) {
text, _ := engine.TokenText(token)
fmt.Print(text)
},
})
go run ./cmd/ds4go prompt --model ./ds4flash.gguf -p "Explain Redis streams in one paragraph."
go run ./cmd/ds4go prompt --model ./ds4flash.gguf
cmd/ds4go prompt and the examples accept the same arguments as the upstream ds4 C programs, parsed with pflag so options take the --option form. cmd/ds4go prompt, examples/simple, and examples/chat mirror the ds4 CLI (ds4_cli.c); examples/openai-compatible mirrors ds4-server (ds4_server.c). Run any of them with --help for the full list.
The only addition with no C equivalent is --lib, which points at the libds4 shared library the pure-Go wrapper loads at runtime. When empty, ds4go searches DS4_LIB, $DS4_DIR/lib (or ~/.ds4/lib), and executable-local paths; it never hands a bare name to the OS loader, whose search includes the working directory on macOS and Windows.
$ ds4go help cheat
ds4go — command cheat sheet
├── completion Generate the autocompletion script for the specified shell
│ ├── bash Generate the autocompletion script for bash
│ ├── fish Generate the autocompletion script for fish
│ ├── powershell Generate the autocompletion script for powershell
│ └── zsh Generate the autocompletion script for zsh
│
├── install Download a prebuilt libds4 shared library
│ └── catalog List available libds4 release assets
│
├── model Browse, download, and manage curated ds4 models
│ ├── status Show loaded engines and in-progress, stalled, or interrupted downloads
│ ├── delete Delete a downloaded model from disk
│ ├── download Download a curated model from Hugging Face
│ ├── info Show details for a curated model
│ ├── list List installed and available models
│ └── set Set the default chat model
│
├── prompt Run prompt or interactive chat inference
│
├── status Find processes holding or using the libds4 shared library
│
├── uninstall Uninstall the installed libds4 shared library
│
├── validate Validate the installed libds4 shared library
│
└── web Test browser-backed web tools
├── search Execute Google search and print Markdown links
└── visit Visit a web page and print extracted Markdown
Run 'ds4go help <command>' for detailed usage.
ds4go includes optional tool packages that can be registered on a
ToolRegistry for model-driven workflows:
workspacetool exposes local workspace tools:
read, more, list, search, view_image (with a vision encoder),
opt-in write / edit, and opt-in shell jobs.scratchtool exposes a session-scoped key/value
scratchpad (scratch_list/get/set/append/delete): durable private working
memory for plans and intermediate notes, persisted under
$DS4_DIR/scratch/<session>/.webtool exposes browser-backed google_search and
visit_page tools, and fetch_image for pulling a web image into a vision
model's context.lsp/lsptool adapts a language-server client as
diagnostics, hover, symbols, and completion tools.Workspace tools are conservative by default: paths are confined to the
configured root, symlink traversal is rejected, writes require AllowWrite, and
shell commands require AllowShell.
go run ./examples/simple --model ./ds4flash.gguf
go run ./examples/chat --model ./ds4flash.gguf
go run ./examples/toolloop --mock
go run ./examples/toolloop --mock --scratch # add the scratchpad tools
go run ./examples/toolloop --model ./ds4flash.gguf --nothink --tokens 512
go run ./examples/openai-compatible --model ./ds4flash.gguf --host 127.0.0.1 --port 8000
The toolloop example registers a Go add tool and exercises model-native
tool-call parsing (DeepSeek DSML or GLM <tool_call> markup), tool dispatch,
tool-result rendering, and replay. Use --mock for a no-model smoke test. The
OpenAI-compatible example exposes POST /v1/chat/completions for a minimal
local test server.
Most users should import the root package ds4 from github.com/NimbleMarkets/ds4go. It provides Go-native runtime policy and convenience helpers on top of the raw API. This includes DetectDefaultBackend(libPath), which queries backend preferences from installation metadata (ds4go-install.json) or falls back to system checks (probes for /dev/nvidia0 or nvidia-smi on Linux to select CUDA; defaults to Metal on macOS arm64, and CPU reference otherwise).
The strict binding layer lives in package ds4api, imported as github.com/NimbleMarkets/ds4go/ds4api. It mirrors the public ds4.h API: engines, sessions, token vectors, chat prompt rendering, tokenization, logprob helpers, MTP metadata, GLM model-family helpers, engine context and placement hints, directional steering options, snapshot/payload save-load, and DS4 context-memory helpers. APIs that take FILE * use the package's opaque ds4api.File wrapper around a C FILE*.
EngineOptions.ContextSize communicates the largest planned session context to
the engine. PlacementCtxHint, PlacementSessionCountHint, and
ShareSessionPrefillWorkspace provide the corresponding GPU-placement planning
hints. The CLI maps --ctx to both ContextSize and PlacementCtxHint.
ds4_log is exposed as LogString, which safely calls it with a fixed "%s" format. Arbitrary C varargs are intentionally not surfaced as a Go variadic API. SetStderr/SetStderrFd redirect libds4's diagnostic stream to a file or descriptor (see below). SetAbortFunc exposes libds4's fatal-invariant hook, which fires immediately before libds4 aborts the process.
libds4 writes its diagnostics — including Metal/CUDA backend messages — to its
own stderr stream. ds4go redirects that stream to a file or descriptor with
SetStderr, SetStderrFd, and DiscardLogs:
f, _ := os.Create("ds4.log")
err := ds4.SetStderr(f) // redirect libds4 diagnostics to f
err = ds4.DiscardLogs() // or send them to the null device
err = ds4.SetStderr(nil) // restore the native stderr
libds4 dups the descriptor internally and writes unbuffered, so you may close
your file once it is no longer the active target. The redirect target is
process-global inside libds4, not per engine, so install it once during startup,
before NewEngine. It is targeted at libds4's own output — not a process-wide
dup2 — so anything other libraries write directly to file descriptor 2 is
unaffected. Diagnostics are redirected as plain text; libds4 uses log levels only
to colorize TTY output, so no per-message level is surfaced to Go.
To capture diagnostics into an in-process io.Writer — a TUI log overlay, a ring
buffer, or an slog adapter — use CaptureStderr, which bridges the descriptor
redirect to a writer with an internal pipe and pump goroutine:
cap, err := ds4.CaptureStderr(myWriter) // libds4 diagnostics stream into myWriter
defer cap.Close() // restore native stderr and drain on exit
Redirection is not supported on Windows: os.File.Fd returns a Win32
HANDLE, which libds4's CRT-based ds4_set_stderr_fd cannot accept, so these
calls return ErrStderrUnsupportedOnWindows there.
For CLI use, you can also redirect stderr with your shell:
ds4go prompt ... 2>ds4.log
ds4go prompt ... 2>/dev/null
Recent libds4 builds expose ds4_abort_set, and ds4go wraps it as
SetAbortFunc. This is a last-chance fatal-invariant hook: libds4 calls it
after logging the fatal message at LogError and immediately before native
abort().
err := ds4.SetAbortFunc(func(msg string) {
crashReporter.Record("libds4 fatal invariant", msg)
})
Returning from the callback does not recover the engine. The native library
still calls abort() because the invariant is already broken. Use the hook for
crash telemetry, flushing logs, or deliberate process termination. Do not call
back into ds4go/libds4 from the callback; it can run from native worker threads
while an FFI call is active.
Do not use signal.NotifyContext around C FFI calls. SIGINT (Ctrl+C) can be delivered to any OS thread, including C worker threads inside libds4 (Metal, CUDA, or CPU). When that happens the C runtime aborts and the process segfaults.
Safe cancellation is programmatic only — pass a context.Context to GenerateOptions.Context and cancel it from Go code. The generator checks ctx.Done() between tokens, so cancellation never interrupts an active FFI call:
ctx, cancel := context.WithTimeout(context.Background(), 30*time.Second)
defer cancel()
_, err = ds4.Generator{Engine: engine, Session: session}.GenerateTokens(prompt, ds4.GenerateOptions{
MaxTokens: 128,
Context: ctx,
OnToken: func(token int) {
text, _ := engine.TokenText(token)
fmt.Print(text)
},
})
This is exactly how examples/openai-compatible handles client disconnects — it wires r.Context() into generation so the engine stops cleanly when the HTTP connection drops.
Bindings are generated by hand against the public ds4 header at https://github.com/antirez/ds4/blob/main/ds4.h.
Inference runs in-process. The Golang wrapper adds FFI calls but does not proxy tokens through a server or copy model weights. Prefill, generation, Metal/CUDA/CPU execution, MTP, KV reuse, and disk KV payload serialization are all handled by the loaded ds4 shared library.
We welcome contributions and feedback. Please adhere to our Code of Conduct when engaging our community.
Thanks to @antirez for his work on ds4 and for his local-LLM advocacy. Thanks to DeepSeek for their public contributions.
Released under the MIT License, see LICENSE.txt.
Copyright (c) 2026 Neomantra Corp.
Made with :heart: and :fire: by the team behind Nimble.Markets.
Go
99.0%