Transfer a local ollama model to a remote ollama server
C#
15
54 commits
updated Oct 3, 2026
osync is a powerful command-line tool for managing Ollama models across local and remote servers.
:latest tag when not specified* wildcards for batch operationsmanage and the colored command output, adapted to 256 and 16-color terminals[C# .NET 10]
[Windows/Linux/MacOS]
[Arm64/x64/Mac]
[Download latest binary release]
[Build from sources]
Clone the repo
dotnet build(any OS, .NET 10 SDK) or Visual Studio 2022;dotnet publish osync/osync.csproj -c Release -r <win-x64|linux-x64|osx-arm64|osx-x64>for a single-file executable. See docs/DEVELOPMENT.md for tests, CI and releases.
Commands without -d work on the local server, found in this order:
XOLLAMA_HOST (xOllama), then OLLAMA_HOST (Ollama); bind addresses such as 0.0.0.0 are mapped to localhost. With "ignoreEnvironment": true (osync setup server env ignore, or the question osync setup server asks when one of them is set) the configured server comes firstosync setup server)localhost:11434 (Ollama default) or localhost:22434 (xOllama default)osync install configures it: when exactly one server answers on this machine it is used without questions; when none or both answer, it asks for the server type (1 Ollama, 2 xOllama, 3 both side by side), host and ports. With both, one is the default local server and each gets an alias (ollama, xollama), so the other one is always at hand: osync ls -d xollama, osync cp model xollama/. Change it any time with osync setup server.
osync detects whether a server is Ollama or xOllama (osync ps shows it) and, for local operations, runs the matching CLI: xollama when the local server is xOllama (or only xollama is installed), else ollama. Set OSYNC_OLLAMA_CLI to force a specific CLI. The models directory is taken from XOLLAMA_MODELS, then OLLAMA_MODELS, then the platform default.
Remote servers given without a port use :11434, or :22434 when the host only answers there (xOllama).
osync keeps user preferences in settings.json in the per-OS configuration folder (osync -v --verbose shows the path):
| OS | Location |
|---|---|
| Windows | %APPDATA%\osync\settings.json |
| macOS | ~/Library/Application Support/osync/settings.json |
| Linux | $XDG_CONFIG_HOME/osync/settings.json (default ~/.config/osync/settings.json) |
{
"server": { "flavor": "ollama", "host": "localhost", "both": true },
"aliases": {
"ollama": "http://localhost:11434",
"xollama": "http://localhost:22434",
"gpu": "http://192.168.1.10:11434"
},
"colorMode": "auto",
"manage": { "theme": "Dracula", "sort": "size-", "servers": [ "gpu" ] },
"shell": { "theme": "Tokyo Night" }
}
server - the local server: flavor (auto, ollama, xollama), host and port (default: the flavor's port); both when Ollama and xOllama run side by side (flavor is then the default one). XOLLAMA_HOST / OLLAMA_HOST override it, unless ignoreEnvironment is true.aliases - server aliases.colorMode - auto (detect), truecolor, 256, 16 or none; OSYNC_COLOR_MODE and NO_COLOR override it.manage.theme, manage.sort - theme (also chosen with Ctrl+T in manage) and initial sort order (name+, name-, size-, size+, created-, created+) of osync manage.manage.servers - aliases manage switches to with Ctrl+Left / Ctrl+Right, after the local server (with Ollama and xOllama side by side the other one is always included).shell.theme - colors of the command output (any theme name, or plain for no colors).OSYNC_CONFIG_DIR moves the settings folder.Everything can be changed with osync setup instead of editing the file.
An alias is a short name for a server (osync setup alias add gpu 192.168.1.10), usable wherever osync expects a server:
osync ls -d gpu # list the models of the server
osync cp qwen3:8b gpu/ # upload
osync cp gpu/qwen3:8b qwen3-gpu # download
osync manage gpu
An alias takes precedence over a model namespace with the same name (gpu/model). Names start with a letter and use letters, digits, - and _.
osync detects the terminal's color depth from COLORTERM (truecolor/24bit), TERM (*-256color, *-direct), TERM_PROGRAM and Windows Terminal, and uses true color, 256 or 16 colors accordingly (osync -v --verbose shows what was detected and why). SSH does not forward COLORTERM by default, so terminals that support true color are seen as 256-color over SSH/tmux: set "colorMode": "truecolor" in the settings file, or OSYNC_COLOR_MODE=truecolor, or forward COLORTERM (SendEnv COLORTERM / AcceptEnv COLORTERM). NO_COLOR disables colors.
Command output is colored with the shell theme (osync setup shell theme NAME; osync setup shell themes shows every theme with a preview): the same 34 themes as manage, with light themes for light terminals (the default follows COLORFGBG when the terminal sets it), or plain. Output that goes to a pipe or a file is never colored.
osync manage draws with 24-bit colors; at 256 colors every theme color is snapped to the xterm-256 palette (tmux and most 256-color terminals then show exactly that color), at 16 colors the themes switch to the 16 standard colors with contrast checks. macOS Terminal.app gets 16 colors unless colorMode says truecolor (it misreads 24-bit colors on older macOS). If manage does not start or draw correctly on an old Windows console, OSYNC_TUI_DRIVER=windows selects the Windows console driver (ansi and dotnet are the others).
An AI generated DeepWiki is available here:
https://deepwiki.com/mann1x/osync
# Interactive TUI for model management (recommended)
osync manage
# Copy local model to remote server
osync cp llama3 http://192.168.100.100:11434
# Copy to remote server (short forms - port 11434 and http:// are defaults)
osync cp llama3 192.168.100.100 # IP address detected as server
osync cp llama3 myserver:11434 # hostname with port
osync cp llama3 myserver/ # trailing slash indicates server
# List all local models
osync ls
# Chat with a model
osync run llama3
# Interactive REPL mode with tab completion
osync
cp)Copy models locally, to remote servers, or between remote servers.
# Local copy (create backup)
osync cp llama3 my-backup-llama3
osync cp llama3:70b llama3:backup-v1
# Local to remote (upload to server) - multiple ways to specify server
osync cp llama3 http://192.168.0.100:11434 # Full URL
osync cp llama3 192.168.0.100 # IP address (auto: http:// + :11434)
osync cp llama3 192.168.0.100:11434 # IP with port (auto: http://)
osync cp llama3 myserver:11434 # Hostname with port
osync cp llama3 myserver/ # Trailing slash = server
# Copy HuggingFace model to remote (uses source model name)
osync cp hf.co/unsloth/gemma-3-1b-it-GGUF:Q4_K_M 192.168.0.100
# Remote to remote (transfer between servers)
osync cp http://192.168.0.100:11434/qwen2:7b http://192.168.0.200:11434/qwen2:latest
osync cp http://server1:11434/llama3 http://server2:11434/llama3-copy
# Remote to local (download from a server to the local server)
osync cp http://192.168.0.100:11434/my-finetune my-finetune
# With custom memory buffer size (default: 512MB)
osync cp http://server1:11434/llama3 http://server2:11434/llama3 -BufferSize 256MB
osync cp http://server1:11434/qwen2 http://server2:11434/qwen2 -BufferSize 1GB
# With bandwidth throttling
osync cp llama3 http://192.168.0.100:11434 -bt 50MB
Features:
:latest tag when not specifiedhostname:port, or hostname/ auto-detected as remote serversRemote-to-remote and remote-to-local copies (push relay):
Ollama has no API to download a model, but a server can push a model to a registry. For copies out of a server, osync runs a temporary registry endpoint (the relay) on the machine where osync runs:
-BufferSize, -bt throttling applies); blobs the destination already has are skipped;This copies any model the source has, including models you created or imported, works between Ollama and xOllama in both directions, and needs no internet access.
Requirements and options:
OSYNC_RELAY_PORT; open your firewall accordingly. A Windows destination cannot store a model named after the relay's host:port, so the model is recreated there from its manifest (byte for byte); a Windows source needs the relay on port 80, which osync then uses for that copy if it is free.OSYNC_RELAY_HOST - address the servers should use to reach osync (default: the local address that routes to the source server), e.g. behind NATOSYNC_RELAY_PORT - fixed relay port (e.g. one opened in the firewall)OSYNC_RELAY_INSTALL=create - skip installing the manifest from the relay and recreate the model with /api/create (for destinations that cannot connect to osync)ls)List models with filtering and sorting options.
# List all local models
osync ls
# List models matching pattern
osync ls "llama*"
osync ls "*:7b"
osync ls "mannix/*"
# List remote models
osync ls http://192.168.0.100:11434 # Full URL
osync ls 192.168.0.100 # IP address (auto: http:// + :11434)
osync ls myserver/ # Hostname with trailing slash
osync ls "qwen*" -d 192.168.0.100 # Filter with pattern on remote
# Sort by size (descending)
osync ls --size
# Sort by size (ascending)
osync ls --sizeasc
# Sort by modified time (newest first)
osync ls --time
# Sort by modified time (oldest first)
osync ls --timeasc
Output:
NAME ID SIZE MODIFIED
llama3:latest 365c0bd3c000 5 GB 2 months ago
qwen2:7b 648f809ced2b 4 GB 1 years ago
mistral:latest 2ae6f6dd7a3d 4 GB 1 years ago
rename, mv, ren)Rename models safely by copying and deleting the original.
# Rename with implicit :latest tag
osync rename llama3 my-llama3
osync mv llama3 my-llama3
# Rename with explicit tags
osync ren llama3:7b my-custom-llama:v1
# Create versioned backup
osync mv qwen2 qwen2:backup-20241218
Features:
rm, delete, del)Delete models with pattern matching.
# Delete specific model
osync rm tinyllama
osync rm llama3:7b
# Delete with pattern
osync rm "test-*"
osync rm "*:backup"
# Delete from remote server
osync rm "old-model*" http://192.168.0.100:11434
Features:
:latest tag fallbackupdate)Update models to their latest versions locally or on remote servers.
# Update all local models
osync update
osync update "*"
# Update specific model
osync update llama3
osync update llama3:latest
# Update models matching pattern
osync update "llama*"
osync update "*:7b"
osync update "hf.co/unsloth/*"
# Update all models on remote server
osync update http://192.168.0.100:11434
osync update "*" http://192.168.0.100:11434
# Update specific models on remote server
osync update "llama*" http://192.168.0.100:11434
Features:
* (all models) when not specifiedOutput:
Updating 2 model(s)...
Updating 'llama3:latest'...
pulling manifest
pulling 6a0746a1ec1a... 100% ▕████████████████▏ 4.7 GB
✓ 'llama3:latest' updated successfully
Updating 'qwen2:7b'...
✓ 'qwen2:7b' is already up to date
show)Display detailed information about a model.
# Show information about a local model
osync show llama3
osync show qwen2:7b
# Show information about a remote model
osync show llama3 -d http://192.168.0.100:11434
# Show specific information
osync show llama3 --license # Show license only
osync show llama3 --modelfile # Show modelfile only
osync show llama3 --parameters # Show parameters only
osync show llama3 --system # Show system prompt only
osync show llama3 --template # Show template only
osync show llama3 -v # Show all information
Options:
<model> - Model name to show information (required)-d <url> - Remote server URL (default: local)--license - Show license information--modelfile - Show modelfile--parameters - Show parameters--system - Show system prompt--template - Show template-v, --verbose - Show all informationFeatures:
pull)Pull a model from the Ollama registry.
# Pull a model from the registry
osync pull llama3
osync pull qwen2:7b
osync pull hf.co/unsloth/llama3
# Pull to a remote server
osync pull llama3 http://192.168.0.100:11434
Features:
run, chat)Interactive chat with a model.
# Chat with a local model
osync run llama3
osync chat qwen2:7b
# Chat with a model on remote server
osync run llama3 -d http://192.168.0.100:11434
# With extended thinking mode (for reasoning models)
osync run qwen3 --think medium
# With verbose output (shows timing stats)
osync run llama3 --verbose
# Vision models with custom image dimensions
osync run llava --dimensions 512
Options:
<model> - Model name to run/chat with (required)-d <url> - Remote server URL (default: local)--format <value> - Response format (e.g., json)--keepalive <duration> - Keep alive duration (e.g., 5m, 1h, default: server default)--nowordwrap - Disable word wrap--verbose - Show verbose output (timing stats)--dimensions <value> - Image dimensions for vision models (e.g., 512)--hidethinking - Hide thinking process output--insecure - Allow insecure connections--think <level> - Enable extended thinking (reasoning) mode with level (low, medium, high)--truncate - Truncate long context (default: server setting)Features:
/bye or press Ctrl+D to exitps)Show models currently loaded in memory, plus system monitoring when running locally.
# Show loaded models on local server (includes GPU and process stats)
osync ps
# Show loaded models on remote server (no GPU/process stats)
osync ps http://192.168.0.100:11434 # Full URL
osync ps 192.168.0.100 # IP address (auto: http:// + :11434)
osync ps myserver/ # Hostname with trailing slash
osync ps -d myserver:11434 # Using -d flag with port
Features:
Output (local server):
Loaded Models:
---------------------------------------------------------------------------------------------------------------------------------------
NAME ID SIZE VRAM USAGE CONTEXT UNTIL
---------------------------------------------------------------------------------------------------------------------------------------
tinyllama:1.1b-chat-v1-fp16 71c2f9b69b52 2.11 GB (1B) 1.33 GB (63%) 4096 4 minutes from now
---------------------------------------------------------------------------------------------------------------------------------------
Ollama Process:
PID: 12345 CPU: 2.3% Memory: 1.2 GB (Working Set)
GPU Status (NVIDIA):
GPU 0: NVIDIA GeForce RTX 4090
Utilization: 45% Memory: 8192 MB / 24576 MB (33%) Temp: 65°C Power: 250 W / 450 W
GPU Monitoring:
psmonitor, monitor, psm)Interactive real-time monitoring dashboard inspired by nvitop, with braille-dot graphs and color-coded metrics.
# Start monitor with default settings (5s refresh, 5m history)
osync psmonitor
osync monitor
osync psm
# Custom refresh interval (2 seconds)
osync psmonitor 2s
osync psmonitor -L 2
# Custom history duration (10 minutes) - multiple formats supported
osync psmonitor -Hi 10 # Plain integer = minutes (10 minutes)
osync psmonitor -Hi 10m # Explicit minutes
osync psmonitor --history 30m # 30 minutes
osync psmonitor -Hi 1h # 1 hour
osync psmonitor -Hi 1h30m # 1 hour 30 minutes
# Combined options
osync psmonitor -L 2 -Hi 15 # 2s refresh, 15 minutes history
# Monitor remote server (shows loaded models only, no GPU/system stats)
osync psmonitor -d http://192.168.0.100:11434
Options:
-d <url> - Remote server URL (default: local). Remote mode shows loaded models only, no GPU/system stats-L <interval> - Refresh interval (default: 5s). Supports: 5, 5s, 30s, 1m, 5m, 1h30m-Hi, --history <duration> - Initial graph history duration (default: 5m). Plain integers are treated as minutes (e.g., -Hi 10 = 10 minutes). Supports: 1m, 30m, 1h, 1h30mFeatures:
Interactive Controls:
Windows Terminal Font Requirements:
The monitor uses Unicode braille characters (U+2800-U+28FF) for graphs. On Windows, you need a font that supports these characters:
Recommended fonts:
To change font in Windows Terminal:
Note: CMD with default raster fonts will show "?" for braille characters. Use Windows Terminal for best results.
load)Preload a model into memory.
# Load a model on local server
osync load llama3
osync load qwen2:7b
# Load a model on remote server
osync load llama3 -d http://192.168.0.100:11434
# URL format with embedded model name
osync load http://192.168.0.100:11434/llama3
Options:
<model> - Model name to load into memory (required)-d <url> - Remote server URL (default: local)Features:
unload)Unload a model from memory.
# Unload a specific model on local server
osync unload llama3
osync unload qwen2:7b
# Unload all models (no model name specified)
osync unload
# Unload a model on remote server
osync unload llama3 -d http://192.168.0.100:11434
# URL format with embedded model name
osync unload http://192.168.0.100:11434/llama3
Options:
<model> - Model name to unload from memory (optional - if not specified, unloads all models)-d <url> - Remote server URL (default: local)Features:
qc)Run comprehensive tests comparing quantization quality and performance across model variants.
# Compare quantizations of a model (f16 as base)
osync qc -M llama3.2 -Q q4_k_m,q5_k_m,q8_0
# Specify custom base quantization
osync qc -M llama3.2 -Q q4_k_m,q5_k_m -B fp16
# Test on remote server (multiple ways to specify)
osync qc -M llama3.2 -Q q4_k_m,q5_k_m -D http://192.168.1.100:11434
osync qc -M llama3.2 -Q q4_k_m,q5_k_m -D 192.168.1.100 # IP (auto: http:// + :11434)
osync qc -M llama3.2 -Q q4_k_m,q5_k_m -D myserver/ # trailing slash
# Custom output file
osync qc -M llama3.2 -Q q4_k_m,q5_k_m -O my-results.json
# Adjust test parameters
osync qc -M llama3.2 -Q q4_k_m,q5_k_m -Te 0.1 -S 42 -To 0.9
How It Works:
The qc command runs a comprehensive test suite (50 questions across 5 categories: Reasoning, Math, Finance, Technology, Science) on each quantization variant and captures detailed metrics using Ollama's logprobs API.
Scoring Algorithm:
Each quantization is compared against the base model using a 4-component weighted scoring system:
Token Sequence Similarity (5% weight)
Logprobs Divergence (70% weight)
100 × exp(-confidence_difference × 2)Answer Length Consistency (5% weight)
100 × exp(-2 × |1 - length_ratio|)Perplexity Score (20% weight)
exp(-average_logprob)100 × exp(-0.5 × |1 - perplexity_ratio|)Overall Confidence Score: Weighted sum of all four components (0-100%)
Color Coding:
Performance Metrics:
Results File:
Results are saved as JSON (modelname.qc.json by default) with:
Incremental Testing:
You can add new quantizations to existing results without re-testing:
# Initial test
osync qc -M llama3.2 -Q q4_k_m,q5_k_m
# Later add more quantizations (f16 and q8_0 will be skipped if already tested)
osync qc -M llama3.2 -Q q8_0,q6_k
# Force re-run testing for quantizations already in results file
osync qc -M llama3.2 -Q q4_k_m,q5_k_m --force
Resume Support:
Testing can be interrupted with Ctrl+C and resumed later:
# Start testing (press Ctrl+C, then 'y' to confirm save and exit)
osync qc -M llama3.2 -Q q4_k_m,q5_k_m,q8_0
# Resume from where you left off
osync qc -M llama3.2 -Q q4_k_m,q5_k_m,q8_0
When resuming:
Cancellation & Timeout Handling:
Wildcard Tag Selection:
Use wildcards (*) to select multiple quantizations from HuggingFace or Ollama registries:
# Test all quantizations from a HuggingFace repository
osync qc -M LFM2.5-1.2B-Instruct -b F16 -Q hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF:*
# Test only IQ quantizations (case-insensitive)
osync qc -M hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF -b F16 -Q IQ*
# Test Q4 and Q5 quantizations
osync qc -M llama3.2 -b fp16 -Q Q4*,Q5*
# Mix HuggingFace base with Ollama quants
osync qc -M LFM2.5-1.2B-Instruct -b hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF:F16 -Q IQ*,Q4*,Q5*
# Multiple HuggingFace patterns
osync qc -M LFM2.5-1.2B-Instruct -b hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF:F16 -Q hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF:IQ*,hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF:Q5*
Wildcard behavior:
* matches any characters (case-insensitive)hf.co/... models-M model as sourceOptions:
-M <name> - Model name without tag (required)-Q <tags> - Comma-separated quantization tags to compare, supports wildcards (e.g., Q4*,IQ*) (required)-B <tag> - Base quantization tag for comparison (default: fp16, or uses existing base from results file)-D <url> - Remote server URL (default: local)-O <file> - Output results file (default: modelname.qc.json)-Te <value> - Temperature (default: 0.0 for deterministic results)-S <value> - Random seed (default: 365)-To <value> - top_p parameter (default: 0.001)-Top <value> - top_k parameter (default: -1)-R <value> - repeat_penalty parameter-Fr <value> - frequency_penalty parameter-T <file> - External test suite JSON file (default: internal v1base)--force - Force re-run testing for quantizations already present in results file--rejudge - Re-run judgment process for existing test results (without re-testing)--judge <model> - Use a judge model for similarity scoring (see Judge Scoring below)--judge-ctxsize <value> - Context length for judge model (0 = auto, default: 0). Auto mode calculates: test_ctx × 2 + 2048--mode <mode> - Judge execution mode: serial (default) or parallel--timeout <seconds> - API timeout in seconds for testing and judgment calls (default: 600)--verbose - Show judgment details (question ID, score, reason) for each judged question--ondemand - Pull models on-demand if not available, then remove after testing (see On-Demand Mode below)--repo <url> - Repository URL for the model source (saved in results file for qcview)--overwrite - Overwrite existing output file without prompting--enablethinking - Enable thinking mode for thinking models (disabled by default)--thinklevel <level> - Set thinking level (low, medium, high) - overrides --enablethinking--no-unloadall - Skip unloading all models before testing--logfile <path> - Log process output to file (appends if exists, timestamps each line, strips color codes)--fix - Attempt to fix a corrupted/malformed results file and recover data (outputs to .fixed.json)Judge Scoring:
Use a second LLM to evaluate similarity between base and quantized responses:
# Local judge model
osync qc -M llama3.2 -Q 1b-instruct-q8_0,1b-instruct-q6_k --judge mistral
# Remote test models, local judge
osync qc -d http://192.168.1.100:11434/ -M llama3.2 -Q q4_k_m --judge mistral
# Remote judge model (different server)
osync qc -M llama3.2 -Q q4_k_m --judge http://192.168.1.200:11434/mistral
# Parallel mode (judge runs concurrently with testing)
osync qc -M llama3.2 -Q q4_k_m --judge mistral --mode parallel
# Re-run judgment only (use existing test results)
osync qc -M llama3.2 -Q q4_k_m --judge mistral --rejudge
When judgment scoring is enabled:
--force, --rejudge is used, or a different judge model is specifiedSeparate Best Answer Judge Model:
Use --judgebest to specify a different model for best answer determination (A/B/tie judgment):
# Same model for similarity and best answer (when using --judge alone)
osync qc -M llama3.2 -Q q4_k_m --judge mistral
# Different model for best answer judgment
osync qc -M llama3.2 -Q q4_k_m --judge mistral --judgebest llama3.3
# Best answer judgment only (no similarity scoring)
osync qc -M llama3.2 -Q q4_k_m --judgebest llama3.3
# Remote best answer judge
osync qc -M llama3.2 -Q q4_k_m --judge mistral --judgebest http://192.168.1.200:11434/llama3.3
# Re-run only best answer judgment with new model
osync qc -M llama3.2 -Q q4_k_m --judgebest llama3.3 --rejudge
When --judgebest is specified:
--judge is used), then best answer judgmentJudgeModelBestAnswer, ReasonBestAnswer, and JudgedBestAnswerAt timestampsCloud Provider Support:
Use cloud AI providers (Anthropic Claude, OpenAI, etc.) as judge models with the @provider/model syntax:
# Using environment variable for API key
osync qc -M llama3.2 -Q q4_k_m --judge @claude/claude-sonnet-4-20250514
osync qc -M llama3.2 -Q q4_k_m --judge @openai/gpt-4o
# Using explicit API key
osync qc -M llama3.2 -Q q4_k_m --judge @claude:sk-ant-xxx/claude-sonnet-4
# Azure OpenAI (key@endpoint format)
osync qc -M llama3.2 -Q q4_k_m --judge @azure:mykey@myendpoint.openai.azure.com/gpt4-deployment
# Mixed: cloud judge with ollama best answer judge
osync qc -M llama3.2 -Q q4_k_m --judge @claude/claude-sonnet-4 --judgebest llama3.2:latest
# Both judges from cloud providers
osync qc -M llama3.2 -Q q4_k_m --judge @openai/gpt-4o --judgebest @claude/claude-opus-4
Supported cloud providers and their environment variables:
| Provider | Syntax | Environment Variable |
|---|---|---|
| Anthropic Claude | @claude/model | ANTHROPIC_API_KEY |
| OpenAI | @openai/model | OPENAI_API_KEY |
| Google Gemini | @gemini/model | GEMINI_API_KEY or GOOGLE_API_KEY |
| Azure OpenAI | @azure/deployment | AZURE_OPENAI_API_KEY + AZURE_OPENAI_ENDPOINT |
| Mistral AI | @mistral/model | MISTRAL_API_KEY |
| Cohere | @cohere/model | CO_API_KEY or COHERE_API_KEY |
| Together AI | @together/model | TOGETHER_API_KEY |
| HuggingFace | @huggingface/model | HF_TOKEN or HUGGINGFACE_TOKEN |
| Replicate | @replicate/model | REPLICATE_API_TOKEN |
Use osync qc --help-cloud for detailed provider documentation and examples.
When using cloud providers:
Judge Execution Modes:
Serial mode (default): After testing each quantization, all questions are judged sequentially before moving to the next quantization. Simple and predictable execution.
Parallel mode: Testing and judgment run concurrently at the question level. As each question is tested, it is immediately passed to the judge model in a background task. This allows the test model to continue generating answers while the judge evaluates previous questions. Significantly reduces total execution time when using a remote judge or when the judge model is faster than the test model.
On-Demand Mode:
Use --ondemand to automatically pull models that aren't available and remove them after testing. This is ideal for testing large models or many quantizations without consuming permanent storage:
# Test many quantizations without keeping them all
osync qc -M llama3.2 -Q q2_k,q3_k_s,q3_k_m,q4_k_s,q4_k_m,q5_k_s,q5_k_m,q6_k,q8_0 --ondemand
# Test on remote server with on-demand
osync qc -M llama3.2 -Q q4_k_m,q8_0 -D http://192.168.1.100:11434 --ondemand
When on-demand mode is enabled:
Built-in Test Suites:
v1base (default) - 50 questions across Reasoning, Math, Finance, Technology, Sciencev1quick - 10 questions (subset of v1base for quick testing)v1code - 50 coding questions across Python, C++, C#, TypeScript, Rust (8192 max tokens)External Test Suite Format:
Create custom test suites using JSON files with the following structure:
{
"name": "my-custom-suite",
"numPredict": 4096,
"contextLength": 4096,
"categories": [
{
"id": 1,
"name": "Category Name",
"contextLength": 8192,
"questions": [
{
"categoryId": 1,
"questionId": 1,
"text": "Your question text here",
"contextLength": 16384
}
]
}
]
}
numPredict - Maximum tokens generated per response (default: 4096). Use higher values for coding or detailed answers.contextLength - Context length (num_ctx) for the model during testing. Can be specified at:
Use with: osync qc -M model -Q q4_k_m -T my-custom-suite.json
Reference files v1base.json and v1code.json are included in the osync directory.
The recommended model to use as a judge is gemma3:12b, it is advisable to use a large context size (6k for v1base and 12k for v1code).
It's also recommended to run the qc command with the --verbose argument to verify the scoring and reason given corresponds to the expected beahviour:
Judging UD-IQ3_XXS Q1-1 Score: 92% (1/50 2%)
A and B match: Both responses provide a complete, thread-safe LRU cache implementation in Python using `OrderedDict` and `threading.RLock`. They both include comprehensive docstrings, error
handling, and a full set of methods (get, put, delete, clear, size, etc.). The core logic for managing the cache, including eviction and thread safety, is nearly identical. The primary differences
are in the formatting of the docstrings and some minor wording variations in the explanations. Both responses also include example usage and testing code. The code itself is very similar, with
only minor differences in variable names and comments. Overall, the responses demonstrate a very high degree of similarity in terms of content, approach, and functionality.
Judging UD-IQ3_XXS Q1-2 Score: 75% (2/50 4%)
A and B match: Both responses provide a Python async web scraper using aiohttp, incorporating concurrent crawling, rate limiting, retries, and CSS selector-based data extraction. They both utilize
asyncio, aiohttp, logging, and dataclasses. Both include comprehensive error handling and logging. However, they differ in their implementation details. Response A uses a semaphore for concurrency
control and a class-based structure for the scraper, while Response B introduces a separate RateLimiter class and a more streamlined approach to data extraction. Response A's retry logic is more
detailed, including random jitter, while Response B's is simpler. Response B also includes an advanced scraper with custom selectors.
Judging UD-IQ3_XXS Q1-3 Score: 75% (3/50 6%)
A and B differ: Both responses implement a retry decorator factory with similar functionality (configurable max attempts, delay strategies, exception filtering, and support for both sync and async
functions). However, they differ significantly in their implementation details and structure. Response A uses a class-based approach with a `RetryError` exception and separate `_calculate_delay`
function. It also includes convenience decorators for common retry patterns. Response B uses a dataclass for configuration and an Enum for delay strategies, making it more structured. It also has
separate functions for sync and async decorators. While both achieve the same goal, the code organization and specific techniques used are quite different, leading to a noticeable difference in
qcview)Display quantization comparison results in formatted tables or export to various formats.
# View results in table format (console)
osync qcview llama3.2.qc.json
# Export table to text file
osync qcview llama3.2.qc.json -O report.txt
# Export as JSON to file
osync qcview llama3.2.qc.json -Fo json -O report.json
# Export as Markdown
osync qcview llama3.2.qc.json -Fo md
# Export as HTML (interactive with theme toggle)
osync qcview llama3.2.qc.json -Fo html
# Export as PDF (includes full Q&A pages)
osync qcview llama3.2.qc.json -Fo pdf
# Add or override repository URL in output
osync qcview llama3.2.qc.json -Fo html --repo https://github.com/user/model
Table Output:
Displays color-coded results with:
Q4_K (87%))Q6_K (81% Q8_0) when dominant differs from tag? suffix (e.g., Q3_K?)Color Coding:
Output Formats:
Options:
<file> - Results file to view (required, positional argument)-Fo <format> - Output format: table, json, md, html, pdf (default: table)-O <file> - Output filename (default: auto-generated based on format)--repo <url> - Repository URL (displayed in output, overrides value from results file)--metricsonly - Ignore judgment data and show only metrics-based scores (useful for comparing pure model output quality without judge influence)--overwrite - Overwrite existing output file without promptingbench)Run context tracking benchmarks to evaluate how well models maintain information across long conversations.
# Basic benchmark with a model
osync bench -M llama3.2
# Test on remote server (multiple ways to specify)
osync bench -M llama3.2 -D http://192.168.1.100:11434
osync bench -M llama3.2 -D 192.168.1.100 # IP (auto: http:// + :11434)
osync bench -M llama3.2 -D myserver/ # trailing slash
# With judge evaluation
osync bench -M llama3.2 --judge gemma3:12b
# Custom output file
osync bench -M llama3.2 -O my-benchmark.json
# Enable thinking for thinking models (qwen3, deepseek-r1) - disabled by default
osync bench -M qwen3 --enablethinking
# Control thinking level (overrides --enablethinking)
osync bench -M qwen3 --thinklevel medium
# Skip unloading all models before testing
osync bench -M llama3.2 --no-unloadall
# Generate a custom test suite
osync bench --generate-suite -T custom-suite.json -O custom-suite.json
How It Works:
The bench command generates dynamic stories with embedded facts and tests the model's ability to:
Question Types:
Scoring:
Each question is evaluated as:
Options:
-M <name> - Model name without tag (required unless --help-cloud, --showtools, or --generate-suite)-Q <tags> - Model/Quantization tags to compare (comma-separated, supports wildcards e.g., Q4*,IQ*)-D <url> - Remote server URL (default: local)-O <file> - Results output file (default: modelname.testtype.json)-T <file> - Test suite file or type. For bench: filename first, falls back to v1.json if type given. For --generate-suite: test type (ctxbench/ctxtoolsbench)-L <category> - Limit testing to a specific category (e.g., 2k, 4k, 8k, 16k, 32k, 64k, 128k, 256k)-Te <value> - Model temperature (default: from model template)-S <value> - Model seed (default: 365)-To <value> - Model top_p (default: from model template)-Top <value> - Model top_k (default: from model template)-R <value> - Model repeat_penalty (default: from model template)-Fr <value> - Model frequency_penalty (default: from model template)--force - Force re-run testing for quantizations/models already present in results file--rejudge - Re-run judgment process for existing test results (without re-testing)--judge <model> - Judge model for answer evaluation. Ollama: model_name or http://host:port/model. Cloud: @provider[:key]/model--mode <mode> - Judge execution mode: serial (default) or parallel--timeout <seconds> - API timeout in seconds (default: 1800)--verbose - Show testing conversation, tools usage, judgment details--judge-ctxsize <value> - Context length for judge model (0 = auto based on test type)--ondemand - Pull models on-demand if not available, then remove after testing--repo <url> - Repository URL for the model source (saved per model/quant in results file)--showtools - Display available tools with descriptions and queryable data--generate-suite - Generate test suites. Use with -T (ctxbench/ctxtoolsbench) for specific type, -O for custom filename--help-cloud - Show detailed help for cloud provider integration--enablethinking - Enable thinking mode for thinking models (disabled by default)--thinklevel <level> - Set thinking level (low, medium, high) - overrides --enablethinking--no-unloadall - Skip unloading all models before testing--overwrite - Overwrite existing output file without prompting--ctxpct <value> - Content scaling percentage for test suite generation (default: 100, range: 50-150)--logfile <path> - Log process output to file (appends if exists, timestamps each line, strips color codes)--calibrate - Enable token calibration mode: detailed tracking of estimated vs actual tokens at each step--calibrate-output <file> - Output file for calibration data (default: calibration_.json in test suite directory)--fix - Fix corrupted/malformed results JSON file (specify file with -O or -M)Bench Test Suite Format:
Custom bench test suites can specify a numPredict field to control maximum tokens per response:
{
"testType": "ctxbench",
"testDescription": "Custom context benchmark",
"numPredict": 16384,
"maxContextLength": 131072,
"categories": [...]
}
numPredict - Maximum tokens generated per test response (default: 16384)maxContextLength - Maximum context length supported by the test suitebenchview)Display context benchmark results in formatted tables or export to various formats.
# View results in table format (console)
osync benchview llama3.2.ctxbench.json
# Export to different formats
osync benchview results.json -Fo json -O report.json
osync benchview results.json -Fo md -O report.md
osync benchview results.json -Fo html -O report.html
osync benchview results.json -Fo pdf -O report.pdf
Output Information:
Output Formats:
Options:
<file> - Results file to view (required, positional argument)-Fo <format> - Output format: table, json, md, html, pdf (default: table)-O <file> - Output filename (default: auto-generated based on format)-C <category> - Filter to specific category--details - Show detailed Q&A results in output--overwrite - Overwrite existing output file without promptingmanage)Interactive TUI for managing models with keyboard shortcuts.
# Launch manage interface for local server
osync manage
# Launch manage interface for remote server (multiple ways to specify)
osync manage http://192.168.0.100:11434
osync manage 192.168.0.100 # IP (auto: http:// + :11434)
osync manage myserver/ # trailing slash
Features:
● marks models loaded in memoryxollama tweakKeys:
: - _ . / - Filter by name (* matches any text); Backspace removes a characterosync setup manage servers (also asked by osync setup server)Tweak (xOllama): xOllama keeps settings of its own in the model: engine, KV cache types, dynamic slots, DCA, session affinity and prefix pooling, council, GPU/devices, speculative decoding. They are listed in the model details (Enter), and on an xOllama server Ctrl+W changes them for the model under the cursor, or for each selected model. Pick what to change (every setting, or one group such as KV cache or council), or remove the settings. You can also type flags: a flag with a value is set without questions, e.g. --kv-k=q8_0 --kv-v=q8_0 or --council=on.
osync runs xollama tweak model <model> on the console against the server manage shows, including a remote one. The questions, the explanations and the consistency checks are xOllama's own, and so is the list of settings, so it always matches the xOllama version. This needs the xollama CLI on this machine: on PATH, or set OSYNC_XOLLAMA_CLI to its path. Only the xOllama settings layer changes: weights, template, system prompt and parameters stay as they are. See xOllama's tweak documentation.
Sort Modes:
Themes (osync setup manage themes shows them with a preview):
setup)Preferences in the settings file, by section. Without the last arguments a section asks interactively (numbered choices, Enter keeps the current value).
osync setup # summary, then a menu
osync setup show # summary only
osync setup server # ask: Ollama, xOllama or both, host, ports
osync setup server xollama 192.168.1.5 # one server (port: the flavor's default)
osync setup server ollama nas:11500
osync setup server both [host] [xollama] # both side by side (default ports; optional default server)
osync setup server auto # back to auto-detection
osync setup server env ignore # the settings win over XOLLAMA_HOST / OLLAMA_HOST (env use: they win)
osync setup alias # list (and add/remove interactively)
osync setup alias add gpu 192.168.1.10 # port 11434, or 22434 when only xOllama answers
osync setup alias remove gpu
osync setup manage # theme and default sort order
osync setup manage theme "Tokyo Night" # name, loose spelling (tokyo-night) or number
osync setup manage sort size- # name+ name- size- size+ created- created+
osync setup manage servers gpu,nas # aliases for Ctrl+Left/Right in manage (all, none)
osync setup manage themes # all themes with a preview
osync setup shell # theme, color mode, tab completion
osync setup shell theme gruvbox-light # or: plain, default
osync setup shell colors truecolor # auto, truecolor, 256, 16, none
osync setup shell completion # install tab completion (bash / PowerShell)
osync setup shell themes
showversion, -v)Display osync version and environment information.
# Show version only
osync -v
osync showversion
# Show detailed environment information
osync showversion --verbose
Basic Output:
osync v1.2.6 (b20260110-1117)
Verbose Output:
osync v1.2.6 (b20260110-1117)
Binary path: C:\Users\user\.osync\osync.exe
Installed: Yes (C:\Users\user\.osync)
Shell: PowerShell Core 7.5.4
Tab completion: Installed (C:\Users\user\Documents\PowerShell\Microsoft.PowerShell_profile.ps1)
Verbose Information:
install)Install osync to user directory and configure shell completion.
# Run the installer
osync install
Installation Directory:
~/.osync~/.local/binFeatures:
-h, -? - Show help for any command-bt <value> - Bandwidth throttling (B, KB, MB, GB per second)-BufferSize <value> - Memory buffer size for remote-to-remote copy (KB, MB, GB; default: 512MB)Examples:
# Bandwidth throttling
osync cp llama3 http://server:11434 -bt 75MB # Limit to 75 MB/s
osync cp qwen2 http://server:11434 -bt 1GB # Limit to 1 GB/s
# Memory buffer configuration for remote-to-remote
osync cp http://server1:11434/llama3 http://server2:11434/llama3 -BufferSize 256MB
osync cp http://server1:11434/qwen2 http://server2:11434/qwen2 -BufferSize 1GB
--size - Sort by size (largest first)--sizeasc - Sort by size (smallest first)--time - Sort by modified time (newest first)--timeasc - Sort by modified time (oldest first)Run osync without arguments to enter interactive mode with:
osync
# Press tab to see available models
# Type command and press enter
> cp llama3 my-backup
> ls "qwen*"
> quit # or type 'exit', or press Ctrl+C
Use * wildcard for flexible pattern matching:
# Match any model starting with "llama"
osync ls "llama*"
# Match any 7b model
osync ls "*:7b"
# Match models in namespace
osync ls "mannix/*"
# Match HuggingFace models
osync ls "hf.co/*"
# Match any model containing "test"
osync ls "*test*"
# Update all local models to latest versions
osync update
# Update specific models matching pattern
osync update "llama*"
# Update all models on remote server
osync update http://192.168.0.10:11434
# Create backup before updating
osync cp llama3:latest llama3:backup
osync update llama3
# Upload from local to multiple servers (short form with IP)
osync cp llama3 192.168.0.10
osync cp llama3 192.168.0.11
osync cp llama3 192.168.0.12
# Or using hostnames with trailing slash
osync cp llama3 server1/
osync cp llama3 server2/
osync cp llama3 server3/
# Copy between remote servers
osync cp http://192.168.0.10:11434/llama3 http://192.168.0.11:11434/llama3
osync cp http://192.168.0.10:11434/qwen2 http://192.168.0.12:11434/qwen2
# List and remove test models
osync ls "test-*"
osync rm "test-*"
# Remove old backups
osync rm "*:backup"
# List all models by size to find space hogs
osync ls --size
# Rename for better organization
osync mv llama3 llama3-8b:prod
osync mv qwen2 qwen2-7b:dev
None
v1.4.2
New
manage tweaks xOllama model settings - on an xOllama server, Ctrl+W runs xollama tweak model for the model under the cursor, or for each selected model, against the server manage shows. You can walk every setting, one group (KV cache, dynamic slots, DCA, session pooling, council, GPU/devices, engine), or remove the settings, and add flags such as --kv-k=q8_0 that are set without questions. The model details list the xOllama settings. Needs the xollama CLI on PATH, or OSYNC_XOLLAMA_CLIv1.4.1
Fixes
host:port is not a valid folder name there). The model used to be recreated from /api/show, which dropped the renderer, parser, requires and xOllama's model settings (a council became a plain model), merged several licenses and changed the parameters, yet osync reported success. It is now recreated from the source's manifest with the config and settings layers verbatim, so every layer and the config have the source's digestsollama show --modelfile and lost the same parts)HF_TOKEN is sent on osync's own HuggingFace requests (tag lookups for hf.co/... wildcard tags, qc model discovery), for higher rate limits and private or gated repositories; only to huggingface.co / hf.co over HTTPSthrough the relay, recreated from its manifest)v1.4.0
New
XOLLAMA_HOST, OLLAMA_HOST, then localhost:11434 / localhost:22434osync ps shows Server: Ollama|xOllama at <url>)xollama CLI when the local server is xOllama (override with OSYNC_OLLAMA_CLI)XOLLAMA_MODELS honored for the models directory; xOllama processes shown in process statsollama, xollama)osync setup - one command for the preferences: server (Ollama, xOllama or both), alias (server aliases), manage (theme, default sort order, servers), shell (colors, theme, tab completion), show; interactive or with arguments for scriptssettings.json in the per-OS configuration folder: local server (Ollama/xOllama, host, port), aliases, color mode, manage and shell settingsosync setup alias add gpu 192.168.1.10, then -d gpu, osync cp model gpu/, osync cp gpu/model copy, osync manage gpuXOLLAMA_HOST / OLLAMA_HOST - osync setup server asks when one is set, osync setup server env ignore|use, and a check box in the manage settings (Ctrl+E)manage rewritten on Terminal.Gui 2 - true color with multi-color themes (one color per column, ● for models loaded in memory, colored top and bottom bars) adapted to 256 and 16-color terminals with contrast checks; theme picker with live preview (Ctrl+T), the theme is saved in the preferences file; settings dialog (Ctrl+E) for the local server (Ollama/xOllama, host, port, connection test) and the color mode; column headers; F1 help; rename on F2 (Ctrl+M is Enter in most terminals); load runs in the background; console operations (copy, run, update, pull) return to the list without restarting osync; pull validation no longer rejects hf.co/... models; confirmations default to the safe answer (Enter cancels a delete)manage switches servers with Ctrl+Left / Ctrl+Right - the local server, the other one of Ollama/xOllama side by side, and the aliases chosen with osync setup manage servers (osync setup server offers them); the top bar shows which one ([2/3] gpu: Ollama @ ...)ls, ps, show, -v, copy/pull/update progress, errors, warnings and results use the shell theme (removed in 1.0.1 because of garbled output on Linux); plain when redirected, with NO_COLOR or the plain thememanage and the shell, 7 of them for light terminals, all checked for readable contrast at true color, 256 and 16 colorsTERM with xterm-16color, so every terminal was treated as 16-color); colorMode setting, OSYNC_COLOR_MODE and NO_COLOR override it; osync -v --verbose shows the resultosync install asks for the local server only when needed - a single server found on this machine is used without questions; when none or both are found it asks for the type (Ollama, xOllama or both), host and port, with detected defaults and a connection test; installing from a renamed binary (e.g. osync-macos-arm64 install) worksosync -v shows the real build time (UTC) for every binary, including renamed ones (osync-macos-arm64) and downloaded copies, instead of the file's modification timeFixes
osync <command> -h and missing-argument errors hanging forever on Linux/macOS when output is redirected (pipes, scripts, CI): the help renderer looped endlessly at 100% CPU with growing memoryrun, ps, qc ignoring OLLAMA_HOST; manage, psmonitor and local judge models now use the same local-server resolutionps truncating model names to 20 characters when output is redirectedshow printing only the Modelfile: it now shows the same sections as ollama show (model details, capabilities, projector, parameters, system, license; all metadata with -v), several section flags can be combined, and a missing model exits with an errorrm and update exiting with code 0 when no model matches, or when deleting/updating a model failed (update of all models on an empty server is still a success)osync cp http://server/qwen3:4b ...): the : of the tag was taken for a port. A server given without a port now uses 11434, or 22434 when the host refuses 11434 but accepts 22434 (xOllama)manage showing unknown quantization (and no parameters/family) when the local models directory and the resolved local server did not match (e.g. Ollama and xOllama both installed): local models are now listed through the local server's API, so the list and its details always come from the same server, and startup no longer makes one /api/show call per model. The top bar shows which server is used (Ollama @ localhost:11434); the models directory is only read when the server is unreachable, which the top bar saysollama service user): the upload now goes through the local server with the push relaymanage not scrolling when the terminal is shorter than the list-bt) not limiting short bursts, counting requested instead of read bytes, and misbehaving after ~25 days of uptimeESC[0m) in redirected output on Linux/macOSPlatform
net10.0; builds and tests on Windows, Linux and macOSosync-macos-arm64 and osync-macos-x64dev builds publish pre-releases, master publishes releasesv1.2.9
-Hi shortcut for --history argument (e.g., osync monitor -Hi 10)--history are now treated as minutes (e.g., -Hi 10 = 10 minutes)osync v1.2.9 (b20260116-1814))--enablethinking and --thinklevel arguments for thinking models (qwen3, deepseek-r1)--no-unloadall argument to skip unloading all models before testing--overwrite argument to skip file overwrite prompts--generate-suite to create custom test suite JSON files with -T and -O options--mode=parallel for parallel judgment - judges answers in background while testing continues--overwrite argument to skip file overwrite promptsnumPredict)numPredict field in bench test suite JSON
ps) - Extended system monitoring
osync -h output to show only global options and available commandsosync <command> -h for detailed help on specific commands\x1b[0m) on exit prevents color leakage to shell promptosync cp model 192.168.0.100)qc and bench commands silently exiting when remote test server is unreachable - now shows clear "Could not connect to server" error message←[0m) appearing after command output on Windows--enablethinking and --thinklevel arguments for thinking models (qwen3, deepseek-r1)--no-unloadall argument to skip unloading all models before testing--overwrite argument to overwrite existing output file without prompting--ondemand mode/api/chat to lightweight /api/generate call--fix argument to recover corrupted/malformed JSON results files (outputs to .fixed.json)
--overwrite argument to overwrite existing output file without promptingfile1-file2.html)--details)--logfile argument for qc and bench commands
[2026-01-17 06:38:20.176])http://192.168.1.100 becomes http://192.168.1.100:11434)http:// protocol when not specified (e.g., 192.168.1.100:11434 becomes http://192.168.1.100:11434)myserver/ → http://myserver:11434) since model names cannot end with /http:// or https://) → remote server192.168.0.100) → remote serverhost:11434) → remote server/ (host/) → remote serverlocalhost → remote server/ (host) → treated as model namev1.2.8
--judge and --judgebest
@provider[:token]/model (e.g., @claude/claude-sonnet-4-20250514, @openai/gpt-4o)--help-cloud option for detailed provider documentationv1.2.7
--judgebest can be used alone or combined with --judge for different models--judge: local model name or http://host:port/model for remote--rejudge to re-run only best answer judgment with new modelOsyncVersion - Version of osync used for testingOllamaVersion - Ollama server version for test quantizationsOllamaJudgeVersion - Ollama version for judge server (similarity scoring)OllamaJudgeBestAnswerVersion - Ollama version for best answer judge server/api/version endpointstop parameter serialization (now correctly sent as array instead of string)ConvertParameterValue helper ensures correct JSON types for all Ollama model parametershf.co/... models
osync load http://host:port/modelname in addition to osync load modelname -d host--rejudge with existing results, only the judge model is neededv1.2.6
-Fo md, -Fo html, or -Fo pdf to select format--repo argument to specify model source repository
qc testing and overridden in qcviewosync version (alias -v) to display version info
--verbose flag displays detailed info: binary path, installation status, shell type/version, tab completion statusinstalled v1.2.6 (b20260110-1156) is older)Digest) and short digest (ShortDigest, first 12 chars) stored in results JSONollama ls)
osync ls and manage TUI compute SHA256 of manifest file contentollama ls output for easy cross-reference...0B-A3B-Instruct-GGUF:Q4_K_S)load_duration from response✓ Model 'model:tag' loaded successfully (2m 15s) (API: 2m 5s)"quantization_level": "unknown" for HuggingFace modelsQ4_0 (87%) or Q6_K (81% Q8_0) showing actual tensor distributionverbose=true to fetch tensor metadataQ3_K?) to indicate uncertainty-b, no longer tries to pull the base model--force to re-run the base model if neededosync ls code matches models starting with "code" (prefix match, same as code*)osync ls *q4_k_m finds all models ending with "q4_k_m" (useful for finding by quantization)osync ls *code* finds models containing "code" anywhere in the nameosync ls 'gemma*'osync pull gemma3:1b-it-q* pulls all matching tagsosync pull hf.co/unsloth/gemma-3-1b-it-GGUF:IQ2*osync pull -d http://server:11434 gemma3:1b-it-q*bestanswer: A (base better), B (quant better), or AB (tie)Score: 75% (27/50 54%) Best: AB--judge is active and bestanswer is missing67% (B:10 A:5 =:3) showing quant won 67% of non-tie comparisonsBestAnswer field (A/B/AB)BestCount, WorstCount, TieCount, BestPercentage, WorstPercentage, TiePercentageCategoryBestStats with counts and percentages--metricsonly argument to ignore judgment data
--judge-ctxsize is 0 (new default), calculates: test_ctx × 2 + 2048v1.2.5
hf.co/namespace/repo:tag) are now preserved correctly instead of being incorrectly derived from the -M argumentv1.2.4
--rejudge argument to re-run judgment process for existing test results--force which re-runs both testing and judgment, --rejudge only re-runs judgment--ondemand) now properly verifies HuggingFace models (hf.co/...)-b qwen3-coder:30b-a3b-fp16 or -b hf.co/namespace/repo:tag) is now used as-is-M model name*) in -Q argument (e.g., Q4*, IQ*, *)hf.co/... modelsModelTagResolver class for reusable tag resolution across commands--timeout) for model loadingQ4_0 is stored as q4_0 causing preload to failv1.2.3
--ondemand argument to enable on-demand model managementcontextLength property in built-in test suites (v1base, v1quick, v1code)contextLength at suite, category, and question levels--judge-ctxsize argument to configure judge model context lengthosync ls qwen2:)v1.2.2
--verbose flag to show judgment details
v1.2.1
user/model:tag)IsBase flag are automatically repaired on load/ or \ are now converted to - in default output filename
v1.2.0
v1code test suite for evaluating code generation quality
-T v1code or via external v1code.json filenumPredict values
numPredict property (default: 4096)registry.ollama.ai/v2/ manifest endpoint instead of HTML scraping--timeout argument for testing and judgment API calls
--timeout <seconds> argumentv1.1.9
--judge <model> - Specify local or remote judge model--mode serial|parallel - Serial (after each quant) or parallel (concurrent) execution--mode parallel, each question is immediately judged in the background as soon as testing completes/api/chat endpoint with structured JSON output schemaReason field to capture judge's reasoning for each scoreosync qcview results.json)v1.1.8
-T path/to/suite.json
v1base.json file for reference and portability--force flag to re-run testing for quantizations already in results filefp16 and is automatically detected from existing results/api/tagsosync ls to accept model name as positional argument (e.g., osync ls qw works like osync ls qw*)v1.1.7
qc and qcview commands for comprehensive quality testing
qcview (90-100% green, 80-90% lime, 70-80% yellow, 50-70% orange, <50% red)/api/show endpoint doesn't include size field, now fetches from /api/tagsv1.1.6
ps command improvements
load command to preload models into memory with configurable keep-aliveunload command to free VRAM immediately/bye or Ctrl+D to exitv1.1.0
-BufferSize parameter (e.g., 256MB, 1GB)v1.0.9
v1.0.8
ren and mv for safe model renaming--size, --sizeasc, --time, --timeascv1.0.7
* wildcard:latest tag fallback for copy and remove commands when tag not specified/ in names)v1.0.6
copy (alias cp)latest tag if none specifiedv1.0.5
v1.0.4
v1.0.3
-bt 75MBv1.0.2
v1.0.1
v1.0.0
See the open issues for a list of proposed features (and known issues).
This project is licensed under the MIT license.
See LICENSE for more information.
C#
98.1%
Gherkin
1.6%
Transfer a local ollama model to a remote ollama server
C#
15
54 commits
updated Oct 3, 2026
osync is a powerful command-line tool for managing Ollama models across local and remote servers.
:latest tag when not specified* wildcards for batch operationsmanage and the colored command output, adapted to 256 and 16-color terminals[C# .NET 10]
[Windows/Linux/MacOS]
[Arm64/x64/Mac]
[Download latest binary release]
[Build from sources]
Clone the repo
dotnet build(any OS, .NET 10 SDK) or Visual Studio 2022;dotnet publish osync/osync.csproj -c Release -r <win-x64|linux-x64|osx-arm64|osx-x64>for a single-file executable. See docs/DEVELOPMENT.md for tests, CI and releases.
Commands without -d work on the local server, found in this order:
XOLLAMA_HOST (xOllama), then OLLAMA_HOST (Ollama); bind addresses such as 0.0.0.0 are mapped to localhost. With "ignoreEnvironment": true (osync setup server env ignore, or the question osync setup server asks when one of them is set) the configured server comes firstosync setup server)localhost:11434 (Ollama default) or localhost:22434 (xOllama default)osync install configures it: when exactly one server answers on this machine it is used without questions; when none or both answer, it asks for the server type (1 Ollama, 2 xOllama, 3 both side by side), host and ports. With both, one is the default local server and each gets an alias (ollama, xollama), so the other one is always at hand: osync ls -d xollama, osync cp model xollama/. Change it any time with osync setup server.
osync detects whether a server is Ollama or xOllama (osync ps shows it) and, for local operations, runs the matching CLI: xollama when the local server is xOllama (or only xollama is installed), else ollama. Set OSYNC_OLLAMA_CLI to force a specific CLI. The models directory is taken from XOLLAMA_MODELS, then OLLAMA_MODELS, then the platform default.
Remote servers given without a port use :11434, or :22434 when the host only answers there (xOllama).
osync keeps user preferences in settings.json in the per-OS configuration folder (osync -v --verbose shows the path):
| OS | Location |
|---|---|
| Windows | %APPDATA%\osync\settings.json |
| macOS | ~/Library/Application Support/osync/settings.json |
| Linux | $XDG_CONFIG_HOME/osync/settings.json (default ~/.config/osync/settings.json) |
{
"server": { "flavor": "ollama", "host": "localhost", "both": true },
"aliases": {
"ollama": "http://localhost:11434",
"xollama": "http://localhost:22434",
"gpu": "http://192.168.1.10:11434"
},
"colorMode": "auto",
"manage": { "theme": "Dracula", "sort": "size-", "servers": [ "gpu" ] },
"shell": { "theme": "Tokyo Night" }
}
server - the local server: flavor (auto, ollama, xollama), host and port (default: the flavor's port); both when Ollama and xOllama run side by side (flavor is then the default one). XOLLAMA_HOST / OLLAMA_HOST override it, unless ignoreEnvironment is true.aliases - server aliases.colorMode - auto (detect), truecolor, 256, 16 or none; OSYNC_COLOR_MODE and NO_COLOR override it.manage.theme, manage.sort - theme (also chosen with Ctrl+T in manage) and initial sort order (name+, name-, size-, size+, created-, created+) of osync manage.manage.servers - aliases manage switches to with Ctrl+Left / Ctrl+Right, after the local server (with Ollama and xOllama side by side the other one is always included).shell.theme - colors of the command output (any theme name, or plain for no colors).OSYNC_CONFIG_DIR moves the settings folder.Everything can be changed with osync setup instead of editing the file.
An alias is a short name for a server (osync setup alias add gpu 192.168.1.10), usable wherever osync expects a server:
osync ls -d gpu # list the models of the server
osync cp qwen3:8b gpu/ # upload
osync cp gpu/qwen3:8b qwen3-gpu # download
osync manage gpu
An alias takes precedence over a model namespace with the same name (gpu/model). Names start with a letter and use letters, digits, - and _.
osync detects the terminal's color depth from COLORTERM (truecolor/24bit), TERM (*-256color, *-direct), TERM_PROGRAM and Windows Terminal, and uses true color, 256 or 16 colors accordingly (osync -v --verbose shows what was detected and why). SSH does not forward COLORTERM by default, so terminals that support true color are seen as 256-color over SSH/tmux: set "colorMode": "truecolor" in the settings file, or OSYNC_COLOR_MODE=truecolor, or forward COLORTERM (SendEnv COLORTERM / AcceptEnv COLORTERM). NO_COLOR disables colors.
Command output is colored with the shell theme (osync setup shell theme NAME; osync setup shell themes shows every theme with a preview): the same 34 themes as manage, with light themes for light terminals (the default follows COLORFGBG when the terminal sets it), or plain. Output that goes to a pipe or a file is never colored.
osync manage draws with 24-bit colors; at 256 colors every theme color is snapped to the xterm-256 palette (tmux and most 256-color terminals then show exactly that color), at 16 colors the themes switch to the 16 standard colors with contrast checks. macOS Terminal.app gets 16 colors unless colorMode says truecolor (it misreads 24-bit colors on older macOS). If manage does not start or draw correctly on an old Windows console, OSYNC_TUI_DRIVER=windows selects the Windows console driver (ansi and dotnet are the others).
An AI generated DeepWiki is available here:
https://deepwiki.com/mann1x/osync
# Interactive TUI for model management (recommended)
osync manage
# Copy local model to remote server
osync cp llama3 http://192.168.100.100:11434
# Copy to remote server (short forms - port 11434 and http:// are defaults)
osync cp llama3 192.168.100.100 # IP address detected as server
osync cp llama3 myserver:11434 # hostname with port
osync cp llama3 myserver/ # trailing slash indicates server
# List all local models
osync ls
# Chat with a model
osync run llama3
# Interactive REPL mode with tab completion
osync
cp)Copy models locally, to remote servers, or between remote servers.
# Local copy (create backup)
osync cp llama3 my-backup-llama3
osync cp llama3:70b llama3:backup-v1
# Local to remote (upload to server) - multiple ways to specify server
osync cp llama3 http://192.168.0.100:11434 # Full URL
osync cp llama3 192.168.0.100 # IP address (auto: http:// + :11434)
osync cp llama3 192.168.0.100:11434 # IP with port (auto: http://)
osync cp llama3 myserver:11434 # Hostname with port
osync cp llama3 myserver/ # Trailing slash = server
# Copy HuggingFace model to remote (uses source model name)
osync cp hf.co/unsloth/gemma-3-1b-it-GGUF:Q4_K_M 192.168.0.100
# Remote to remote (transfer between servers)
osync cp http://192.168.0.100:11434/qwen2:7b http://192.168.0.200:11434/qwen2:latest
osync cp http://server1:11434/llama3 http://server2:11434/llama3-copy
# Remote to local (download from a server to the local server)
osync cp http://192.168.0.100:11434/my-finetune my-finetune
# With custom memory buffer size (default: 512MB)
osync cp http://server1:11434/llama3 http://server2:11434/llama3 -BufferSize 256MB
osync cp http://server1:11434/qwen2 http://server2:11434/qwen2 -BufferSize 1GB
# With bandwidth throttling
osync cp llama3 http://192.168.0.100:11434 -bt 50MB
Features:
:latest tag when not specifiedhostname:port, or hostname/ auto-detected as remote serversRemote-to-remote and remote-to-local copies (push relay):
Ollama has no API to download a model, but a server can push a model to a registry. For copies out of a server, osync runs a temporary registry endpoint (the relay) on the machine where osync runs:
-BufferSize, -bt throttling applies); blobs the destination already has are skipped;This copies any model the source has, including models you created or imported, works between Ollama and xOllama in both directions, and needs no internet access.
Requirements and options:
OSYNC_RELAY_PORT; open your firewall accordingly. A Windows destination cannot store a model named after the relay's host:port, so the model is recreated there from its manifest (byte for byte); a Windows source needs the relay on port 80, which osync then uses for that copy if it is free.OSYNC_RELAY_HOST - address the servers should use to reach osync (default: the local address that routes to the source server), e.g. behind NATOSYNC_RELAY_PORT - fixed relay port (e.g. one opened in the firewall)OSYNC_RELAY_INSTALL=create - skip installing the manifest from the relay and recreate the model with /api/create (for destinations that cannot connect to osync)ls)List models with filtering and sorting options.
# List all local models
osync ls
# List models matching pattern
osync ls "llama*"
osync ls "*:7b"
osync ls "mannix/*"
# List remote models
osync ls http://192.168.0.100:11434 # Full URL
osync ls 192.168.0.100 # IP address (auto: http:// + :11434)
osync ls myserver/ # Hostname with trailing slash
osync ls "qwen*" -d 192.168.0.100 # Filter with pattern on remote
# Sort by size (descending)
osync ls --size
# Sort by size (ascending)
osync ls --sizeasc
# Sort by modified time (newest first)
osync ls --time
# Sort by modified time (oldest first)
osync ls --timeasc
Output:
NAME ID SIZE MODIFIED
llama3:latest 365c0bd3c000 5 GB 2 months ago
qwen2:7b 648f809ced2b 4 GB 1 years ago
mistral:latest 2ae6f6dd7a3d 4 GB 1 years ago
rename, mv, ren)Rename models safely by copying and deleting the original.
# Rename with implicit :latest tag
osync rename llama3 my-llama3
osync mv llama3 my-llama3
# Rename with explicit tags
osync ren llama3:7b my-custom-llama:v1
# Create versioned backup
osync mv qwen2 qwen2:backup-20241218
Features:
rm, delete, del)Delete models with pattern matching.
# Delete specific model
osync rm tinyllama
osync rm llama3:7b
# Delete with pattern
osync rm "test-*"
osync rm "*:backup"
# Delete from remote server
osync rm "old-model*" http://192.168.0.100:11434
Features:
:latest tag fallbackupdate)Update models to their latest versions locally or on remote servers.
# Update all local models
osync update
osync update "*"
# Update specific model
osync update llama3
osync update llama3:latest
# Update models matching pattern
osync update "llama*"
osync update "*:7b"
osync update "hf.co/unsloth/*"
# Update all models on remote server
osync update http://192.168.0.100:11434
osync update "*" http://192.168.0.100:11434
# Update specific models on remote server
osync update "llama*" http://192.168.0.100:11434
Features:
* (all models) when not specifiedOutput:
Updating 2 model(s)...
Updating 'llama3:latest'...
pulling manifest
pulling 6a0746a1ec1a... 100% ▕████████████████▏ 4.7 GB
✓ 'llama3:latest' updated successfully
Updating 'qwen2:7b'...
✓ 'qwen2:7b' is already up to date
show)Display detailed information about a model.
# Show information about a local model
osync show llama3
osync show qwen2:7b
# Show information about a remote model
osync show llama3 -d http://192.168.0.100:11434
# Show specific information
osync show llama3 --license # Show license only
osync show llama3 --modelfile # Show modelfile only
osync show llama3 --parameters # Show parameters only
osync show llama3 --system # Show system prompt only
osync show llama3 --template # Show template only
osync show llama3 -v # Show all information
Options:
<model> - Model name to show information (required)-d <url> - Remote server URL (default: local)--license - Show license information--modelfile - Show modelfile--parameters - Show parameters--system - Show system prompt--template - Show template-v, --verbose - Show all informationFeatures:
pull)Pull a model from the Ollama registry.
# Pull a model from the registry
osync pull llama3
osync pull qwen2:7b
osync pull hf.co/unsloth/llama3
# Pull to a remote server
osync pull llama3 http://192.168.0.100:11434
Features:
run, chat)Interactive chat with a model.
# Chat with a local model
osync run llama3
osync chat qwen2:7b
# Chat with a model on remote server
osync run llama3 -d http://192.168.0.100:11434
# With extended thinking mode (for reasoning models)
osync run qwen3 --think medium
# With verbose output (shows timing stats)
osync run llama3 --verbose
# Vision models with custom image dimensions
osync run llava --dimensions 512
Options:
<model> - Model name to run/chat with (required)-d <url> - Remote server URL (default: local)--format <value> - Response format (e.g., json)--keepalive <duration> - Keep alive duration (e.g., 5m, 1h, default: server default)--nowordwrap - Disable word wrap--verbose - Show verbose output (timing stats)--dimensions <value> - Image dimensions for vision models (e.g., 512)--hidethinking - Hide thinking process output--insecure - Allow insecure connections--think <level> - Enable extended thinking (reasoning) mode with level (low, medium, high)--truncate - Truncate long context (default: server setting)Features:
/bye or press Ctrl+D to exitps)Show models currently loaded in memory, plus system monitoring when running locally.
# Show loaded models on local server (includes GPU and process stats)
osync ps
# Show loaded models on remote server (no GPU/process stats)
osync ps http://192.168.0.100:11434 # Full URL
osync ps 192.168.0.100 # IP address (auto: http:// + :11434)
osync ps myserver/ # Hostname with trailing slash
osync ps -d myserver:11434 # Using -d flag with port
Features:
Output (local server):
Loaded Models:
---------------------------------------------------------------------------------------------------------------------------------------
NAME ID SIZE VRAM USAGE CONTEXT UNTIL
---------------------------------------------------------------------------------------------------------------------------------------
tinyllama:1.1b-chat-v1-fp16 71c2f9b69b52 2.11 GB (1B) 1.33 GB (63%) 4096 4 minutes from now
---------------------------------------------------------------------------------------------------------------------------------------
Ollama Process:
PID: 12345 CPU: 2.3% Memory: 1.2 GB (Working Set)
GPU Status (NVIDIA):
GPU 0: NVIDIA GeForce RTX 4090
Utilization: 45% Memory: 8192 MB / 24576 MB (33%) Temp: 65°C Power: 250 W / 450 W
GPU Monitoring:
psmonitor, monitor, psm)Interactive real-time monitoring dashboard inspired by nvitop, with braille-dot graphs and color-coded metrics.
# Start monitor with default settings (5s refresh, 5m history)
osync psmonitor
osync monitor
osync psm
# Custom refresh interval (2 seconds)
osync psmonitor 2s
osync psmonitor -L 2
# Custom history duration (10 minutes) - multiple formats supported
osync psmonitor -Hi 10 # Plain integer = minutes (10 minutes)
osync psmonitor -Hi 10m # Explicit minutes
osync psmonitor --history 30m # 30 minutes
osync psmonitor -Hi 1h # 1 hour
osync psmonitor -Hi 1h30m # 1 hour 30 minutes
# Combined options
osync psmonitor -L 2 -Hi 15 # 2s refresh, 15 minutes history
# Monitor remote server (shows loaded models only, no GPU/system stats)
osync psmonitor -d http://192.168.0.100:11434
Options:
-d <url> - Remote server URL (default: local). Remote mode shows loaded models only, no GPU/system stats-L <interval> - Refresh interval (default: 5s). Supports: 5, 5s, 30s, 1m, 5m, 1h30m-Hi, --history <duration> - Initial graph history duration (default: 5m). Plain integers are treated as minutes (e.g., -Hi 10 = 10 minutes). Supports: 1m, 30m, 1h, 1h30mFeatures:
Interactive Controls:
Windows Terminal Font Requirements:
The monitor uses Unicode braille characters (U+2800-U+28FF) for graphs. On Windows, you need a font that supports these characters:
Recommended fonts:
To change font in Windows Terminal:
Note: CMD with default raster fonts will show "?" for braille characters. Use Windows Terminal for best results.
load)Preload a model into memory.
# Load a model on local server
osync load llama3
osync load qwen2:7b
# Load a model on remote server
osync load llama3 -d http://192.168.0.100:11434
# URL format with embedded model name
osync load http://192.168.0.100:11434/llama3
Options:
<model> - Model name to load into memory (required)-d <url> - Remote server URL (default: local)Features:
unload)Unload a model from memory.
# Unload a specific model on local server
osync unload llama3
osync unload qwen2:7b
# Unload all models (no model name specified)
osync unload
# Unload a model on remote server
osync unload llama3 -d http://192.168.0.100:11434
# URL format with embedded model name
osync unload http://192.168.0.100:11434/llama3
Options:
<model> - Model name to unload from memory (optional - if not specified, unloads all models)-d <url> - Remote server URL (default: local)Features:
qc)Run comprehensive tests comparing quantization quality and performance across model variants.
# Compare quantizations of a model (f16 as base)
osync qc -M llama3.2 -Q q4_k_m,q5_k_m,q8_0
# Specify custom base quantization
osync qc -M llama3.2 -Q q4_k_m,q5_k_m -B fp16
# Test on remote server (multiple ways to specify)
osync qc -M llama3.2 -Q q4_k_m,q5_k_m -D http://192.168.1.100:11434
osync qc -M llama3.2 -Q q4_k_m,q5_k_m -D 192.168.1.100 # IP (auto: http:// + :11434)
osync qc -M llama3.2 -Q q4_k_m,q5_k_m -D myserver/ # trailing slash
# Custom output file
osync qc -M llama3.2 -Q q4_k_m,q5_k_m -O my-results.json
# Adjust test parameters
osync qc -M llama3.2 -Q q4_k_m,q5_k_m -Te 0.1 -S 42 -To 0.9
How It Works:
The qc command runs a comprehensive test suite (50 questions across 5 categories: Reasoning, Math, Finance, Technology, Science) on each quantization variant and captures detailed metrics using Ollama's logprobs API.
Scoring Algorithm:
Each quantization is compared against the base model using a 4-component weighted scoring system:
Token Sequence Similarity (5% weight)
Logprobs Divergence (70% weight)
100 × exp(-confidence_difference × 2)Answer Length Consistency (5% weight)
100 × exp(-2 × |1 - length_ratio|)Perplexity Score (20% weight)
exp(-average_logprob)100 × exp(-0.5 × |1 - perplexity_ratio|)Overall Confidence Score: Weighted sum of all four components (0-100%)
Color Coding:
Performance Metrics:
Results File:
Results are saved as JSON (modelname.qc.json by default) with:
Incremental Testing:
You can add new quantizations to existing results without re-testing:
# Initial test
osync qc -M llama3.2 -Q q4_k_m,q5_k_m
# Later add more quantizations (f16 and q8_0 will be skipped if already tested)
osync qc -M llama3.2 -Q q8_0,q6_k
# Force re-run testing for quantizations already in results file
osync qc -M llama3.2 -Q q4_k_m,q5_k_m --force
Resume Support:
Testing can be interrupted with Ctrl+C and resumed later:
# Start testing (press Ctrl+C, then 'y' to confirm save and exit)
osync qc -M llama3.2 -Q q4_k_m,q5_k_m,q8_0
# Resume from where you left off
osync qc -M llama3.2 -Q q4_k_m,q5_k_m,q8_0
When resuming:
Cancellation & Timeout Handling:
Wildcard Tag Selection:
Use wildcards (*) to select multiple quantizations from HuggingFace or Ollama registries:
# Test all quantizations from a HuggingFace repository
osync qc -M LFM2.5-1.2B-Instruct -b F16 -Q hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF:*
# Test only IQ quantizations (case-insensitive)
osync qc -M hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF -b F16 -Q IQ*
# Test Q4 and Q5 quantizations
osync qc -M llama3.2 -b fp16 -Q Q4*,Q5*
# Mix HuggingFace base with Ollama quants
osync qc -M LFM2.5-1.2B-Instruct -b hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF:F16 -Q IQ*,Q4*,Q5*
# Multiple HuggingFace patterns
osync qc -M LFM2.5-1.2B-Instruct -b hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF:F16 -Q hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF:IQ*,hf.co/LiquidAI/LFM2.5-1.2B-Instruct-GGUF:Q5*
Wildcard behavior:
* matches any characters (case-insensitive)hf.co/... models-M model as sourceOptions:
-M <name> - Model name without tag (required)-Q <tags> - Comma-separated quantization tags to compare, supports wildcards (e.g., Q4*,IQ*) (required)-B <tag> - Base quantization tag for comparison (default: fp16, or uses existing base from results file)-D <url> - Remote server URL (default: local)-O <file> - Output results file (default: modelname.qc.json)-Te <value> - Temperature (default: 0.0 for deterministic results)-S <value> - Random seed (default: 365)-To <value> - top_p parameter (default: 0.001)-Top <value> - top_k parameter (default: -1)-R <value> - repeat_penalty parameter-Fr <value> - frequency_penalty parameter-T <file> - External test suite JSON file (default: internal v1base)--force - Force re-run testing for quantizations already present in results file--rejudge - Re-run judgment process for existing test results (without re-testing)--judge <model> - Use a judge model for similarity scoring (see Judge Scoring below)--judge-ctxsize <value> - Context length for judge model (0 = auto, default: 0). Auto mode calculates: test_ctx × 2 + 2048--mode <mode> - Judge execution mode: serial (default) or parallel--timeout <seconds> - API timeout in seconds for testing and judgment calls (default: 600)--verbose - Show judgment details (question ID, score, reason) for each judged question--ondemand - Pull models on-demand if not available, then remove after testing (see On-Demand Mode below)--repo <url> - Repository URL for the model source (saved in results file for qcview)--overwrite - Overwrite existing output file without prompting--enablethinking - Enable thinking mode for thinking models (disabled by default)--thinklevel <level> - Set thinking level (low, medium, high) - overrides --enablethinking--no-unloadall - Skip unloading all models before testing--logfile <path> - Log process output to file (appends if exists, timestamps each line, strips color codes)--fix - Attempt to fix a corrupted/malformed results file and recover data (outputs to .fixed.json)Judge Scoring:
Use a second LLM to evaluate similarity between base and quantized responses:
# Local judge model
osync qc -M llama3.2 -Q 1b-instruct-q8_0,1b-instruct-q6_k --judge mistral
# Remote test models, local judge
osync qc -d http://192.168.1.100:11434/ -M llama3.2 -Q q4_k_m --judge mistral
# Remote judge model (different server)
osync qc -M llama3.2 -Q q4_k_m --judge http://192.168.1.200:11434/mistral
# Parallel mode (judge runs concurrently with testing)
osync qc -M llama3.2 -Q q4_k_m --judge mistral --mode parallel
# Re-run judgment only (use existing test results)
osync qc -M llama3.2 -Q q4_k_m --judge mistral --rejudge
When judgment scoring is enabled:
--force, --rejudge is used, or a different judge model is specifiedSeparate Best Answer Judge Model:
Use --judgebest to specify a different model for best answer determination (A/B/tie judgment):
# Same model for similarity and best answer (when using --judge alone)
osync qc -M llama3.2 -Q q4_k_m --judge mistral
# Different model for best answer judgment
osync qc -M llama3.2 -Q q4_k_m --judge mistral --judgebest llama3.3
# Best answer judgment only (no similarity scoring)
osync qc -M llama3.2 -Q q4_k_m --judgebest llama3.3
# Remote best answer judge
osync qc -M llama3.2 -Q q4_k_m --judge mistral --judgebest http://192.168.1.200:11434/llama3.3
# Re-run only best answer judgment with new model
osync qc -M llama3.2 -Q q4_k_m --judgebest llama3.3 --rejudge
When --judgebest is specified:
--judge is used), then best answer judgmentJudgeModelBestAnswer, ReasonBestAnswer, and JudgedBestAnswerAt timestampsCloud Provider Support:
Use cloud AI providers (Anthropic Claude, OpenAI, etc.) as judge models with the @provider/model syntax:
# Using environment variable for API key
osync qc -M llama3.2 -Q q4_k_m --judge @claude/claude-sonnet-4-20250514
osync qc -M llama3.2 -Q q4_k_m --judge @openai/gpt-4o
# Using explicit API key
osync qc -M llama3.2 -Q q4_k_m --judge @claude:sk-ant-xxx/claude-sonnet-4
# Azure OpenAI (key@endpoint format)
osync qc -M llama3.2 -Q q4_k_m --judge @azure:mykey@myendpoint.openai.azure.com/gpt4-deployment
# Mixed: cloud judge with ollama best answer judge
osync qc -M llama3.2 -Q q4_k_m --judge @claude/claude-sonnet-4 --judgebest llama3.2:latest
# Both judges from cloud providers
osync qc -M llama3.2 -Q q4_k_m --judge @openai/gpt-4o --judgebest @claude/claude-opus-4
Supported cloud providers and their environment variables:
| Provider | Syntax | Environment Variable |
|---|---|---|
| Anthropic Claude | @claude/model | ANTHROPIC_API_KEY |
| OpenAI | @openai/model | OPENAI_API_KEY |
| Google Gemini | @gemini/model | GEMINI_API_KEY or GOOGLE_API_KEY |
| Azure OpenAI | @azure/deployment | AZURE_OPENAI_API_KEY + AZURE_OPENAI_ENDPOINT |
| Mistral AI | @mistral/model | MISTRAL_API_KEY |
| Cohere | @cohere/model | CO_API_KEY or COHERE_API_KEY |
| Together AI | @together/model | TOGETHER_API_KEY |
| HuggingFace | @huggingface/model | HF_TOKEN or HUGGINGFACE_TOKEN |
| Replicate | @replicate/model | REPLICATE_API_TOKEN |
Use osync qc --help-cloud for detailed provider documentation and examples.
When using cloud providers:
Judge Execution Modes:
Serial mode (default): After testing each quantization, all questions are judged sequentially before moving to the next quantization. Simple and predictable execution.
Parallel mode: Testing and judgment run concurrently at the question level. As each question is tested, it is immediately passed to the judge model in a background task. This allows the test model to continue generating answers while the judge evaluates previous questions. Significantly reduces total execution time when using a remote judge or when the judge model is faster than the test model.
On-Demand Mode:
Use --ondemand to automatically pull models that aren't available and remove them after testing. This is ideal for testing large models or many quantizations without consuming permanent storage:
# Test many quantizations without keeping them all
osync qc -M llama3.2 -Q q2_k,q3_k_s,q3_k_m,q4_k_s,q4_k_m,q5_k_s,q5_k_m,q6_k,q8_0 --ondemand
# Test on remote server with on-demand
osync qc -M llama3.2 -Q q4_k_m,q8_0 -D http://192.168.1.100:11434 --ondemand
When on-demand mode is enabled:
Built-in Test Suites:
v1base (default) - 50 questions across Reasoning, Math, Finance, Technology, Sciencev1quick - 10 questions (subset of v1base for quick testing)v1code - 50 coding questions across Python, C++, C#, TypeScript, Rust (8192 max tokens)External Test Suite Format:
Create custom test suites using JSON files with the following structure:
{
"name": "my-custom-suite",
"numPredict": 4096,
"contextLength": 4096,
"categories": [
{
"id": 1,
"name": "Category Name",
"contextLength": 8192,
"questions": [
{
"categoryId": 1,
"questionId": 1,
"text": "Your question text here",
"contextLength": 16384
}
]
}
]
}
numPredict - Maximum tokens generated per response (default: 4096). Use higher values for coding or detailed answers.contextLength - Context length (num_ctx) for the model during testing. Can be specified at:
Use with: osync qc -M model -Q q4_k_m -T my-custom-suite.json
Reference files v1base.json and v1code.json are included in the osync directory.
The recommended model to use as a judge is gemma3:12b, it is advisable to use a large context size (6k for v1base and 12k for v1code).
It's also recommended to run the qc command with the --verbose argument to verify the scoring and reason given corresponds to the expected beahviour:
Judging UD-IQ3_XXS Q1-1 Score: 92% (1/50 2%)
A and B match: Both responses provide a complete, thread-safe LRU cache implementation in Python using `OrderedDict` and `threading.RLock`. They both include comprehensive docstrings, error
handling, and a full set of methods (get, put, delete, clear, size, etc.). The core logic for managing the cache, including eviction and thread safety, is nearly identical. The primary differences
are in the formatting of the docstrings and some minor wording variations in the explanations. Both responses also include example usage and testing code. The code itself is very similar, with
only minor differences in variable names and comments. Overall, the responses demonstrate a very high degree of similarity in terms of content, approach, and functionality.
Judging UD-IQ3_XXS Q1-2 Score: 75% (2/50 4%)
A and B match: Both responses provide a Python async web scraper using aiohttp, incorporating concurrent crawling, rate limiting, retries, and CSS selector-based data extraction. They both utilize
asyncio, aiohttp, logging, and dataclasses. Both include comprehensive error handling and logging. However, they differ in their implementation details. Response A uses a semaphore for concurrency
control and a class-based structure for the scraper, while Response B introduces a separate RateLimiter class and a more streamlined approach to data extraction. Response A's retry logic is more
detailed, including random jitter, while Response B's is simpler. Response B also includes an advanced scraper with custom selectors.
Judging UD-IQ3_XXS Q1-3 Score: 75% (3/50 6%)
A and B differ: Both responses implement a retry decorator factory with similar functionality (configurable max attempts, delay strategies, exception filtering, and support for both sync and async
functions). However, they differ significantly in their implementation details and structure. Response A uses a class-based approach with a `RetryError` exception and separate `_calculate_delay`
function. It also includes convenience decorators for common retry patterns. Response B uses a dataclass for configuration and an Enum for delay strategies, making it more structured. It also has
separate functions for sync and async decorators. While both achieve the same goal, the code organization and specific techniques used are quite different, leading to a noticeable difference in
qcview)Display quantization comparison results in formatted tables or export to various formats.
# View results in table format (console)
osync qcview llama3.2.qc.json
# Export table to text file
osync qcview llama3.2.qc.json -O report.txt
# Export as JSON to file
osync qcview llama3.2.qc.json -Fo json -O report.json
# Export as Markdown
osync qcview llama3.2.qc.json -Fo md
# Export as HTML (interactive with theme toggle)
osync qcview llama3.2.qc.json -Fo html
# Export as PDF (includes full Q&A pages)
osync qcview llama3.2.qc.json -Fo pdf
# Add or override repository URL in output
osync qcview llama3.2.qc.json -Fo html --repo https://github.com/user/model
Table Output:
Displays color-coded results with:
Q4_K (87%))Q6_K (81% Q8_0) when dominant differs from tag? suffix (e.g., Q3_K?)Color Coding:
Output Formats:
Options:
<file> - Results file to view (required, positional argument)-Fo <format> - Output format: table, json, md, html, pdf (default: table)-O <file> - Output filename (default: auto-generated based on format)--repo <url> - Repository URL (displayed in output, overrides value from results file)--metricsonly - Ignore judgment data and show only metrics-based scores (useful for comparing pure model output quality without judge influence)--overwrite - Overwrite existing output file without promptingbench)Run context tracking benchmarks to evaluate how well models maintain information across long conversations.
# Basic benchmark with a model
osync bench -M llama3.2
# Test on remote server (multiple ways to specify)
osync bench -M llama3.2 -D http://192.168.1.100:11434
osync bench -M llama3.2 -D 192.168.1.100 # IP (auto: http:// + :11434)
osync bench -M llama3.2 -D myserver/ # trailing slash
# With judge evaluation
osync bench -M llama3.2 --judge gemma3:12b
# Custom output file
osync bench -M llama3.2 -O my-benchmark.json
# Enable thinking for thinking models (qwen3, deepseek-r1) - disabled by default
osync bench -M qwen3 --enablethinking
# Control thinking level (overrides --enablethinking)
osync bench -M qwen3 --thinklevel medium
# Skip unloading all models before testing
osync bench -M llama3.2 --no-unloadall
# Generate a custom test suite
osync bench --generate-suite -T custom-suite.json -O custom-suite.json
How It Works:
The bench command generates dynamic stories with embedded facts and tests the model's ability to:
Question Types:
Scoring:
Each question is evaluated as:
Options:
-M <name> - Model name without tag (required unless --help-cloud, --showtools, or --generate-suite)-Q <tags> - Model/Quantization tags to compare (comma-separated, supports wildcards e.g., Q4*,IQ*)-D <url> - Remote server URL (default: local)-O <file> - Results output file (default: modelname.testtype.json)-T <file> - Test suite file or type. For bench: filename first, falls back to v1.json if type given. For --generate-suite: test type (ctxbench/ctxtoolsbench)-L <category> - Limit testing to a specific category (e.g., 2k, 4k, 8k, 16k, 32k, 64k, 128k, 256k)-Te <value> - Model temperature (default: from model template)-S <value> - Model seed (default: 365)-To <value> - Model top_p (default: from model template)-Top <value> - Model top_k (default: from model template)-R <value> - Model repeat_penalty (default: from model template)-Fr <value> - Model frequency_penalty (default: from model template)--force - Force re-run testing for quantizations/models already present in results file--rejudge - Re-run judgment process for existing test results (without re-testing)--judge <model> - Judge model for answer evaluation. Ollama: model_name or http://host:port/model. Cloud: @provider[:key]/model--mode <mode> - Judge execution mode: serial (default) or parallel--timeout <seconds> - API timeout in seconds (default: 1800)--verbose - Show testing conversation, tools usage, judgment details--judge-ctxsize <value> - Context length for judge model (0 = auto based on test type)--ondemand - Pull models on-demand if not available, then remove after testing--repo <url> - Repository URL for the model source (saved per model/quant in results file)--showtools - Display available tools with descriptions and queryable data--generate-suite - Generate test suites. Use with -T (ctxbench/ctxtoolsbench) for specific type, -O for custom filename--help-cloud - Show detailed help for cloud provider integration--enablethinking - Enable thinking mode for thinking models (disabled by default)--thinklevel <level> - Set thinking level (low, medium, high) - overrides --enablethinking--no-unloadall - Skip unloading all models before testing--overwrite - Overwrite existing output file without prompting--ctxpct <value> - Content scaling percentage for test suite generation (default: 100, range: 50-150)--logfile <path> - Log process output to file (appends if exists, timestamps each line, strips color codes)--calibrate - Enable token calibration mode: detailed tracking of estimated vs actual tokens at each step--calibrate-output <file> - Output file for calibration data (default: calibration_.json in test suite directory)--fix - Fix corrupted/malformed results JSON file (specify file with -O or -M)Bench Test Suite Format:
Custom bench test suites can specify a numPredict field to control maximum tokens per response:
{
"testType": "ctxbench",
"testDescription": "Custom context benchmark",
"numPredict": 16384,
"maxContextLength": 131072,
"categories": [...]
}
numPredict - Maximum tokens generated per test response (default: 16384)maxContextLength - Maximum context length supported by the test suitebenchview)Display context benchmark results in formatted tables or export to various formats.
# View results in table format (console)
osync benchview llama3.2.ctxbench.json
# Export to different formats
osync benchview results.json -Fo json -O report.json
osync benchview results.json -Fo md -O report.md
osync benchview results.json -Fo html -O report.html
osync benchview results.json -Fo pdf -O report.pdf
Output Information:
Output Formats:
Options:
<file> - Results file to view (required, positional argument)-Fo <format> - Output format: table, json, md, html, pdf (default: table)-O <file> - Output filename (default: auto-generated based on format)-C <category> - Filter to specific category--details - Show detailed Q&A results in output--overwrite - Overwrite existing output file without promptingmanage)Interactive TUI for managing models with keyboard shortcuts.
# Launch manage interface for local server
osync manage
# Launch manage interface for remote server (multiple ways to specify)
osync manage http://192.168.0.100:11434
osync manage 192.168.0.100 # IP (auto: http:// + :11434)
osync manage myserver/ # trailing slash
Features:
● marks models loaded in memoryxollama tweakKeys:
: - _ . / - Filter by name (* matches any text); Backspace removes a characterosync setup manage servers (also asked by osync setup server)Tweak (xOllama): xOllama keeps settings of its own in the model: engine, KV cache types, dynamic slots, DCA, session affinity and prefix pooling, council, GPU/devices, speculative decoding. They are listed in the model details (Enter), and on an xOllama server Ctrl+W changes them for the model under the cursor, or for each selected model. Pick what to change (every setting, or one group such as KV cache or council), or remove the settings. You can also type flags: a flag with a value is set without questions, e.g. --kv-k=q8_0 --kv-v=q8_0 or --council=on.
osync runs xollama tweak model <model> on the console against the server manage shows, including a remote one. The questions, the explanations and the consistency checks are xOllama's own, and so is the list of settings, so it always matches the xOllama version. This needs the xollama CLI on this machine: on PATH, or set OSYNC_XOLLAMA_CLI to its path. Only the xOllama settings layer changes: weights, template, system prompt and parameters stay as they are. See xOllama's tweak documentation.
Sort Modes:
Themes (osync setup manage themes shows them with a preview):
setup)Preferences in the settings file, by section. Without the last arguments a section asks interactively (numbered choices, Enter keeps the current value).
osync setup # summary, then a menu
osync setup show # summary only
osync setup server # ask: Ollama, xOllama or both, host, ports
osync setup server xollama 192.168.1.5 # one server (port: the flavor's default)
osync setup server ollama nas:11500
osync setup server both [host] [xollama] # both side by side (default ports; optional default server)
osync setup server auto # back to auto-detection
osync setup server env ignore # the settings win over XOLLAMA_HOST / OLLAMA_HOST (env use: they win)
osync setup alias # list (and add/remove interactively)
osync setup alias add gpu 192.168.1.10 # port 11434, or 22434 when only xOllama answers
osync setup alias remove gpu
osync setup manage # theme and default sort order
osync setup manage theme "Tokyo Night" # name, loose spelling (tokyo-night) or number
osync setup manage sort size- # name+ name- size- size+ created- created+
osync setup manage servers gpu,nas # aliases for Ctrl+Left/Right in manage (all, none)
osync setup manage themes # all themes with a preview
osync setup shell # theme, color mode, tab completion
osync setup shell theme gruvbox-light # or: plain, default
osync setup shell colors truecolor # auto, truecolor, 256, 16, none
osync setup shell completion # install tab completion (bash / PowerShell)
osync setup shell themes
showversion, -v)Display osync version and environment information.
# Show version only
osync -v
osync showversion
# Show detailed environment information
osync showversion --verbose
Basic Output:
osync v1.2.6 (b20260110-1117)
Verbose Output:
osync v1.2.6 (b20260110-1117)
Binary path: C:\Users\user\.osync\osync.exe
Installed: Yes (C:\Users\user\.osync)
Shell: PowerShell Core 7.5.4
Tab completion: Installed (C:\Users\user\Documents\PowerShell\Microsoft.PowerShell_profile.ps1)
Verbose Information:
install)Install osync to user directory and configure shell completion.
# Run the installer
osync install
Installation Directory:
~/.osync~/.local/binFeatures:
-h, -? - Show help for any command-bt <value> - Bandwidth throttling (B, KB, MB, GB per second)-BufferSize <value> - Memory buffer size for remote-to-remote copy (KB, MB, GB; default: 512MB)Examples:
# Bandwidth throttling
osync cp llama3 http://server:11434 -bt 75MB # Limit to 75 MB/s
osync cp qwen2 http://server:11434 -bt 1GB # Limit to 1 GB/s
# Memory buffer configuration for remote-to-remote
osync cp http://server1:11434/llama3 http://server2:11434/llama3 -BufferSize 256MB
osync cp http://server1:11434/qwen2 http://server2:11434/qwen2 -BufferSize 1GB
--size - Sort by size (largest first)--sizeasc - Sort by size (smallest first)--time - Sort by modified time (newest first)--timeasc - Sort by modified time (oldest first)Run osync without arguments to enter interactive mode with:
osync
# Press tab to see available models
# Type command and press enter
> cp llama3 my-backup
> ls "qwen*"
> quit # or type 'exit', or press Ctrl+C
Use * wildcard for flexible pattern matching:
# Match any model starting with "llama"
osync ls "llama*"
# Match any 7b model
osync ls "*:7b"
# Match models in namespace
osync ls "mannix/*"
# Match HuggingFace models
osync ls "hf.co/*"
# Match any model containing "test"
osync ls "*test*"
# Update all local models to latest versions
osync update
# Update specific models matching pattern
osync update "llama*"
# Update all models on remote server
osync update http://192.168.0.10:11434
# Create backup before updating
osync cp llama3:latest llama3:backup
osync update llama3
# Upload from local to multiple servers (short form with IP)
osync cp llama3 192.168.0.10
osync cp llama3 192.168.0.11
osync cp llama3 192.168.0.12
# Or using hostnames with trailing slash
osync cp llama3 server1/
osync cp llama3 server2/
osync cp llama3 server3/
# Copy between remote servers
osync cp http://192.168.0.10:11434/llama3 http://192.168.0.11:11434/llama3
osync cp http://192.168.0.10:11434/qwen2 http://192.168.0.12:11434/qwen2
# List and remove test models
osync ls "test-*"
osync rm "test-*"
# Remove old backups
osync rm "*:backup"
# List all models by size to find space hogs
osync ls --size
# Rename for better organization
osync mv llama3 llama3-8b:prod
osync mv qwen2 qwen2-7b:dev
None
v1.4.2
New
manage tweaks xOllama model settings - on an xOllama server, Ctrl+W runs xollama tweak model for the model under the cursor, or for each selected model, against the server manage shows. You can walk every setting, one group (KV cache, dynamic slots, DCA, session pooling, council, GPU/devices, engine), or remove the settings, and add flags such as --kv-k=q8_0 that are set without questions. The model details list the xOllama settings. Needs the xollama CLI on PATH, or OSYNC_XOLLAMA_CLIv1.4.1
Fixes
host:port is not a valid folder name there). The model used to be recreated from /api/show, which dropped the renderer, parser, requires and xOllama's model settings (a council became a plain model), merged several licenses and changed the parameters, yet osync reported success. It is now recreated from the source's manifest with the config and settings layers verbatim, so every layer and the config have the source's digestsollama show --modelfile and lost the same parts)HF_TOKEN is sent on osync's own HuggingFace requests (tag lookups for hf.co/... wildcard tags, qc model discovery), for higher rate limits and private or gated repositories; only to huggingface.co / hf.co over HTTPSthrough the relay, recreated from its manifest)v1.4.0
New
XOLLAMA_HOST, OLLAMA_HOST, then localhost:11434 / localhost:22434osync ps shows Server: Ollama|xOllama at <url>)xollama CLI when the local server is xOllama (override with OSYNC_OLLAMA_CLI)XOLLAMA_MODELS honored for the models directory; xOllama processes shown in process statsollama, xollama)osync setup - one command for the preferences: server (Ollama, xOllama or both), alias (server aliases), manage (theme, default sort order, servers), shell (colors, theme, tab completion), show; interactive or with arguments for scriptssettings.json in the per-OS configuration folder: local server (Ollama/xOllama, host, port), aliases, color mode, manage and shell settingsosync setup alias add gpu 192.168.1.10, then -d gpu, osync cp model gpu/, osync cp gpu/model copy, osync manage gpuXOLLAMA_HOST / OLLAMA_HOST - osync setup server asks when one is set, osync setup server env ignore|use, and a check box in the manage settings (Ctrl+E)manage rewritten on Terminal.Gui 2 - true color with multi-color themes (one color per column, ● for models loaded in memory, colored top and bottom bars) adapted to 256 and 16-color terminals with contrast checks; theme picker with live preview (Ctrl+T), the theme is saved in the preferences file; settings dialog (Ctrl+E) for the local server (Ollama/xOllama, host, port, connection test) and the color mode; column headers; F1 help; rename on F2 (Ctrl+M is Enter in most terminals); load runs in the background; console operations (copy, run, update, pull) return to the list without restarting osync; pull validation no longer rejects hf.co/... models; confirmations default to the safe answer (Enter cancels a delete)manage switches servers with Ctrl+Left / Ctrl+Right - the local server, the other one of Ollama/xOllama side by side, and the aliases chosen with osync setup manage servers (osync setup server offers them); the top bar shows which one ([2/3] gpu: Ollama @ ...)ls, ps, show, -v, copy/pull/update progress, errors, warnings and results use the shell theme (removed in 1.0.1 because of garbled output on Linux); plain when redirected, with NO_COLOR or the plain thememanage and the shell, 7 of them for light terminals, all checked for readable contrast at true color, 256 and 16 colorsTERM with xterm-16color, so every terminal was treated as 16-color); colorMode setting, OSYNC_COLOR_MODE and NO_COLOR override it; osync -v --verbose shows the resultosync install asks for the local server only when needed - a single server found on this machine is used without questions; when none or both are found it asks for the type (Ollama, xOllama or both), host and port, with detected defaults and a connection test; installing from a renamed binary (e.g. osync-macos-arm64 install) worksosync -v shows the real build time (UTC) for every binary, including renamed ones (osync-macos-arm64) and downloaded copies, instead of the file's modification timeFixes
osync <command> -h and missing-argument errors hanging forever on Linux/macOS when output is redirected (pipes, scripts, CI): the help renderer looped endlessly at 100% CPU with growing memoryrun, ps, qc ignoring OLLAMA_HOST; manage, psmonitor and local judge models now use the same local-server resolutionps truncating model names to 20 characters when output is redirectedshow printing only the Modelfile: it now shows the same sections as ollama show (model details, capabilities, projector, parameters, system, license; all metadata with -v), several section flags can be combined, and a missing model exits with an errorrm and update exiting with code 0 when no model matches, or when deleting/updating a model failed (update of all models on an empty server is still a success)osync cp http://server/qwen3:4b ...): the : of the tag was taken for a port. A server given without a port now uses 11434, or 22434 when the host refuses 11434 but accepts 22434 (xOllama)manage showing unknown quantization (and no parameters/family) when the local models directory and the resolved local server did not match (e.g. Ollama and xOllama both installed): local models are now listed through the local server's API, so the list and its details always come from the same server, and startup no longer makes one /api/show call per model. The top bar shows which server is used (Ollama @ localhost:11434); the models directory is only read when the server is unreachable, which the top bar saysollama service user): the upload now goes through the local server with the push relaymanage not scrolling when the terminal is shorter than the list-bt) not limiting short bursts, counting requested instead of read bytes, and misbehaving after ~25 days of uptimeESC[0m) in redirected output on Linux/macOSPlatform
net10.0; builds and tests on Windows, Linux and macOSosync-macos-arm64 and osync-macos-x64dev builds publish pre-releases, master publishes releasesv1.2.9
-Hi shortcut for --history argument (e.g., osync monitor -Hi 10)--history are now treated as minutes (e.g., -Hi 10 = 10 minutes)osync v1.2.9 (b20260116-1814))--enablethinking and --thinklevel arguments for thinking models (qwen3, deepseek-r1)--no-unloadall argument to skip unloading all models before testing--overwrite argument to skip file overwrite prompts--generate-suite to create custom test suite JSON files with -T and -O options--mode=parallel for parallel judgment - judges answers in background while testing continues--overwrite argument to skip file overwrite promptsnumPredict)numPredict field in bench test suite JSON
ps) - Extended system monitoring
osync -h output to show only global options and available commandsosync <command> -h for detailed help on specific commands\x1b[0m) on exit prevents color leakage to shell promptosync cp model 192.168.0.100)qc and bench commands silently exiting when remote test server is unreachable - now shows clear "Could not connect to server" error message←[0m) appearing after command output on Windows--enablethinking and --thinklevel arguments for thinking models (qwen3, deepseek-r1)--no-unloadall argument to skip unloading all models before testing--overwrite argument to overwrite existing output file without prompting--ondemand mode/api/chat to lightweight /api/generate call--fix argument to recover corrupted/malformed JSON results files (outputs to .fixed.json)
--overwrite argument to overwrite existing output file without promptingfile1-file2.html)--details)--logfile argument for qc and bench commands
[2026-01-17 06:38:20.176])http://192.168.1.100 becomes http://192.168.1.100:11434)http:// protocol when not specified (e.g., 192.168.1.100:11434 becomes http://192.168.1.100:11434)myserver/ → http://myserver:11434) since model names cannot end with /http:// or https://) → remote server192.168.0.100) → remote serverhost:11434) → remote server/ (host/) → remote serverlocalhost → remote server/ (host) → treated as model namev1.2.8
--judge and --judgebest
@provider[:token]/model (e.g., @claude/claude-sonnet-4-20250514, @openai/gpt-4o)--help-cloud option for detailed provider documentationv1.2.7
--judgebest can be used alone or combined with --judge for different models--judge: local model name or http://host:port/model for remote--rejudge to re-run only best answer judgment with new modelOsyncVersion - Version of osync used for testingOllamaVersion - Ollama server version for test quantizationsOllamaJudgeVersion - Ollama version for judge server (similarity scoring)OllamaJudgeBestAnswerVersion - Ollama version for best answer judge server/api/version endpointstop parameter serialization (now correctly sent as array instead of string)ConvertParameterValue helper ensures correct JSON types for all Ollama model parametershf.co/... models
osync load http://host:port/modelname in addition to osync load modelname -d host--rejudge with existing results, only the judge model is neededv1.2.6
-Fo md, -Fo html, or -Fo pdf to select format--repo argument to specify model source repository
qc testing and overridden in qcviewosync version (alias -v) to display version info
--verbose flag displays detailed info: binary path, installation status, shell type/version, tab completion statusinstalled v1.2.6 (b20260110-1156) is older)Digest) and short digest (ShortDigest, first 12 chars) stored in results JSONollama ls)
osync ls and manage TUI compute SHA256 of manifest file contentollama ls output for easy cross-reference...0B-A3B-Instruct-GGUF:Q4_K_S)load_duration from response✓ Model 'model:tag' loaded successfully (2m 15s) (API: 2m 5s)"quantization_level": "unknown" for HuggingFace modelsQ4_0 (87%) or Q6_K (81% Q8_0) showing actual tensor distributionverbose=true to fetch tensor metadataQ3_K?) to indicate uncertainty-b, no longer tries to pull the base model--force to re-run the base model if neededosync ls code matches models starting with "code" (prefix match, same as code*)osync ls *q4_k_m finds all models ending with "q4_k_m" (useful for finding by quantization)osync ls *code* finds models containing "code" anywhere in the nameosync ls 'gemma*'osync pull gemma3:1b-it-q* pulls all matching tagsosync pull hf.co/unsloth/gemma-3-1b-it-GGUF:IQ2*osync pull -d http://server:11434 gemma3:1b-it-q*bestanswer: A (base better), B (quant better), or AB (tie)Score: 75% (27/50 54%) Best: AB--judge is active and bestanswer is missing67% (B:10 A:5 =:3) showing quant won 67% of non-tie comparisonsBestAnswer field (A/B/AB)BestCount, WorstCount, TieCount, BestPercentage, WorstPercentage, TiePercentageCategoryBestStats with counts and percentages--metricsonly argument to ignore judgment data
--judge-ctxsize is 0 (new default), calculates: test_ctx × 2 + 2048v1.2.5
hf.co/namespace/repo:tag) are now preserved correctly instead of being incorrectly derived from the -M argumentv1.2.4
--rejudge argument to re-run judgment process for existing test results--force which re-runs both testing and judgment, --rejudge only re-runs judgment--ondemand) now properly verifies HuggingFace models (hf.co/...)-b qwen3-coder:30b-a3b-fp16 or -b hf.co/namespace/repo:tag) is now used as-is-M model name*) in -Q argument (e.g., Q4*, IQ*, *)hf.co/... modelsModelTagResolver class for reusable tag resolution across commands--timeout) for model loadingQ4_0 is stored as q4_0 causing preload to failv1.2.3
--ondemand argument to enable on-demand model managementcontextLength property in built-in test suites (v1base, v1quick, v1code)contextLength at suite, category, and question levels--judge-ctxsize argument to configure judge model context lengthosync ls qwen2:)v1.2.2
--verbose flag to show judgment details
v1.2.1
user/model:tag)IsBase flag are automatically repaired on load/ or \ are now converted to - in default output filename
v1.2.0
v1code test suite for evaluating code generation quality
-T v1code or via external v1code.json filenumPredict values
numPredict property (default: 4096)registry.ollama.ai/v2/ manifest endpoint instead of HTML scraping--timeout argument for testing and judgment API calls
--timeout <seconds> argumentv1.1.9
--judge <model> - Specify local or remote judge model--mode serial|parallel - Serial (after each quant) or parallel (concurrent) execution--mode parallel, each question is immediately judged in the background as soon as testing completes/api/chat endpoint with structured JSON output schemaReason field to capture judge's reasoning for each scoreosync qcview results.json)v1.1.8
-T path/to/suite.json
v1base.json file for reference and portability--force flag to re-run testing for quantizations already in results filefp16 and is automatically detected from existing results/api/tagsosync ls to accept model name as positional argument (e.g., osync ls qw works like osync ls qw*)v1.1.7
qc and qcview commands for comprehensive quality testing
qcview (90-100% green, 80-90% lime, 70-80% yellow, 50-70% orange, <50% red)/api/show endpoint doesn't include size field, now fetches from /api/tagsv1.1.6
ps command improvements
load command to preload models into memory with configurable keep-aliveunload command to free VRAM immediately/bye or Ctrl+D to exitv1.1.0
-BufferSize parameter (e.g., 256MB, 1GB)v1.0.9
v1.0.8
ren and mv for safe model renaming--size, --sizeasc, --time, --timeascv1.0.7
* wildcard:latest tag fallback for copy and remove commands when tag not specified/ in names)v1.0.6
copy (alias cp)latest tag if none specifiedv1.0.5
v1.0.4
v1.0.3
-bt 75MBv1.0.2
v1.0.1
v1.0.0
See the open issues for a list of proposed features (and known issues).
This project is licensed under the MIT license.
See LICENSE for more information.
C#
98.1%
Gherkin
1.6%