PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.
See the codePowdered Metal — An ML SDK, framework, and application suite for Apple Silicon, written in Rust.
PMetal is a complete machine learning platform for Apple Silicon — from low-level Metal GPU kernels and Apple Neural Engine integration to high-level training APIs, a terminal TUI, and a full desktop GUI. Ship fine-tuned models without leaving the Apple ecosystem.
A full Tauri + Svelte desktop application for visual model management, training, and inference.
cd crates/pmetal-gui
bun install && bun tauri dev
19 pages: Dashboard, Training, GRPO, Distillation, Pretrain, Inference, DFlash, Models, Datasets, Merging, Quantize, Embed Train, RLKD, Ollama, Serve, Bench, Eval, Jobs, and Settings. Download models from HuggingFace, configure LoRA training with live loss metrics, chat with models, merge weights, and quantize — all from the GUI. Training, inference, distillation and GRPO run in-process with real-time progress updates; the remaining pages drive the pmetal CLI as a subprocess, which the app bundles.
A full-featured terminal control center with 20 tabs.
pmetal tui
| Tab | Description |
|---|---|
| Device | GPU/ANE info, Metal feature detection, memory gauge, kernel tuning, UltraFusion topology |
| Models | Browse cached models, HuggingFace Hub search (S), memory fit estimation, download |
| Datasets | Scan and preview local datasets (JSONL, Parquet, CSV) with line counts |
| Tokenize | Tokenize a text corpus into binary shards for pretraining |
| Training | Configure and launch SFT/LoRA/QLoRA training runs with sectioned parameter forms |
| Embed Train | Train a sentence-embedding (encoder-only) model with contrastive losses |
| Pretrain | Full-parameter pretraining from scratch |
| Distillation | Configure knowledge distillation (online, offline, progressive) |
| RLKD | Reinforcement learning with knowledge distillation |
| GRPO | Configure GRPO/DAPO reasoning training with reward functions and sampling params |
| Dashboard | Live loss curves (braille), LR schedule, throughput sparklines, timing breakdown gauges |
| Inference | Interactive chat interface with markdown rendering and generation settings sidebar |
| DFlash | Block-diffusion speculative decoding |
| Serve | OpenAI-compatible server control |
| Quantize | GGUF and MLX quantization with bit/method selection |
| Merge | SLERP, TIES, DARE and linear model merging |
| Bench | Training and inference benchmarking |
| Eval | Perplexity evaluation against a dataset |
| Ollama | Modelfile generation and Ollama export |
| Jobs | Training run history with log viewer, status tracking, and metadata |
Keybindings: Ctrl+P to jump to any tab (type to filter), Tab/Shift+Tab to cycle, Alt+1-9 (or Ctrl+1-9) for the first nine, L to adjust learning rate mid-run, ? for contextual help, q to quit.
# LoRA fine-tuning with sequence packing (default)
pmetal train \
--model Qwen/Qwen3-0.6B \
--dataset train.jsonl \
--output ./output \
--lora-r 16 --batch-size 4 --learning-rate 2e-4
# Inference with LoRA adapter
pmetal infer \
--model Qwen/Qwen3-0.6B \
--lora ./output/lora_weights.safetensors \
--prompt "Explain quantum entanglement" \
--chat
# Train and use a Qwen3Next/Qwen3.6 MTP predictor
pmetal tokenize --input train.jsonl --output ./tok --tokenizer Qwen/Qwen3.6-30B-A3B-Instruct
pmetal train-mtp \
--model Qwen/Qwen3.6-30B-A3B-Instruct \
--family qwen3-next \
--shards ./tok/shard_00000.bin \
--output ./qwen-mtp
pmetal infer \
--model Qwen/Qwen3.6-30B-A3B-Instruct \
--mtp --mtp-model ./qwen-mtp \
--prompt "Explain monotonic queues."
# Knowledge distillation
pmetal distill \
--teacher Qwen/Qwen3-4B \
--student Qwen/Qwen3.5-0.8B-Base \
--dataset train.jsonl
# GRPO reasoning training
pmetal grpo \
--model Qwen/Qwen3-0.6B \
--dataset reasoning.jsonl \
--reasoning-rewards
# HuggingFace model search with memory fit
pmetal search "qwen 0.6b" --detailed
# Merge models with SLERP
pmetal merge \
--model-a model-a --model-b model-b \
--method slerp --t 0.5
# Quantize to GGUF
pmetal quantize \
--model ./output \
--output model.gguf --method q4_k_m
# Fuse LoRA into base model
pmetal fuse \
--model Qwen/Qwen3-0.6B \
--lora ./output/lora_weights.safetensors
# Evaluate perplexity
pmetal eval \
--model Qwen/Qwen3-0.6B \
--dataset eval.jsonl
# Start OpenAI-compatible server
# (in the prebuilt binary and the brew formula; add --features serve if you build it yourself)
pmetal serve --model Qwen/Qwen3-0.6B --port 8080
| Command | Description |
|---|---|
train | Fine-tune with LoRA/QLoRA (SFT) |
train-mtp | Train Gemma 4 assistant or Qwen3Next/Qwen3.6 MTP predictor checkpoints |
train-draft | Train DFlash block-diffusion draft checkpoints |
train-diffusion | LoRA/QLoRA fine-tune a DiffusionGemma block-diffusion model |
pretrain | Pretrain a model from scratch (full-parameter, no LoRA) |
infer | Interactive inference with chat, tool use, and thinking mode |
distill | Knowledge distillation (online, offline, progressive) |
grpo | GRPO/DAPO reasoning training (VLM, speculative, async rewards) |
rlkd | Reinforcement Learning with Knowledge Distillation |
embed-train | Sentence-transformer fine-tuning (InfoNCE, Triplet, CoSENT) |
search | Search HuggingFace Hub with memory fit estimation |
download | Download a model from HuggingFace Hub |
merge | Merge two models (12 strategies) |
quantize | GGUF quantization (24 methods) |
fuse | Fuse LoRA adapter weights into base model |
eval | Evaluate model perplexity on a dataset |
serve | OpenAI- and Anthropic-compatible inference server |
tui | Full TUI control center (20 tabs) |
dashboard | Real-time training metrics visualization |
dataset | Dataset utilities: analyze, download, convert |
ollama | Ollama integration: modelfile, create, templates |
info | Show device info (GPU, ANE, bandwidth, NAX) |
memory | Show memory usage and available capacity |
init | Generate a sample configuration file |
bench | Benchmark training performance |
bench-gen | Benchmark generation loop timing |
bench-ffi | Benchmark FFI overhead |
bench-workload | Benchmark real cached inference/training workloads |
bench-corpus | Structured kernel benchmarking with JSON reporting |
bench-gdn | Benchmark Qwen3.5 GDN backends on real layer shapes |
tokenize | Tokenize a text corpus into binary shards for pretraining |
pack-experts | Pack expert weights for SSD-offloaded MoE inference |
dflash | Block-diffusion speculative decoding |
mcp | Start MCP server over stdio (51 tools for Claude Desktop / Claude Code) |
cluster | Multi-Mac cluster: discover peers, train across machines, run all-reduce / pipeline benchmarks |
Connect two or more Apple Silicon Macs into a "home cluster" for distributed training and inference. PMetal auto-detects every NIC on the box (Thunderbolt-Bridge, Ethernet, Wi-Fi), advertises them via mDNS, and forms a ring biased toward the fastest fabric — Thunderbolt cables are picked over Ethernet, Ethernet over Wi-Fi, all without configuration.
# 1. Connect the Macs (Thunderbolt-4/5 cable recommended; Ethernet works too).
# 2. On every Mac:
pmetal cluster status # Show local NICs + any peers already announcing.
pmetal cluster up # Join the cluster, hold connection open.
# 3. On every Mac at the same time:
pmetal cluster bench --mb 64 --iters 10 # All-reduce throughput per fabric.
pmetal cluster pipeline-bench --tokens 16 --layers 32 # Pipeline activation transport bench.
# 4. Distributed training (each Mac runs the same command simultaneously):
pmetal train --model Qwen/Qwen3-0.6B \
--dataset train.jsonl \
--distributed-auto # mDNS-discovers peers, all-reduce gradients.
pmetal cluster status example output:
Local peer: 12D3KooW…XyZ (rank 0/2)
Thunderbolt ring: yes
Local interfaces:
bridge0 thunderbolt 169.254.42.1
en0 ethernet 192.168.1.10
lo0 loopback 127.0.0.1, ::1
Cluster peers:
peer-id local primary-addr fabric paths
12D3KooW…XyZ yes 169.254.42.1:52416 thunderbolt 2
12D3KooW…AbC no 169.254.42.2:52416 thunderbolt 2
What's wired today: gradient all-reduce (multi-machine training, real), fabric-aware ring formation with Thunderbolt > Ethernet > Wi-Fi priority, automatic fabric fallback when a cable is unplugged mid-job, gradient compression (TopK, FP16/BF16/INT8 quantization, error feedback), and a transport-tested pipeline harness with multi-process integration tests. Per-architecture partial-layer execution (the prerequisite for serving a model that doesn't fit on one Mac) is the next step — the harness is ready, the model API is the bottleneck.
PMetal is an embeddable SDK — integrate training, inference, and model operations into your own Rust applications. pmetal re-exports every sub-crate, so one dependency gets you the whole framework.
orchestrator::run_training is the one-call training entry point the CLI itself uses:
use pmetal::trainer::orchestrator::{TrainingJobConfig, run_training};
let config = TrainingJobConfig {
model_id: "Qwen/Qwen3-0.6B".to_string(),
dataset: "train.jsonl".to_string(),
output_dir: "./output".to_string(),
..Default::default()
};
let result = run_training(config, None, Vec::new()).await?;
println!("final loss {:.4} over {} steps", result.final_loss, result.total_steps);
Inference goes through the model dispatcher:
use pmetal::models::generate;
use pmetal::prelude::*;
let model_dir = pmetal::hub::resolve_model_path("Qwen/Qwen3-0.6B", None, None).await?;
let mut model = DynamicModel::load(&model_dir)?;
let tokenizer = Tokenizer::from_model_dir(&model_dir)?;
let input_ids = tokenizer.encode_with_special_tokens("What is 2+2?")?;
let output = generate(
|input| model.forward(input, None),
&input_ids,
GenerationConfig::sampling(256, 0.7),
)?;
println!("{}", tokenizer.decode(&output.token_ids[input_ids.len()..])?);
Preference optimization (DpoTrainer, SimpoTrainer, OrpoTrainer, KtoTrainer) and TAID
distillation are library-only for now — there is no CLI subcommand for them yet.
For step-by-step control, use the crates directly: pmetal_trainer::TrainingLoop,
pmetal_models::DynamicModel, pmetal_lora::DynamicLoraModel, pmetal_distill::Distiller. Every
code block in each crate's README is compiled as a doctest, so they stay honest. See
examples/ for complete working programs, including manual training-loop
orchestration and ANE-specific workflows.
PMetal exposes a Python extension module via PyO3. Install with maturin develop from crates/pmetal-py.
import pmetal
# Fine-tune with sensible defaults
result = pmetal.finetune(
"Qwen/Qwen3-0.6B",
"train.jsonl",
lora_r=16,
learning_rate=2e-4,
epochs=3,
)
print(f"Loss: {result['final_loss']}, Steps: {result['total_steps']}")
# Inference
text = pmetal.infer("Qwen/Qwen3-0.6B", "What is 2+2?")
print(text)
# Inference with LoRA adapter
text = pmetal.infer(
"Qwen/Qwen3-0.6B",
"Explain quantum entanglement",
lora="./output/lora_weights.safetensors",
)
import pmetal
# Configure training components
lora_config = pmetal.LoraConfig(r=16, alpha=32.0)
training_config = pmetal.TrainingConfig(
learning_rate=2e-4,
num_epochs=3,
batch_size=4,
max_seq_len=2048,
)
# Create trainer
trainer = pmetal.Trainer(
model_id="Qwen/Qwen3-0.6B",
lora_config=lora_config,
training_config=training_config,
dataset_path="train.jsonl",
)
trainer.add_callback(pmetal.ProgressCallback())
result = trainer.train()
# Load model for inference
model = pmetal.Model.load("Qwen/Qwen3-0.6B")
print(model.generate("Hello world", temperature=0.7))
Prebuilt signed binaries are available on the Releases page.
Crates are available on crates.io.
Build from source:
git clone https://github.com/epistates/pmetal.git && cd pmetal
cargo build --release # CLI + TUI
cd crates/pmetal-gui && bun install && bun tauri build # GUI (optional)
PMetal automatically detects Apple Silicon capabilities at startup and tunes kernel parameters accordingly.
| Chip Family | GPU Family | NAX | ANE | UltraFusion | Status |
|---|---|---|---|---|---|
| M1 / Pro / Max / Ultra | Apple7 | - | 16 cores | Ultra: 2-die | Fully supported |
| M2 / Pro / Max / Ultra | Apple8 | - | 16 cores | Ultra: 2-die | Fully supported |
| M3 / Pro / Max / Ultra | Apple9 | - | 16 cores | Ultra: 2-die | Fully supported |
| M4 / Pro / Max / Ultra | Apple9 | - | 16 cores | Ultra: 2-die | Fully supported |
| M5 / Pro / Max / Ultra | Apple10 | Yes | 16 cores | Ultra: 2-die | Fully supported |
Auto-detected features: GPU family, device tier, core counts, memory bandwidth, dynamic caching, mesh shaders, NAX (M5+), UltraFusion topology (via sysctl hw.packages), ANE availability.
Tier-based kernel tuning: Matrix tile sizes, FlashAttention block sizes, fused kernel threadgroup sizes, and batch multipliers are automatically selected based on device tier (Base/Pro/Max/Ultra) and GPU family. See docs/hardware-support.md for the full tuning matrix.
PMetal is organized as a Rust workspace with 20 specialized crates:
pmetal/
├── pmetal-bridge # Zero-allocation MLX C++ bridge (inline array FFI)
├── pmetal-core # Foundation: configs, traits, types, error handling
├── pmetal-core-derive # Derive macros for the core traits
├── pmetal-metal # Custom Metal GPU kernels + ANE runtime
├── pmetal-mlx # MLX backend integration (KV cache, RoPE, etc.)
├── pmetal-models # LLM architectures (Llama, Qwen, DeepSeek, etc.)
├── pmetal-lora # LoRA/QLoRA training implementations
├── pmetal-trainer # Training loops (SFT, DPO, SimPO, ORPO, KTO, GRPO, etc.)
├── pmetal-data # Dataset loading, chat templates, tokenization
├── pmetal-hub # HuggingFace Hub integration + model fit estimation
├── pmetal-distill # Knowledge distillation losses, offline caches, and TAID
├── pmetal-merge # Model merging (15 strategies)
├── pmetal-gguf # GGUF format with imatrix quantization
├── pmetal-mhc # Manifold-Constrained Hyper-Connections
├── pmetal-distributed # Distributed training (mDNS, Ring All-Reduce)
├── pmetal-vocoder # BigVGAN neural vocoder
├── pmetal-serve # OpenAI- and Anthropic-compatible inference server
├── pmetal-mcp # MCP server (51 tools for Claude Desktop)
├── pmetal-py # Python bindings (maturin/PyO3)
├── pmetal # Umbrella crate: CLI binary + TUI control center
└── pmetal-gui # Desktop GUI (Tauri + Svelte + TailwindCSS)
The pmetal crate is both the umbrella library — it re-exports every sub-crate behind a feature flag — and the binary that ships the CLI and TUI.
DynamicModel dispatcher)All causal language models below can be loaded from HuggingFace Hub or local safetensors and used for generation via the CLI, TUI, GUI, or SDK.
| Family | Architecture | Variants | model_type values |
|---|---|---|---|
| Llama | Llama | 2, 3, 3.1, 3.2, 3.3 | llama, llama3 |
| Llama 4 | Llama4 | Scout, Maverick | llama4 |
| Qwen 2 | Qwen2 | 2, 2.5 | qwen2, qwen2_5 |
| Qwen 3 | Qwen3 | 3 | qwen3 |
| Qwen 3 MoE | Qwen3MoE | 3-MoE | qwen3_moe |
| Qwen 3.5 / 3.6 | Qwen3Next | 3.5 (Next), 3.6 | qwen3_next, qwen3_5, qwen3_6 |
| DeepSeek | DeepSeek | V3, V3.2, V3.2-Speciale | deepseek, deepseek_v3 |
| Mistral | Mistral | 7B, Mixtral 8x7B | mistral, mixtral |
| Gemma | Gemma | 2, 3 | gemma, gemma2, gemma3 |
| Phi 3 | Phi | 3, 3.5 | phi, phi3 |
| Phi 4 | Phi4 | 4 | phi4 |
| Cohere | Cohere | Command R | cohere, command_r |
| Granite | Granite | 3.0, 3.1, Hybrid MoE | granite, granitehybrid |
| NemotronH | NemotronH | Hybrid (Mamba+Attention) | nemotron_h |
| GPT-OSS | GptOss | 20B, 120B | gpt_oss, gpt-oss |
| Gemma 4 | Gemma4 | 4 | gemma4, gemma4_text |
| Llama 3.2 Vision | Mllama | 11B, 90B | mllama, mllama_text_model |
| DiffusionGemma | DiffusionGemma | dense, MoE | diffusion_gemma, diffusion_gemma_text |
Gemma 4 MTP assistant checkpoints (model_type = "gemma4_assistant") are supported as
draft assistants via pmetal infer --draft-model <assistant>. The standard sampling
controls (--temperature, --top-k, --top-p, --min-p, and penalties) are preserved
through speculative verification, so MTP does not change generation quality.
Qwen3Next/Qwen3.6 checkpoints with bundled mtp.* weights can use exact speculative
decoding via pmetal infer --mtp --mtp-draft-tokens 3. This path preserves the same
sampling controls, supports bundled multi-predictor MTP heads, works with FP8 target/MTP
weights, packed expert offload, and LoRA-merged targets, and verifies every drafted token
against the target model. LoRA and packed expert offload are separate modes; fuse the adapter
first if you need both together. MTP inference prints draft acceptance metrics, including
accepted/attempted draft tokens, accepted tokens per verify step, and target bonus/correction
tokens.
Custom draft checkpoint creation is wired through pmetal train-mtp and pmetal train-draft.
train-mtp exports HF-compatible Gemma 4 assistant checkpoints and Qwen mtp.* predictor
checkpoints; load the latter with pmetal infer --mtp --mtp-model ./qwen-mtp. train-draft
exports DFlash draft checkpoints for the dedicated pmetal dflash runtime.
| Family | Architecture | Variants | model_type values |
|---|---|---|---|
| BERT | Bert | BERT, RoBERTa, DistilBERT, XLM-RoBERTa | bert, roberta, distilbert, xlm-roberta, xlm_roberta |
LoRA training is supported for models that have implementations in DynamicLoraModel. Architecture detection is automatic — just point pmetal train at a model directory or HuggingFace ID.
| Architecture | LoRA | QLoRA | Notes |
|---|---|---|---|
| Llama | Yes | Yes | Covers Llama 2, 3, 3.1, 3.2, 3.3. Gradient checkpointing supported. |
| Llama 4 | Yes | Yes | Scout/Maverick support via DynamicLoraModel. |
| Qwen 2 | Yes | Yes | Uses Qwen3 LoRA implementation internally. |
| Qwen 3 | Yes | Yes | Gradient checkpointing supported. |
| Qwen 3 MoE | Yes | Yes | Sparse MoE support. |
| Qwen 3.5 / 3.6 (Next) | Yes | Yes | Hybrid architecture with nested text_config handling and bundled MTP inference. |
| Gemma | Yes | Yes | GeGLU activation, special RMSNorm. |
| Gemma 4 | Yes | Yes | Multimodal-era Gemma text path with MTP assistant inference support. |
| Mistral | Yes | Yes | Sliding window attention support. |
| Phi 3/4 | Yes | Yes | Partial RoPE, fused gate_up projection. |
| DeepSeek | Yes | Yes | V3-family support. |
| Cohere | Yes | Yes | Command R support. |
| Granite | Yes | Yes | Dense and hybrid variants. |
| NemotronH | Yes | Yes | Hybrid architecture support. |
| GPT-OSS | Yes | Yes | MoE variants. |
| DiffusionGemma | Yes | Yes | Block-diffusion; trained via pmetal train-diffusion (add --qlora). |
The following architectures have implementations in pmetal-models but are not wired into the DynamicModel dispatcher and cannot be loaded via the CLI or DynamicModel::load():
| Family | Module | Notes |
|---|---|---|
| Pixtral | pixtral | 12B vision-language model |
| Qwen2-VL | qwen2_vl | 2B, 7B vision-language model |
| CLIP | clip | ViT-L/14 vision encoder |
| Whisper | whisper | Base, Small, Medium, Large speech models |
| T5 | t5 | Encoder-decoder architecture |
These modules can be used directly via their Rust types (e.g., pmetal_models::architectures::pixtral::Pixtral) but require manual weight loading.
| Family | Variants | Status |
|---|---|---|
| Flux | 1-dev, 1-schnell | Dispatcher + pipeline implemented |
All training methods support callback-based cancellation (should_stop()), metrics JSONL logging, and adaptive learning rate control.
| Method | CLI | GUI | TUI | Library |
|---|---|---|---|---|
| SFT (Supervised Fine-Tuning) | train | Yes | Yes | orchestrator::run_training() |
| LoRA | train | Yes | Yes | orchestrator::run_training() |
| QLoRA (4-bit) | train --quantization nf4 | Yes | Yes | orchestrator::run_training() |
| DoRA | — | — | — | LoraConfig { use_dora: true } |
| DPO (Direct Preference) | — | — | — | DpoTrainer |
| SimPO (Simple Preference) | — | — | — | SimpoTrainer |
| ORPO (Odds-Ratio Preference) | — | — | — | OrpoTrainer |
| KTO (Kahneman-Tversky) | — | — | — | KtoTrainer |
| GRPO (Reasoning) | grpo | Yes | Yes | GrpoTrainer |
| DAPO (Decoupled GRPO) | grpo --dapo | Yes | Yes | GrpoTrainer DAPO mode |
| Knowledge Distillation | distill | Yes | Yes | Distiller |
| TAID (Temporally Adaptive) | — | — | — | TaidDistiller |
| ANE Training | train (auto) | — | Yes | AneTrainingLoop |
| RLKD (RL + Distillation) | rlkd | Yes | Yes | RlkdTrainer |
| Embedding Training | embed-train | Yes | Yes | EmbeddingTrainer |
| Block-Diffusion (DiffusionGemma) | train-diffusion | — | — | DiffusionTrainingLoop |
| Gemma/Qwen MTP Predictor Training | train-mtp | — | — | pmetal_trainer::mtp_training |
| DFlash Draft Training | train-draft | — | — | pmetal_trainer::mtp_training |
Additional methods available via the library only: GSPO (GspoTrainer), PPO (PpoTrainer), Online DPO (OnlineDpoTrainer).
Custom Metal shaders provide significant speedups:
lora-metal-fused feature)Native ANE integration for power-efficient training and inference:
Near-optimal KV cache compression for long-context inference:
q3_5 (near-lossless), q2_5 (6.4x compression)TrainingCallback trait with lifecycle hooks (on_step_start, on_step_end, should_stop) for metrics logging, progress reporting, and clean cancellationAuto-detected training data formats:
{"conversations": [{"from": "human", "value": "..."}, ...]}{"instruction": "...", "input": "...", "output": "..."}{"messages": [{"role": "user", "content": "..."}, ...]}{"problem": "...", "thinking": "...", "solution": "..."}{"text": "..."}Custom columns: Use --text-column for arbitrary field names, --text-columns col1,col2 to concatenate multiple columns, and --prompt-column/--response-column for SFT loss masking. All training commands (train, distill, grpo, rlkd) support column flags uniformly.
The pmetal dataset subcommand provides utilities for analysis, download from HuggingFace, and format conversion (Parquet, JSON, JSONL, CSV, ShareGPT, Alpaca).
HuggingFace Hub Search: pmetal search with memory fit estimation and download
Model Merging (15 strategies via MergeConfig, 12 via CLI):
| CLI | Library | Description |
|---|---|---|
linear | LinearMerge | Simple weighted averaging |
slerp | SlerpMerge | Spherical linear interpolation |
ties | TiesMerge | Task arithmetic with sparsification and sign consensus |
dare_ties | DareMerge | Random pruning with rescaling (TIES variant) |
dare_linear | DareMerge | Random pruning with rescaling (linear variant) |
task_arithmetic | TaskArithmeticMerge | Task vector arithmetic |
della | DellaMerge | Adaptive magnitude-based pruning |
della_linear | DellaMerge | Adaptive magnitude pruning (linear variant) |
breadcrumbs | BreadcrumbsMerge | Breadcrumbs merge strategy |
model_stock | ModelStockMerge | Geometric interpolation based on task vector similarity |
nearswap | NearswapMerge | Near-swap merge strategy |
passthrough | PassthroughMerge | Layer passthrough composition |
| — | RamMerge | RAM merge strategy |
| — | SouperMerge | Souper merge strategy |
| — | MultiSlerpMerge | Multi-model SLERP |
GPU-Accelerated Merging: Metal-based merge operations for large models
FP8-Aware Merging: Merge with FP8 quantization for memory efficiency
Async Merge Pipeline: Double-buffered streaming merge for large models
LoRA Fusing: Merge LoRA adapters into base weights (standard and accurate modes)
GGUF Quantization (13 format options):
| Format | Description |
|---|---|
dynamic | Auto-select per layer |
q8_0 | 8-bit quantization |
q6k | 6-bit k-quant |
q5km | 5-bit k-quant (medium) |
q5ks | 5-bit k-quant (small) |
q4km | 4-bit k-quant (medium) |
q4ks | 4-bit k-quant (small) |
q3km | 3-bit k-quant (medium) |
q3ks | 3-bit k-quant (small) |
q3kl | 3-bit k-quant (large) |
q2k | 2-bit k-quant |
f16 | Float16 |
f32 | Float32 |
Supports importance matrix (--imatrix) for improved quantization quality. KL-calibrated quantization (--kl-calibrate) selects per-tensor quantization types via NRMSE + cosine distance, with optional --target-bpw for budget-constrained quantization.
FP8 Runtime Quantization: Convert to FP8 (E4M3) at inference time for ~2x memory reduction
Multiple distillation methods and loss functions:
TaidDistillerpmetal train Parameters| Parameter | Default | Description |
|---|---|---|
--lora-r | 16 | LoRA rank |
--lora-alpha | 32.0 | LoRA scaling factor (2x rank) |
--batch-size | 1 | Micro-batch size |
--learning-rate | 2e-4 | Learning rate |
--max-seq-len | 0 | Max seq len (0 = auto-detect) |
--epochs | 1 | Number of training epochs |
--max-grad-norm | 1.0 | Gradient clipping |
--quantization | none | QLoRA method (nf4, fp4, int8) |
--gradient-accumulation-steps | 4 | Gradient accumulation steps |
--embedding-lr | None | Separate LR for embeddings |
--no-metal-fused-optimizer | false | Disable Metal fused optimizer |
--lr-schedule | cosine | Schedule type (constant, linear, cosine, cosine_with_restarts, polynomial, wsd) |
--no-gradient-checkpointing | false | Disable gradient checkpointing (enabled by default) |
--gradient-checkpointing-layers | 4 | Number of layers per checkpoint block |
--warmup-steps | 100 | Learning rate warmup steps |
--weight-decay | 0.01 | AdamW weight decay coefficient |
--no-sequence-packing | false | Disable sequence packing |
--cut-cross-entropy | false | Memory-efficient loss (avoids full logit materialization) |
--text-column | — | Custom JSONL column name for training text |
--text-columns | — | Multi-column concat (comma-separated, e.g. thinking,solution) |
--prompt-column | — | Column for prompt (enables SFT loss masking) |
--response-column | — | Column for response (with prompt masking) |
--column-separator | \n\n | Separator for --text-columns |
--config | — | Path to YAML configuration file |
pmetal infer Parameters| Parameter | Default | Description |
|---|---|---|
--temperature | Model default | Sampling temperature |
--top-k | Model default | Top-k sampling |
--top-p | Model default | Nucleus sampling |
--min-p | Model default | Min-p dynamic sampling |
--max-tokens | 256 | Maximum generation length |
--repetition-penalty | 1.0 | Repetition penalty |
--frequency-penalty | 0.0 | Frequency penalty |
--presence-penalty | 0.0 | Presence penalty |
--chat | false | Apply chat template |
--draft-model | — | Gemma 4 MTP assistant checkpoint |
--mtp | false | Enable Qwen3Next/Qwen3.6 exact speculative MTP |
--mtp-model | bundled mtp.* | Optional external Qwen MTP checkpoint from train-mtp |
--mtp-draft-tokens | 3 | Qwen MTP draft tokens per verification step |
--fp8 | false | Use FP8 weights (~2x mem reduction) |
--compiled | false | Use JIT-compiled sampling |
--ane-max-seq-len | 1024 | Max ANE kernel sequence length |
--tools | — | Tool/function definitions file (OpenAI format) |
--system | — | System message |
Defaults are cli, dashboard, trainer, lora, merge, ane and distributed; the rest are pulled in transitively.
| Feature | Default | Crate | Description |
|---|---|---|---|
cli | Yes | — | The pmetal binary and its CLI-only dependencies |
core | Yes* | pmetal-core | Foundation types, configs, traits |
gguf | Yes* | pmetal-gguf | GGUF format support |
metal | Yes* | pmetal-metal | Metal GPU kernels |
hub | Yes* | pmetal-hub | HuggingFace Hub integration |
mlx | Yes* | pmetal-mlx | MLX backend |
models | Yes* | pmetal-models | LLM architectures |
lora | Yes | pmetal-lora | LoRA/QLoRA |
trainer | Yes | pmetal-trainer | Training loops (pulls in data, distill) |
data | Yes* | pmetal-data | Dataset loading (*via cli and trainer) |
distill | Yes* | pmetal-distill | Knowledge distillation (*via trainer) |
merge | Yes | pmetal-merge | Model merging strategies |
distributed | Yes | pmetal-distributed | Distributed training and the cluster subcommand |
ane | Yes | — | Apple Neural Engine |
dashboard | Yes | — | TUI control center |
native-only | No | pmetal-bridge | Bridge-only build with no mlx-rs/mlx-sys |
lora-metal-fused | No | — | ~2x LoRA training speedup via fused Metal kernels |
vocoder | No | pmetal-vocoder | BigVGAN neural vocoder |
mhc | No | pmetal-mhc | Manifold-Constrained Hyper-Connections |
serve | No | pmetal-serve | OpenAI- and Anthropic-compatible inference server |
mcp | No | pmetal-mcp | MCP server (51 tools for Claude Desktop) |
full | No | — | All sub-crate features (not cli, serve or mcp) |
serve and mcp stay out of the default set so library consumers don't inherit axum and rmcp. The prebuilt binary and the Homebrew formula both build with --features serve,mcp, so pmetal serve and pmetal mcp are there if you installed either way. Building yourself, add the flag: cargo install pmetal --features serve,mcp.
# Release build (default features: ANE + Dashboard)
cargo build --release
# Build without ANE
cargo build --release --no-default-features --features dashboard
# Run tests (single-threaded for Metal compatibility)
just test
# Build GUI
cd crates/pmetal-gui && bun install && bun tauri build
# cargo-kani proofs for ring all-reduce and topology
just kani-verify
Licensed under either of MIT or Apache-2.0.
Rust
89.9%
Metal
3.9%
C++
2.7%
Svelte
2.2%
PMetal: high-performance Apple Silicon framework for local LLM inference, LoRA/QLoRA fine-tuning, serving, quantization, and MLX/Metal acceleration.
See the codePowdered Metal — An ML SDK, framework, and application suite for Apple Silicon, written in Rust.
PMetal is a complete machine learning platform for Apple Silicon — from low-level Metal GPU kernels and Apple Neural Engine integration to high-level training APIs, a terminal TUI, and a full desktop GUI. Ship fine-tuned models without leaving the Apple ecosystem.
A full Tauri + Svelte desktop application for visual model management, training, and inference.
cd crates/pmetal-gui
bun install && bun tauri dev
19 pages: Dashboard, Training, GRPO, Distillation, Pretrain, Inference, DFlash, Models, Datasets, Merging, Quantize, Embed Train, RLKD, Ollama, Serve, Bench, Eval, Jobs, and Settings. Download models from HuggingFace, configure LoRA training with live loss metrics, chat with models, merge weights, and quantize — all from the GUI. Training, inference, distillation and GRPO run in-process with real-time progress updates; the remaining pages drive the pmetal CLI as a subprocess, which the app bundles.
A full-featured terminal control center with 20 tabs.
pmetal tui
| Tab | Description |
|---|---|
| Device | GPU/ANE info, Metal feature detection, memory gauge, kernel tuning, UltraFusion topology |
| Models | Browse cached models, HuggingFace Hub search (S), memory fit estimation, download |
| Datasets | Scan and preview local datasets (JSONL, Parquet, CSV) with line counts |
| Tokenize | Tokenize a text corpus into binary shards for pretraining |
| Training | Configure and launch SFT/LoRA/QLoRA training runs with sectioned parameter forms |
| Embed Train | Train a sentence-embedding (encoder-only) model with contrastive losses |
| Pretrain | Full-parameter pretraining from scratch |
| Distillation | Configure knowledge distillation (online, offline, progressive) |
| RLKD | Reinforcement learning with knowledge distillation |
| GRPO | Configure GRPO/DAPO reasoning training with reward functions and sampling params |
| Dashboard | Live loss curves (braille), LR schedule, throughput sparklines, timing breakdown gauges |
| Inference | Interactive chat interface with markdown rendering and generation settings sidebar |
| DFlash | Block-diffusion speculative decoding |
| Serve | OpenAI-compatible server control |
| Quantize | GGUF and MLX quantization with bit/method selection |
| Merge | SLERP, TIES, DARE and linear model merging |
| Bench | Training and inference benchmarking |
| Eval | Perplexity evaluation against a dataset |
| Ollama | Modelfile generation and Ollama export |
| Jobs | Training run history with log viewer, status tracking, and metadata |
Keybindings: Ctrl+P to jump to any tab (type to filter), Tab/Shift+Tab to cycle, Alt+1-9 (or Ctrl+1-9) for the first nine, L to adjust learning rate mid-run, ? for contextual help, q to quit.
# LoRA fine-tuning with sequence packing (default)
pmetal train \
--model Qwen/Qwen3-0.6B \
--dataset train.jsonl \
--output ./output \
--lora-r 16 --batch-size 4 --learning-rate 2e-4
# Inference with LoRA adapter
pmetal infer \
--model Qwen/Qwen3-0.6B \
--lora ./output/lora_weights.safetensors \
--prompt "Explain quantum entanglement" \
--chat
# Train and use a Qwen3Next/Qwen3.6 MTP predictor
pmetal tokenize --input train.jsonl --output ./tok --tokenizer Qwen/Qwen3.6-30B-A3B-Instruct
pmetal train-mtp \
--model Qwen/Qwen3.6-30B-A3B-Instruct \
--family qwen3-next \
--shards ./tok/shard_00000.bin \
--output ./qwen-mtp
pmetal infer \
--model Qwen/Qwen3.6-30B-A3B-Instruct \
--mtp --mtp-model ./qwen-mtp \
--prompt "Explain monotonic queues."
# Knowledge distillation
pmetal distill \
--teacher Qwen/Qwen3-4B \
--student Qwen/Qwen3.5-0.8B-Base \
--dataset train.jsonl
# GRPO reasoning training
pmetal grpo \
--model Qwen/Qwen3-0.6B \
--dataset reasoning.jsonl \
--reasoning-rewards
# HuggingFace model search with memory fit
pmetal search "qwen 0.6b" --detailed
# Merge models with SLERP
pmetal merge \
--model-a model-a --model-b model-b \
--method slerp --t 0.5
# Quantize to GGUF
pmetal quantize \
--model ./output \
--output model.gguf --method q4_k_m
# Fuse LoRA into base model
pmetal fuse \
--model Qwen/Qwen3-0.6B \
--lora ./output/lora_weights.safetensors
# Evaluate perplexity
pmetal eval \
--model Qwen/Qwen3-0.6B \
--dataset eval.jsonl
# Start OpenAI-compatible server
# (in the prebuilt binary and the brew formula; add --features serve if you build it yourself)
pmetal serve --model Qwen/Qwen3-0.6B --port 8080
| Command | Description |
|---|---|
train | Fine-tune with LoRA/QLoRA (SFT) |
train-mtp | Train Gemma 4 assistant or Qwen3Next/Qwen3.6 MTP predictor checkpoints |
train-draft | Train DFlash block-diffusion draft checkpoints |
train-diffusion | LoRA/QLoRA fine-tune a DiffusionGemma block-diffusion model |
pretrain | Pretrain a model from scratch (full-parameter, no LoRA) |
infer | Interactive inference with chat, tool use, and thinking mode |
distill | Knowledge distillation (online, offline, progressive) |
grpo | GRPO/DAPO reasoning training (VLM, speculative, async rewards) |
rlkd | Reinforcement Learning with Knowledge Distillation |
embed-train | Sentence-transformer fine-tuning (InfoNCE, Triplet, CoSENT) |
search | Search HuggingFace Hub with memory fit estimation |
download | Download a model from HuggingFace Hub |
merge | Merge two models (12 strategies) |
quantize | GGUF quantization (24 methods) |
fuse | Fuse LoRA adapter weights into base model |
eval | Evaluate model perplexity on a dataset |
serve | OpenAI- and Anthropic-compatible inference server |
tui | Full TUI control center (20 tabs) |
dashboard | Real-time training metrics visualization |
dataset | Dataset utilities: analyze, download, convert |
ollama | Ollama integration: modelfile, create, templates |
info | Show device info (GPU, ANE, bandwidth, NAX) |
memory | Show memory usage and available capacity |
init | Generate a sample configuration file |
bench | Benchmark training performance |
bench-gen | Benchmark generation loop timing |
bench-ffi | Benchmark FFI overhead |
bench-workload | Benchmark real cached inference/training workloads |
bench-corpus | Structured kernel benchmarking with JSON reporting |
bench-gdn | Benchmark Qwen3.5 GDN backends on real layer shapes |
tokenize | Tokenize a text corpus into binary shards for pretraining |
pack-experts | Pack expert weights for SSD-offloaded MoE inference |
dflash | Block-diffusion speculative decoding |
mcp | Start MCP server over stdio (51 tools for Claude Desktop / Claude Code) |
cluster | Multi-Mac cluster: discover peers, train across machines, run all-reduce / pipeline benchmarks |
Connect two or more Apple Silicon Macs into a "home cluster" for distributed training and inference. PMetal auto-detects every NIC on the box (Thunderbolt-Bridge, Ethernet, Wi-Fi), advertises them via mDNS, and forms a ring biased toward the fastest fabric — Thunderbolt cables are picked over Ethernet, Ethernet over Wi-Fi, all without configuration.
# 1. Connect the Macs (Thunderbolt-4/5 cable recommended; Ethernet works too).
# 2. On every Mac:
pmetal cluster status # Show local NICs + any peers already announcing.
pmetal cluster up # Join the cluster, hold connection open.
# 3. On every Mac at the same time:
pmetal cluster bench --mb 64 --iters 10 # All-reduce throughput per fabric.
pmetal cluster pipeline-bench --tokens 16 --layers 32 # Pipeline activation transport bench.
# 4. Distributed training (each Mac runs the same command simultaneously):
pmetal train --model Qwen/Qwen3-0.6B \
--dataset train.jsonl \
--distributed-auto # mDNS-discovers peers, all-reduce gradients.
pmetal cluster status example output:
Local peer: 12D3KooW…XyZ (rank 0/2)
Thunderbolt ring: yes
Local interfaces:
bridge0 thunderbolt 169.254.42.1
en0 ethernet 192.168.1.10
lo0 loopback 127.0.0.1, ::1
Cluster peers:
peer-id local primary-addr fabric paths
12D3KooW…XyZ yes 169.254.42.1:52416 thunderbolt 2
12D3KooW…AbC no 169.254.42.2:52416 thunderbolt 2
What's wired today: gradient all-reduce (multi-machine training, real), fabric-aware ring formation with Thunderbolt > Ethernet > Wi-Fi priority, automatic fabric fallback when a cable is unplugged mid-job, gradient compression (TopK, FP16/BF16/INT8 quantization, error feedback), and a transport-tested pipeline harness with multi-process integration tests. Per-architecture partial-layer execution (the prerequisite for serving a model that doesn't fit on one Mac) is the next step — the harness is ready, the model API is the bottleneck.
PMetal is an embeddable SDK — integrate training, inference, and model operations into your own Rust applications. pmetal re-exports every sub-crate, so one dependency gets you the whole framework.
orchestrator::run_training is the one-call training entry point the CLI itself uses:
use pmetal::trainer::orchestrator::{TrainingJobConfig, run_training};
let config = TrainingJobConfig {
model_id: "Qwen/Qwen3-0.6B".to_string(),
dataset: "train.jsonl".to_string(),
output_dir: "./output".to_string(),
..Default::default()
};
let result = run_training(config, None, Vec::new()).await?;
println!("final loss {:.4} over {} steps", result.final_loss, result.total_steps);
Inference goes through the model dispatcher:
use pmetal::models::generate;
use pmetal::prelude::*;
let model_dir = pmetal::hub::resolve_model_path("Qwen/Qwen3-0.6B", None, None).await?;
let mut model = DynamicModel::load(&model_dir)?;
let tokenizer = Tokenizer::from_model_dir(&model_dir)?;
let input_ids = tokenizer.encode_with_special_tokens("What is 2+2?")?;
let output = generate(
|input| model.forward(input, None),
&input_ids,
GenerationConfig::sampling(256, 0.7),
)?;
println!("{}", tokenizer.decode(&output.token_ids[input_ids.len()..])?);
Preference optimization (DpoTrainer, SimpoTrainer, OrpoTrainer, KtoTrainer) and TAID
distillation are library-only for now — there is no CLI subcommand for them yet.
For step-by-step control, use the crates directly: pmetal_trainer::TrainingLoop,
pmetal_models::DynamicModel, pmetal_lora::DynamicLoraModel, pmetal_distill::Distiller. Every
code block in each crate's README is compiled as a doctest, so they stay honest. See
examples/ for complete working programs, including manual training-loop
orchestration and ANE-specific workflows.
PMetal exposes a Python extension module via PyO3. Install with maturin develop from crates/pmetal-py.
import pmetal
# Fine-tune with sensible defaults
result = pmetal.finetune(
"Qwen/Qwen3-0.6B",
"train.jsonl",
lora_r=16,
learning_rate=2e-4,
epochs=3,
)
print(f"Loss: {result['final_loss']}, Steps: {result['total_steps']}")
# Inference
text = pmetal.infer("Qwen/Qwen3-0.6B", "What is 2+2?")
print(text)
# Inference with LoRA adapter
text = pmetal.infer(
"Qwen/Qwen3-0.6B",
"Explain quantum entanglement",
lora="./output/lora_weights.safetensors",
)
import pmetal
# Configure training components
lora_config = pmetal.LoraConfig(r=16, alpha=32.0)
training_config = pmetal.TrainingConfig(
learning_rate=2e-4,
num_epochs=3,
batch_size=4,
max_seq_len=2048,
)
# Create trainer
trainer = pmetal.Trainer(
model_id="Qwen/Qwen3-0.6B",
lora_config=lora_config,
training_config=training_config,
dataset_path="train.jsonl",
)
trainer.add_callback(pmetal.ProgressCallback())
result = trainer.train()
# Load model for inference
model = pmetal.Model.load("Qwen/Qwen3-0.6B")
print(model.generate("Hello world", temperature=0.7))
Prebuilt signed binaries are available on the Releases page.
Crates are available on crates.io.
Build from source:
git clone https://github.com/epistates/pmetal.git && cd pmetal
cargo build --release # CLI + TUI
cd crates/pmetal-gui && bun install && bun tauri build # GUI (optional)
PMetal automatically detects Apple Silicon capabilities at startup and tunes kernel parameters accordingly.
| Chip Family | GPU Family | NAX | ANE | UltraFusion | Status |
|---|---|---|---|---|---|
| M1 / Pro / Max / Ultra | Apple7 | - | 16 cores | Ultra: 2-die | Fully supported |
| M2 / Pro / Max / Ultra | Apple8 | - | 16 cores | Ultra: 2-die | Fully supported |
| M3 / Pro / Max / Ultra | Apple9 | - | 16 cores | Ultra: 2-die | Fully supported |
| M4 / Pro / Max / Ultra | Apple9 | - | 16 cores | Ultra: 2-die | Fully supported |
| M5 / Pro / Max / Ultra | Apple10 | Yes | 16 cores | Ultra: 2-die | Fully supported |
Auto-detected features: GPU family, device tier, core counts, memory bandwidth, dynamic caching, mesh shaders, NAX (M5+), UltraFusion topology (via sysctl hw.packages), ANE availability.
Tier-based kernel tuning: Matrix tile sizes, FlashAttention block sizes, fused kernel threadgroup sizes, and batch multipliers are automatically selected based on device tier (Base/Pro/Max/Ultra) and GPU family. See docs/hardware-support.md for the full tuning matrix.
PMetal is organized as a Rust workspace with 20 specialized crates:
pmetal/
├── pmetal-bridge # Zero-allocation MLX C++ bridge (inline array FFI)
├── pmetal-core # Foundation: configs, traits, types, error handling
├── pmetal-core-derive # Derive macros for the core traits
├── pmetal-metal # Custom Metal GPU kernels + ANE runtime
├── pmetal-mlx # MLX backend integration (KV cache, RoPE, etc.)
├── pmetal-models # LLM architectures (Llama, Qwen, DeepSeek, etc.)
├── pmetal-lora # LoRA/QLoRA training implementations
├── pmetal-trainer # Training loops (SFT, DPO, SimPO, ORPO, KTO, GRPO, etc.)
├── pmetal-data # Dataset loading, chat templates, tokenization
├── pmetal-hub # HuggingFace Hub integration + model fit estimation
├── pmetal-distill # Knowledge distillation losses, offline caches, and TAID
├── pmetal-merge # Model merging (15 strategies)
├── pmetal-gguf # GGUF format with imatrix quantization
├── pmetal-mhc # Manifold-Constrained Hyper-Connections
├── pmetal-distributed # Distributed training (mDNS, Ring All-Reduce)
├── pmetal-vocoder # BigVGAN neural vocoder
├── pmetal-serve # OpenAI- and Anthropic-compatible inference server
├── pmetal-mcp # MCP server (51 tools for Claude Desktop)
├── pmetal-py # Python bindings (maturin/PyO3)
├── pmetal # Umbrella crate: CLI binary + TUI control center
└── pmetal-gui # Desktop GUI (Tauri + Svelte + TailwindCSS)
The pmetal crate is both the umbrella library — it re-exports every sub-crate behind a feature flag — and the binary that ships the CLI and TUI.
DynamicModel dispatcher)All causal language models below can be loaded from HuggingFace Hub or local safetensors and used for generation via the CLI, TUI, GUI, or SDK.
| Family | Architecture | Variants | model_type values |
|---|---|---|---|
| Llama | Llama | 2, 3, 3.1, 3.2, 3.3 | llama, llama3 |
| Llama 4 | Llama4 | Scout, Maverick | llama4 |
| Qwen 2 | Qwen2 | 2, 2.5 | qwen2, qwen2_5 |
| Qwen 3 | Qwen3 | 3 | qwen3 |
| Qwen 3 MoE | Qwen3MoE | 3-MoE | qwen3_moe |
| Qwen 3.5 / 3.6 | Qwen3Next | 3.5 (Next), 3.6 | qwen3_next, qwen3_5, qwen3_6 |
| DeepSeek | DeepSeek | V3, V3.2, V3.2-Speciale | deepseek, deepseek_v3 |
| Mistral | Mistral | 7B, Mixtral 8x7B | mistral, mixtral |
| Gemma | Gemma | 2, 3 | gemma, gemma2, gemma3 |
| Phi 3 | Phi | 3, 3.5 | phi, phi3 |
| Phi 4 | Phi4 | 4 | phi4 |
| Cohere | Cohere | Command R | cohere, command_r |
| Granite | Granite | 3.0, 3.1, Hybrid MoE | granite, granitehybrid |
| NemotronH | NemotronH | Hybrid (Mamba+Attention) | nemotron_h |
| GPT-OSS | GptOss | 20B, 120B | gpt_oss, gpt-oss |
| Gemma 4 | Gemma4 | 4 | gemma4, gemma4_text |
| Llama 3.2 Vision | Mllama | 11B, 90B | mllama, mllama_text_model |
| DiffusionGemma | DiffusionGemma | dense, MoE | diffusion_gemma, diffusion_gemma_text |
Gemma 4 MTP assistant checkpoints (model_type = "gemma4_assistant") are supported as
draft assistants via pmetal infer --draft-model <assistant>. The standard sampling
controls (--temperature, --top-k, --top-p, --min-p, and penalties) are preserved
through speculative verification, so MTP does not change generation quality.
Qwen3Next/Qwen3.6 checkpoints with bundled mtp.* weights can use exact speculative
decoding via pmetal infer --mtp --mtp-draft-tokens 3. This path preserves the same
sampling controls, supports bundled multi-predictor MTP heads, works with FP8 target/MTP
weights, packed expert offload, and LoRA-merged targets, and verifies every drafted token
against the target model. LoRA and packed expert offload are separate modes; fuse the adapter
first if you need both together. MTP inference prints draft acceptance metrics, including
accepted/attempted draft tokens, accepted tokens per verify step, and target bonus/correction
tokens.
Custom draft checkpoint creation is wired through pmetal train-mtp and pmetal train-draft.
train-mtp exports HF-compatible Gemma 4 assistant checkpoints and Qwen mtp.* predictor
checkpoints; load the latter with pmetal infer --mtp --mtp-model ./qwen-mtp. train-draft
exports DFlash draft checkpoints for the dedicated pmetal dflash runtime.
| Family | Architecture | Variants | model_type values |
|---|---|---|---|
| BERT | Bert | BERT, RoBERTa, DistilBERT, XLM-RoBERTa | bert, roberta, distilbert, xlm-roberta, xlm_roberta |
LoRA training is supported for models that have implementations in DynamicLoraModel. Architecture detection is automatic — just point pmetal train at a model directory or HuggingFace ID.
| Architecture | LoRA | QLoRA | Notes |
|---|---|---|---|
| Llama | Yes | Yes | Covers Llama 2, 3, 3.1, 3.2, 3.3. Gradient checkpointing supported. |
| Llama 4 | Yes | Yes | Scout/Maverick support via DynamicLoraModel. |
| Qwen 2 | Yes | Yes | Uses Qwen3 LoRA implementation internally. |
| Qwen 3 | Yes | Yes | Gradient checkpointing supported. |
| Qwen 3 MoE | Yes | Yes | Sparse MoE support. |
| Qwen 3.5 / 3.6 (Next) | Yes | Yes | Hybrid architecture with nested text_config handling and bundled MTP inference. |
| Gemma | Yes | Yes | GeGLU activation, special RMSNorm. |
| Gemma 4 | Yes | Yes | Multimodal-era Gemma text path with MTP assistant inference support. |
| Mistral | Yes | Yes | Sliding window attention support. |
| Phi 3/4 | Yes | Yes | Partial RoPE, fused gate_up projection. |
| DeepSeek | Yes | Yes | V3-family support. |
| Cohere | Yes | Yes | Command R support. |
| Granite | Yes | Yes | Dense and hybrid variants. |
| NemotronH | Yes | Yes | Hybrid architecture support. |
| GPT-OSS | Yes | Yes | MoE variants. |
| DiffusionGemma | Yes | Yes | Block-diffusion; trained via pmetal train-diffusion (add --qlora). |
The following architectures have implementations in pmetal-models but are not wired into the DynamicModel dispatcher and cannot be loaded via the CLI or DynamicModel::load():
| Family | Module | Notes |
|---|---|---|
| Pixtral | pixtral | 12B vision-language model |
| Qwen2-VL | qwen2_vl | 2B, 7B vision-language model |
| CLIP | clip | ViT-L/14 vision encoder |
| Whisper | whisper | Base, Small, Medium, Large speech models |
| T5 | t5 | Encoder-decoder architecture |
These modules can be used directly via their Rust types (e.g., pmetal_models::architectures::pixtral::Pixtral) but require manual weight loading.
| Family | Variants | Status |
|---|---|---|
| Flux | 1-dev, 1-schnell | Dispatcher + pipeline implemented |
All training methods support callback-based cancellation (should_stop()), metrics JSONL logging, and adaptive learning rate control.
| Method | CLI | GUI | TUI | Library |
|---|---|---|---|---|
| SFT (Supervised Fine-Tuning) | train | Yes | Yes | orchestrator::run_training() |
| LoRA | train | Yes | Yes | orchestrator::run_training() |
| QLoRA (4-bit) | train --quantization nf4 | Yes | Yes | orchestrator::run_training() |
| DoRA | — | — | — | LoraConfig { use_dora: true } |
| DPO (Direct Preference) | — | — | — | DpoTrainer |
| SimPO (Simple Preference) | — | — | — | SimpoTrainer |
| ORPO (Odds-Ratio Preference) | — | — | — | OrpoTrainer |
| KTO (Kahneman-Tversky) | — | — | — | KtoTrainer |
| GRPO (Reasoning) | grpo | Yes | Yes | GrpoTrainer |
| DAPO (Decoupled GRPO) | grpo --dapo | Yes | Yes | GrpoTrainer DAPO mode |
| Knowledge Distillation | distill | Yes | Yes | Distiller |
| TAID (Temporally Adaptive) | — | — | — | TaidDistiller |
| ANE Training | train (auto) | — | Yes | AneTrainingLoop |
| RLKD (RL + Distillation) | rlkd | Yes | Yes | RlkdTrainer |
| Embedding Training | embed-train | Yes | Yes | EmbeddingTrainer |
| Block-Diffusion (DiffusionGemma) | train-diffusion | — | — | DiffusionTrainingLoop |
| Gemma/Qwen MTP Predictor Training | train-mtp | — | — | pmetal_trainer::mtp_training |
| DFlash Draft Training | train-draft | — | — | pmetal_trainer::mtp_training |
Additional methods available via the library only: GSPO (GspoTrainer), PPO (PpoTrainer), Online DPO (OnlineDpoTrainer).
Custom Metal shaders provide significant speedups:
lora-metal-fused feature)Native ANE integration for power-efficient training and inference:
Near-optimal KV cache compression for long-context inference:
q3_5 (near-lossless), q2_5 (6.4x compression)TrainingCallback trait with lifecycle hooks (on_step_start, on_step_end, should_stop) for metrics logging, progress reporting, and clean cancellationAuto-detected training data formats:
{"conversations": [{"from": "human", "value": "..."}, ...]}{"instruction": "...", "input": "...", "output": "..."}{"messages": [{"role": "user", "content": "..."}, ...]}{"problem": "...", "thinking": "...", "solution": "..."}{"text": "..."}Custom columns: Use --text-column for arbitrary field names, --text-columns col1,col2 to concatenate multiple columns, and --prompt-column/--response-column for SFT loss masking. All training commands (train, distill, grpo, rlkd) support column flags uniformly.
The pmetal dataset subcommand provides utilities for analysis, download from HuggingFace, and format conversion (Parquet, JSON, JSONL, CSV, ShareGPT, Alpaca).
HuggingFace Hub Search: pmetal search with memory fit estimation and download
Model Merging (15 strategies via MergeConfig, 12 via CLI):
| CLI | Library | Description |
|---|---|---|
linear | LinearMerge | Simple weighted averaging |
slerp | SlerpMerge | Spherical linear interpolation |
ties | TiesMerge | Task arithmetic with sparsification and sign consensus |
dare_ties | DareMerge | Random pruning with rescaling (TIES variant) |
dare_linear | DareMerge | Random pruning with rescaling (linear variant) |
task_arithmetic | TaskArithmeticMerge | Task vector arithmetic |
della | DellaMerge | Adaptive magnitude-based pruning |
della_linear | DellaMerge | Adaptive magnitude pruning (linear variant) |
breadcrumbs | BreadcrumbsMerge | Breadcrumbs merge strategy |
model_stock | ModelStockMerge | Geometric interpolation based on task vector similarity |
nearswap | NearswapMerge | Near-swap merge strategy |
passthrough | PassthroughMerge | Layer passthrough composition |
| — | RamMerge | RAM merge strategy |
| — | SouperMerge | Souper merge strategy |
| — | MultiSlerpMerge | Multi-model SLERP |
GPU-Accelerated Merging: Metal-based merge operations for large models
FP8-Aware Merging: Merge with FP8 quantization for memory efficiency
Async Merge Pipeline: Double-buffered streaming merge for large models
LoRA Fusing: Merge LoRA adapters into base weights (standard and accurate modes)
GGUF Quantization (13 format options):
| Format | Description |
|---|---|
dynamic | Auto-select per layer |
q8_0 | 8-bit quantization |
q6k | 6-bit k-quant |
q5km | 5-bit k-quant (medium) |
q5ks | 5-bit k-quant (small) |
q4km | 4-bit k-quant (medium) |
q4ks | 4-bit k-quant (small) |
q3km | 3-bit k-quant (medium) |
q3ks | 3-bit k-quant (small) |
q3kl | 3-bit k-quant (large) |
q2k | 2-bit k-quant |
f16 | Float16 |
f32 | Float32 |
Supports importance matrix (--imatrix) for improved quantization quality. KL-calibrated quantization (--kl-calibrate) selects per-tensor quantization types via NRMSE + cosine distance, with optional --target-bpw for budget-constrained quantization.
FP8 Runtime Quantization: Convert to FP8 (E4M3) at inference time for ~2x memory reduction
Multiple distillation methods and loss functions:
TaidDistillerpmetal train Parameters| Parameter | Default | Description |
|---|---|---|
--lora-r | 16 | LoRA rank |
--lora-alpha | 32.0 | LoRA scaling factor (2x rank) |
--batch-size | 1 | Micro-batch size |
--learning-rate | 2e-4 | Learning rate |
--max-seq-len | 0 | Max seq len (0 = auto-detect) |
--epochs | 1 | Number of training epochs |
--max-grad-norm | 1.0 | Gradient clipping |
--quantization | none | QLoRA method (nf4, fp4, int8) |
--gradient-accumulation-steps | 4 | Gradient accumulation steps |
--embedding-lr | None | Separate LR for embeddings |
--no-metal-fused-optimizer | false | Disable Metal fused optimizer |
--lr-schedule | cosine | Schedule type (constant, linear, cosine, cosine_with_restarts, polynomial, wsd) |
--no-gradient-checkpointing | false | Disable gradient checkpointing (enabled by default) |
--gradient-checkpointing-layers | 4 | Number of layers per checkpoint block |
--warmup-steps | 100 | Learning rate warmup steps |
--weight-decay | 0.01 | AdamW weight decay coefficient |
--no-sequence-packing | false | Disable sequence packing |
--cut-cross-entropy | false | Memory-efficient loss (avoids full logit materialization) |
--text-column | — | Custom JSONL column name for training text |
--text-columns | — | Multi-column concat (comma-separated, e.g. thinking,solution) |
--prompt-column | — | Column for prompt (enables SFT loss masking) |
--response-column | — | Column for response (with prompt masking) |
--column-separator | \n\n | Separator for --text-columns |
--config | — | Path to YAML configuration file |
pmetal infer Parameters| Parameter | Default | Description |
|---|---|---|
--temperature | Model default | Sampling temperature |
--top-k | Model default | Top-k sampling |
--top-p | Model default | Nucleus sampling |
--min-p | Model default | Min-p dynamic sampling |
--max-tokens | 256 | Maximum generation length |
--repetition-penalty | 1.0 | Repetition penalty |
--frequency-penalty | 0.0 | Frequency penalty |
--presence-penalty | 0.0 | Presence penalty |
--chat | false | Apply chat template |
--draft-model | — | Gemma 4 MTP assistant checkpoint |
--mtp | false | Enable Qwen3Next/Qwen3.6 exact speculative MTP |
--mtp-model | bundled mtp.* | Optional external Qwen MTP checkpoint from train-mtp |
--mtp-draft-tokens | 3 | Qwen MTP draft tokens per verification step |
--fp8 | false | Use FP8 weights (~2x mem reduction) |
--compiled | false | Use JIT-compiled sampling |
--ane-max-seq-len | 1024 | Max ANE kernel sequence length |
--tools | — | Tool/function definitions file (OpenAI format) |
--system | — | System message |
Defaults are cli, dashboard, trainer, lora, merge, ane and distributed; the rest are pulled in transitively.
| Feature | Default | Crate | Description |
|---|---|---|---|
cli | Yes | — | The pmetal binary and its CLI-only dependencies |
core | Yes* | pmetal-core | Foundation types, configs, traits |
gguf | Yes* | pmetal-gguf | GGUF format support |
metal | Yes* | pmetal-metal | Metal GPU kernels |
hub | Yes* | pmetal-hub | HuggingFace Hub integration |
mlx | Yes* | pmetal-mlx | MLX backend |
models | Yes* | pmetal-models | LLM architectures |
lora | Yes | pmetal-lora | LoRA/QLoRA |
trainer | Yes | pmetal-trainer | Training loops (pulls in data, distill) |
data | Yes* | pmetal-data | Dataset loading (*via cli and trainer) |
distill | Yes* | pmetal-distill | Knowledge distillation (*via trainer) |
merge | Yes | pmetal-merge | Model merging strategies |
distributed | Yes | pmetal-distributed | Distributed training and the cluster subcommand |
ane | Yes | — | Apple Neural Engine |
dashboard | Yes | — | TUI control center |
native-only | No | pmetal-bridge | Bridge-only build with no mlx-rs/mlx-sys |
lora-metal-fused | No | — | ~2x LoRA training speedup via fused Metal kernels |
vocoder | No | pmetal-vocoder | BigVGAN neural vocoder |
mhc | No | pmetal-mhc | Manifold-Constrained Hyper-Connections |
serve | No | pmetal-serve | OpenAI- and Anthropic-compatible inference server |
mcp | No | pmetal-mcp | MCP server (51 tools for Claude Desktop) |
full | No | — | All sub-crate features (not cli, serve or mcp) |
serve and mcp stay out of the default set so library consumers don't inherit axum and rmcp. The prebuilt binary and the Homebrew formula both build with --features serve,mcp, so pmetal serve and pmetal mcp are there if you installed either way. Building yourself, add the flag: cargo install pmetal --features serve,mcp.
# Release build (default features: ANE + Dashboard)
cargo build --release
# Build without ANE
cargo build --release --no-default-features --features dashboard
# Run tests (single-threaded for Metal compatibility)
just test
# Build GUI
cd crates/pmetal-gui && bun install && bun tauri build
# cargo-kani proofs for ring all-reduce and topology
just kani-verify
Licensed under either of MIT or Apache-2.0.
Rust
89.9%
Metal
3.9%
C++
2.7%
Svelte
2.2%