31 repos
srvsngh99/Krill
Fast local LLM inference CLI for Apple Silicon. 1.57x faster than Ollama, 58% less memory.
4
372 commits
vllm-project/vllm-metal
Community maintained hardware plugin for vLLM on Apple Silicon
1,744
560 commits
apocryphx/Apertura
A from-scratch Objective-C/MLX rebuild of Google's Gemma-4 for Apple Silicon — built to be…
2
106 commits
joshuarossi/mlx-bun
MLX inference for TypeScript/Bun on Apple Silicon. Native library and signed…
1
755 commits
carloslfu/slotstream
Run a 105 GB AI model on a 48 GB Mac. Qwen3.8-Flash-Next (125B mixture of experts) streams its…
373
46 commits
nanguoyu/minirun-app
Run Kimi K3 (2.8 T parameters,1.56 TB) on iPhone 16 Pro? Weights stream from your SSD through a…
58
13 commits
jundot/omlx
LLM inference server with continuous batching & SSD caching for Apple Silicon — managed from the…
21,845
2,517 commits
nielspeter/mlx-ts
A TypeScript MLX SDK over Apple's mlx-c via FFI. Runs on Bun, Deno and Node; output matches Apple's…
0
111 commits
walter-grace/expert-sniper
Run MoE models bigger than your RAM — SSD expert streaming on one Mac, or an Expert Network of…
75 commits
PipeNetwork/inkling-mlx
Run Thinking Machines Lab's Inkling (975B-A41B MoE, multimodal) on Apple Silicon with MLX —…
6
19 commits
youssofal/MTPLX
The fastest way to run Qwen 3.8 Flash Next and Qwen 3.8 27B on a Mac: 125 tok/s in OpenCode on an…
2,353
1,184 commits
jonready/vllm-ios
vLLM-style continuous batching for iPhone. Native Swift on MLX, no Python. The fastest multi-agent…
3
14 commits
raullenchai/Rapid-MLX
The fastest local AI engine for Apple Silicon. 4.2x faster than Ollama, 0.08s cached TTFT, 100%…
3,772
2,017 commits
szibis/mlx-flash
Run AI models too large for your Mac's memory — at near-full speed. Intelligent expert caching,…
7
117 commits
Keith-CY/melix
Local-first AI runtime for Apple Silicon with CLI and macOS operator workflows for LoRA training,…
2,930 commits
ddalcu/mlx-serve
Native LLM inference server for Apple Silicon. OpenAI + Anthropic API compatible. No Python.…
1,369
451 commits
sethdford/gemma-realtime
Personalize Gemma 4 and make it real-time on Apple Silicon. Fine-tune on your conversations, serve…
9
7 commits
waybarrios/vllm-mlx
High-performance OpenAI and Anthropic compatible LLM inference server for Apple Silicon. Native…
1,578
597 commits
Greninja9257/LabLLM
A native macOS lab for teaching tiny language models to think — build the architecture, train the…
75
45 commits
Blaizzy/mlx-vlm
MLX-VLM is a package for inference and fine-tuning of Vision Language Models (VLMs) on your Mac…
5,507
1,420 commits
cryptopoly/ChaosEngineAI
Local AI workstation — discover, run, chat, benchmark, and generate images from open-weight models.…
25
578 commits
dthinkr/mlx-lens
Mechanistic interpretability on Apple Silicon: steering vectors, residual capture, and SAE analysis…
1 commits
ivanfioravanti/vlm-bakeoff
VLM bake-off — MLX vs GGUF: identical vision benchmarks across mlx-vlm and llama.cpp on Apple…
24 commits
jjang-ai/jangq
JANG — GGUF for MLX. YOU MUST USE JANG_Q RUNTIME. Adaptive Mixed-Precision Quantization + Runtime…
227
654 commits
Hemeskyo/TimesFM-3-MLX
A pure-MLX port of Google Research's TimesFM-3, a 330M-parameter foundation model for time-series…
6 commits
sebastien-burel/KaozKit
JavaScript LLM agents on the XS engine, embedded in Swift. Snapshots, resident agents, confined…
98 commits
yanun0323/Whallm
DeepSeek-V4-Flash-0731 284B inference in ~25 GB of RAM / Qwen3.8-Next-Flash-FP8 inference in ~18…
101
40 commits
drumih/turbo-fieldfare
Gemma 4 26B-A4B inference in ~2 GB of RAM on any M-series MacBook
6,766
44 commits
mlx-node/mlx-node
No description
157
167 commits
argonautlabsai/argodrive
Layout, balancer and instruments for running mixture-of-experts models from SSDs. Three models, two…
NightMean/OlliteRT
Turn your Android phone into an OpenAI-compatible LLM inference server - Fully local, private and…
351
1,202 commits