WeeLM is a compact language-model runtime built in Northern Ireland for .NET. It runs conversational transformer weights entirely in managed C#, without Python, ONNX, or LibTorch. The default 360M model pack is designed for responsive local chat.
dotnet run --project src/WeeLM.Cli -c Release -- fetch
dotnet run --project src/WeeLM.Cli -c Release -- chat
Inside chat, use /clear, /system, /stats, /help, or /exit. Model files are downloaded into the gitignored models/ directory and are not redistributed with WeeLM.
Use the larger unquantized WeeLM 1.7B quality tier:
dotnet run --project src/WeeLM.Cli -c Release -- fetch --model wee-1.7b
dotnet run --project src/WeeLM.Cli -c Release -- chat --quality quality
For substantially stronger general reasoning and coding, add the text-only Qwen3.5-4B tier:
dotnet run --project src/WeeLM.Cli -c Release -- fetch --model qwen3.5-4b
dotnet run --project src/WeeLM.Cli -c Release -- quantize --model models/Qwen3.5-4B
dotnet run --project src/WeeLM.Cli -c Release -- chat --quality best --max-tokens 512
Qwen3.5 is optional and does not replace the responsive 360M or 1.7B tiers. Its official checkpoint is roughly 9.3 GB because both shards also contain vision weights; WeeLM downloads the unchanged shards but loads only language tensors. A 64-bit process and ample memory are required, and CPU decoding is materially slower than the smaller tiers.
The Qwen tier uses native reasoning internally. WeeLM suppresses its <think>
section and stores only the final answer. If the output budget ends before a
final answer appears, increase --max-tokens.
WeeLM keeps the official BF16 weights unchanged and can create an optional
groupwise INT8 sidecar for faster, lower-memory CPU inference. --weights auto
uses a valid sidecar when present; use --weights bf16 for the exact original
path or --weights int8 to require quantized weights. The sidecar is validated
against the source checkpoint and never modifies its Safetensors files. The
lazy, copy-on-write KV cache stores keys and values in BF16.
The native 8,192-token mode remains the default. Enable the experimental 16K YaRN profile at load time:
dotnet run --project src/WeeLM.Cli -c Release -- chat --context 16384
--rope-scaling auto is the default: it preserves standard RoPE at the trained context length and selects YaRN only beyond it. --rope-scaling standard|yarn can be used to make the choice explicit. Extending positional encoding does not guarantee 16K retrieval quality by itself, so evaluate it on your own long-context tasks.
Qwen3.5 also defaults to 8,192 tokens but uses native partial RoPE. Explicit Qwen contexts up to 32,768 tokens are supported; YaRN and larger Qwen windows are rejected in this CPU-first release.
Long chats use a bounded rolling context rather than silently discarding everything old. WeeLM protects the initial system prompt as an attention sink, keeps the most recent turns, scores older turns for durable facts and constraints, and folds evicted turns into a structured memory with these sections:
/stats reports the active context and BF16 KV allocation. When compaction occurs, chat reports how many turns were summarized and how many important older turns survived eviction.
Chat uses confidence-aware min-p sampling through the balanced profile. Switch behavior at runtime:
/profile precise
/profile creative
/mode deliberate
deliberate prefills once, generates three independent replies, and selects the strongest completed, non-repetitive candidate. It is slower and displays only the winner. fast streams one reply immediately.
Compare the modes over WeeLM's built-in general-quality prompt suite:
dotnet run --project src/WeeLM.Cli -c Release -- evaluate-quality --max-tokens 48
Measure cold loading, prompt ingestion, decoding, allocations, and memory:
dotnet run --project src/WeeLM.Cli -c Release -- benchmark --quality quality --prompt-tokens 128 --decode-tokens 32
The benchmark also reports precision, worker count, Qwen prefill chunk size,
and average CPU cores used. Tune managed execution with --workers and
--prefill-chunk; Qwen defaults to 32-token layer-wise prefill chunks.
Chat also includes safe local calculator and clock tools. Try calculate sqrt(81) + 7 or current UTC time; toggle them with /tools on|off.
WeeLM can train and load an output-projection LoRA entirely in C#. Input can be a UTF-8 text file or JSONL with one {"text":"..."} document per line:
dotnet run --project src/WeeLM.Cli -c Release -- train-lora `
--model models/WeeLM-360M `
--data training.jsonl `
--output runs/belfast.weelora `
--context 16384 `
--sequence-length 16384 `
--rank 8 `
--max-training-tokens 2000
dotnet run --project src/WeeLM.Cli -c Release -- chat `
--model models/WeeLM-360M `
--context 16384 `
--adapter runs/belfast.weelora
The backbone remains frozen while the full transformer produces context-dependent hidden states and AdamW trains the low-rank language-model-head delta. This is suitable for vocabulary, domain, and response-style calibration. It is not attention-projection LongLoRA: training Q/K/V adapters would require full transformer backpropagation and substantially more compute and memory.
Qwen3.5 vision/video, MTP speculative decoding and LoRA are not enabled. Attention-projection LoRA and an Intel Arc compute backend are future milestones.
C#
92.7%
HTML
4.2%
CSS
3.0%
WeeLM is a compact language-model runtime built in Northern Ireland for .NET. It runs conversational transformer weights entirely in managed C#, without Python, ONNX, or LibTorch. The default 360M model pack is designed for responsive local chat.
dotnet run --project src/WeeLM.Cli -c Release -- fetch
dotnet run --project src/WeeLM.Cli -c Release -- chat
Inside chat, use /clear, /system, /stats, /help, or /exit. Model files are downloaded into the gitignored models/ directory and are not redistributed with WeeLM.
Use the larger unquantized WeeLM 1.7B quality tier:
dotnet run --project src/WeeLM.Cli -c Release -- fetch --model wee-1.7b
dotnet run --project src/WeeLM.Cli -c Release -- chat --quality quality
For substantially stronger general reasoning and coding, add the text-only Qwen3.5-4B tier:
dotnet run --project src/WeeLM.Cli -c Release -- fetch --model qwen3.5-4b
dotnet run --project src/WeeLM.Cli -c Release -- quantize --model models/Qwen3.5-4B
dotnet run --project src/WeeLM.Cli -c Release -- chat --quality best --max-tokens 512
Qwen3.5 is optional and does not replace the responsive 360M or 1.7B tiers. Its official checkpoint is roughly 9.3 GB because both shards also contain vision weights; WeeLM downloads the unchanged shards but loads only language tensors. A 64-bit process and ample memory are required, and CPU decoding is materially slower than the smaller tiers.
The Qwen tier uses native reasoning internally. WeeLM suppresses its <think>
section and stores only the final answer. If the output budget ends before a
final answer appears, increase --max-tokens.
WeeLM keeps the official BF16 weights unchanged and can create an optional
groupwise INT8 sidecar for faster, lower-memory CPU inference. --weights auto
uses a valid sidecar when present; use --weights bf16 for the exact original
path or --weights int8 to require quantized weights. The sidecar is validated
against the source checkpoint and never modifies its Safetensors files. The
lazy, copy-on-write KV cache stores keys and values in BF16.
The native 8,192-token mode remains the default. Enable the experimental 16K YaRN profile at load time:
dotnet run --project src/WeeLM.Cli -c Release -- chat --context 16384
--rope-scaling auto is the default: it preserves standard RoPE at the trained context length and selects YaRN only beyond it. --rope-scaling standard|yarn can be used to make the choice explicit. Extending positional encoding does not guarantee 16K retrieval quality by itself, so evaluate it on your own long-context tasks.
Qwen3.5 also defaults to 8,192 tokens but uses native partial RoPE. Explicit Qwen contexts up to 32,768 tokens are supported; YaRN and larger Qwen windows are rejected in this CPU-first release.
Long chats use a bounded rolling context rather than silently discarding everything old. WeeLM protects the initial system prompt as an attention sink, keeps the most recent turns, scores older turns for durable facts and constraints, and folds evicted turns into a structured memory with these sections:
/stats reports the active context and BF16 KV allocation. When compaction occurs, chat reports how many turns were summarized and how many important older turns survived eviction.
Chat uses confidence-aware min-p sampling through the balanced profile. Switch behavior at runtime:
/profile precise
/profile creative
/mode deliberate
deliberate prefills once, generates three independent replies, and selects the strongest completed, non-repetitive candidate. It is slower and displays only the winner. fast streams one reply immediately.
Compare the modes over WeeLM's built-in general-quality prompt suite:
dotnet run --project src/WeeLM.Cli -c Release -- evaluate-quality --max-tokens 48
Measure cold loading, prompt ingestion, decoding, allocations, and memory:
dotnet run --project src/WeeLM.Cli -c Release -- benchmark --quality quality --prompt-tokens 128 --decode-tokens 32
The benchmark also reports precision, worker count, Qwen prefill chunk size,
and average CPU cores used. Tune managed execution with --workers and
--prefill-chunk; Qwen defaults to 32-token layer-wise prefill chunks.
Chat also includes safe local calculator and clock tools. Try calculate sqrt(81) + 7 or current UTC time; toggle them with /tools on|off.
WeeLM can train and load an output-projection LoRA entirely in C#. Input can be a UTF-8 text file or JSONL with one {"text":"..."} document per line:
dotnet run --project src/WeeLM.Cli -c Release -- train-lora `
--model models/WeeLM-360M `
--data training.jsonl `
--output runs/belfast.weelora `
--context 16384 `
--sequence-length 16384 `
--rank 8 `
--max-training-tokens 2000
dotnet run --project src/WeeLM.Cli -c Release -- chat `
--model models/WeeLM-360M `
--context 16384 `
--adapter runs/belfast.weelora
The backbone remains frozen while the full transformer produces context-dependent hidden states and AdamW trains the low-rank language-model-head delta. This is suitable for vocabulary, domain, and response-style calibration. It is not attention-projection LongLoRA: training Q/K/V adapters would require full transformer backpropagation and substantially more compute and memory.
Qwen3.5 vision/video, MTP speculative decoding and LoRA are not enabled. Attention-projection LoRA and an Intel Arc compute backend are future milestones.
C#
92.7%
HTML
4.2%
CSS
3.0%