BenjaminEdgar/FreeLM

C#

0

1 commits

updated Jul 19, 2026

See the code

README

WeeLM

WeeLM is a compact language-model runtime built in Northern Ireland for .NET. It runs conversational transformer weights entirely in managed C#, without Python, ONNX, or LibTorch. The default 360M model pack is designed for responsive local chat.

Quick start

dotnet run --project src/WeeLM.Cli -c Release -- fetch
dotnet run --project src/WeeLM.Cli -c Release -- chat

Inside chat, use /clear, /system, /stats, /help, or /exit. Model files are downloaded into the gitignored models/ directory and are not redistributed with WeeLM.

Use the larger unquantized WeeLM 1.7B quality tier:

dotnet run --project src/WeeLM.Cli -c Release -- fetch --model wee-1.7b
dotnet run --project src/WeeLM.Cli -c Release -- chat --quality quality

For substantially stronger general reasoning and coding, add the text-only Qwen3.5-4B tier:

dotnet run --project src/WeeLM.Cli -c Release -- fetch --model qwen3.5-4b
dotnet run --project src/WeeLM.Cli -c Release -- quantize --model models/Qwen3.5-4B
dotnet run --project src/WeeLM.Cli -c Release -- chat --quality best --max-tokens 512

Qwen3.5 is optional and does not replace the responsive 360M or 1.7B tiers. Its official checkpoint is roughly 9.3 GB because both shards also contain vision weights; WeeLM downloads the unchanged shards but loads only language tensors. A 64-bit process and ample memory are required, and CPU decoding is materially slower than the smaller tiers.

The Qwen tier uses native reasoning internally. WeeLM suppresses its <think> section and stores only the final answer. If the output budget ends before a final answer appears, increase --max-tokens.

WeeLM keeps the official BF16 weights unchanged and can create an optional groupwise INT8 sidecar for faster, lower-memory CPU inference. --weights auto uses a valid sidecar when present; use --weights bf16 for the exact original path or --weights int8 to require quantized weights. The sidecar is validated against the source checkpoint and never modifies its Safetensors files. The lazy, copy-on-write KV cache stores keys and values in BF16.

Long context and memory

The native 8,192-token mode remains the default. Enable the experimental 16K YaRN profile at load time:

dotnet run --project src/WeeLM.Cli -c Release -- chat --context 16384

--rope-scaling auto is the default: it preserves standard RoPE at the trained context length and selects YaRN only beyond it. --rope-scaling standard|yarn can be used to make the choice explicit. Extending positional encoding does not guarantee 16K retrieval quality by itself, so evaluate it on your own long-context tasks.

Qwen3.5 also defaults to 8,192 tokens but uses native partial RoPE. Explicit Qwen contexts up to 32,768 tokens are supported; YaRN and larger Qwen windows are rejected in this CPU-first release.

Long chats use a bounded rolling context rather than silently discarding everything old. WeeLM protects the initial system prompt as an attention sink, keeps the most recent turns, scores older turns for durable facts and constraints, and folds evicted turns into a structured memory with these sections:

  • User facts
  • Decisions and constraints
  • Open tasks
  • Relevant context

/stats reports the active context and BF16 KV allocation. When compaction occurs, chat reports how many turns were summarized and how many important older turns survived eviction.

Chat uses confidence-aware min-p sampling through the balanced profile. Switch behavior at runtime:

/profile precise
/profile creative
/mode deliberate

deliberate prefills once, generates three independent replies, and selects the strongest completed, non-repetitive candidate. It is slower and displays only the winner. fast streams one reply immediately.

Compare the modes over WeeLM's built-in general-quality prompt suite:

dotnet run --project src/WeeLM.Cli -c Release -- evaluate-quality --max-tokens 48

Measure cold loading, prompt ingestion, decoding, allocations, and memory:

dotnet run --project src/WeeLM.Cli -c Release -- benchmark --quality quality --prompt-tokens 128 --decode-tokens 32

The benchmark also reports precision, worker count, Qwen prefill chunk size, and average CPU cores used. Tune managed execution with --workers and --prefill-chunk; Qwen defaults to 32-token layer-wise prefill chunks.

Chat also includes safe local calculator and clock tools. Try calculate sqrt(81) + 7 or current UTC time; toggle them with /tools on|off.

Long-context LoRA

WeeLM can train and load an output-projection LoRA entirely in C#. Input can be a UTF-8 text file or JSONL with one {"text":"..."} document per line:

dotnet run --project src/WeeLM.Cli -c Release -- train-lora `
  --model models/WeeLM-360M `
  --data training.jsonl `
  --output runs/belfast.weelora `
  --context 16384 `
  --sequence-length 16384 `
  --rank 8 `
  --max-training-tokens 2000

dotnet run --project src/WeeLM.Cli -c Release -- chat `
  --model models/WeeLM-360M `
  --context 16384 `
  --adapter runs/belfast.weelora

The backbone remains frozen while the full transformer produces context-dependent hidden states and AdamW trains the low-rank language-model-head delta. This is suitable for vocabulary, domain, and response-style calibration. It is not attention-projection LongLoRA: training Q/K/V adapters would require full transformer backpropagation and substantially more compute and memory.

Current scope

  • Native F32/BF16 Safetensors, optional groupwise INT8 sidecars, and Hugging Face byte-level BPE loading
  • Single-file and indexed multi-shard Safetensors loading
  • Llama decoder math: RMSNorm, RoPE, grouped-query attention, SwiGLU and tied embeddings
  • Qwen3.5 text inference: gated DeltaNet, gated grouped-query attention, partial RoPE and FP32 recurrent state
  • ChatML formatting and persistent multi-turn KV cache
  • Optional 16K YaRN RoPE, rolling attention-sink context, importance eviction and structured memory
  • BF16 paged KV storage with copy-on-write forks
  • Portable output-projection LoRA loading and long-sequence C# training
  • Greedy, temperature, top-k/top-p, repetition, frequency and presence sampling controls
  • Managed AVX2 BF16/INT8 kernels, persistent inference workers, and chunked Qwen prefill

Qwen3.5 vision/video, MTP speculative decoding and LoRA are not enabled. Attention-projection LoRA and an Intel Arc compute backend are future milestones.

BenjaminEdgar/FreeLM

C#

0

1 commits

updated Jul 19, 2026

See the code

README

WeeLM

WeeLM is a compact language-model runtime built in Northern Ireland for .NET. It runs conversational transformer weights entirely in managed C#, without Python, ONNX, or LibTorch. The default 360M model pack is designed for responsive local chat.

Quick start

dotnet run --project src/WeeLM.Cli -c Release -- fetch
dotnet run --project src/WeeLM.Cli -c Release -- chat

Inside chat, use /clear, /system, /stats, /help, or /exit. Model files are downloaded into the gitignored models/ directory and are not redistributed with WeeLM.

Use the larger unquantized WeeLM 1.7B quality tier:

dotnet run --project src/WeeLM.Cli -c Release -- fetch --model wee-1.7b
dotnet run --project src/WeeLM.Cli -c Release -- chat --quality quality

For substantially stronger general reasoning and coding, add the text-only Qwen3.5-4B tier:

dotnet run --project src/WeeLM.Cli -c Release -- fetch --model qwen3.5-4b
dotnet run --project src/WeeLM.Cli -c Release -- quantize --model models/Qwen3.5-4B
dotnet run --project src/WeeLM.Cli -c Release -- chat --quality best --max-tokens 512

Qwen3.5 is optional and does not replace the responsive 360M or 1.7B tiers. Its official checkpoint is roughly 9.3 GB because both shards also contain vision weights; WeeLM downloads the unchanged shards but loads only language tensors. A 64-bit process and ample memory are required, and CPU decoding is materially slower than the smaller tiers.

The Qwen tier uses native reasoning internally. WeeLM suppresses its <think> section and stores only the final answer. If the output budget ends before a final answer appears, increase --max-tokens.

WeeLM keeps the official BF16 weights unchanged and can create an optional groupwise INT8 sidecar for faster, lower-memory CPU inference. --weights auto uses a valid sidecar when present; use --weights bf16 for the exact original path or --weights int8 to require quantized weights. The sidecar is validated against the source checkpoint and never modifies its Safetensors files. The lazy, copy-on-write KV cache stores keys and values in BF16.

Long context and memory

The native 8,192-token mode remains the default. Enable the experimental 16K YaRN profile at load time:

dotnet run --project src/WeeLM.Cli -c Release -- chat --context 16384

--rope-scaling auto is the default: it preserves standard RoPE at the trained context length and selects YaRN only beyond it. --rope-scaling standard|yarn can be used to make the choice explicit. Extending positional encoding does not guarantee 16K retrieval quality by itself, so evaluate it on your own long-context tasks.

Qwen3.5 also defaults to 8,192 tokens but uses native partial RoPE. Explicit Qwen contexts up to 32,768 tokens are supported; YaRN and larger Qwen windows are rejected in this CPU-first release.

Long chats use a bounded rolling context rather than silently discarding everything old. WeeLM protects the initial system prompt as an attention sink, keeps the most recent turns, scores older turns for durable facts and constraints, and folds evicted turns into a structured memory with these sections:

  • User facts
  • Decisions and constraints
  • Open tasks
  • Relevant context

/stats reports the active context and BF16 KV allocation. When compaction occurs, chat reports how many turns were summarized and how many important older turns survived eviction.

Chat uses confidence-aware min-p sampling through the balanced profile. Switch behavior at runtime:

/profile precise
/profile creative
/mode deliberate

deliberate prefills once, generates three independent replies, and selects the strongest completed, non-repetitive candidate. It is slower and displays only the winner. fast streams one reply immediately.

Compare the modes over WeeLM's built-in general-quality prompt suite:

dotnet run --project src/WeeLM.Cli -c Release -- evaluate-quality --max-tokens 48

Measure cold loading, prompt ingestion, decoding, allocations, and memory:

dotnet run --project src/WeeLM.Cli -c Release -- benchmark --quality quality --prompt-tokens 128 --decode-tokens 32

The benchmark also reports precision, worker count, Qwen prefill chunk size, and average CPU cores used. Tune managed execution with --workers and --prefill-chunk; Qwen defaults to 32-token layer-wise prefill chunks.

Chat also includes safe local calculator and clock tools. Try calculate sqrt(81) + 7 or current UTC time; toggle them with /tools on|off.

Long-context LoRA

WeeLM can train and load an output-projection LoRA entirely in C#. Input can be a UTF-8 text file or JSONL with one {"text":"..."} document per line:

dotnet run --project src/WeeLM.Cli -c Release -- train-lora `
  --model models/WeeLM-360M `
  --data training.jsonl `
  --output runs/belfast.weelora `
  --context 16384 `
  --sequence-length 16384 `
  --rank 8 `
  --max-training-tokens 2000

dotnet run --project src/WeeLM.Cli -c Release -- chat `
  --model models/WeeLM-360M `
  --context 16384 `
  --adapter runs/belfast.weelora

The backbone remains frozen while the full transformer produces context-dependent hidden states and AdamW trains the low-rank language-model-head delta. This is suitable for vocabulary, domain, and response-style calibration. It is not attention-projection LongLoRA: training Q/K/V adapters would require full transformer backpropagation and substantially more compute and memory.

Current scope

  • Native F32/BF16 Safetensors, optional groupwise INT8 sidecars, and Hugging Face byte-level BPE loading
  • Single-file and indexed multi-shard Safetensors loading
  • Llama decoder math: RMSNorm, RoPE, grouped-query attention, SwiGLU and tied embeddings
  • Qwen3.5 text inference: gated DeltaNet, gated grouped-query attention, partial RoPE and FP32 recurrent state
  • ChatML formatting and persistent multi-turn KV cache
  • Optional 16K YaRN RoPE, rolling attention-sink context, importance eviction and structured memory
  • BF16 paged KV storage with copy-on-write forks
  • Portable output-projection LoRA loading and long-sequence C# training
  • Greedy, temperature, top-k/top-p, repetition, frequency and presence sampling controls
  • Managed AVX2 BF16/INT8 kernels, persistent inference workers, and chunked Qwen prefill

Qwen3.5 vision/video, MTP speculative decoding and LoRA are not enabled. Attention-projection LoRA and an Intel Arc compute backend are future milestones.

Languages

C#

92.7%

HTML

4.2%

CSS

3.0%