library_name: pytorch tags:
AGILLM4.1 is the promoted AGILLM4 mainline evolved from the AGILLM3.5 prototype, and it is larger than AGILLM3/AGILLM3.5. Resumed checkpoints are the source of truth for the exact architecture, with AGILLM4-sized presets available for fresh starts.
The mainline runnable artifact is agillm41.py. The historical implementation file remains agillm35.py for compatibility with existing worker paths, checkpoints, and automation. The helper modules are folded into the single file so the runtime can be cloned, inspected, and launched without restoring the whole AGILLM4 source tree.
The live architecture supports AR, SAT, and NAT objectives/heads. Distributed
inference is AR-first today; monolithic runtime inference supports --mode ar,
--mode sat, and --mode nat.
Live coordinator (Scott's network):
https://join.opentransformers.onlinepublic_join/agillm41_network_host.py starts a signed-lease HTTPS coordinator for people who want to run their own network.
public_join/agillm41_join_worker.py is an outbound-only worker for untrusted joiners. It requests short-lived leases, verifies package hashes, runs a local worker command, and submits results to quarantine rather than exposing SSH or writing directly into the master merge path.
public_join/README.md documents the two intended public paths:
distributed_infer/agillm41_distributed_infer.py is a single-file distributed AR inference harness for the real AGILLM4.1 transformer. It splits contiguous transformer/DiffusionBlock layer ranges across local or HTTP worker stages, using the actual Block implementation and MoE FFNs from the checkpoint config.
Plan layer ranges:
python distributed_infer/agillm41_distributed_infer.py plan \
--agillm41-path ./agillm41.py \
--ckpt /path/to/master.pt \
--dblock-blocks 8
Start a worker for one layer range:
AGILLM41_INFER_TOKEN='change-me' python distributed_infer/agillm41_distributed_infer.py worker \
--agillm41-path ./agillm41.py \
--ckpt /path/to/master.pt \
--start-layer 0 \
--end-layer 12 \
--host 0.0.0.0 \
--port 9100
Run the coordinator:
AGILLM41_INFER_TOKEN='change-me' python distributed_infer/agillm41_distributed_infer.py infer \
--agillm41-path ./agillm41.py \
--ckpt /path/to/master.pt \
--prompt "Hello" \
--max-new 32 \
--cache-mode kv \
--stage https://worker-a.example:9100,0,12 \
--stage local:12:24
Network tensor payloads use a small raw tensor wire format rather than unpickling remote worker responses. Use TLS plus a bearer token for workers exposed beyond localhost. --cache-mode kv is the default and keeps per-session KV state on each worker after the prompt prefill, so decode steps send only the new hidden token through the pipeline. --cache-mode full is kept for comparison/debugging. SAT/NAT distributed decoding is a later phase.
For checkpoint sharing, export inference-slim or split-stage artifacts with the
scripts in agillm4/ops/; full training checkpoints are intentionally not kept
in this code repository.
deepseek-ai/DeepSeek-V4-Proagillm4_floor, agillm4_main, agillm4_biglarge (d=1024, layers=24, heads=16, rank=128)agillm35.py or TOKENIZER_ID=deepseek-ai/DeepSeek-V3.2 ... --agillm3_compat--dblock--async_update_dir; side workers never block the master looppython agillm41.py --help
python agillm41.py status --ckpt /path/to/pretrain_step00051081.pt
python agillm41.py infer --mode ar --ckpt /path/to/pretrain_step00051081.pt --prompt "Hello"
python agillm41.py train \
--preset agillm4_floor \
--resume /path/to/agillm41_master.pt \
--block 1122 \
--batch_size 4 \
--source HuggingFaceFW/fineweb-edu \
--save_dir ckpts \
--dblock \
--dblock_blocks 8 \
--async_update_dir ckpts/side_updates/incoming \
--async_update_every_steps 100
This repository contains code only, not AGILLM checkpoint weights.
DiffusionBlock logs report raw CE-style loss plus the actual EDM-weighted training objective as weighted. The weighted value is the optimization target; the raw value is the sanity-check number to compare with ordinary AR/SAT loss.
The Linux smoke test compiles the single file and completes a one-step synthetic training save. The full AGILLM4.1 continuation run is managed separately by the disaggregated Hetzner worker setup. Legacy agillm35.py, AGILLM35_*, and --agillm35-path names remain supported as compatibility aliases.
55 commits
Python
86.7%
Shell
10.9%
PowerShell
2.5%
library_name: pytorch tags:
AGILLM4.1 is the promoted AGILLM4 mainline evolved from the AGILLM3.5 prototype, and it is larger than AGILLM3/AGILLM3.5. Resumed checkpoints are the source of truth for the exact architecture, with AGILLM4-sized presets available for fresh starts.
The mainline runnable artifact is agillm41.py. The historical implementation file remains agillm35.py for compatibility with existing worker paths, checkpoints, and automation. The helper modules are folded into the single file so the runtime can be cloned, inspected, and launched without restoring the whole AGILLM4 source tree.
The live architecture supports AR, SAT, and NAT objectives/heads. Distributed
inference is AR-first today; monolithic runtime inference supports --mode ar,
--mode sat, and --mode nat.
Live coordinator (Scott's network):
https://join.opentransformers.onlinepublic_join/agillm41_network_host.py starts a signed-lease HTTPS coordinator for people who want to run their own network.
public_join/agillm41_join_worker.py is an outbound-only worker for untrusted joiners. It requests short-lived leases, verifies package hashes, runs a local worker command, and submits results to quarantine rather than exposing SSH or writing directly into the master merge path.
public_join/README.md documents the two intended public paths:
distributed_infer/agillm41_distributed_infer.py is a single-file distributed AR inference harness for the real AGILLM4.1 transformer. It splits contiguous transformer/DiffusionBlock layer ranges across local or HTTP worker stages, using the actual Block implementation and MoE FFNs from the checkpoint config.
Plan layer ranges:
python distributed_infer/agillm41_distributed_infer.py plan \
--agillm41-path ./agillm41.py \
--ckpt /path/to/master.pt \
--dblock-blocks 8
Start a worker for one layer range:
AGILLM41_INFER_TOKEN='change-me' python distributed_infer/agillm41_distributed_infer.py worker \
--agillm41-path ./agillm41.py \
--ckpt /path/to/master.pt \
--start-layer 0 \
--end-layer 12 \
--host 0.0.0.0 \
--port 9100
Run the coordinator:
AGILLM41_INFER_TOKEN='change-me' python distributed_infer/agillm41_distributed_infer.py infer \
--agillm41-path ./agillm41.py \
--ckpt /path/to/master.pt \
--prompt "Hello" \
--max-new 32 \
--cache-mode kv \
--stage https://worker-a.example:9100,0,12 \
--stage local:12:24
Network tensor payloads use a small raw tensor wire format rather than unpickling remote worker responses. Use TLS plus a bearer token for workers exposed beyond localhost. --cache-mode kv is the default and keeps per-session KV state on each worker after the prompt prefill, so decode steps send only the new hidden token through the pipeline. --cache-mode full is kept for comparison/debugging. SAT/NAT distributed decoding is a later phase.
For checkpoint sharing, export inference-slim or split-stage artifacts with the
scripts in agillm4/ops/; full training checkpoints are intentionally not kept
in this code repository.
deepseek-ai/DeepSeek-V4-Proagillm4_floor, agillm4_main, agillm4_biglarge (d=1024, layers=24, heads=16, rank=128)agillm35.py or TOKENIZER_ID=deepseek-ai/DeepSeek-V3.2 ... --agillm3_compat--dblock--async_update_dir; side workers never block the master looppython agillm41.py --help
python agillm41.py status --ckpt /path/to/pretrain_step00051081.pt
python agillm41.py infer --mode ar --ckpt /path/to/pretrain_step00051081.pt --prompt "Hello"
python agillm41.py train \
--preset agillm4_floor \
--resume /path/to/agillm41_master.pt \
--block 1122 \
--batch_size 4 \
--source HuggingFaceFW/fineweb-edu \
--save_dir ckpts \
--dblock \
--dblock_blocks 8 \
--async_update_dir ckpts/side_updates/incoming \
--async_update_every_steps 100
This repository contains code only, not AGILLM checkpoint weights.
DiffusionBlock logs report raw CE-style loss plus the actual EDM-weighted training objective as weighted. The weighted value is the optimization target; the raw value is the sanity-check number to compare with ordinary AR/SAT loss.
The Linux smoke test compiles the single file and completes a one-step synthetic training save. The full AGILLM4.1 continuation run is managed separately by the disaggregated Hetzner worker setup. Legacy agillm35.py, AGILLM35_*, and --agillm35-path names remain supported as compatibility aliases.
55 commits
Python
86.7%
Shell
10.9%
PowerShell
2.5%