An industrial-grade, off-policy synthetic data distillation engine for LLMs.
Turn YAML curriculum blueprints into gold-standard task benchmarks, and mill verifiable reasoning traces and multi-turn agent trajectories from any model with an OpenAI-compatible endpoint.
user <> assistant)sftmill is a high-throughput, modular data engine designed to generate custom Supervised Fine-Tuning (SFT) datasets through off-policy model distillation. Instead of relying on expensive human annotations or brittle web scrapers, sftmill allows researchers and engineers to define high-level educational curricula in declarative YAML and distill rich training corpora directly from any frontier or open-weights teacher model (such as DeepSeek-R1, Qwen 2.5, OpenAI o1/o3-mini/GPT-4o, or locally served vLLM/Ollama instances).
sftmill natively produces:
user <> assistant conversations capturing both latent internal chain-of-thought (reasoning_content) and polished conversational answers (content).verify: python3 check.py) and exact post-edit filesystem states (expect_files).| Feature | Typical Synthetic Data Scripts | sftmill Data Mill |
|---|---|---|
| Pipeline Model | One-off messy scripts with prompt leaks | Decoupled 2-stage architecture: Spec ➔ Benchmark ➔ Rollout |
| Reasoning Distillation | Flattened text or stripped thinking | Native Hybrid Reasoning: preserves reasoning_content across turns |
| Tool Execution | Mocked regex or unexecuted hallucinations | Real hermetic sandbox: isolated subshells, timeouts, path protection |
| Verification | Blind trust in LLM output | Deterministic grading (exact, contains, check.py, gold file trees) |
| Tool Format Lock-in | Hardcoded to one provider | Multi-harness translation: OpenAI, Claude Code XML, Alice, custom |
| Evaluation Leakage | Contaminated test splits | Group Hashing: sibling harnesses never cross train/val splits |
| Production Scale | Memory-heavy arrays | Auto-partitioned sharded JSONL (part-000000.jsonl) with live streaming |
┌────────────────────────────────────────────────────────┐
│ YAML Curriculum Blueprint │
│ Define skills, counts, templates, and match rules │
└───────────────────────────┬────────────────────────────┘
│
▼ Stage 1: sftmill tasks
┌────────────────────────────────────────────────────────┐
│ Task Benchmark Synthesis │
│ Synthesis teacher invents diverse questions, tests, │
│ seed files, and multi-turn user prompts │
└───────────────────────────┬────────────────────────────┘
│ Output: tasks.jsonl
▼
┌────────────────────────────────────────────────────────┐
│ Stage 2: sftmill generate │
│ Rollout & Milling │
└─────────────┬────────────────────────────┬─────────────┘
│ │
[kind: trace] [kind: trajectory]
Multi-Turn CoT Sandboxed Tools
User-Assistant Turns Execution Loop
│ │
▼ ▼
Reasoning + Answer Capture Workspace Actions:
(reasoning_content + content) list_dir, read_file,
edit_file, bash, python
│ │
└─────────────┬──────────────┘
│
▼
┌────────────────────────────────────────────────────────┐
│ Acceptance & Grading │
│ - Exact answer match / substring contains │
│ - python3 check.py passes in workspace │
│ - expect_files tree matches gold state │
│ - Substantive prose verification (anti-stub filter) │
└───────────────────────────┬────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────┐
│ Production Sharded SFT Dataset │
│ data/part-000000.jsonl ... part-00000N.jsonl │
│ Ready for Axolotl, TRL, or Unsloth │
└────────────────────────────────────────────────────────┘
sftmill is lightweight and only requires Python 3.10+ and pyyaml.
# Clone the repository
git clone https://github.com/jackjusko/sftmill.git
cd sftmill
# Install in editable mode
pip install -e .
# Or using uv (recommended for ultra-fast setup)
uv pip install -e ".[dev]"
Use a teacher model to expand a curriculum specification into a concrete tasks.jsonl benchmark:
sftmill tasks \
--curriculum configs/curriculum/code_agent_test.yaml \
--out tasks.jsonl \
--base-url https://api.openai.com/v1 \
--model gpt-4o \
--api-key $OPENAI_API_KEY \
--jobs 4
Run generation rollouts over the synthesized tasks. Trajectories run tools in sandboxes and traces distill hybrid reasoning:
sftmill generate \
--tasks tasks.jsonl \
--out sft_output/ \
--base-url https://api.deepseek.com/v1 \
--model deepseek-reasoner \
--api-key $DEEPSEEK_API_KEY \
--shard-size 1000 \
--jobs 4
Compatible with any OpenAI-style backend: Works out of the box with vLLM, Ollama, DeepSeek, Groq, OpenAI, SGLang, or LiteLLM. Just configure
--base-urland--model.
user <> assistant)A core strength of sftmill is distilling conversational datasets where every assistant response contains both an internal chain-of-thought scratchpad (reasoning_content) and user-facing conversational prose (content).
When training smaller models (3B, 7B, 14B, 32B), teaching the model to formulate a structured thought process before answering follow-up queries or clarifying constraints significantly elevates downstream multi-turn performance.
Set kind: trace, match: open, and user_turns: 2 (or 3):
name: code_instruct
categories:
- id: clarify_requirements
count: 100
kind: trace
match: open
user_turns: 3
template: |
JSON key: turns, a list of exactly 3 user messages and no assistant messages.
The first turn is a programming task missing one critical technical constraint.
The second turn answers with that missing constraint.
The third turn asks for the final complete implementation and unit tests.
No tools, no files, no answer key.
example:
turns:
- "We need a rate limiter in Python for our public API endpoints."
- "Use Redis token-bucket algorithm, limiting to 60 requests per minute per IP."
- "Great. Please provide the complete production implementation with async redis-py."
sftmill Generates Multi-Turn Tracessftmill feeds turn[0] to the teacher model.delta.reasoning_content) is captured separately from its visible output (delta.content).user ➔ assistant with reasoning and text) is passed back for turn[1].sftmill.generate.filters.prose_ok automatically rejects one-word responses, lazy apologies, or stubs.The resulting JSONL dataset row is ready for standard fine-tuning pipelines:
{
"kind": "trace",
"task_id": "clarify_requirements-000042",
"messages": [
{
"role": "user",
"content": "We need a rate limiter in Python for our public API endpoints."
},
{
"role": "assistant",
"reasoning_content": "The user wants an API rate limiter in Python, but has not specified the concurrency model (FastAPI, Flask, Django), the storage backend (in-memory, Redis, Memcached), or the algorithm (token bucket, leaky bucket, sliding window). I should identify the missing details and ask for clarification.",
"content": "I would be happy to help implement a rate limiter. To tailor it to your architecture, could you clarify:\n1. Which backend or cache you plan to use (e.g., Redis, in-memory)?\n2. Which algorithm you prefer (e.g., token bucket, sliding window log)?\n3. What are the rate limits and identifier (e.g., 60 req/min per IP)?"
},
{
"role": "user",
"content": "Use Redis token-bucket algorithm, limiting to 60 requests per minute per IP."
},
{
"role": "assistant",
"reasoning_content": "The parameters are now clear: Redis backend, token bucket algorithm, 60 req/min per IP. The user hasn't asked for code yet, so I should confirm understanding and ask if they prefer async (redis-py async) or synchronous execution.",
"content": "Understood: Redis-backed token bucket allowing 60 requests per minute per IP. Do you need this implemented for an asynchronous framework (like FastAPI with `redis-py` async) or a WSGI application (like Flask)?"
},
{
"role": "user",
"content": "Great. Please provide the complete production implementation with async redis-py."
},
{
"role": "assistant",
"reasoning_content": "Requirements finalized: Async implementation using redis.asyncio, token bucket with capacity=60 and refill_rate=1 token/sec, atomic Lua script for concurrency safety, and a reusable dependency/middleware. Writing clean, robust code.",
"content": "Here is a complete, production-ready asynchronous token bucket rate limiter using `redis-py` with an atomic Lua script:\n\n```python\nimport time\nfrom redis import asyncio as aioredis\n\nTOKEN_BUCKET_LUA = \"\"\"\nlocal key = KEYS[1]\nlocal capacity = tonumber(ARGV[1])\nlocal refill_rate = tonumber(ARGV[2])\nlocal now = tonumber(ARGV[3])\nlocal requested = tonumber(ARGV[4])\n...\n```"
}
]
}
📖 For complete fine-tuning recipes with Axolotl, Unsloth, and TRL, read the Multi-Turn Hybrid-Reasoning Guide.
For agent training (kind: trajectory), sftmill executes real tool calls inside isolated temporary directory sandboxes.
read_file(path, offset, limit): Read file lines with pagination.write_file(path, content): Create or overwrite workspace files.edit_file(path, old, new): Precise single-occurrence string replacement.search(pattern, path): Literal recursive search across the workspace (path:line:text).list_dir(path): List directory contents.bash(command): Execute subshell commands with sanitized environment variables.python(code): Execute Python code with a 5-second timeout.verify: python3 check.py): An automated test script is executed after the agent declares completion. If check.py fails or exits non-zero, the trajectory is discarded.expect_files): The final workspace filesystem tree is compared against ground-truth files.require_observation: PASSED): Trajectories are rejected unless the agent witnessed a successful test run during its rollout... traversals are blocked).📖 Read the Agent Trajectories & Tool Use Guide for details.
A curriculum YAML controls how sftmill tasks synthesizes problems:
name: code_agent_benchmark
categories:
- id: bug_fix_caching
count: 25 # Number of rows to synthesize
kind: trajectory # trace | trajectory
match: exact # exact | contains | open
max_steps: 16 # Tool step budget per episode
toolset: workspace # workspace | custom
require_observation: PASSED # Substring that must appear in tool output
verify: python3 check.py # Command run in workspace to grade trajectory
harnesses: [openai, claude_code] # Envelopes to generate
template: |
JSON keys: question, answer, files, expect_files.
files includes app.py with a subtle bug, and check.py with assertions.
The check prints PASSED only when app.py is fixed.
expect_files contains the corrected app.py.
answer is the single token ok.
exact: Graded answer must match the gold string exactly.contains: Expected token or needle must appear in the assistant's final response or reasoning.open: Open-ended conversational prose (requires user_turns: 2 or 3 and validates multi-sentence responses).📖 See the full Curriculum Specification Reference.
The teacher generates canonical actions, and sftmill automatically translates them into platform-specific envelopes:
openai: Native function-calling JSON ({"name": "...", "arguments": {...}}).claude_code: Anthropic XML blocks (<tool_call><tool_name>...</tool_name></tool_call>).alice: Jinja-delimited native envelopes (<tool>\n...\n</tool>).novel: Out-of-distribution evaluation envelopes ([[tool]] ... [[/tool]]).When a problem is expanded across multiple harnesses (e.g. openai and claude_code), both share a group_id. sftmill includes a SHA-256 partitioner ensuring all variants of a problem always land on the same side of the train/val split:
from sftmill.dataset import iter_jsonl_dataset, split_by_group
# Load and split
all_rows = list(iter_jsonl_dataset("sft_output/"))
train_data, val_data = split_by_group(all_rows, val_fraction=0.1, seed=42)
print(f"Train rows: {len(train_data)} | Val rows: {len(val_data)}")
sftmill tasksSynthesize task definitions from a curriculum YAML file.
sftmill tasks \
--curriculum <path/to/curriculum.yaml> \
--out <path/to/tasks.jsonl> \
--base-url <url> \
--model <model_name> \
[--api-key <key>] \
[--temperature 0.7] \
[--batch-size 5] \
[--jobs 3] \
[--progress-interval 30.0]
| Flag | Default | Description |
|---|---|---|
--curriculum | required | Path to curriculum YAML spec. |
--out | required | Destination path for synthesized tasks JSONL. |
--base-url | required | Base URL of the OpenAI-compatible endpoint. |
--model | required | Teacher model identifier. |
--api-key | None | API key (or omit for local models). |
--jobs | 3 | Parallel categories in flight. |
--batch-size | 5 | Tasks requested per synthesis prompt. |
sftmill generateExecute task rollouts to produce final SFT training data.
sftmill generate \
--tasks <path/to/tasks.jsonl> \
--out <directory_or_file> \
--base-url <url> \
--model <model_name> \
[--api-key <key>] \
[--kind trace|trajectory|both] \
[--temperature 1.0] \
[--max-steps 4] \
[--jobs 3] \
[--shard-size 1000] \
[--progress-interval 30.0]
| Flag | Default | Description |
|---|---|---|
--tasks | required | Path to tasks JSONL file. |
--out | required | Output path (file .jsonl or directory for auto-sharding). |
--kind | both | Filter by trace, trajectory, or both. |
--max-steps | 4 | Maximum tool-call rounds for trajectories. |
--jobs | 3 | Concurrent worker threads. |
--shard-size | 1000 | Rows per shard file (part-000000.jsonl) when --out is a directory. |
--progress-interval | 30.0 | Seconds between stderr progress status lines. |
# 1. Stream sharded JSONL datasets
from sftmill.dataset import iter_jsonl_dataset, count_jsonl_dataset
total = count_jsonl_dataset("data/sft_output")
for row in iter_jsonl_dataset("data/sft_output"):
print(row["task_id"], len(row["messages"]))
# 2. Thread-safe sharded writing
from sftmill.dataset import JsonlDatasetWriter
with JsonlDatasetWriter("data/output_dir", shard_size=500) as writer:
writer.write({"messages": [...]})
# 3. Direct teacher completions
from sftmill.teachers.openai_compat import OpenAICompatibleTeacher
teacher = OpenAICompatibleTeacher(
base_url="http://localhost:8000/v1",
model="deepseek-ai/DeepSeek-R1-Distill-Qwen-32B",
stream=True
)
response = teacher.complete([{"role": "user", "content": "Explain Paxos consensus."}])
print("Thoughts:", response.get("reasoning_content"))
print("Answer:", response.get("content"))
sftmill/
├── assets/ # Visual assets and banners
│ └── sftmill-banner.png
├── configs/
│ ├── curriculum/ # Ready-to-use curriculum templates
│ │ ├── code_agent.yaml # Full agent curriculum (16 categories, tool use)
│ │ ├── code_agent_test.yaml # Fast test curriculum (32 sample tasks)
│ │ └── code_instruct.yaml # Multi-turn hybrid-reasoning instruct curriculum
│ └── tasks/
│ ├── examples.jsonl # Minimal reference tasks
│ └── code_agent_test.jsonl # Pre-synthesized task benchmark
├── docs/ # In-depth architectural & technical documentation
│ ├── spec.md # Curriculum YAML specification
│ ├── multi_turn_hybrid_reasoning.md # In-depth guide on multi-turn CoT distillation
│ ├── agent_trajectories.md # Sandboxing, workspace tools & verification
│ └── architecture.md # Concurrency, data flow & system design
├── src/sftmill/
│ ├── cli.py # Command-line interface
│ ├── dataset.py # Sharded writer, dataset iterator, group splitter
│ ├── harness.py # Multi-harness envelope translation
│ ├── progress.py # Live background status reporting
│ ├── schema.py # Message and tool validation
│ ├── generate/
│ │ ├── filters.py # Exact/contains/prose acceptance filters
│ │ ├── interpret.py # Reasoning extraction & call normalization
│ │ ├── synthesize_tasks.py # Stage 1 synthesis engine
│ │ ├── traces.py # Stage 2 reasoning traces & multi-turn loop
│ │ └── trajectories.py # Stage 2 tool execution & workspace sandbox
│ ├── teachers/
│ │ └── openai_compat.py # OpenAI-compatible streaming SSE client
│ └── tools/
│ ├── python_sandbox.py # Isolated timeout-capped Python runner
│ └── workspace.py # Hermetic file/bash/search workspace
├── tests/ # Comprehensive test suite (79 tests)
│ ├── test_cli.py
│ ├── test_dataset.py
│ ├── test_generate.py
│ ├── test_progress.py
│ ├── test_sandbox.py
│ ├── test_synthesize_tasks.py
│ └── test_workspace_tools.py
├── pyproject.toml # Project packaging & dependency specifications
├── CONTRIBUTING.md # Contribution and testing guidelines
└── LICENSE # MIT License
check.py), and harness envelopes.sftmill includes a test suite covering the entire pipeline. No GPU is required:
# Run all tests
pytest
# Or via uv
uv run --with pytest pytest
============================== 79 passed in 2.90s ==============================
Distributed under the MIT License. See LICENSE for details.
Python
100.0%
An industrial-grade, off-policy synthetic data distillation engine for LLMs.
Turn YAML curriculum blueprints into gold-standard task benchmarks, and mill verifiable reasoning traces and multi-turn agent trajectories from any model with an OpenAI-compatible endpoint.
user <> assistant)sftmill is a high-throughput, modular data engine designed to generate custom Supervised Fine-Tuning (SFT) datasets through off-policy model distillation. Instead of relying on expensive human annotations or brittle web scrapers, sftmill allows researchers and engineers to define high-level educational curricula in declarative YAML and distill rich training corpora directly from any frontier or open-weights teacher model (such as DeepSeek-R1, Qwen 2.5, OpenAI o1/o3-mini/GPT-4o, or locally served vLLM/Ollama instances).
sftmill natively produces:
user <> assistant conversations capturing both latent internal chain-of-thought (reasoning_content) and polished conversational answers (content).verify: python3 check.py) and exact post-edit filesystem states (expect_files).| Feature | Typical Synthetic Data Scripts | sftmill Data Mill |
|---|---|---|
| Pipeline Model | One-off messy scripts with prompt leaks | Decoupled 2-stage architecture: Spec ➔ Benchmark ➔ Rollout |
| Reasoning Distillation | Flattened text or stripped thinking | Native Hybrid Reasoning: preserves reasoning_content across turns |
| Tool Execution | Mocked regex or unexecuted hallucinations | Real hermetic sandbox: isolated subshells, timeouts, path protection |
| Verification | Blind trust in LLM output | Deterministic grading (exact, contains, check.py, gold file trees) |
| Tool Format Lock-in | Hardcoded to one provider | Multi-harness translation: OpenAI, Claude Code XML, Alice, custom |
| Evaluation Leakage | Contaminated test splits | Group Hashing: sibling harnesses never cross train/val splits |
| Production Scale | Memory-heavy arrays | Auto-partitioned sharded JSONL (part-000000.jsonl) with live streaming |
┌────────────────────────────────────────────────────────┐
│ YAML Curriculum Blueprint │
│ Define skills, counts, templates, and match rules │
└───────────────────────────┬────────────────────────────┘
│
▼ Stage 1: sftmill tasks
┌────────────────────────────────────────────────────────┐
│ Task Benchmark Synthesis │
│ Synthesis teacher invents diverse questions, tests, │
│ seed files, and multi-turn user prompts │
└───────────────────────────┬────────────────────────────┘
│ Output: tasks.jsonl
▼
┌────────────────────────────────────────────────────────┐
│ Stage 2: sftmill generate │
│ Rollout & Milling │
└─────────────┬────────────────────────────┬─────────────┘
│ │
[kind: trace] [kind: trajectory]
Multi-Turn CoT Sandboxed Tools
User-Assistant Turns Execution Loop
│ │
▼ ▼
Reasoning + Answer Capture Workspace Actions:
(reasoning_content + content) list_dir, read_file,
edit_file, bash, python
│ │
└─────────────┬──────────────┘
│
▼
┌────────────────────────────────────────────────────────┐
│ Acceptance & Grading │
│ - Exact answer match / substring contains │
│ - python3 check.py passes in workspace │
│ - expect_files tree matches gold state │
│ - Substantive prose verification (anti-stub filter) │
└───────────────────────────┬────────────────────────────┘
│
▼
┌────────────────────────────────────────────────────────┐
│ Production Sharded SFT Dataset │
│ data/part-000000.jsonl ... part-00000N.jsonl │
│ Ready for Axolotl, TRL, or Unsloth │
└────────────────────────────────────────────────────────┘
sftmill is lightweight and only requires Python 3.10+ and pyyaml.
# Clone the repository
git clone https://github.com/jackjusko/sftmill.git
cd sftmill
# Install in editable mode
pip install -e .
# Or using uv (recommended for ultra-fast setup)
uv pip install -e ".[dev]"
Use a teacher model to expand a curriculum specification into a concrete tasks.jsonl benchmark:
sftmill tasks \
--curriculum configs/curriculum/code_agent_test.yaml \
--out tasks.jsonl \
--base-url https://api.openai.com/v1 \
--model gpt-4o \
--api-key $OPENAI_API_KEY \
--jobs 4
Run generation rollouts over the synthesized tasks. Trajectories run tools in sandboxes and traces distill hybrid reasoning:
sftmill generate \
--tasks tasks.jsonl \
--out sft_output/ \
--base-url https://api.deepseek.com/v1 \
--model deepseek-reasoner \
--api-key $DEEPSEEK_API_KEY \
--shard-size 1000 \
--jobs 4
Compatible with any OpenAI-style backend: Works out of the box with vLLM, Ollama, DeepSeek, Groq, OpenAI, SGLang, or LiteLLM. Just configure
--base-urland--model.
user <> assistant)A core strength of sftmill is distilling conversational datasets where every assistant response contains both an internal chain-of-thought scratchpad (reasoning_content) and user-facing conversational prose (content).
When training smaller models (3B, 7B, 14B, 32B), teaching the model to formulate a structured thought process before answering follow-up queries or clarifying constraints significantly elevates downstream multi-turn performance.
Set kind: trace, match: open, and user_turns: 2 (or 3):
name: code_instruct
categories:
- id: clarify_requirements
count: 100
kind: trace
match: open
user_turns: 3
template: |
JSON key: turns, a list of exactly 3 user messages and no assistant messages.
The first turn is a programming task missing one critical technical constraint.
The second turn answers with that missing constraint.
The third turn asks for the final complete implementation and unit tests.
No tools, no files, no answer key.
example:
turns:
- "We need a rate limiter in Python for our public API endpoints."
- "Use Redis token-bucket algorithm, limiting to 60 requests per minute per IP."
- "Great. Please provide the complete production implementation with async redis-py."
sftmill Generates Multi-Turn Tracessftmill feeds turn[0] to the teacher model.delta.reasoning_content) is captured separately from its visible output (delta.content).user ➔ assistant with reasoning and text) is passed back for turn[1].sftmill.generate.filters.prose_ok automatically rejects one-word responses, lazy apologies, or stubs.The resulting JSONL dataset row is ready for standard fine-tuning pipelines:
{
"kind": "trace",
"task_id": "clarify_requirements-000042",
"messages": [
{
"role": "user",
"content": "We need a rate limiter in Python for our public API endpoints."
},
{
"role": "assistant",
"reasoning_content": "The user wants an API rate limiter in Python, but has not specified the concurrency model (FastAPI, Flask, Django), the storage backend (in-memory, Redis, Memcached), or the algorithm (token bucket, leaky bucket, sliding window). I should identify the missing details and ask for clarification.",
"content": "I would be happy to help implement a rate limiter. To tailor it to your architecture, could you clarify:\n1. Which backend or cache you plan to use (e.g., Redis, in-memory)?\n2. Which algorithm you prefer (e.g., token bucket, sliding window log)?\n3. What are the rate limits and identifier (e.g., 60 req/min per IP)?"
},
{
"role": "user",
"content": "Use Redis token-bucket algorithm, limiting to 60 requests per minute per IP."
},
{
"role": "assistant",
"reasoning_content": "The parameters are now clear: Redis backend, token bucket algorithm, 60 req/min per IP. The user hasn't asked for code yet, so I should confirm understanding and ask if they prefer async (redis-py async) or synchronous execution.",
"content": "Understood: Redis-backed token bucket allowing 60 requests per minute per IP. Do you need this implemented for an asynchronous framework (like FastAPI with `redis-py` async) or a WSGI application (like Flask)?"
},
{
"role": "user",
"content": "Great. Please provide the complete production implementation with async redis-py."
},
{
"role": "assistant",
"reasoning_content": "Requirements finalized: Async implementation using redis.asyncio, token bucket with capacity=60 and refill_rate=1 token/sec, atomic Lua script for concurrency safety, and a reusable dependency/middleware. Writing clean, robust code.",
"content": "Here is a complete, production-ready asynchronous token bucket rate limiter using `redis-py` with an atomic Lua script:\n\n```python\nimport time\nfrom redis import asyncio as aioredis\n\nTOKEN_BUCKET_LUA = \"\"\"\nlocal key = KEYS[1]\nlocal capacity = tonumber(ARGV[1])\nlocal refill_rate = tonumber(ARGV[2])\nlocal now = tonumber(ARGV[3])\nlocal requested = tonumber(ARGV[4])\n...\n```"
}
]
}
📖 For complete fine-tuning recipes with Axolotl, Unsloth, and TRL, read the Multi-Turn Hybrid-Reasoning Guide.
For agent training (kind: trajectory), sftmill executes real tool calls inside isolated temporary directory sandboxes.
read_file(path, offset, limit): Read file lines with pagination.write_file(path, content): Create or overwrite workspace files.edit_file(path, old, new): Precise single-occurrence string replacement.search(pattern, path): Literal recursive search across the workspace (path:line:text).list_dir(path): List directory contents.bash(command): Execute subshell commands with sanitized environment variables.python(code): Execute Python code with a 5-second timeout.verify: python3 check.py): An automated test script is executed after the agent declares completion. If check.py fails or exits non-zero, the trajectory is discarded.expect_files): The final workspace filesystem tree is compared against ground-truth files.require_observation: PASSED): Trajectories are rejected unless the agent witnessed a successful test run during its rollout... traversals are blocked).📖 Read the Agent Trajectories & Tool Use Guide for details.
A curriculum YAML controls how sftmill tasks synthesizes problems:
name: code_agent_benchmark
categories:
- id: bug_fix_caching
count: 25 # Number of rows to synthesize
kind: trajectory # trace | trajectory
match: exact # exact | contains | open
max_steps: 16 # Tool step budget per episode
toolset: workspace # workspace | custom
require_observation: PASSED # Substring that must appear in tool output
verify: python3 check.py # Command run in workspace to grade trajectory
harnesses: [openai, claude_code] # Envelopes to generate
template: |
JSON keys: question, answer, files, expect_files.
files includes app.py with a subtle bug, and check.py with assertions.
The check prints PASSED only when app.py is fixed.
expect_files contains the corrected app.py.
answer is the single token ok.
exact: Graded answer must match the gold string exactly.contains: Expected token or needle must appear in the assistant's final response or reasoning.open: Open-ended conversational prose (requires user_turns: 2 or 3 and validates multi-sentence responses).📖 See the full Curriculum Specification Reference.
The teacher generates canonical actions, and sftmill automatically translates them into platform-specific envelopes:
openai: Native function-calling JSON ({"name": "...", "arguments": {...}}).claude_code: Anthropic XML blocks (<tool_call><tool_name>...</tool_name></tool_call>).alice: Jinja-delimited native envelopes (<tool>\n...\n</tool>).novel: Out-of-distribution evaluation envelopes ([[tool]] ... [[/tool]]).When a problem is expanded across multiple harnesses (e.g. openai and claude_code), both share a group_id. sftmill includes a SHA-256 partitioner ensuring all variants of a problem always land on the same side of the train/val split:
from sftmill.dataset import iter_jsonl_dataset, split_by_group
# Load and split
all_rows = list(iter_jsonl_dataset("sft_output/"))
train_data, val_data = split_by_group(all_rows, val_fraction=0.1, seed=42)
print(f"Train rows: {len(train_data)} | Val rows: {len(val_data)}")
sftmill tasksSynthesize task definitions from a curriculum YAML file.
sftmill tasks \
--curriculum <path/to/curriculum.yaml> \
--out <path/to/tasks.jsonl> \
--base-url <url> \
--model <model_name> \
[--api-key <key>] \
[--temperature 0.7] \
[--batch-size 5] \
[--jobs 3] \
[--progress-interval 30.0]
| Flag | Default | Description |
|---|---|---|
--curriculum | required | Path to curriculum YAML spec. |
--out | required | Destination path for synthesized tasks JSONL. |
--base-url | required | Base URL of the OpenAI-compatible endpoint. |
--model | required | Teacher model identifier. |
--api-key | None | API key (or omit for local models). |
--jobs | 3 | Parallel categories in flight. |
--batch-size | 5 | Tasks requested per synthesis prompt. |
sftmill generateExecute task rollouts to produce final SFT training data.
sftmill generate \
--tasks <path/to/tasks.jsonl> \
--out <directory_or_file> \
--base-url <url> \
--model <model_name> \
[--api-key <key>] \
[--kind trace|trajectory|both] \
[--temperature 1.0] \
[--max-steps 4] \
[--jobs 3] \
[--shard-size 1000] \
[--progress-interval 30.0]
| Flag | Default | Description |
|---|---|---|
--tasks | required | Path to tasks JSONL file. |
--out | required | Output path (file .jsonl or directory for auto-sharding). |
--kind | both | Filter by trace, trajectory, or both. |
--max-steps | 4 | Maximum tool-call rounds for trajectories. |
--jobs | 3 | Concurrent worker threads. |
--shard-size | 1000 | Rows per shard file (part-000000.jsonl) when --out is a directory. |
--progress-interval | 30.0 | Seconds between stderr progress status lines. |
# 1. Stream sharded JSONL datasets
from sftmill.dataset import iter_jsonl_dataset, count_jsonl_dataset
total = count_jsonl_dataset("data/sft_output")
for row in iter_jsonl_dataset("data/sft_output"):
print(row["task_id"], len(row["messages"]))
# 2. Thread-safe sharded writing
from sftmill.dataset import JsonlDatasetWriter
with JsonlDatasetWriter("data/output_dir", shard_size=500) as writer:
writer.write({"messages": [...]})
# 3. Direct teacher completions
from sftmill.teachers.openai_compat import OpenAICompatibleTeacher
teacher = OpenAICompatibleTeacher(
base_url="http://localhost:8000/v1",
model="deepseek-ai/DeepSeek-R1-Distill-Qwen-32B",
stream=True
)
response = teacher.complete([{"role": "user", "content": "Explain Paxos consensus."}])
print("Thoughts:", response.get("reasoning_content"))
print("Answer:", response.get("content"))
sftmill/
├── assets/ # Visual assets and banners
│ └── sftmill-banner.png
├── configs/
│ ├── curriculum/ # Ready-to-use curriculum templates
│ │ ├── code_agent.yaml # Full agent curriculum (16 categories, tool use)
│ │ ├── code_agent_test.yaml # Fast test curriculum (32 sample tasks)
│ │ └── code_instruct.yaml # Multi-turn hybrid-reasoning instruct curriculum
│ └── tasks/
│ ├── examples.jsonl # Minimal reference tasks
│ └── code_agent_test.jsonl # Pre-synthesized task benchmark
├── docs/ # In-depth architectural & technical documentation
│ ├── spec.md # Curriculum YAML specification
│ ├── multi_turn_hybrid_reasoning.md # In-depth guide on multi-turn CoT distillation
│ ├── agent_trajectories.md # Sandboxing, workspace tools & verification
│ └── architecture.md # Concurrency, data flow & system design
├── src/sftmill/
│ ├── cli.py # Command-line interface
│ ├── dataset.py # Sharded writer, dataset iterator, group splitter
│ ├── harness.py # Multi-harness envelope translation
│ ├── progress.py # Live background status reporting
│ ├── schema.py # Message and tool validation
│ ├── generate/
│ │ ├── filters.py # Exact/contains/prose acceptance filters
│ │ ├── interpret.py # Reasoning extraction & call normalization
│ │ ├── synthesize_tasks.py # Stage 1 synthesis engine
│ │ ├── traces.py # Stage 2 reasoning traces & multi-turn loop
│ │ └── trajectories.py # Stage 2 tool execution & workspace sandbox
│ ├── teachers/
│ │ └── openai_compat.py # OpenAI-compatible streaming SSE client
│ └── tools/
│ ├── python_sandbox.py # Isolated timeout-capped Python runner
│ └── workspace.py # Hermetic file/bash/search workspace
├── tests/ # Comprehensive test suite (79 tests)
│ ├── test_cli.py
│ ├── test_dataset.py
│ ├── test_generate.py
│ ├── test_progress.py
│ ├── test_sandbox.py
│ ├── test_synthesize_tasks.py
│ └── test_workspace_tools.py
├── pyproject.toml # Project packaging & dependency specifications
├── CONTRIBUTING.md # Contribution and testing guidelines
└── LICENSE # MIT License
check.py), and harness envelopes.sftmill includes a test suite covering the entire pipeline. No GPU is required:
# Run all tests
pytest
# Or via uv
uv run --with pytest pytest
============================== 79 passed in 2.90s ==============================
Distributed under the MIT License. See LICENSE for details.
Python
100.0%