amazingvince/chess-reasoning-llm

0

stars

0

commits

Python

primary language

Jul 5, 2026

updated

README

Chess LLM

This repository contains the package, tools, and runbooks for training and evaluating chess-focused language models. The current loop is:

generate SFT data -> train -> evaluate -> judge failures -> generate targeted data

The canonical implementation lives under src/chess_llm/. Console scripts are declared in pyproject.toml and should be preferred over path-based Python entry points.

Current Shape

  • src/chess_llm/ contains the data, training, evaluation, artifact, Autodata, and shared chess utilities.
  • src/chess_llm_ui/ plus ui/ contain the workbench backend and React UI.
  • docs/ contains current architecture, runbooks, schemas, research notes, and archived historical material.
  • sft/training/ is now only operational scaffolding: Docker files, PowerShell launchers, requirements files, and local environment templates.
  • polyglot_opening_books/ stores local Polyglot source archives and extracted .bin books. Both are ignored and can be provisioned per machine.
  • data/syzygy/ is the ignored local tablebase cache used by package defaults.
  • chess_sft_data/, chess_sft_checkpoints/, artifacts/, .runtime/, and logs/ are local generated/runtime roots and are ignored.

The old Python compatibility shims under sft/make_data/ and sft/training/ have been removed. Use chess_llm.* imports and chess-llm-* commands.

Docs

Start with docs/index.md.

Useful current docs:

Historical plans, agent execution plans, and dated reviews live under docs/archive/. They are useful context, but package settings and current runbooks are the source of truth.

Install

Install the root package in editable mode before using the CLIs:

python -m pip install -e ".[dev]"

Install optional extras for the workflow you are running:

python -m pip install -e ".[data]"
python -m pip install -e ".[train]"
python -m pip install -e ".[eval]"
python -m pip install -e ".[ui]"

For one WSL/Linux GPU environment that generates data, trains, and evaluates:

python -m pip install -e ".[data,train,eval]"

Stockfish remains an external binary. Pass --stockfish-path where supported or set STOCKFISH_PATH.

Main Commands

Data generation and validation:

chess-llm-preflight
chess-llm-extract-polyglot-books --polyglot-dir polyglot_opening_books
chess-llm-download-tablebases --pieces 3,4,5 --dry-run
chess-llm-make-data --volume 100 --all
chess-llm-make-data --tier 1 2
chess-llm-validate-outputs --output-dir chess_sft_data/output --blocklist chess_sft_data/eval_splits/blocklist.txt
chess-llm-upload-data --all --dry-run

Evaluation:

chess-llm-run-eval-split --volume 100
chess-llm-run-eval-harness --split-dir chess_sft_data/eval_splits
chess-llm-freeze-benchmark --split-dir chess_sft_data/eval_splits --output-dir chess_sft_data/benchmark
chess-llm-run-benchmark --benchmark-dir chess_sft_data/benchmark --predictions predictions.jsonl
chess-llm-batch-judge --benchmark-dir chess_sft_data/benchmark --predictions predictions.jsonl --output-dir artifacts --model-id Qwen/Qwen3-4B
chess-llm-sft-refresh --prompts artifacts/prompts.jsonl --rollouts artifacts/rollouts.jsonl --judgments artifacts/judgments.jsonl --output-dir chess_sft_data/autodata_refresh

Training:

chess-llm-train --phase a --dry-run
chess-llm-train --phase a --smoke-run --wandb-project chess-sft --run-name phase-a-smoke
chess-llm-evaluate --model chess_sft_checkpoints/phase_a/best --benchmark-dir chess_sft_data/benchmark --output predictions.jsonl --phase a
chess-llm-run-curriculum --start-phase a --end-phase c --wandb-project chess-sft --run-prefix full-curriculum

The python -m chess_llm... module form remains available for package modules, but path-based legacy commands are retired.

Current Curriculum Totals

The source of truth is src/chess_llm/sft/settings.py.

TierFocusTasksTarget Rows
1Perception, state, material mechanics191,560,000
2Rules and legal-move decomposition12740,000
3Tactics9450,000
4Evaluation3150,000
5Openings330,000
6Endgames4140,000
7Planning and verification7310,000
Total573,380,000

Phase A uses tiers 1-2: 31 tasks and 2,300,000 target rows before task upsampling.

Tests

Run root package tests:

python -m pytest tests -q

Run UI checks:

cd ui
npm test
npm run build

amazingvince/chess-reasoning-llm

0

stars

0

commits

Python

primary language

Jul 5, 2026

updated

README

Chess LLM

This repository contains the package, tools, and runbooks for training and evaluating chess-focused language models. The current loop is:

generate SFT data -> train -> evaluate -> judge failures -> generate targeted data

The canonical implementation lives under src/chess_llm/. Console scripts are declared in pyproject.toml and should be preferred over path-based Python entry points.

Current Shape

  • src/chess_llm/ contains the data, training, evaluation, artifact, Autodata, and shared chess utilities.
  • src/chess_llm_ui/ plus ui/ contain the workbench backend and React UI.
  • docs/ contains current architecture, runbooks, schemas, research notes, and archived historical material.
  • sft/training/ is now only operational scaffolding: Docker files, PowerShell launchers, requirements files, and local environment templates.
  • polyglot_opening_books/ stores local Polyglot source archives and extracted .bin books. Both are ignored and can be provisioned per machine.
  • data/syzygy/ is the ignored local tablebase cache used by package defaults.
  • chess_sft_data/, chess_sft_checkpoints/, artifacts/, .runtime/, and logs/ are local generated/runtime roots and are ignored.

The old Python compatibility shims under sft/make_data/ and sft/training/ have been removed. Use chess_llm.* imports and chess-llm-* commands.

Docs

Start with docs/index.md.

Useful current docs:

Historical plans, agent execution plans, and dated reviews live under docs/archive/. They are useful context, but package settings and current runbooks are the source of truth.

Install

Install the root package in editable mode before using the CLIs:

python -m pip install -e ".[dev]"

Install optional extras for the workflow you are running:

python -m pip install -e ".[data]"
python -m pip install -e ".[train]"
python -m pip install -e ".[eval]"
python -m pip install -e ".[ui]"

For one WSL/Linux GPU environment that generates data, trains, and evaluates:

python -m pip install -e ".[data,train,eval]"

Stockfish remains an external binary. Pass --stockfish-path where supported or set STOCKFISH_PATH.

Main Commands

Data generation and validation:

chess-llm-preflight
chess-llm-extract-polyglot-books --polyglot-dir polyglot_opening_books
chess-llm-download-tablebases --pieces 3,4,5 --dry-run
chess-llm-make-data --volume 100 --all
chess-llm-make-data --tier 1 2
chess-llm-validate-outputs --output-dir chess_sft_data/output --blocklist chess_sft_data/eval_splits/blocklist.txt
chess-llm-upload-data --all --dry-run

Evaluation:

chess-llm-run-eval-split --volume 100
chess-llm-run-eval-harness --split-dir chess_sft_data/eval_splits
chess-llm-freeze-benchmark --split-dir chess_sft_data/eval_splits --output-dir chess_sft_data/benchmark
chess-llm-run-benchmark --benchmark-dir chess_sft_data/benchmark --predictions predictions.jsonl
chess-llm-batch-judge --benchmark-dir chess_sft_data/benchmark --predictions predictions.jsonl --output-dir artifacts --model-id Qwen/Qwen3-4B
chess-llm-sft-refresh --prompts artifacts/prompts.jsonl --rollouts artifacts/rollouts.jsonl --judgments artifacts/judgments.jsonl --output-dir chess_sft_data/autodata_refresh

Training:

chess-llm-train --phase a --dry-run
chess-llm-train --phase a --smoke-run --wandb-project chess-sft --run-name phase-a-smoke
chess-llm-evaluate --model chess_sft_checkpoints/phase_a/best --benchmark-dir chess_sft_data/benchmark --output predictions.jsonl --phase a
chess-llm-run-curriculum --start-phase a --end-phase c --wandb-project chess-sft --run-prefix full-curriculum

The python -m chess_llm... module form remains available for package modules, but path-based legacy commands are retired.

Current Curriculum Totals

The source of truth is src/chess_llm/sft/settings.py.

TierFocusTasksTarget Rows
1Perception, state, material mechanics191,560,000
2Rules and legal-move decomposition12740,000
3Tactics9450,000
4Evaluation3150,000
5Openings330,000
6Endgames4140,000
7Planning and verification7310,000
Total573,380,000

Phase A uses tiers 1-2: 31 tasks and 2,300,000 target rows before task upsampling.

Tests

Run root package tests:

python -m pytest tests -q

Run UI checks:

cd ui
npm test
npm run build

Languages

Python

94.3%

TypeScript

3.6%

PowerShell

1.4%