Pedagogical semantic chunking and RAG dataset preparation from OpenStax textbooks
Python
0
18 commits
updated Oct 2, 2026
Pedagogical semantic chunking, RAG dataset preparation, and LLM fine-tuning pipelines from OpenStax textbooks.
Built on top of openstax-md, openstax-llm
transforms OpenStax college textbooks into structured, citation-aware, formula-safe datasets
for vector search (RAG) and model fine-tuning.
This repository ships three things that share one core:
| Artifact | What it is | Entry point |
|---|---|---|
openstax-llm | Python library and CLI | openstax-llm |
openstax-llm-mcp | Model Context Protocol server | openstax-llm-mcp |
skills/openstax-llm | Agent skill for coding assistants | /skill:openstax-llm |
Generic chunkers (simple character or recursive token splitters) break down on technical academic textbooks:
$x^2 + \dots$ split from \dots + y^2$).openstax-llm provides:
Example 1.1), problem sets,
definitions, and summaries are kept whole.$...$ or
$$...$$ block, or inside a fenced code block.max_words is actually honoured; oversized paragraphs are
split at sentence boundaries that lie outside math.chunk_ids, compatible with
Chroma, Qdrant, Pinecone, LanceDB, LlamaIndex, LangChain, and Hugging Face datasets.# Library + CLI
uv add openstax-llm
# CLI as a standalone tool
uv tool install openstax-llm
# MCP server, one-shot via uvx
uvx openstax-llm-mcp
Or run the container:
docker build -t openstax-llm-mcp .
To track an unreleased commit instead, install from git:
uv add "git+https://github.com/michaelnavazhylau/openstax-llm.git#subdirectory=packages/openstax-llm"
uvx --from "git+https://github.com/michaelnavazhylau/openstax-llm.git#subdirectory=packages/openstax-llm-mcp" openstax-llm-mcp
# Search the catalog for available textbooks
openstax-llm search physics
# Inspect a textbook's chunk statistics and section index
openstax-llm info astronomy-2e
# Compile and chunk a textbook into a JSONL dataset
openstax-llm prepare astronomy-2e -o datasets/astronomy-2e.jsonl
# Verify an export before using it downstream
openstax-llm validate datasets/astronomy-2e.jsonl
validate checks formula integrity, chunk_id uniqueness, and provenance, and exits
non-zero on structural errors. Add --strict to also fail on warnings such as front matter
that has no section number. Always run it before loading a dataset into an index: a split
formula that reaches an embedding store is very hard to detect afterwards.
from openstax_llm import DocumentChunker, TextBookDataset
chunker = DocumentChunker(target_words=400, max_words=600, overlap_words=50)
dataset = TextBookDataset.from_textbook("calculus-volume-1", chunker=chunker)
dataset.to_jsonl("calculus.jsonl")
print(dataset.summary())
# {'book_slug': 'calculus-volume-1', 'total_chunks': 1445, 'total_words': 262184, ...}
DocumentChunker also works on arbitrary markdown:
from openstax_llm import DocumentChunker
chunks = DocumentChunker().chunk_markdown(
"## 1.2 Functions\n\nAn inline formula $f(x)=x^2$ stays intact.\n",
book_slug="my-notes",
section="1.2",
section_title="Functions",
)
Exposes the same three operations to any MCP client over stdio or streamable HTTP.
# Register with Pi
pi mcp add openstax-llm -- uvx openstax-llm-mcp
# Or serve over HTTP
docker run --rm -p 8765:8765 openstax-llm-mcp
| Tool | Purpose |
|---|---|
search_catalog(query, limit) | Resolve a subject to a canonical slug (offline) |
inspect_textbook(target) | Chunk totals plus a per-section index |
prepare_textbook(target, out) | Export JSONL, sandboxed to the server's output directory |
Resources: textbook://<slug> for an overview and textbook://<slug>/<section> for every
chunk in one section. See packages/openstax-llm-mcp/README.md
for options, tool schemas, and failure modes.
npx skills add michaelnavazhylau/openstax-llm
Also on skills.sh — that directory is populated
from anonymous install telemetry, so the listing appears only after the first
npx skills add (the command above is always the canonical path).
The skill teaches an agent when and how to reach for these tools: resolving slugs instead of guessing titles, verifying exports before indexing, tuning chunk sizes, and loading the result into Chroma, Qdrant, Pinecone, or Hugging Face. It deliberately contains no chunking logic of its own.
| Field | Type | Description |
|---|---|---|
chunk_id | string | Unique within a book: <section>-c<index>, e.g. 1.2-c003 |
text | string | Markdown with intact LaTeX math |
book_slug | string | Canonical OpenStax slug |
book_title | string | Human-readable title |
chapter | string | Chapter number from the section hierarchy |
section | string | Section number, e.g. 1.2 |
section_title | string | Module title |
chunk_type | string | prose, example, exercise, definition, summary |
word_count | integer | Whitespace-delimited word count |
token_est | integer | Heuristic estimate, words * 1.3 |
metadata | object | Carries module_id; free for downstream use |
Front matter — prefaces, formula tables, chapter introductions — has no section number. Its
chunk_id is prefixed with the module id and its provenance lives in metadata.module_id.
The machine-readable schema is at
skills/openstax-llm/assets/chunk.schema.json.
uv sync --all-groups --all-packages
uv run pytest -v
uv run ruff check . && uv run ruff format --check .
uv run mypy
# Real-textbook integration tests (clones a book, needs network)
OPENSTAX_LLM_NETWORK_TESTS=1 uv run pytest tests/test_integration.py -v
The repository is a uv workspace: the root pyproject.toml declares members and owns the
shared tool configuration, while each package under packages/ is independently
distributable. See AGENTS.md for the architecture and engineering invariants,
and docs/PUBLISHING.md for the release process.
MIT
Python
85.4%
Shell
12.3%
Dockerfile
2.3%
Pedagogical semantic chunking and RAG dataset preparation from OpenStax textbooks
Python
0
18 commits
updated Oct 2, 2026
Pedagogical semantic chunking, RAG dataset preparation, and LLM fine-tuning pipelines from OpenStax textbooks.
Built on top of openstax-md, openstax-llm
transforms OpenStax college textbooks into structured, citation-aware, formula-safe datasets
for vector search (RAG) and model fine-tuning.
This repository ships three things that share one core:
| Artifact | What it is | Entry point |
|---|---|---|
openstax-llm | Python library and CLI | openstax-llm |
openstax-llm-mcp | Model Context Protocol server | openstax-llm-mcp |
skills/openstax-llm | Agent skill for coding assistants | /skill:openstax-llm |
Generic chunkers (simple character or recursive token splitters) break down on technical academic textbooks:
$x^2 + \dots$ split from \dots + y^2$).openstax-llm provides:
Example 1.1), problem sets,
definitions, and summaries are kept whole.$...$ or
$$...$$ block, or inside a fenced code block.max_words is actually honoured; oversized paragraphs are
split at sentence boundaries that lie outside math.chunk_ids, compatible with
Chroma, Qdrant, Pinecone, LanceDB, LlamaIndex, LangChain, and Hugging Face datasets.# Library + CLI
uv add openstax-llm
# CLI as a standalone tool
uv tool install openstax-llm
# MCP server, one-shot via uvx
uvx openstax-llm-mcp
Or run the container:
docker build -t openstax-llm-mcp .
To track an unreleased commit instead, install from git:
uv add "git+https://github.com/michaelnavazhylau/openstax-llm.git#subdirectory=packages/openstax-llm"
uvx --from "git+https://github.com/michaelnavazhylau/openstax-llm.git#subdirectory=packages/openstax-llm-mcp" openstax-llm-mcp
# Search the catalog for available textbooks
openstax-llm search physics
# Inspect a textbook's chunk statistics and section index
openstax-llm info astronomy-2e
# Compile and chunk a textbook into a JSONL dataset
openstax-llm prepare astronomy-2e -o datasets/astronomy-2e.jsonl
# Verify an export before using it downstream
openstax-llm validate datasets/astronomy-2e.jsonl
validate checks formula integrity, chunk_id uniqueness, and provenance, and exits
non-zero on structural errors. Add --strict to also fail on warnings such as front matter
that has no section number. Always run it before loading a dataset into an index: a split
formula that reaches an embedding store is very hard to detect afterwards.
from openstax_llm import DocumentChunker, TextBookDataset
chunker = DocumentChunker(target_words=400, max_words=600, overlap_words=50)
dataset = TextBookDataset.from_textbook("calculus-volume-1", chunker=chunker)
dataset.to_jsonl("calculus.jsonl")
print(dataset.summary())
# {'book_slug': 'calculus-volume-1', 'total_chunks': 1445, 'total_words': 262184, ...}
DocumentChunker also works on arbitrary markdown:
from openstax_llm import DocumentChunker
chunks = DocumentChunker().chunk_markdown(
"## 1.2 Functions\n\nAn inline formula $f(x)=x^2$ stays intact.\n",
book_slug="my-notes",
section="1.2",
section_title="Functions",
)
Exposes the same three operations to any MCP client over stdio or streamable HTTP.
# Register with Pi
pi mcp add openstax-llm -- uvx openstax-llm-mcp
# Or serve over HTTP
docker run --rm -p 8765:8765 openstax-llm-mcp
| Tool | Purpose |
|---|---|
search_catalog(query, limit) | Resolve a subject to a canonical slug (offline) |
inspect_textbook(target) | Chunk totals plus a per-section index |
prepare_textbook(target, out) | Export JSONL, sandboxed to the server's output directory |
Resources: textbook://<slug> for an overview and textbook://<slug>/<section> for every
chunk in one section. See packages/openstax-llm-mcp/README.md
for options, tool schemas, and failure modes.
npx skills add michaelnavazhylau/openstax-llm
Also on skills.sh — that directory is populated
from anonymous install telemetry, so the listing appears only after the first
npx skills add (the command above is always the canonical path).
The skill teaches an agent when and how to reach for these tools: resolving slugs instead of guessing titles, verifying exports before indexing, tuning chunk sizes, and loading the result into Chroma, Qdrant, Pinecone, or Hugging Face. It deliberately contains no chunking logic of its own.
| Field | Type | Description |
|---|---|---|
chunk_id | string | Unique within a book: <section>-c<index>, e.g. 1.2-c003 |
text | string | Markdown with intact LaTeX math |
book_slug | string | Canonical OpenStax slug |
book_title | string | Human-readable title |
chapter | string | Chapter number from the section hierarchy |
section | string | Section number, e.g. 1.2 |
section_title | string | Module title |
chunk_type | string | prose, example, exercise, definition, summary |
word_count | integer | Whitespace-delimited word count |
token_est | integer | Heuristic estimate, words * 1.3 |
metadata | object | Carries module_id; free for downstream use |
Front matter — prefaces, formula tables, chapter introductions — has no section number. Its
chunk_id is prefixed with the module id and its provenance lives in metadata.module_id.
The machine-readable schema is at
skills/openstax-llm/assets/chunk.schema.json.
uv sync --all-groups --all-packages
uv run pytest -v
uv run ruff check . && uv run ruff format --check .
uv run mypy
# Real-textbook integration tests (clones a book, needs network)
OPENSTAX_LLM_NETWORK_TESTS=1 uv run pytest tests/test_integration.py -v
The repository is a uv workspace: the root pyproject.toml declares members and owns the
shared tool configuration, while each package under packages/ is independently
distributable. See AGENTS.md for the architecture and engineering invariants,
and docs/PUBLISHING.md for the release process.
MIT
Python
85.4%
Shell
12.3%
Dockerfile
2.3%