High-throughput synthetic and pretraining dataset sifter for TypeSafe Jev. Rust streaming core, Parquet and JSONL I/O, typed Choice/Score/Noul judgments, speculative fan-out, 24.0 rows/sec measured.
See the codejev-curateSupport: fuel the next build —
High-Throughput Synthetic & Pretraining Dataset Sifter Powered by TypeSafe AI (Jev)
Beta: v0.1.1 is experimental. Expect rough edges. Please contribute by opening an issue or PR.
Live: jev-curate.vercel.app (measured 24.0 rows/sec single-node on local mock bench examples/bench_mock.rs, 1,500+ cluster target)
Why jev-curate • Quickstart • CLI Reference • Python API • Architecture • Non-Goals • Ecosystem
jev-curate?Cleaning 10M to 1B rows of synthetic reasoning data, instruction tuning pairs, or web-scraped corpora is an economic and technical nightmare:
jev-curate solves this by piping Apache Arrow and Parquet streams through TypeSafe AI's Jev model (jev-1.13.0):
Noul), ordinal rubrics (Score on a 0 to 4 scale, sent as an ordered list), and categorical choices (Choice), eliminating generative text slop.# Install via Cargo
cargo install jev-curate
# Set your TypeSafe AI key
export TYPESAFE_API_KEY="your-api-key"
# Filter a Parquet dataset using the math reasoning preset:
jev-curate filter train.parquet \
--preset reasoning-math \
--out ./output/ \
--concurrency 32
PyO3 bindings (build from source with the python cargo feature):
git clone https://github.com/AkashPriyadarshii/jev-curate
cd jev-curate
maturin develop # python feature auto-enabled via pyproject.toml
from jev_curate import PyJevCurator
curator = PyJevCurator(api_key="your-api-key", preset="reasoning-math")
PyJevCurator currently wraps the same filter pipeline as the CLI (see src/filter.rs); constructor-only for now. Use the CLI for row-level sifting.
jev-curate filter [OPTIONS] <INPUT_PATH>
| Flag | Default | Description |
|---|---|---|
<INPUT_PATH> | Required | Path to input .parquet or .jsonl file. |
-p, --preset | reasoning-math | Pre-built rubric (reasoning-math, anti-sycophancy, code-correctness). |
-o, --out | ./curated/ | Destination folder for clean.jsonl and rejected.jsonl. |
-c, --concurrency | 32 | Worker concurrency (adaptive token bucket prevents 429 rate limits). |
--dry-run | false | Offline evaluation simulation with host pre-filtering and zero API calls (no TYPESAFE_API_KEY needed). |
--endpoint | None | Custom API endpoint URL for offline mock testing (or set TYPESAFE_ENDPOINT). |
| Preset | Primitives Evaluated | Target Problem Solved |
|---|---|---|
reasoning-math | has_circular_reasoning (Noul)reasoning_depth (Score 0-4, keep 2.0 or higher) | Drops ungrounded math derivations and repetitive circular proofs. |
anti-sycophancy | is_sycophantic (Noul)has_ai_disclaimer (Noul) | Eliminates "As an AI...", ungrounded flattery, and conversational filler. |
code-correctness | has_stub_placeholders (Noul)code_quality (Score 0-4, keep 2.0 or higher) | Drops incomplete code blocks and unrunnable pseudo-code mocks. |
# Rust core
cargo build --release
cargo test
# Python bindings (via maturin)
maturin develop
pytest
All tests run against an in-process mock server with zero live API credits in CI.
jev-curate/
├── Cargo.toml # Rust core manifest (arrow, parquet, tokio, pyo3, clap, reqwest, serde)
├── pyproject.toml # Maturin Python package manifest
├── src/
│ ├── lib.rs # PyO3 module bindings & crate entry
│ ├── main.rs # Standalone CLI binary entrypoint
│ ├── client.rs # TypeSafe AI HTTP client (speculative fan-out)
│ ├── filter.rs # Host-side sanity pruning & Jev pipeline
│ ├── parquet_io.rs # Streaming Parquet/Arrow reader and writer
│ ├── rate_limiter.rs # Adaptive token-bucket with auto 429 backoff
│ └── presets.rs # Pre-built post-training evaluation rubrics
└── tests/
└── mock_test.rs # In-process mock tests via typesafe-rs-mock (100% offline)
jev-curate never paraphrases or re-generates text. Data is kept 100% verbatim.
Built with high-performance Rust for the TypeSafe AI System One (Jev) ecosystem.
Keywords: TypeSafe AI, Jev, api.typesafe.ai, System One, Choice, Score, Noul, dataset curation, synthetic data filtering, pretraining datasets, Parquet streaming, arrow, rust.
Rust
100.0%
High-throughput synthetic and pretraining dataset sifter for TypeSafe Jev. Rust streaming core, Parquet and JSONL I/O, typed Choice/Score/Noul judgments, speculative fan-out, 24.0 rows/sec measured.
See the codejev-curateSupport: fuel the next build —
High-Throughput Synthetic & Pretraining Dataset Sifter Powered by TypeSafe AI (Jev)
Beta: v0.1.1 is experimental. Expect rough edges. Please contribute by opening an issue or PR.
Live: jev-curate.vercel.app (measured 24.0 rows/sec single-node on local mock bench examples/bench_mock.rs, 1,500+ cluster target)
Why jev-curate • Quickstart • CLI Reference • Python API • Architecture • Non-Goals • Ecosystem
jev-curate?Cleaning 10M to 1B rows of synthetic reasoning data, instruction tuning pairs, or web-scraped corpora is an economic and technical nightmare:
jev-curate solves this by piping Apache Arrow and Parquet streams through TypeSafe AI's Jev model (jev-1.13.0):
Noul), ordinal rubrics (Score on a 0 to 4 scale, sent as an ordered list), and categorical choices (Choice), eliminating generative text slop.# Install via Cargo
cargo install jev-curate
# Set your TypeSafe AI key
export TYPESAFE_API_KEY="your-api-key"
# Filter a Parquet dataset using the math reasoning preset:
jev-curate filter train.parquet \
--preset reasoning-math \
--out ./output/ \
--concurrency 32
PyO3 bindings (build from source with the python cargo feature):
git clone https://github.com/AkashPriyadarshii/jev-curate
cd jev-curate
maturin develop # python feature auto-enabled via pyproject.toml
from jev_curate import PyJevCurator
curator = PyJevCurator(api_key="your-api-key", preset="reasoning-math")
PyJevCurator currently wraps the same filter pipeline as the CLI (see src/filter.rs); constructor-only for now. Use the CLI for row-level sifting.
jev-curate filter [OPTIONS] <INPUT_PATH>
| Flag | Default | Description |
|---|---|---|
<INPUT_PATH> | Required | Path to input .parquet or .jsonl file. |
-p, --preset | reasoning-math | Pre-built rubric (reasoning-math, anti-sycophancy, code-correctness). |
-o, --out | ./curated/ | Destination folder for clean.jsonl and rejected.jsonl. |
-c, --concurrency | 32 | Worker concurrency (adaptive token bucket prevents 429 rate limits). |
--dry-run | false | Offline evaluation simulation with host pre-filtering and zero API calls (no TYPESAFE_API_KEY needed). |
--endpoint | None | Custom API endpoint URL for offline mock testing (or set TYPESAFE_ENDPOINT). |
| Preset | Primitives Evaluated | Target Problem Solved |
|---|---|---|
reasoning-math | has_circular_reasoning (Noul)reasoning_depth (Score 0-4, keep 2.0 or higher) | Drops ungrounded math derivations and repetitive circular proofs. |
anti-sycophancy | is_sycophantic (Noul)has_ai_disclaimer (Noul) | Eliminates "As an AI...", ungrounded flattery, and conversational filler. |
code-correctness | has_stub_placeholders (Noul)code_quality (Score 0-4, keep 2.0 or higher) | Drops incomplete code blocks and unrunnable pseudo-code mocks. |
# Rust core
cargo build --release
cargo test
# Python bindings (via maturin)
maturin develop
pytest
All tests run against an in-process mock server with zero live API credits in CI.
jev-curate/
├── Cargo.toml # Rust core manifest (arrow, parquet, tokio, pyo3, clap, reqwest, serde)
├── pyproject.toml # Maturin Python package manifest
├── src/
│ ├── lib.rs # PyO3 module bindings & crate entry
│ ├── main.rs # Standalone CLI binary entrypoint
│ ├── client.rs # TypeSafe AI HTTP client (speculative fan-out)
│ ├── filter.rs # Host-side sanity pruning & Jev pipeline
│ ├── parquet_io.rs # Streaming Parquet/Arrow reader and writer
│ ├── rate_limiter.rs # Adaptive token-bucket with auto 429 backoff
│ └── presets.rs # Pre-built post-training evaluation rubrics
└── tests/
└── mock_test.rs # In-process mock tests via typesafe-rs-mock (100% offline)
jev-curate never paraphrases or re-generates text. Data is kept 100% verbatim.
Built with high-performance Rust for the TypeSafe AI System One (Jev) ecosystem.
Keywords: TypeSafe AI, Jev, api.typesafe.ai, System One, Choice, Score, Noul, dataset curation, synthetic data filtering, pretraining datasets, Parquet streaming, arrow, rust.
Rust
100.0%