High-performance Crystal Reports XML parser built on the rypipe columnar ingestion engine.
See the code
Stream Crystal Reports XML at memory bandwidth.
High-performance Crystal Reports XML → Arrow/DataFrame engine for Python.
Parse, filter, rename, cast, and project Crystal Reports XML directly into
columnar data, with Rust execution, parallel parsing, bounded-memory
processing, and automatic query fusion.
from crxml import CrystalXMLSource
source = CrystalXMLSource("report.xml", row_tag="Details")
# Row iteration: yields dicts lazily
for row in source:
print(row["invoice"], row["amount"])
# DataFrame (auto-routes to parallel engine)
df = source.to_dataframe()
print(df.head())
That is it. df is a pandas DataFrame with zero-copy ArrowDtype strings,
built in under a second for a 100 MB file.
For maximum throughput, declare the schema upfront:
schema = ["Level","Section","Field22","Field23","Field38","Field39",
"Field61","Field73","FieldG","Text20"]
src = CrystalXMLSource("report.xml", row_tag="Details", schema=schema)
batches = src.iter_record_batches(memory="64MB", threads=16)
With pipeline stages fused into the Rust parse loop:
from crxml.stages import RenameFields, DropFields
pipeline = source | RenameFields({"f1": "invoice"}) | DropFields(["temp_id"])
df = pipeline.to_dataframe()
Performance tip: When your column set is known, pass
schema=[...]to skip column discovery and enable the fast path. On production data this is the single largest performance lever (crxml goes from 4.2 GB/s to 7.6 GB/s on a 533 MB report), and therow_satisfiedprojection skip reaches 11 GB/s on benchmarks.
This library was originally inspired by carlosplanchon/xmlstreamer.
Crystal Reports XML exports are deeply nested: <Group> wraps <GroupHeader>
wraps <Section> wraps <Details> wraps <Field>/<Text>/<FormattedValue>/
<Value>/<TextValue>. Standard XML libraries (ElementTree, SAX, lxml)
spend most of their CPU time descending into children you do not need.
crxml skips the nesting:
memchr scanner (src/crxml_core/src/xml/scanner.rs, scan_one_row scanner.rs:119 via RowSink src/crxml_core/src/lib.rs:603) and yields flat dicts: 508 MB/s 100 MB.splitter.rs:57 find_split_points), and parses each chunk on its own thread into Arrow buffers directly (no dicts): up to 4.2 GB/s on high-cardinality production reports (533 MB real par128 4231) and 4.2 GB/s on uniform exports (1 GB par128 4158) via rypipe (rypipe-core Vec<ColumnBuilder>+field_index engine.rs:16, row_dirty engine.rs:26).| Task | stream (single) | parallel (full RAM) | parallel streaming (bounded) |
|---|---|---|---|
| Row iteration | Yields dicts lazily | Arrow table first, then dicts (slower) | Yields RecordBatches incrementally, stable schema |
| DataFrame / Table output | Collects dicts, converts | Direct Arrow buffers, zero-copy | Same, incremental + bounded |
| 533 MB real export (Table) | 953 MB/s (single) / 723 MB/s (1 MB) | 4231 MB/s par128 (4.16 MB) | 3828 auto / 7630 explicit schema=[...] (2 MB) |
| 1 GB (Table) | 940 MB/s | 4158 MB/s par128 | 3782 auto / ~4900 explicit |
| Peak RssAnon (533 MB) | 24 MB (1 MB) | 137 MB | 88 MB (auto or explicit) |
| Pipeline fusion | No (dict path) | Yes (Rust BuildPlan) | Yes (same plan, streamed) |
ParquetWriter | N/A | N/A | write_batch succeeds (batches share schema schema.rs:14) |
Auto discovery (16x2 MiB windows for >128 MB) adds ~15% (19 ms on 533 MB) so auto is -10% vs par128 (3828 vs 4231) but still bounded and incremental. Explicit schema=[...] (FrozenSchema::from_plan) avoids Discovery and is +80% vs par128 (7630 vs 4231). Fastest bounded mode needs explicit schema; auto is safe and bounded but slightly slower. Use iter_record_batches(memory="64MB", threads=16, schema=[...]) for the fast path.
Full benchmark details: like-for-like Table vs Vec, chunk-per-cell, fixed-chunk isolation, and schema cost.
pip install crxml
The columnar and parallel engines are included by default. For performance
profiling counters: pip install -e . --config-settings=--features=profile.
| Category | What crxml handles |
|---|---|
| Stream engine | Row-by-row XML parsing, yields dict[str, str], GIL-released batching |
| Columnar engine | Single-threaded Arrow table output, zero-copy string columns |
| Parallel engine | Multi-threaded (rayon), file split at row boundaries, off-GIL parse |
| Bounded mode | memory="500MB" splits into chunks; RSS independent of file size |
| Pipeline fusion | RenameFields, DropFields, CastTypes, FilterRows compile into Rust BuildPlan |
| mmap | Memory-maps input files (default, zero-copy) |
| prefault | MADV_WILLNEED vs MADV_SEQUENTIAL for RSS/speed trade-off |
| Arrow sinks | to_arrow(), to_pandas() (ArrowDtype), to_polars(), to_parquet() |
| Auto-dict encoding | auto_dict=True encodes low-cardinality string columns |
| Field typing | field_types={"amount": "float64"} coerces at parse time |
| Filter pushdown | filter={"field": "Status", "op": "==", "value": "Active"} in Rust |
| Correctness | All engines validated byte-identical against stream oracle (29 test cases + 465k-row real cross-check) |
| Engine / API | When to use | Throughput 533 MB / 1 GB | RssAnon |
|---|---|---|---|
stream (for row in source) | Row-by-row dict iteration | 723 MB/s 1 MB budget (24 MB anon) | 24 MB |
columnar (single) | Single-threaded Arrow Table | 953 / 940 MB/s | 134 MB |
parallel (par128 full RAM, 4 MB) | Fastest full-RAM Table | 4231 / 4158 MB/s | 137 MB |
iter_record_batches(..., threads=16, schema=[...]) (explicit schema) | Fastest bounded, stable schema, yields RecordBatches | 7630 / — MB/s | 88 MB |
iter_record_batches(memory="64MB", threads=16) auto | Bounded + incremental, stable schema | 3828 / 3782 MB/s (-14% vs par, +15% Discovery) | 88 MB |
bounded (memory="64MB" single) | Single-thread bounded | 645 / 546 MB/s | 133 MB |
Pass engine= explicitly, or let auto select per call. auto stays "parallel if it fits" (blocked: auto discovery adds 15% and would make auto slower until cheaper). Streaming is opt-in via iter_record_batches(..., threads=16), keeping 4 MB for par (src/crxml/source.py:164), 2 MB via budget/(threads*2) for streaming. Provide schema= for the fast path.
# Recommended bounded paths
from crxml import CrystalXMLSource
import pyarrow as pa, pyarrow.parquet as pq
src = CrystalXMLSource("report.xml", row_tag="Details")
# explicit schema: fastest, no Discovery, writer succeeds
schema = ["Level","Section","Field22","Field23","Field38","Field39","Field61","Field73","FieldG","Text20"]
src = CrystalXMLSource("report.xml", row_tag="Details", schema=schema)
batches = src.iter_record_batches(memory="64MB", threads=16)
# auto: stable but pays 15% Discovery (16×2 MiB windows for >128 MB)
batches = src.iter_record_batches(memory="64MB", threads=16)
# ParquetWriter (now works; batches share schema)
it = src.iter_record_batches(memory="64MB", threads=16)
first = next(it)
w = pq.ParquetWriter("out.parquet", first.schema)
w.write_batch(first)
for b in it: w.write_batch(b)
w.close()
| Framework | Integration |
|---|---|
| FastAPI / Starlette / Litestar | Parse in route handler, return DataFrame or Arrow table directly |
| Django / Flask | Call source.to_dataframe() in view; pass to template or response |
| Pandas / Polars | source.to_dataframe() / source.to_polars() for zero-copy analysis |
| Airflow / Prefect | Parse in task, write to parquet with source.to_parquet() |
| CLI / ETL scripts | Use to_csv() sink or iterate rows for line-by-line processing |
.gz/.zst files must be decompressed before parsing.Full docs at crxml.emiliano-go.com covering:
CrystalXMLSource parametersMIT
186 commits
Python
57.4%
Rust
42.6%
High-performance Crystal Reports XML parser built on the rypipe columnar ingestion engine.
See the code
Stream Crystal Reports XML at memory bandwidth.
High-performance Crystal Reports XML → Arrow/DataFrame engine for Python.
Parse, filter, rename, cast, and project Crystal Reports XML directly into
columnar data, with Rust execution, parallel parsing, bounded-memory
processing, and automatic query fusion.
from crxml import CrystalXMLSource
source = CrystalXMLSource("report.xml", row_tag="Details")
# Row iteration: yields dicts lazily
for row in source:
print(row["invoice"], row["amount"])
# DataFrame (auto-routes to parallel engine)
df = source.to_dataframe()
print(df.head())
That is it. df is a pandas DataFrame with zero-copy ArrowDtype strings,
built in under a second for a 100 MB file.
For maximum throughput, declare the schema upfront:
schema = ["Level","Section","Field22","Field23","Field38","Field39",
"Field61","Field73","FieldG","Text20"]
src = CrystalXMLSource("report.xml", row_tag="Details", schema=schema)
batches = src.iter_record_batches(memory="64MB", threads=16)
With pipeline stages fused into the Rust parse loop:
from crxml.stages import RenameFields, DropFields
pipeline = source | RenameFields({"f1": "invoice"}) | DropFields(["temp_id"])
df = pipeline.to_dataframe()
Performance tip: When your column set is known, pass
schema=[...]to skip column discovery and enable the fast path. On production data this is the single largest performance lever (crxml goes from 4.2 GB/s to 7.6 GB/s on a 533 MB report), and therow_satisfiedprojection skip reaches 11 GB/s on benchmarks.
This library was originally inspired by carlosplanchon/xmlstreamer.
Crystal Reports XML exports are deeply nested: <Group> wraps <GroupHeader>
wraps <Section> wraps <Details> wraps <Field>/<Text>/<FormattedValue>/
<Value>/<TextValue>. Standard XML libraries (ElementTree, SAX, lxml)
spend most of their CPU time descending into children you do not need.
crxml skips the nesting:
memchr scanner (src/crxml_core/src/xml/scanner.rs, scan_one_row scanner.rs:119 via RowSink src/crxml_core/src/lib.rs:603) and yields flat dicts: 508 MB/s 100 MB.splitter.rs:57 find_split_points), and parses each chunk on its own thread into Arrow buffers directly (no dicts): up to 4.2 GB/s on high-cardinality production reports (533 MB real par128 4231) and 4.2 GB/s on uniform exports (1 GB par128 4158) via rypipe (rypipe-core Vec<ColumnBuilder>+field_index engine.rs:16, row_dirty engine.rs:26).| Task | stream (single) | parallel (full RAM) | parallel streaming (bounded) |
|---|---|---|---|
| Row iteration | Yields dicts lazily | Arrow table first, then dicts (slower) | Yields RecordBatches incrementally, stable schema |
| DataFrame / Table output | Collects dicts, converts | Direct Arrow buffers, zero-copy | Same, incremental + bounded |
| 533 MB real export (Table) | 953 MB/s (single) / 723 MB/s (1 MB) | 4231 MB/s par128 (4.16 MB) | 3828 auto / 7630 explicit schema=[...] (2 MB) |
| 1 GB (Table) | 940 MB/s | 4158 MB/s par128 | 3782 auto / ~4900 explicit |
| Peak RssAnon (533 MB) | 24 MB (1 MB) | 137 MB | 88 MB (auto or explicit) |
| Pipeline fusion | No (dict path) | Yes (Rust BuildPlan) | Yes (same plan, streamed) |
ParquetWriter | N/A | N/A | write_batch succeeds (batches share schema schema.rs:14) |
Auto discovery (16x2 MiB windows for >128 MB) adds ~15% (19 ms on 533 MB) so auto is -10% vs par128 (3828 vs 4231) but still bounded and incremental. Explicit schema=[...] (FrozenSchema::from_plan) avoids Discovery and is +80% vs par128 (7630 vs 4231). Fastest bounded mode needs explicit schema; auto is safe and bounded but slightly slower. Use iter_record_batches(memory="64MB", threads=16, schema=[...]) for the fast path.
Full benchmark details: like-for-like Table vs Vec, chunk-per-cell, fixed-chunk isolation, and schema cost.
pip install crxml
The columnar and parallel engines are included by default. For performance
profiling counters: pip install -e . --config-settings=--features=profile.
| Category | What crxml handles |
|---|---|
| Stream engine | Row-by-row XML parsing, yields dict[str, str], GIL-released batching |
| Columnar engine | Single-threaded Arrow table output, zero-copy string columns |
| Parallel engine | Multi-threaded (rayon), file split at row boundaries, off-GIL parse |
| Bounded mode | memory="500MB" splits into chunks; RSS independent of file size |
| Pipeline fusion | RenameFields, DropFields, CastTypes, FilterRows compile into Rust BuildPlan |
| mmap | Memory-maps input files (default, zero-copy) |
| prefault | MADV_WILLNEED vs MADV_SEQUENTIAL for RSS/speed trade-off |
| Arrow sinks | to_arrow(), to_pandas() (ArrowDtype), to_polars(), to_parquet() |
| Auto-dict encoding | auto_dict=True encodes low-cardinality string columns |
| Field typing | field_types={"amount": "float64"} coerces at parse time |
| Filter pushdown | filter={"field": "Status", "op": "==", "value": "Active"} in Rust |
| Correctness | All engines validated byte-identical against stream oracle (29 test cases + 465k-row real cross-check) |
| Engine / API | When to use | Throughput 533 MB / 1 GB | RssAnon |
|---|---|---|---|
stream (for row in source) | Row-by-row dict iteration | 723 MB/s 1 MB budget (24 MB anon) | 24 MB |
columnar (single) | Single-threaded Arrow Table | 953 / 940 MB/s | 134 MB |
parallel (par128 full RAM, 4 MB) | Fastest full-RAM Table | 4231 / 4158 MB/s | 137 MB |
iter_record_batches(..., threads=16, schema=[...]) (explicit schema) | Fastest bounded, stable schema, yields RecordBatches | 7630 / — MB/s | 88 MB |
iter_record_batches(memory="64MB", threads=16) auto | Bounded + incremental, stable schema | 3828 / 3782 MB/s (-14% vs par, +15% Discovery) | 88 MB |
bounded (memory="64MB" single) | Single-thread bounded | 645 / 546 MB/s | 133 MB |
Pass engine= explicitly, or let auto select per call. auto stays "parallel if it fits" (blocked: auto discovery adds 15% and would make auto slower until cheaper). Streaming is opt-in via iter_record_batches(..., threads=16), keeping 4 MB for par (src/crxml/source.py:164), 2 MB via budget/(threads*2) for streaming. Provide schema= for the fast path.
# Recommended bounded paths
from crxml import CrystalXMLSource
import pyarrow as pa, pyarrow.parquet as pq
src = CrystalXMLSource("report.xml", row_tag="Details")
# explicit schema: fastest, no Discovery, writer succeeds
schema = ["Level","Section","Field22","Field23","Field38","Field39","Field61","Field73","FieldG","Text20"]
src = CrystalXMLSource("report.xml", row_tag="Details", schema=schema)
batches = src.iter_record_batches(memory="64MB", threads=16)
# auto: stable but pays 15% Discovery (16×2 MiB windows for >128 MB)
batches = src.iter_record_batches(memory="64MB", threads=16)
# ParquetWriter (now works; batches share schema)
it = src.iter_record_batches(memory="64MB", threads=16)
first = next(it)
w = pq.ParquetWriter("out.parquet", first.schema)
w.write_batch(first)
for b in it: w.write_batch(b)
w.close()
| Framework | Integration |
|---|---|
| FastAPI / Starlette / Litestar | Parse in route handler, return DataFrame or Arrow table directly |
| Django / Flask | Call source.to_dataframe() in view; pass to template or response |
| Pandas / Polars | source.to_dataframe() / source.to_polars() for zero-copy analysis |
| Airflow / Prefect | Parse in task, write to parquet with source.to_parquet() |
| CLI / ETL scripts | Use to_csv() sink or iterate rows for line-by-line processing |
.gz/.zst files must be decompressed before parsing.Full docs at crxml.emiliano-go.com covering:
CrystalXMLSource parametersMIT
186 commits
Python
57.4%
Rust
42.6%