English | 中文
Billion-scale embedded vector database built entirely on Parquet and Arrow.
Website | Browser Demo | Quick Start | Status | Documentation
ParqDB is an embedded vector database for larger-than-memory search and analytics on billion-scale multimodal data, with Parquet storage and Arrow-native execution.
Try the live browser demo →
IVF-LVQ8 over HTTP Range · Parquet · WebAssembly · no query server
⭐ If ParqDB is useful, star the repo to help more people find it.
Key Features
Install ParqDB:
python -m pip install parqdb
From a new working directory, build a source-encoded IVF index over the dataset included in the package and run a filtered vector query:
import parqdb
session = parqdb.connect("./parqdb-data")
session.register_parquet("documents", parqdb.datasets.uri("documents"))
documents = session.table("documents")
documents.create_index(
"documents_embedding",
column="embedding",
key=["document_id"],
config=parqdb.IVF(nlist=3),
)
documents.wait_for_index("documents_embedding")
query = (
documents.search([0.2, 0.0], column="embedding")
.where("tenant_id = 42 AND status = 'published'")
.nprobes(3)
.limit(3)
.select(["document_id", "title", "category"])
)
print(session.collect(query).to_pylist())
Vector search remains relational rather than becoming a terminal service call. Compile it as a SQL subquery and compose it with the rest of the analysis:
session.register_parquet(
"document_stats",
parqdb.datasets.uri("document_stats"),
)
search_sql = session.to_sql(query)
summary = session.sql(f"""
SELECT
h.category,
COUNT(*) AS matches,
AVG(h._distance) AS avg_distance,
MAX(s.popularity) AS max_popularity
FROM ({search_sql}) AS h
JOIN document_stats AS s USING (document_id)
GROUP BY h.category
ORDER BY h.category
""")
print(summary.to_pydict())
The packaged dataset makes this example self-contained. The getting-started guide covers persistent tables, existing indexes, query inspection, and source schema requirements.
Publish a source table and an immutable browser index in one command. For raw
text, ParqDB uses the pinned MiniLM ONNX model for both offline embeddings and
browser parity metadata, then builds hierarchical IVF-LVQ8 and uploads every
object before exposing the single top-level manifest.json:
python -m pip install "parqdb[publish]"
parqdb publish \
--source documents.parquet \
--key chunk_id \
--text-column title \
--text-column section \
--text-column text \
--nlist 4096 \
--destination s3://my-bucket/kb/v1 \
--s3-endpoint https://ACCOUNT_ID.r2.cloudflarestorage.com \
--s3-region auto \
--public-url https://data.example.com/kb/v1 \
--include-source \
--include-model
By default only the index artifact is uploaded. --include-source enables
browser payload lookup, while --include-model publishes the pinned text
embedding model. Credentials come from the standard AWS_ACCESS_KEY_ID and
AWS_SECRET_ACCESS_KEY environment variables. If documents.parquet already
contains embeddings, replace the three --text-column options with
--vector-column embedding. Publication refuses to overwrite an existing
prefix and verifies public HTTP Range and CORS behavior before succeeding.
For a complete document-to-GitHub-Pages knowledge base, including token-aware
chunking and the search UI, see
parqdb-knowledgebase.
| Runtime | Storage | Current capability | Status |
|---|---|---|---|
| Embedded DataFusion | Parquet | Build and query IVF, IVF-LVQ4, and IVF-LVQ8 indexes | Supported |
| Browser/WASM | Public HTTPS object storage | Query immutable IVF-LVQ4 and IVF-LVQ8 indexes over HTTP Range | Experimental |
| Client/server | Authorized Parquet sources | Build and query through the HTTP API | Experimental |
The first supported product surface is the embedded DataFusion runtime. The index specification remains independent of that runtime; distributed engine adapters are no longer bundled into the Python package.
See the embedded guide for installation and configuration.
The experimental HTTP server is documented in the server guide.
Public documentation is maintained at parqdb.io/docs.
Use the version selector there to switch between latest and release snapshots.
TEngineDB-V: An OLAP-Native Vector Search System for Large-k Workloads at Tencent is Tencent's production system for large-k vector search. On a 10-billion-vector deployment, its deep integration with TEngineDB delivers up to a 52x speedup over the legacy system.
ParqDB shares the idea, not the implementation. It rebuilds table-native vector search around open index formats and existing SQL engines, aiming for TEngineDB-V-class performance without requiring a proprietary engine.
If you use ParqDB in your research, please cite our VLDB 2026 Industry Track paper:
@misc{wu2026tenginedbvolapnativevectorsearch,
title = {{TEngineDB-V}: An {OLAP}-Native Vector Search System for Large-$k$ Workloads at Tencent},
author = {Xufei Wu and Pengcheng Zhang and Yitong Song and Xiaobo Zhang and Anqi Liang and Kai Wang and Jijun Du and Yidi Xiong and Guangxu Cheng and Zhe Chen and Peng Chen and Guoliang Li and Xuanhe Zhou and Fan Wu},
year = {2026},
eprint = {2608.00650},
archivePrefix = {arXiv},
primaryClass = {cs.DB},
url = {https://arxiv.org/abs/2608.00650},
}
ParqDB's next phase is being designed in public. We welcome concrete use cases, benchmark results, design feedback, and implementation help:
If you are working on RAG, agent trajectory storage, Parquet performance, or embedded lakehouse systems, share your workload and requirements in the relevant issue. Comment before starting a large change so that scope and interfaces can be agreed on first.
ParqDB uses uv, Maturin, Cargo, and a small Makefile orchestration layer:
make sync
make develop
make check
See CONTRIBUTING.md for quality gates, fixtures, benchmarks, and contribution guidelines.
Made with contrib.rocks.
ParqDB's original code is available under the MIT License. Wheels include the vendored DataFusion Python binding under Apache-2.0; see the third-party notices.
ParqDB builds on work from LanceDB, DataFusion, DuckDB, StarRocks, Apache Spark, and Apache Iceberg, with gratitude to their contributors and communities.
Python
50.5%
Rust
46.0%
TypeScript
2.4%
English | 中文
Billion-scale embedded vector database built entirely on Parquet and Arrow.
Website | Browser Demo | Quick Start | Status | Documentation
ParqDB is an embedded vector database for larger-than-memory search and analytics on billion-scale multimodal data, with Parquet storage and Arrow-native execution.
Try the live browser demo →
IVF-LVQ8 over HTTP Range · Parquet · WebAssembly · no query server
⭐ If ParqDB is useful, star the repo to help more people find it.
Key Features
Install ParqDB:
python -m pip install parqdb
From a new working directory, build a source-encoded IVF index over the dataset included in the package and run a filtered vector query:
import parqdb
session = parqdb.connect("./parqdb-data")
session.register_parquet("documents", parqdb.datasets.uri("documents"))
documents = session.table("documents")
documents.create_index(
"documents_embedding",
column="embedding",
key=["document_id"],
config=parqdb.IVF(nlist=3),
)
documents.wait_for_index("documents_embedding")
query = (
documents.search([0.2, 0.0], column="embedding")
.where("tenant_id = 42 AND status = 'published'")
.nprobes(3)
.limit(3)
.select(["document_id", "title", "category"])
)
print(session.collect(query).to_pylist())
Vector search remains relational rather than becoming a terminal service call. Compile it as a SQL subquery and compose it with the rest of the analysis:
session.register_parquet(
"document_stats",
parqdb.datasets.uri("document_stats"),
)
search_sql = session.to_sql(query)
summary = session.sql(f"""
SELECT
h.category,
COUNT(*) AS matches,
AVG(h._distance) AS avg_distance,
MAX(s.popularity) AS max_popularity
FROM ({search_sql}) AS h
JOIN document_stats AS s USING (document_id)
GROUP BY h.category
ORDER BY h.category
""")
print(summary.to_pydict())
The packaged dataset makes this example self-contained. The getting-started guide covers persistent tables, existing indexes, query inspection, and source schema requirements.
Publish a source table and an immutable browser index in one command. For raw
text, ParqDB uses the pinned MiniLM ONNX model for both offline embeddings and
browser parity metadata, then builds hierarchical IVF-LVQ8 and uploads every
object before exposing the single top-level manifest.json:
python -m pip install "parqdb[publish]"
parqdb publish \
--source documents.parquet \
--key chunk_id \
--text-column title \
--text-column section \
--text-column text \
--nlist 4096 \
--destination s3://my-bucket/kb/v1 \
--s3-endpoint https://ACCOUNT_ID.r2.cloudflarestorage.com \
--s3-region auto \
--public-url https://data.example.com/kb/v1 \
--include-source \
--include-model
By default only the index artifact is uploaded. --include-source enables
browser payload lookup, while --include-model publishes the pinned text
embedding model. Credentials come from the standard AWS_ACCESS_KEY_ID and
AWS_SECRET_ACCESS_KEY environment variables. If documents.parquet already
contains embeddings, replace the three --text-column options with
--vector-column embedding. Publication refuses to overwrite an existing
prefix and verifies public HTTP Range and CORS behavior before succeeding.
For a complete document-to-GitHub-Pages knowledge base, including token-aware
chunking and the search UI, see
parqdb-knowledgebase.
| Runtime | Storage | Current capability | Status |
|---|---|---|---|
| Embedded DataFusion | Parquet | Build and query IVF, IVF-LVQ4, and IVF-LVQ8 indexes | Supported |
| Browser/WASM | Public HTTPS object storage | Query immutable IVF-LVQ4 and IVF-LVQ8 indexes over HTTP Range | Experimental |
| Client/server | Authorized Parquet sources | Build and query through the HTTP API | Experimental |
The first supported product surface is the embedded DataFusion runtime. The index specification remains independent of that runtime; distributed engine adapters are no longer bundled into the Python package.
See the embedded guide for installation and configuration.
The experimental HTTP server is documented in the server guide.
Public documentation is maintained at parqdb.io/docs.
Use the version selector there to switch between latest and release snapshots.
TEngineDB-V: An OLAP-Native Vector Search System for Large-k Workloads at Tencent is Tencent's production system for large-k vector search. On a 10-billion-vector deployment, its deep integration with TEngineDB delivers up to a 52x speedup over the legacy system.
ParqDB shares the idea, not the implementation. It rebuilds table-native vector search around open index formats and existing SQL engines, aiming for TEngineDB-V-class performance without requiring a proprietary engine.
If you use ParqDB in your research, please cite our VLDB 2026 Industry Track paper:
@misc{wu2026tenginedbvolapnativevectorsearch,
title = {{TEngineDB-V}: An {OLAP}-Native Vector Search System for Large-$k$ Workloads at Tencent},
author = {Xufei Wu and Pengcheng Zhang and Yitong Song and Xiaobo Zhang and Anqi Liang and Kai Wang and Jijun Du and Yidi Xiong and Guangxu Cheng and Zhe Chen and Peng Chen and Guoliang Li and Xuanhe Zhou and Fan Wu},
year = {2026},
eprint = {2608.00650},
archivePrefix = {arXiv},
primaryClass = {cs.DB},
url = {https://arxiv.org/abs/2608.00650},
}
ParqDB's next phase is being designed in public. We welcome concrete use cases, benchmark results, design feedback, and implementation help:
If you are working on RAG, agent trajectory storage, Parquet performance, or embedded lakehouse systems, share your workload and requirements in the relevant issue. Comment before starting a large change so that scope and interfaces can be agreed on first.
ParqDB uses uv, Maturin, Cargo, and a small Makefile orchestration layer:
make sync
make develop
make check
See CONTRIBUTING.md for quality gates, fixtures, benchmarks, and contribution guidelines.
Made with contrib.rocks.
ParqDB's original code is available under the MIT License. Wheels include the vendored DataFusion Python binding under Apache-2.0; see the third-party notices.
ParqDB builds on work from LanceDB, DataFusion, DuckDB, StarRocks, Apache Spark, and Apache Iceberg, with gratitude to their contributors and communities.
Python
50.5%
Rust
46.0%
TypeScript
2.4%