Bharat Guide: screened Indian public information documents
10
5 commits
updated Oct 1, 2026
This snapshot contains 45,041 distinct normalized text bodies and 52,701 source records. Generated 2026-10-01T11:59:53.737883+00:00.
One Parquet row represents one normalized source body, with original extracted text, source title, URL, extraction quality, observed retrieval time, rule version, topic labels and all current-body source aliases. Duplicate URLs/editions are preserved as provenance, not counted as additional distinct documents. The canonical record is the lexicographically smallest eligible document identifier. Topic labels overlap.
Selection requires an eligible assessment whose body checksum still matches the current document, with full_text or ocr_text extraction. This is heuristic screening, not manual factual validation. Older and newer screening versions coexist and are reported in manifest.json. PTI secondary news, directory-only entries, summary-only records, rejected text and OCR awaiting review are not qualifying rows. A reviewed alias can be retained inside an eligible row's provenance, explicitly labelled. The text exports contain no raw response bytes or account credentials. The combined CSV labels screened and review documents explicitly; failure records are separate. The separate SQLite backup retains original evidence and operational records; its manifest records its own snapshot date.
This is not a complete census of Indian websites, laws, schemes or facilities. Government-adjacent institutional publications are labelled separately from government sources. Retrieval time is not publication time. Source validity, amendments, legal force, policy deadlines and numerical OCR accuracy require checking the original. Exact normalized-body deduplication does not establish semantic equivalence.
Rights remain with each original publisher. Government provenance does not establish that every source is public domain or has one common reuse licence. No blanket licence is granted over collected source text. Consult each source's terms and applicable rights before redistribution or other reuse. This repository is initially private; publication scope can be changed by its owner after source-rights review.
manifest.json gives checksums and row counts. publication.json tracks confirmed publication content hashes and the additional-document threshold. Each update is committed atomically with its data and manifest; the next release is due after at least 10,000 newly qualifying distinct bodies, not 10,000 fetched URLs or chunks.
documents.csv.gz is the main download: one deduplicated, nonempty extracted text
body per row, combining screened and review content. The status column is good
or review. Only good rows count toward the screened-document milestone.
Review rows may include thin text, directory entries, news or incomplete extraction;
keep them separate when evaluating quality. Duplicate source aliases are preserved
in provenance/sources.csv.gz without repeating their text in the main CSV.
The CSV contains: content_hash, document_id, title, text, url, host,
kind, trust, quality, state, category, fetched_at, metadata, sha256,
status, assessment_version, assessment_status, assessment_reasons, topics.
text is the full stored extracted body, not a summary. metadata preserves the
representative document's SQLite JSON metadata. topics is a JSON list, and labels
may overlap; unclassified documents are explicitly labelled. Fetch time is not
publication time. A row represents a document body, not every relational SQLite row.
topics/document_topics.csv.gz maps topics to content hashes and good/review status;
topics/summary.json provides counts. This index avoids copying full text into many
folders. Topic labels are heuristic, not manual factual or legal validation.
fails/urls.csv.gz contains failure, robots-blocked, retry and OCR-pending records
with their exact statuses; not every entry is a permanent failure.
fails/empty_documents.csv.gz lists records without usable text. Neither is counted
as a distinct text document. Unattempted queued URLs and seed inventories are not
exported. No crawler handoff package is included.
CSV files use UTF-8, gzip compression and standard quoting for multiline text. Decompress to obtain an ordinary CSV, or read directly:
import pandas as pd
df = pd.read_csv("documents.csv.gz")
good = df[df["status"] == "good"]
Optional data/train-*.parquet files contain screened text in larger batches for
the Hugging Face viewer. The CSV contains additional review content, so its row
count is higher than the qualifying-document count. Source values are preserved;
treat spreadsheet formulas in source text as untrusted, not executable instructions.
SQLITE_SCHEMA.md lists every table and field; sqlite_schema.sql
contains the reference schema. documents stores text and metadata;
document_assessments stores screening decisions; document_topics stores labels.
Original PDFs and responses are in raw_blobs.payload, linked through
fetch_receipts.raw_sha256, with codec describing zlib compression or raw bytes.
ocr_pages stores page-level results. These binary files and operational tables
are not embedded in the CSV.
The private recovery files backups/bharat.sqlite and
backups/sqlite-manifest.json retain their own snapshot date and checksum, which
may precede the text release. The full SQLite backup includes crawl queues and
source inventories even though this CSV release excludes them. Keep that backup
private when publishing text-only data.
Bharat Guide: screened Indian public information documents
10
5 commits
updated Oct 1, 2026
This snapshot contains 45,041 distinct normalized text bodies and 52,701 source records. Generated 2026-10-01T11:59:53.737883+00:00.
One Parquet row represents one normalized source body, with original extracted text, source title, URL, extraction quality, observed retrieval time, rule version, topic labels and all current-body source aliases. Duplicate URLs/editions are preserved as provenance, not counted as additional distinct documents. The canonical record is the lexicographically smallest eligible document identifier. Topic labels overlap.
Selection requires an eligible assessment whose body checksum still matches the current document, with full_text or ocr_text extraction. This is heuristic screening, not manual factual validation. Older and newer screening versions coexist and are reported in manifest.json. PTI secondary news, directory-only entries, summary-only records, rejected text and OCR awaiting review are not qualifying rows. A reviewed alias can be retained inside an eligible row's provenance, explicitly labelled. The text exports contain no raw response bytes or account credentials. The combined CSV labels screened and review documents explicitly; failure records are separate. The separate SQLite backup retains original evidence and operational records; its manifest records its own snapshot date.
This is not a complete census of Indian websites, laws, schemes or facilities. Government-adjacent institutional publications are labelled separately from government sources. Retrieval time is not publication time. Source validity, amendments, legal force, policy deadlines and numerical OCR accuracy require checking the original. Exact normalized-body deduplication does not establish semantic equivalence.
Rights remain with each original publisher. Government provenance does not establish that every source is public domain or has one common reuse licence. No blanket licence is granted over collected source text. Consult each source's terms and applicable rights before redistribution or other reuse. This repository is initially private; publication scope can be changed by its owner after source-rights review.
manifest.json gives checksums and row counts. publication.json tracks confirmed publication content hashes and the additional-document threshold. Each update is committed atomically with its data and manifest; the next release is due after at least 10,000 newly qualifying distinct bodies, not 10,000 fetched URLs or chunks.
documents.csv.gz is the main download: one deduplicated, nonempty extracted text
body per row, combining screened and review content. The status column is good
or review. Only good rows count toward the screened-document milestone.
Review rows may include thin text, directory entries, news or incomplete extraction;
keep them separate when evaluating quality. Duplicate source aliases are preserved
in provenance/sources.csv.gz without repeating their text in the main CSV.
The CSV contains: content_hash, document_id, title, text, url, host,
kind, trust, quality, state, category, fetched_at, metadata, sha256,
status, assessment_version, assessment_status, assessment_reasons, topics.
text is the full stored extracted body, not a summary. metadata preserves the
representative document's SQLite JSON metadata. topics is a JSON list, and labels
may overlap; unclassified documents are explicitly labelled. Fetch time is not
publication time. A row represents a document body, not every relational SQLite row.
topics/document_topics.csv.gz maps topics to content hashes and good/review status;
topics/summary.json provides counts. This index avoids copying full text into many
folders. Topic labels are heuristic, not manual factual or legal validation.
fails/urls.csv.gz contains failure, robots-blocked, retry and OCR-pending records
with their exact statuses; not every entry is a permanent failure.
fails/empty_documents.csv.gz lists records without usable text. Neither is counted
as a distinct text document. Unattempted queued URLs and seed inventories are not
exported. No crawler handoff package is included.
CSV files use UTF-8, gzip compression and standard quoting for multiline text. Decompress to obtain an ordinary CSV, or read directly:
import pandas as pd
df = pd.read_csv("documents.csv.gz")
good = df[df["status"] == "good"]
Optional data/train-*.parquet files contain screened text in larger batches for
the Hugging Face viewer. The CSV contains additional review content, so its row
count is higher than the qualifying-document count. Source values are preserved;
treat spreadsheet formulas in source text as untrusted, not executable instructions.
SQLITE_SCHEMA.md lists every table and field; sqlite_schema.sql
contains the reference schema. documents stores text and metadata;
document_assessments stores screening decisions; document_topics stores labels.
Original PDFs and responses are in raw_blobs.payload, linked through
fetch_receipts.raw_sha256, with codec describing zlib compression or raw bytes.
ocr_pages stores page-level results. These binary files and operational tables
are not embedded in the CSV.
The private recovery files backups/bharat.sqlite and
backups/sqlite-manifest.json retain their own snapshot date and checksum, which
may precede the text release. The full SQLite backup includes crawl queues and
source inventories even though this CSV release excludes them. Keep that backup
private when publishing text-only data.