ankitjh4/bharat-government-documents

Dataset

Bharat Guide: screened Indian public information documents

10

5 commits

updated Oct 1, 2026

See the code

README

Bharat Guide: screened Indian public information documents

This snapshot contains 45,041 distinct normalized text bodies and 52,701 source records. Generated 2026-10-01T11:59:53.737883+00:00.

Contents and provenance

One Parquet row represents one normalized source body, with original extracted text, source title, URL, extraction quality, observed retrieval time, rule version, topic labels and all current-body source aliases. Duplicate URLs/editions are preserved as provenance, not counted as additional distinct documents. The canonical record is the lexicographically smallest eligible document identifier. Topic labels overlap.

Selection requires an eligible assessment whose body checksum still matches the current document, with full_text or ocr_text extraction. This is heuristic screening, not manual factual validation. Older and newer screening versions coexist and are reported in manifest.json. PTI secondary news, directory-only entries, summary-only records, rejected text and OCR awaiting review are not qualifying rows. A reviewed alias can be retained inside an eligible row's provenance, explicitly labelled. The text exports contain no raw response bytes or account credentials. The combined CSV labels screened and review documents explicitly; failure records are separate. The separate SQLite backup retains original evidence and operational records; its manifest records its own snapshot date.

Limitations

This is not a complete census of Indian websites, laws, schemes or facilities. Government-adjacent institutional publications are labelled separately from government sources. Retrieval time is not publication time. Source validity, amendments, legal force, policy deadlines and numerical OCR accuracy require checking the original. Exact normalized-body deduplication does not establish semantic equivalence.

Rights and use

Rights remain with each original publisher. Government provenance does not establish that every source is public domain or has one common reuse licence. No blanket licence is granted over collected source text. Consult each source's terms and applicable rights before redistribution or other reuse. This repository is initially private; publication scope can be changed by its owner after source-rights review.

Releases

manifest.json gives checksums and row counts. publication.json tracks confirmed publication content hashes and the additional-document threshold. Each update is committed atomically with its data and manifest; the next release is due after at least 10,000 newly qualifying distinct bodies, not 10,000 fetched URLs or chunks.

Download layout and CSV fields

documents.csv.gz is the main download: one deduplicated, nonempty extracted text body per row, combining screened and review content. The status column is good or review. Only good rows count toward the screened-document milestone. Review rows may include thin text, directory entries, news or incomplete extraction; keep them separate when evaluating quality. Duplicate source aliases are preserved in provenance/sources.csv.gz without repeating their text in the main CSV.

The CSV contains: content_hash, document_id, title, text, url, host, kind, trust, quality, state, category, fetched_at, metadata, sha256, status, assessment_version, assessment_status, assessment_reasons, topics. text is the full stored extracted body, not a summary. metadata preserves the representative document's SQLite JSON metadata. topics is a JSON list, and labels may overlap; unclassified documents are explicitly labelled. Fetch time is not publication time. A row represents a document body, not every relational SQLite row.

topics/document_topics.csv.gz maps topics to content hashes and good/review status; topics/summary.json provides counts. This index avoids copying full text into many folders. Topic labels are heuristic, not manual factual or legal validation.

fails/urls.csv.gz contains failure, robots-blocked, retry and OCR-pending records with their exact statuses; not every entry is a permanent failure. fails/empty_documents.csv.gz lists records without usable text. Neither is counted as a distinct text document. Unattempted queued URLs and seed inventories are not exported. No crawler handoff package is included.

CSV files use UTF-8, gzip compression and standard quoting for multiline text. Decompress to obtain an ordinary CSV, or read directly:

import pandas as pd
df = pd.read_csv("documents.csv.gz")
good = df[df["status"] == "good"]

Optional data/train-*.parquet files contain screened text in larger batches for the Hugging Face viewer. The CSV contains additional review content, so its row count is higher than the qualifying-document count. Source values are preserved; treat spreadsheet formulas in source text as untrusted, not executable instructions.

SQLite reference

SQLITE_SCHEMA.md lists every table and field; sqlite_schema.sql contains the reference schema. documents stores text and metadata; document_assessments stores screening decisions; document_topics stores labels. Original PDFs and responses are in raw_blobs.payload, linked through fetch_receipts.raw_sha256, with codec describing zlib compression or raw bytes. ocr_pages stores page-level results. These binary files and operational tables are not embedded in the CSV.

The private recovery files backups/bharat.sqlite and backups/sqlite-manifest.json retain their own snapshot date and checksum, which may precede the text release. The full SQLite backup includes crawl queues and source inventories even though this CSV release excludes them. Keep that backup private when publishing text-only data.

ankitjh4/bharat-government-documents

Dataset

Bharat Guide: screened Indian public information documents

10

5 commits

updated Oct 1, 2026

See the code

README

Bharat Guide: screened Indian public information documents

This snapshot contains 45,041 distinct normalized text bodies and 52,701 source records. Generated 2026-10-01T11:59:53.737883+00:00.

Contents and provenance

One Parquet row represents one normalized source body, with original extracted text, source title, URL, extraction quality, observed retrieval time, rule version, topic labels and all current-body source aliases. Duplicate URLs/editions are preserved as provenance, not counted as additional distinct documents. The canonical record is the lexicographically smallest eligible document identifier. Topic labels overlap.

Selection requires an eligible assessment whose body checksum still matches the current document, with full_text or ocr_text extraction. This is heuristic screening, not manual factual validation. Older and newer screening versions coexist and are reported in manifest.json. PTI secondary news, directory-only entries, summary-only records, rejected text and OCR awaiting review are not qualifying rows. A reviewed alias can be retained inside an eligible row's provenance, explicitly labelled. The text exports contain no raw response bytes or account credentials. The combined CSV labels screened and review documents explicitly; failure records are separate. The separate SQLite backup retains original evidence and operational records; its manifest records its own snapshot date.

Limitations

This is not a complete census of Indian websites, laws, schemes or facilities. Government-adjacent institutional publications are labelled separately from government sources. Retrieval time is not publication time. Source validity, amendments, legal force, policy deadlines and numerical OCR accuracy require checking the original. Exact normalized-body deduplication does not establish semantic equivalence.

Rights and use

Rights remain with each original publisher. Government provenance does not establish that every source is public domain or has one common reuse licence. No blanket licence is granted over collected source text. Consult each source's terms and applicable rights before redistribution or other reuse. This repository is initially private; publication scope can be changed by its owner after source-rights review.

Releases

manifest.json gives checksums and row counts. publication.json tracks confirmed publication content hashes and the additional-document threshold. Each update is committed atomically with its data and manifest; the next release is due after at least 10,000 newly qualifying distinct bodies, not 10,000 fetched URLs or chunks.

Download layout and CSV fields

documents.csv.gz is the main download: one deduplicated, nonempty extracted text body per row, combining screened and review content. The status column is good or review. Only good rows count toward the screened-document milestone. Review rows may include thin text, directory entries, news or incomplete extraction; keep them separate when evaluating quality. Duplicate source aliases are preserved in provenance/sources.csv.gz without repeating their text in the main CSV.

The CSV contains: content_hash, document_id, title, text, url, host, kind, trust, quality, state, category, fetched_at, metadata, sha256, status, assessment_version, assessment_status, assessment_reasons, topics. text is the full stored extracted body, not a summary. metadata preserves the representative document's SQLite JSON metadata. topics is a JSON list, and labels may overlap; unclassified documents are explicitly labelled. Fetch time is not publication time. A row represents a document body, not every relational SQLite row.

topics/document_topics.csv.gz maps topics to content hashes and good/review status; topics/summary.json provides counts. This index avoids copying full text into many folders. Topic labels are heuristic, not manual factual or legal validation.

fails/urls.csv.gz contains failure, robots-blocked, retry and OCR-pending records with their exact statuses; not every entry is a permanent failure. fails/empty_documents.csv.gz lists records without usable text. Neither is counted as a distinct text document. Unattempted queued URLs and seed inventories are not exported. No crawler handoff package is included.

CSV files use UTF-8, gzip compression and standard quoting for multiline text. Decompress to obtain an ordinary CSV, or read directly:

import pandas as pd
df = pd.read_csv("documents.csv.gz")
good = df[df["status"] == "good"]

Optional data/train-*.parquet files contain screened text in larger batches for the Hugging Face viewer. The CSV contains additional review content, so its row count is higher than the qualifying-document count. Source values are preserved; treat spreadsheet formulas in source text as untrusted, not executable instructions.

SQLite reference

SQLITE_SCHEMA.md lists every table and field; sqlite_schema.sql contains the reference schema. documents stores text and metadata; document_assessments stores screening decisions; document_topics stores labels. Original PDFs and responses are in raw_blobs.payload, linked through fetch_receipts.raw_sha256, with codec describing zlib compression or raw bytes. ocr_pages stores page-level results. These binary files and operational tables are not embedded in the CSV.

The private recovery files backups/bharat.sqlite and backups/sqlite-manifest.json retain their own snapshot date and checksum, which may precede the text release. The full SQLite backup includes crawl queues and source inventories even though this CSV release excludes them. Keep that backup private when publishing text-only data.