वागर्थाविव संपृक्तौ वागर्थप्रतिपत्तये ।
जगतः पितरौ वन्दे पार्वतीपरमेश्वरौ ॥"United as word and meaning are united, I bow to the parents of the world,
Pārvatī and Parameśvara, that I may attain an understanding of word and meaning."— Kālidāsa, Raghuvaṃśa 1.1
Vāgartha — vāk (word) and artha (meaning) — is a corpus of 217,959 Sanskrit verses, each paired with a detailed, structured explanation in English. The name is taken from the invocation above, in which Kālidāsa talks about the inseparability of word and meaning; the dataset's two substantive columns.
The verses span roughly three millennia of Sanskrit literature: the Mahābhārata and Rāmāyaṇa, the eighteen Mahāpurāṇas, the Vedas and Brāhmaṇas, the Upaniṣads and Darśanas, Āyurveda and the Nāṭyaśāstra, kāvya and nāṭaka, lexicons, and Jain, Buddhist and Sikh scripture.
| Field | Type | Description |
|---|---|---|
source | string | Provenance of the verse, as a /-delimited taxonomy path or a text name (see Source naming) |
shloka | string | The verse itself, in Devanāgarī. Mean 91 characters |
explanation | string | A structured explanation in Markdown, in English. Mean 4,041 characters |
Single train split. 217,959 rows, ~901M characters, 562 MB as Parquet.
explanationEvery explanation has three things in order, so the field has a consistent internal structure:
source: mahabharata
shloka: तथैव स गिरिभूर्यः प्रपुष्पितलताद्रुमः।
सपक्षिगणसंघुष्टः सश्वापदसरीसृपः॥
### 1. Word by word translation
* **तथैव** (tathaiva): In that very way; Likewise; Just so.
(A compound of `tathā` "thus" + `eva` "indeed/very").
* **स** (sa): That.
* **गिरिभूर्यः** (giribhūryaḥ): The great mountain; the chief of mountains.
(A compound of `giri` "mountain" + `bhūryaḥ` "chief/great").
* **प्रपुष्पितलताद्रुमः** (prapuṣpitalatādrumaḥ): (One) whose creepers (`latā`)
and trees (`druma`) were in full bloom (`prapuṣpita`).
...
### 2. Exact semantic translation
"In that very way, that great mountain had its creepers and trees in full bloom,
was resounding with the calls of flocks of birds, and was inhabited by wild
beasts and reptiles."
### 3. Expanded translation with additional background and details
... The verse uses a series of long compound adjectives (known as *Bahuvrīhi*
compounds in Sanskrit grammar) to paint a holistic and vivid picture of the
mountain. Each adjective adds a layer to the description, engaging different
senses. ...
| Category | Rows | Share | Sources |
|---|---|---|---|
| Mahābhārata | 73,431 | 33.7% | 3 |
| Purāṇas (18 Mahāpurāṇas) | 70,795 | 32.5% | 34 |
| Vālmīki Rāmāyaṇa | 18,280 | 8.4% | 1 |
| Upavedas (Āyurveda, Nāṭyaśāstra, …) | 14,833 | 6.8% | 48 |
| Upapurāṇas | 9,419 | 4.3% | 11 |
| Vedas (Saṃhitās) | 5,384 | 2.5% | 36 |
| Darśanas (philosophy) | 4,591 | 2.1% | 42 |
| Gītās | 4,133 | 1.9% | 84 |
| Kāvya & Nāṭaka | 4,080 | 1.9% | 66 |
| Scientific literature | 3,535 | 1.6% | 7 |
| Lexicons (koṣa) | 2,967 | 1.4% | 9 |
| Brāhmaṇas | 1,729 | 0.8% | 19 |
| Vedāṅgas (grammar, śikṣā, nirukta) | 1,487 | 0.7% | 37 |
| Smṛtis | 1,237 | 0.6% | 3 |
| Upaniṣads | 841 | 0.4% | 30 |
| Āraṇyakas | 505 | 0.2% | 5 |
| Jain / Buddhist / Sikh scripture | 382 | 0.2% | 15 |
| Vedic-studies papers & misc. | 330 | 0.2% | 80 |
| Total | 217,959 | 530 |
The distribution is heavily skewed: the Mahābhārata and the Mahāpurāṇas together account for two thirds of all rows. Weight or subsample accordingly.
source values are normalised. 371 of the 530 sources are /-delimited taxonomy
paths, from broad category down to a specific chapter or volume:
upaveda/natyashastra/translation-with-chandrika-notes-by-dr-sudhakar-malaviya/vol-iii
upanishad/main-upanishad/kena/translation-based-on-sanakaras-commentary/by-swami-gambhirananda
vedas/yajur-veda/krishna-yajur-veda/taittiriya/vol-iii/part-i
puranas-18-puranas-mahapurana/garud-puran/garuda-vol-2
The remaining 159 are flat names for standalone texts (mahabharata,
valmiki-ramayana, Charakasamhita, Taittiriya_Brahmana_Vol_I), including the
Vedic-studies seminar papers, which are named by their title in the language they
were written in (DHARANA_A_YOGIC_SCIENCE, वैदिकं_विज्ञानम्).
This makes coarse filtering a prefix match:
puranas = ds.filter(lambda r: r["source"].startswith("puranas-18-puranas-mahapurana/"))
ayurveda = ds.filter(lambda r: r["source"].startswith("upaveda/ayurveda/"))
Verses were extracted from digitised editions of the source texts. Explanations were then generated verse-by-verse with gemini-2.5-pro. We run rudimentary checks of the translations, but errors may exist.
from datasets import load_dataset
ds = load_dataset("sarvamai/vagartha", split="train")
print(ds[0]["shloka"])
print(ds[0]["explanation"])
Streaming, for the 562 MB you may not want to download:
ds = load_dataset("sarvamai/vagartha", split="train", streaming=True)
for row in ds.take(5):
print(row["source"], row["shloka"])
Please read this section before training on the corpus.
source labels are not always right. Verse attribution is inherited from the
extraction pipeline, and the pipeline is imperfect. In the very example quoted
above, source says mahabharata while the explanation identifies the verse
as Rāmāyaṇa, Kiṣkindhā Kāṇḍa — and the explanation is correct. Treat source
as a strong hint, not a citation.BhG 2.47), only the source path. Verse order within a source is not
guaranteed to be reading order.Released under CC BY 4.0.
The underlying verses are classical works long in the public domain. The explanations are machine-generated derivative text, released under the same terms. Individual source editions and their modern commentaries may carry their own rights; this release covers the verse text and the generated explanations only.
@misc{vagartha2026,
title = {V\={a}gartha: Sanskrit Verses with Structured Explanations},
author = {Sarvam AI},
year = {2026},
url = {https://huggingface.co/datasets/sarvamai/vagartha}
}
5 commits
वागर्थाविव संपृक्तौ वागर्थप्रतिपत्तये ।
जगतः पितरौ वन्दे पार्वतीपरमेश्वरौ ॥"United as word and meaning are united, I bow to the parents of the world,
Pārvatī and Parameśvara, that I may attain an understanding of word and meaning."— Kālidāsa, Raghuvaṃśa 1.1
Vāgartha — vāk (word) and artha (meaning) — is a corpus of 217,959 Sanskrit verses, each paired with a detailed, structured explanation in English. The name is taken from the invocation above, in which Kālidāsa talks about the inseparability of word and meaning; the dataset's two substantive columns.
The verses span roughly three millennia of Sanskrit literature: the Mahābhārata and Rāmāyaṇa, the eighteen Mahāpurāṇas, the Vedas and Brāhmaṇas, the Upaniṣads and Darśanas, Āyurveda and the Nāṭyaśāstra, kāvya and nāṭaka, lexicons, and Jain, Buddhist and Sikh scripture.
| Field | Type | Description |
|---|---|---|
source | string | Provenance of the verse, as a /-delimited taxonomy path or a text name (see Source naming) |
shloka | string | The verse itself, in Devanāgarī. Mean 91 characters |
explanation | string | A structured explanation in Markdown, in English. Mean 4,041 characters |
Single train split. 217,959 rows, ~901M characters, 562 MB as Parquet.
explanationEvery explanation has three things in order, so the field has a consistent internal structure:
source: mahabharata
shloka: तथैव स गिरिभूर्यः प्रपुष्पितलताद्रुमः।
सपक्षिगणसंघुष्टः सश्वापदसरीसृपः॥
### 1. Word by word translation
* **तथैव** (tathaiva): In that very way; Likewise; Just so.
(A compound of `tathā` "thus" + `eva` "indeed/very").
* **स** (sa): That.
* **गिरिभूर्यः** (giribhūryaḥ): The great mountain; the chief of mountains.
(A compound of `giri` "mountain" + `bhūryaḥ` "chief/great").
* **प्रपुष्पितलताद्रुमः** (prapuṣpitalatādrumaḥ): (One) whose creepers (`latā`)
and trees (`druma`) were in full bloom (`prapuṣpita`).
...
### 2. Exact semantic translation
"In that very way, that great mountain had its creepers and trees in full bloom,
was resounding with the calls of flocks of birds, and was inhabited by wild
beasts and reptiles."
### 3. Expanded translation with additional background and details
... The verse uses a series of long compound adjectives (known as *Bahuvrīhi*
compounds in Sanskrit grammar) to paint a holistic and vivid picture of the
mountain. Each adjective adds a layer to the description, engaging different
senses. ...
| Category | Rows | Share | Sources |
|---|---|---|---|
| Mahābhārata | 73,431 | 33.7% | 3 |
| Purāṇas (18 Mahāpurāṇas) | 70,795 | 32.5% | 34 |
| Vālmīki Rāmāyaṇa | 18,280 | 8.4% | 1 |
| Upavedas (Āyurveda, Nāṭyaśāstra, …) | 14,833 | 6.8% | 48 |
| Upapurāṇas | 9,419 | 4.3% | 11 |
| Vedas (Saṃhitās) | 5,384 | 2.5% | 36 |
| Darśanas (philosophy) | 4,591 | 2.1% | 42 |
| Gītās | 4,133 | 1.9% | 84 |
| Kāvya & Nāṭaka | 4,080 | 1.9% | 66 |
| Scientific literature | 3,535 | 1.6% | 7 |
| Lexicons (koṣa) | 2,967 | 1.4% | 9 |
| Brāhmaṇas | 1,729 | 0.8% | 19 |
| Vedāṅgas (grammar, śikṣā, nirukta) | 1,487 | 0.7% | 37 |
| Smṛtis | 1,237 | 0.6% | 3 |
| Upaniṣads | 841 | 0.4% | 30 |
| Āraṇyakas | 505 | 0.2% | 5 |
| Jain / Buddhist / Sikh scripture | 382 | 0.2% | 15 |
| Vedic-studies papers & misc. | 330 | 0.2% | 80 |
| Total | 217,959 | 530 |
The distribution is heavily skewed: the Mahābhārata and the Mahāpurāṇas together account for two thirds of all rows. Weight or subsample accordingly.
source values are normalised. 371 of the 530 sources are /-delimited taxonomy
paths, from broad category down to a specific chapter or volume:
upaveda/natyashastra/translation-with-chandrika-notes-by-dr-sudhakar-malaviya/vol-iii
upanishad/main-upanishad/kena/translation-based-on-sanakaras-commentary/by-swami-gambhirananda
vedas/yajur-veda/krishna-yajur-veda/taittiriya/vol-iii/part-i
puranas-18-puranas-mahapurana/garud-puran/garuda-vol-2
The remaining 159 are flat names for standalone texts (mahabharata,
valmiki-ramayana, Charakasamhita, Taittiriya_Brahmana_Vol_I), including the
Vedic-studies seminar papers, which are named by their title in the language they
were written in (DHARANA_A_YOGIC_SCIENCE, वैदिकं_विज्ञानम्).
This makes coarse filtering a prefix match:
puranas = ds.filter(lambda r: r["source"].startswith("puranas-18-puranas-mahapurana/"))
ayurveda = ds.filter(lambda r: r["source"].startswith("upaveda/ayurveda/"))
Verses were extracted from digitised editions of the source texts. Explanations were then generated verse-by-verse with gemini-2.5-pro. We run rudimentary checks of the translations, but errors may exist.
from datasets import load_dataset
ds = load_dataset("sarvamai/vagartha", split="train")
print(ds[0]["shloka"])
print(ds[0]["explanation"])
Streaming, for the 562 MB you may not want to download:
ds = load_dataset("sarvamai/vagartha", split="train", streaming=True)
for row in ds.take(5):
print(row["source"], row["shloka"])
Please read this section before training on the corpus.
source labels are not always right. Verse attribution is inherited from the
extraction pipeline, and the pipeline is imperfect. In the very example quoted
above, source says mahabharata while the explanation identifies the verse
as Rāmāyaṇa, Kiṣkindhā Kāṇḍa — and the explanation is correct. Treat source
as a strong hint, not a citation.BhG 2.47), only the source path. Verse order within a source is not
guaranteed to be reading order.Released under CC BY 4.0.
The underlying verses are classical works long in the public domain. The explanations are machine-generated derivative text, released under the same terms. Individual source editions and their modern commentaries may carry their own rights; this release covers the verse text and the generated explanations only.
@misc{vagartha2026,
title = {V\={a}gartha: Sanskrit Verses with Structured Explanations},
author = {Sarvam AI},
year = {2026},
url = {https://huggingface.co/datasets/sarvamai/vagartha}
}
5 commits