Importance Matrix Calibration Datasets
This repository provides calibration datasets used to generate importance matrices (imatrix), which are required to minimize errors when quantizing models with LLaMA C++.
The llama-imatrix program cannot handle parquet files directly and thus requires them to be converted into text format first. There are many ways to do this but a simple approach is to use DuckDB with the following command: duckdb -noheader -ascii -c "SELECT content FROM 'tools_micro.parquet';" > tools_micro.txt
Code calibration datasets
This dataset consists of cleaned and de-duplicated code prompts and is available in six sizes, ranging from huge (~ 200,000 lines equivalent to approx. 14.5M tokens), to micro (~ 6,200 lines and 2.4M tokens avg).
Original data sourced from Vezora/Open-Critic-GPT, OpenCoder-LLM/opc-sft-stage2, ise-uiuc/Magicoder-Evol-Instruct-110K, and Multilingual-Multimodal-NLP/McEval-Instruct
Math calibration datasets
This dataset consists of cleaned and de-duplicated math prompts and is available in six sizes, ranging from huge (~ 200,000 lines equivalent to approx. 6 million), to micro (~ 6,250 lines and 0.9 million tokens avg).
Original data sourced from nvidia/OpenMathInstruct-2
This dataset consists of cleaned and de-duplicated tool prompts and is available in six sizes, ranging from huge (~ 100,000 lines equivalent to approx. 10 million tokens), to micro (~ 3,100 lines and 1 million tokens).
Original data sourced from BitAgent/tool_calling and JungHun/Efficient_ToolCalling
Language calibration datasets
This dataset consists of cleaned and de-duplicated text prompts for 18 different languages. Each language file is available in five sizes, ranging from large (~ 25,000 lines equivalent to approx. 725K tokens), to micro (~ 1,600 lines and 125K tokens avg).
Original data sourced from HuggingFaceFW/fineweb, HuggingFaceFW/fineweb-2, and Common Crawl
| File | Language | Lines |
|---|
| text_ar_large | Arabic | 25,000 |
| text_ar_medium | Arabic | 12,500 |
| text_ar_small | Arabic | 6,250 |
| text_ar_tiny | Arabic | 3,125 |
| text_ar_micro | Arabic | 1,562 |
| text_cn_large | Chinese | 25,000 |
| text_cn_medium | Chinese | 12,500 |
| text_cn_small | Chinese | 6,250 |
| text_cn_tiny | Chinese | 3,125 |
| text_cn_micro | Chinese | 1,562 |
| text_de_large | German | 25,000 |
| text_de_medium | German | 12,500 |
| text_de_small | German | 6,250 |
| text_de_tiny | German | 3,125 |
| text_de_micro | German | 1,562 |
| text_en_large | English | 25,000 |
| text_en_medium | English | 12,500 |
| text_en_small | English | 6,250 |
| text_en_tiny | English | 3,125 |
| text_en_micro | English | 1,562 |
| text_es_large | Spanish | 25,000 |
| text_es_medium | Spanish | 12,500 |
| text_es_small | Spanish | 6,250 |
| text_es_tiny | Spanish | 3,125 |
| text_es_micro | Spanish | 1,562 |
| text_fr_large | French | 25,000 |
| text_fr_medium | French | 12,500 |
| text_fr_small | French | 6,250 |
| text_fr_tiny | French | 3,125 |
| text_fr_micro | French | 1,562 |
| text_hi_large | Hindi | 25,000 |
| text_hi_medium | Hindi | 12,500 |
| text_hi_small | Hindi | 6,250 |
| text_hi_tiny | Hindi | 3,125 |
| text_hi_micro | Hindi | 1,562 |
| text_id_large | Indonesian | 24,999 |
| text_id_medium | Indonesian | 12,500 |
| text_id_small | Indonesian | 6,250 |
| text_id_tiny | Indonesian | 3,125 |
| text_id_micro | Indonesian | 1,562 |
| text_it_large | Italian | 25,000 |
| text_it_medium | Italian | 12,500 |
| text_it_small | Italian | 6,250 |
| text_it_tiny | Italian | 3,125 |
| text_it_micro | Italian | 1,562 |
| text_jp_large | Japanese | 25,000 |
| text_jp_medium | Japanese | 12,500 |
| text_jp_small | Japanese | 6,250 |
| text_jp_tiny | Japanese | 3,125 |
| text_jp_micro | Japanese | 1,562 |
| text_mm_large | Burmese | 25,000 |
| text_mm_medium | Burmese | 12,500 |
| text_mm_small | Burmese | 6,250 |
| text_mm_tiny | Burmese | 3,125 |
| text_mm_micro | Burmese | 1,562 |
| text_nl_large | Dutch | 25,000 |
| text_nl_medium | Dutch | 12,500 |
| text_nl_small | Dutch | 6,250 |
| text_nl_tiny | Dutch | 3,125 |
| text_nl_micro | Dutch | 1,562 |
| text_ph_large | Filipino | 25,000 |
| text_ph_medium | Filipino | 12,500 |
| text_ph_small | Filipino | 6,250 |
| text_ph_tiny | Filipino | 3,125 |
| text_ph_micro | Filipino | 1,562 |
| text_pl_large | Polish | 25,000 |
| text_pl_medium | Polish | 12,500 |
| text_pl_small | Polish | 6,250 |
| text_pl_tiny | Polish | 3,125 |
| text_pl_micro | Polish | 1,562 |
| text_pt_large | Portuguese | 25,000 |
| text_pt_medium | Portuguese | 12,500 |
| text_pt_small | Portuguese | 6,250 |
| text_pt_tiny | Portuguese | 3,125 |
| text_pt_micro | Portuguese | 1,562 |
| text_ru_large | Russian | 25,000 |
| text_ru_medium | Russian | 12,500 |
| text_ru_small | Russian | 6,250 |
| text_ru_tiny | Russian | 3,125 |
| text_ru_micro | Russian | 1,562 |
| text_th_large | Thai | 25,000 |
| text_th_medium | Thai | 12,500 |
| text_th_small | Thai | 6,250 |
| text_th_tiny | Thai | 3,125 |
| text_th_micro | Thai | 1,562 |
| text_vn_large | Vietnamese | 25,000 |
| text_vn_medium | Vietnamese | 12,500 |
| text_vn_small | Vietnamese | 6,250 |
| text_vn_tiny | Vietnamese | 3,125 |
| text_vn_micro | Vietnamese | 1,562 |
Language groups
In addition to single language files, the dataset includes randomized and files by language family/region and all languages in dataset
All languages (all)
European languages: English, French, German, Italian, Portuguese & Spanish (eur)
Germanic languages: Dutch, English & German (gem)
Romance languages: French, Italian, Portuguese & Spanish (roa)
Rest of World: Arabic, Chinese, Hindi & Japanese (row)
Southeast Asia languages: Burmese, Filipino, Indonesian, Thai & Vietnamese (sea)
Slavic languages: Polish & Russian (sla)
Math & Code calibration datasets
This dataset combines math and code prompts into single calibration files.
This dataset combines tool, math, code and language prompts into single calibration files.
| File | Language | Lines |
|---|
| combined_ar_huge | Arabic | 100,000 |
| combined_ar_large | Arabic | 50,000 |
| combined_ar_medium | Arabic | 25,000 |
| combined_ar_small | Arabic | 12,500 |
| combined_ar_tiny | Arabic | 6,248 |
| combined_ar_micro | Arabic | 3,124 |
| combined_cn_huge | Chinese | 100,000 |
| combined_cn_large | Chinese | 50,000 |
| combined_cn_medium | Chinese | 25,000 |
| combined_cn_small | Chinese | 12,500 |
| combined_cn_tiny | Chinese | 6,248 |
| combined_cn_micro | Chinese | 3,124 |
| combined_de_huge | German | 100,000 |
| combined_de_large | German | 50,000 |
| combined_de_medium | German | 25,000 |
| combined_de_small | German | 12,500 |
| combined_de_tiny | German | 6,248 |
| combined_de_micro | German | 3,124 |
| combined_en_huge | English | 99,999 |
| combined_en_large | English | 50,000 |
| combined_en_medium | English | 25,000 |
| combined_en_small | English | 12,500 |
| combined_en_tiny | English | 6,248 |
| combined_en_micro | English | 3,124 |
| combined_es_huge | Spanish | 100,000 |
| combined_es_large | Spanish | 50,000 |
| combined_es_medium | Spanish | 25,000 |
| combined_es_small | Spanish | 12,500 |
| combined_es_tiny | Spanish | 6,248 |
| combined_es_micro | Spanish | 3,124 |
| combined_fr_huge | French | 100,000 |
| combined_fr_large | French | 50,000 |
| combined_fr_medium | French | 25,000 |
| combined_fr_small | French | 12,500 |
| combined_fr_tiny | French | 6,248 |
| combined_fr_micro | French | 3,124 |
| combined_hi_huge | Hindi | 100,000 |
| combined_hi_large | Hindi | 50,000 |
| combined_hi_medium | Hindi | 25,000 |
| combined_hi_small | Hindi | 12,500 |
| combined_hi_tiny | Hindi | 6,248 |
| combined_hi_micro | Hindi | 3,124 |
| combined_id_huge | Indonesian | 99,999 |
| combined_id_large | Indonesian | 50,000 |
| combined_id_medium | Indonesian | 25,000 |
| combined_id_small | Indonesian | 12,500 |
| combined_id_tiny | Indonesian | 6,248 |
| combined_id_micro | Indonesian | 3,124 |
| combined_it_huge | Italian | 100,000 |
| combined_it_large | Italian | 50,000 |
| combined_it_medium | Italian | 25,000 |
| combined_it_small | Italian | 12,500 |
| combined_it_tiny | Italian | 6,248 |
| combined_it_micro | Italian | 3,124 |
| combined_jp_huge | Japanese | 100,000 |
| combined_jp_large | Japanese | 50,000 |
| combined_jp_medium | Japanese | 25,000 |
| combined_jp_small | Japanese | 12,500 |
| combined_jp_tiny | Japanese | 6,248 |
| combined_jp_micro | Japanese | 3,124 |
| combined_mm_huge | Burmese | 100,000 |
| combined_mm_large | Burmese | 50,000 |
| combined_mm_medium | Burmese | 25,000 |
| combined_mm_small | Burmese | 12,500 |
| combined_mm_tiny | Burmese | 6,248 |
| combined_mm_micro | Burmese | 3,124 |
| combined_nl_huge | Dutch | 100,000 |
| combined_nl_large | Dutch | 50,000 |
| combined_nl_medium | Dutch | 25,000 |
| combined_nl_small | Dutch | 12,500 |
| combined_nl_tiny | Dutch | 6,248 |
| combined_nl_micro | Dutch | 3,124 |
| combined_ph_huge | Filipino | 100,000 |
| combined_ph_large | Filipino | 49,999 |
| combined_ph_medium | Filipino | 25,000 |
| combined_ph_small | Filipino | 12,500 |
| combined_ph_tiny | Filipino | 6,248 |
| combined_ph_micro | Filipino | 3,124 |
| combined_pl_huge | Polish | 100,000 |
| combined_pl_large | Polish | 50,000 |
| combined_pl_medium | Polish | 25,000 |
| combined_pl_small | Polish | 12,500 |
| combined_pl_tiny | Polish | 6,248 |
| combined_pl_micro | Polish | 3,124 |
| combined_pt_huge | Portuguese | 100,000 |
| combined_pt_large | Portuguese | 50,000 |
| combined_pt_medium | Portuguese | 25,000 |
| combined_pt_small | Portuguese | 12,500 |
| combined_pt_tiny | Portuguese | 6,248 |
| combined_pt_micro | Portuguese | 3,124 |
| combined_ru_huge | Russian | 99,999 |
| combined_ru_large | Russian | 50,000 |
| combined_ru_medium | Russian | 25,000 |
| combined_ru_small | Russian | 12,500 |
| combined_ru_tiny | Russian | 6,248 |
| combined_ru_micro | Russian | 3,124 |
| combined_th_huge | Thai | 100,000 |
| combined_th_large | Thai | 50,000 |
| combined_th_medium | Thai | 25,000 |
| combined_th_small | Thai | 12,500 |
| combined_th_tiny | Thai | 6,248 |
| combined_th_micro | Thai | 3,124 |
| combined_vn_huge | Vietnamese | 99,999 |
| combined_vn_large | Vietnamese | 50,000 |
| combined_vn_medium | Vietnamese | 25,000 |
| combined_vn_small | Vietnamese | 12,499 |
| combined_vn_tiny | Vietnamese | 6,248 |
| combined_vn_micro | Vietnamese | 3,124 |
In addition to single tool, math, code and language files, the dataset includes combined and randomized files by language family/region and all languages in dataset
All languages (all)
European languages: English, French, German, Italian, Portuguese & Spanish (eur)
Germanic languages: Dutch, English & German (gem)
Romance languages: French, Italian, Portuguese & Spanish (roa)
Rest of World: Arabic, Chinese, Hindi & Japanese (row)
Southeast Asia languages: Burmese, Filipino, Indonesian, Thai & Vietnamese (sea)
Slavic languages: Polish & Russian (sla)