Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset

Dataset

๐Ÿ“– The Open Distillation Codex

207

500 commits

updated Aug 29, 2026

See the code

README

Version Storage Sources License Samples Cybersecurity



๐Ÿ“– The Open Distillation Codex

๐ŸŒŒ The Ultimate Open-Source Distillation Dataset โ€” No Skip, Full, with Attack & Defense ๐ŸŒŒ

Where 73 open-source minds converge into one unified stream of intelligence

18M+ Distilled Signals ยท 7,090 Raw GitHub Repositories ยท 8 Curated Categories ยท ~76 GB+


"We did not write this dataset. We assembled it. Every line is an echo โ€” of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three sources. Eight categories. Zero gatekeeping. No skipping. Fully processed. Now fortified with real-world cybersecurity confrontations."



๐Ÿ“Œ Table of Contents

#SectionDescription
1๐Ÿ“Š Dataset SummaryHigh-level overview & value proposition
2๐Ÿ—‚๏ธ Directory StructureASCII tree + folder explanation
3๐ŸŒ Data SourcesAll 73 sources with full attribution
4๐Ÿ›ก๏ธ Cybersecurity Deep Dive: Attack & DefenseImportance, attack traces, defense, exploit analysis
5๐Ÿ› ๏ธ How to Use & TrainLoading, streaming, training scripts
6๐Ÿ” Licensing & LimitationsLicense, intended use, limitations
7๐Ÿ“œ ChangelogVersion history

๐Ÿ“Š Dataset Summary

๐ŸŽฏ The Numbers That Matter

MetricValueStatus
Total Storage76 GB+โœ… Verified
JSONL Data Shards516โœ… Verified
Archive Files (tar.gz)7,090โœ… Verified
Source Datasets73โœ… Verified
Categories8โœ… Verified
Total Samples18M+โœ… Verified
Largest Source8.15M (Vibe-Coding-Instruct-V2)โœ…
Archive Size~64 GB (compressed GitHub repos)โœ…
Cybersecurity Sources6โœ…
Cybersecurity Data Size~2.6 GBโœ…

๐ŸŒŸ Why "Ultimate Distilled"?

This dataset is not a raw scrape. Every sample has been distilled through a unified extraction pipeline:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    UNIFIED EXTRACTION PIPELINE              โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                                             โ”‚
โ”‚  73 Upstream Sources (ALL FULLY PROCESSED, NO SKIP)         โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ”          โ”‚
โ”‚  โ”‚ HF  โ”‚ โ”‚ HF  โ”‚ โ”‚ HF  โ”‚ โ”‚ GH  โ”‚ โ”‚ HF  โ”‚ โ”‚ ... โ”‚          โ”‚
โ”‚  โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜          โ”‚
โ”‚     โ”‚       โ”‚       โ”‚       โ”‚       โ”‚       โ”‚               โ”‚
โ”‚     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜               โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚ EXTRACT โ”‚ โ† Field normalization        โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜   (instruction/response)      โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚CATEGORIZEโ”‚ โ† 8 semantic categories      โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚  SHARD  โ”‚ โ† 20K samples per shard       โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚ UPLOAD  โ”‚ โ† Batch commits to HF         โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                                                             โ”‚
โ”‚  STATUS: ALL 73 SOURCES COMPLETE. NO SKIPPING. 18M+ ROWS.   โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿ’Ž Value to the Open-Source AI Community

๐ŸŽฏ For...๐Ÿ“ฆ This dataset provides...
Model TrainersSingle load_dataset() call to stream 18M+ SFT-ready samples
Coding Agent Researchers11M+ agentic coding traces from Fable-5, Vibe-Coding, Royal Ghost, Kimi, DeepSeek
Code Pretraining7,090 full GitHub repository snapshots (64 GB compressed)
Reasoning Researchers2.7M+ distilled reasoning traces from Claude, Gemini, Grok, GPT-5.5, Opus 4.8
Domain Specialists25K-sample sweeps across 29 disciplines
Cybersecurity ResearchersDedicated cybersecurity category with attack/defense/exploit traces, red/blue team dialogues, and incident reports
Red Team / Blue Team TrainersRealistic attack scenarios, defense strategies, exploit code, and post-mortem analysis

๐Ÿ—‚๏ธ Directory Structure

๐Ÿ“‚ Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset/
โ”‚
โ”œโ”€โ”€ ๐Ÿ“ฆ archives/                          # ~64 GB โ€” 7,090 compressed GitHub repos
โ”‚   โ”œโ”€โ”€ 0-chi__sonaure-lp.tar.gz
โ”‚   โ”œโ”€โ”€ 00MB__bitcoin_trading_bot.tar.gz
โ”‚   โ”œโ”€โ”€ 0101-agents__plugins.tar.gz
โ”‚   โ”œโ”€โ”€ ... (7,090 files total)
โ”‚   โ””โ”€โ”€ zznmg1__playable-survivor-ad.tar.gz
โ”‚
โ”œโ”€โ”€ ๐Ÿ“ data/                              # ~12 GB โ€” 516 JSONL shards (18M+ samples)
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ’ป coding/                        # 28 sources ยท ~11M+ samples
โ”‚   โ”‚   โ”œโ”€โ”€ vibe_instruct_v2/             # 8,152,510 samples
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_2m/                    # 2,006,487 samples
โ”‚   โ”‚   โ”œโ”€โ”€ vibe_instruct_v1/             # 1,100,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ vibe_coding/                  # 1,100,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ royal_ghost_1m/               # 1,000,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ citation_ground/              # 980,064 samples
โ”‚   โ”‚   โ”œโ”€โ”€ royal_ghost_501k/             # 703,449 samples
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_repos_full/            # 7,090 archive pointers
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_agentic_sft/           # 159,972 samples
โ”‚   โ”‚   โ”œโ”€โ”€ gpt55_codex/                  # 119,436 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ alpca_gpt55/                  # 49,099 samples
โ”‚   โ”‚   โ”œโ”€โ”€ deepseek_v4_pro_agent/        # 96,597 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_traces/                # 49,544 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code_100k/            # 68,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code/                 # 49,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ kimi_coding/                  # 9,014 samples
โ”‚   โ”‚   โ”œโ”€โ”€ mimo_claude_code_traces/      # 15,046 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ kimi_k26_claude_code_traces/  # 7,438 samples
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code_10k/             # 9,800 samples
โ”‚   โ”‚   โ”œโ”€โ”€ legend_python/                # 5,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ autonomy/                     # 10,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code_demo/            # 1,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ god_coder/                    # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ python_god_coder/             # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ elite_god_coder/              # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ omega_genesis/                # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ open_tool_trace/              # 48 samples
โ”‚   โ”‚   โ””โ”€โ”€ genesis_v11/                  # partial recovery
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿงฎ math/                          # 2 sources
โ”‚   โ”‚   โ”œโ”€โ”€ math_25k/
โ”‚   โ”‚   โ””โ”€โ”€ deepseek_prover_v1/           # 27,503 Lean theorem proofs
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ”ฌ science/                       # 7 sources
โ”‚   โ”‚   โ”œโ”€โ”€ science_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ physics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ chemistry_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ biology_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ medical_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ cs_25k/
โ”‚   โ”‚   โ””โ”€โ”€ biology_r2med/                # โญ NEW
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ โš™๏ธ applied/                       # 8 sources
โ”‚   โ”‚   โ”œโ”€โ”€ robotics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ nano_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ materials_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ earth_climate_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ renewable_energy_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ evolution_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ universe_25k/
โ”‚   โ”‚   โ””โ”€โ”€ kardashev_25k/
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ“š humanities/                    # 8 sources
โ”‚   โ”‚   โ”œโ”€โ”€ psychology_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ economics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ law_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ statistics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ sports_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ human_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ conscience_25k/
โ”‚   โ”‚   โ””โ”€โ”€ supernatural_25k/
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿง  distilled/                     # 9 sources ยท frontier distillations
โ”‚   โ”‚   โ”œโ”€โ”€ claude_mythos/
โ”‚   โ”‚   โ”œโ”€โ”€ gemini35/
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_cleaned/
โ”‚   โ”‚   โ”œโ”€โ”€ grok44/
โ”‚   โ”‚   โ”œโ”€โ”€ gemini_pro32/
โ”‚   โ”‚   โ”œโ”€โ”€ gpt55_thinking/
โ”‚   โ”‚   โ”œโ”€โ”€ gpt55_distilled/
โ”‚   โ”‚   โ”œโ”€โ”€ claude_opus_48_distill/       # โญ NEW
โ”‚   โ”‚   โ””โ”€โ”€ claude_opus_48_max_thinking/  # โญ NEW
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ“ instruction/                   # 3 sources
โ”‚   โ”‚   โ”œโ”€โ”€ alpaca/                       # 52,002 samples
โ”‚   โ”‚   โ”œโ”€โ”€ oasst/                        # 32,141 samples
โ”‚   โ”‚   โ””โ”€โ”€ dolly/                        # 15,011 samples
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ”’ cybersecurity/                 # 6 sources
โ”‚   โ”‚   โ”œโ”€โ”€ high_quality_cybersecurity/
โ”‚   โ”‚   โ”œโ”€โ”€ heimdall_v1_1/                # โญ NEW โ€” 78 MB conversations
โ”‚   โ”‚   โ”œโ”€โ”€ fenrir_v2_1/                  # โญ NEW โ€” 411 MB (2.1M+ entries)
โ”‚   โ”‚   โ”œโ”€โ”€ clydeiii_cybersecurity/       # โญ NEW โ€” 20 MB yearly corpus
โ”‚   โ”‚   โ”œโ”€โ”€ precinct6_cybersecurity/      # โญ NEW โ€” 2.1 GB (graph+signals+ref)
โ”‚   โ”‚   โ””โ”€โ”€ savani_cyber_attack/          # โญ NEW โ€” 17 MB attack CSV
โ”‚   โ”‚
โ”‚   โ””โ”€โ”€ ๐Ÿ“‡ index/                         # 2 sources
โ”‚       โ”œโ”€โ”€ species_25k/
โ”‚       โ””โ”€โ”€ transport_25k/
โ”‚
โ”œโ”€โ”€ ๐Ÿ“„ README.md
โ””โ”€โ”€ ๐Ÿ“„ dataset_info.json

๐Ÿค” Why is archives/ kept compressed?

ReasonExplanation
๐Ÿ’พ Space EfficiencyUncompressed would exceed 200+ GB. Compressed = 64 GB (3ร— saving)
๐ŸŽฏ On-Demand AccessDownload only specific repositories you need
๐Ÿ” Preservation Fidelitytar.gz preserves exact file permissions, directory structure, binaries

๐Ÿ’ก Tip: For training on code content, use data/coding/fable5_repos_full/ (475K samples, each a file extracted from archives, capped at 4KB). For full untruncated file access, stream directly from archives/.


๐ŸŒ Data Sources & Provenance

๐Ÿ—บ๏ธ 73 Sources Across 8 Categories

CategorySourcesSamplesDescription
๐Ÿ’ป coding28~11M+Agentic traces, code repos, coder distillations
๐Ÿง  distilled9~200KFrontier model distillations
โš™๏ธ applied8~200KRobotics, nano, materials, climate, energy
๐Ÿ“š humanities8~200KPsychology, economics, law, statistics
๐Ÿ”ฌ science7~175KPhysics, chemistry, biology, medical, CS
๐Ÿ“ instruction3~99KClassic instruction (alpaca, oasst, dolly)
๐Ÿ“‡ index2~50KSpecies index, transport
๐Ÿ”’ cybersecurity6~2.6 GBHigh-quality attack, defense, exploit traces
๐Ÿงฎ math2~52KMath + Lean theorem proofs

๐Ÿ’ป Coding Category (28 sources โ€” ALL FULLY PROCESSED โญ)

Source SlugUpstream DatasetTypeSamples
vibe_instruct_v2CodeDevX/Vibe-Coding-Instruct-V2Agentic coding8,152,510
fable5_2mCrownelius/Complete-FABLE.5-traces-2MFable-5 traces2,006,487
vibe_instruct_v1CodeDevX/Vibe-Coding-InstructAgentic coding1,100,000
vibe_codingattentionAllYouNeed/Vibe-Coding-Claude-Fable-5Claude coding1,100,000
royal_ghost_1mWithinUsAI/Royal_Ghost_Coder_1MGhost coder1,000,000
citation_groundWithinUsAI/CitationGround-1MCitation-grounded980,064
royal_ghost_501kWithinUsAI/Royal_Ghost_Coder_501kGhost coder703,449
fable5_repos_fullnotune/fable5-repos7,090 repo pointers7,090
fable5_agentic_sftNexlab/fable5-agentic-coding-sftAgentic SFT159,972
gpt55_codexAletheiaResearch/GPT-5.5-CodexGPT-5.5 Codex119,436
alpca_gpt55GabrielFreeze-2/alpca-mlt-gpt-5.5_chatmlGPT-5.5 chatml49,099
deepseek_v4_pro_agentTeichAI/DeepSeek-v4-Pro-AgentDeepSeek v496,597
fable5_tracesGlint-Research/Fable-5-tracesFable-5 traces49,544
genesis_code_100kWithinUsAI/Genesis_AI_Code_100kGenesis code68,000
genesis_codeWithinUsAI/Genesis_AI_Code_50kGenesis code49,000
kimi_codingtrjxter/Kimi-K2.7-CodingTraces-9000xKimi K2.79,014
mimo_claude_code_traceschoucsan/mimo-claude-code-traces-1kMimo Claude15,046
kimi_k26_claude_code_tracesarmand0e/kimi-k2.6-claude-code-tracesKimi K2.67,438
genesis_code_10kWithinUsAI/Genesis_AI_Code_10kGenesis code9,800
legend_pythonWithinUsAI/Legend_Python_CoderV.1Python coder5,000
autonomyWithinUsAI/The_Autonomy_From_WithIn_10kAutonomy10,000
genesis_code_demoWithinUsAI/Genesis_AI_Code_1k_DemoGenesis demo1,000
god_coderWithinUsAI/GOD_Coder_100kGOD coderFULL โญ
python_god_coderWithinUsAI/python_GOD_coder_100kPython GODFULL โญ
elite_god_coderWithinUsAI/Elite_GOD_Coder_100kElite GODFULL โญ
omega_genesisWithinUsAI/Omega_Genesis_Coder_100kOmega GenesisFULL โญ
open_tool_traceWithinUsAI/OpenToolTrace-XTool traces48
genesis_v11WithinUsAI/Genesis_v1_1_Update...Genesis v1.1partial

๐Ÿง  Distilled Category (9 sources)

SourceUpstreamDistilled From
claude_mythosWithinUsAI/claude_mythos_distilled_25kClaude
gemini35WithinUsAI/gemini_3.5_flash_distilled_25kGemini 3.5 Flash
fable5_cleanedWithinUsAI/fable_5_distillation_merged_cleaned_25kFable-5
grok44WithinUsAI/Grok4.4_heavy_max_distill_god_seed_25kGrok 4.4
gemini_pro32WithinUsAI/GeminiPro3.2_max_distill_god_seed_25kGemini Pro 3.2
gpt55_thinkingWithinUsAI/GPT5.5_thinking_max_distill_god_seed_25KGPT-5.5
gpt55_distilledWithinUsAI/GPT_5.5_DistilledGPT-5.5
claude_opus_48_distill11-47/claude_opus_4.8_distill_5kClaude Opus 4.8 โญ
claude_opus_48_max_thinking11-47/claude_opus_4.8_max_thinking_5k_v2Opus 4.8 Max โญ

๐Ÿ”ฌ Science ยท โš™๏ธ Applied ยท ๐Ÿ“š Humanities ยท ๐Ÿงฎ Math ยท ๐Ÿ“ Instruction ยท ๐Ÿ”’ Cybersecurity ยท ๐Ÿ“‡ Index

๐Ÿ“– Click to expand all other categories

๐Ÿ”ฌ Science (7 sources): science_25k, physics_25k, chemistry_25k, biology_25k, medical_25k, cs_25k, biology_r2med (R2MED/Biology)

โš™๏ธ Applied (8 sources): robotics_25k, nano_25k, materials_25k, earth_climate_25k, renewable_energy_25k, evolution_25k, universe_25k, kardashev_25k

๐Ÿ“š Humanities (8 sources): psychology_25k, economics_25k, law_25k, statistics_25k, sports_25k, human_25k, conscience_25k, supernatural_25k

๐Ÿงฎ Math (2 sources): math_25k, deepseek_prover_v1 (27,503 Lean proofs)

๐Ÿ“ Instruction (3 sources): alpaca (52K), oasst (32K), dolly (15K)

๐Ÿ”’ Cybersecurity (6 sources): high_quality_cybersecurity, heimdall_v1_1, fenrir_v2_1, clydeiii_cybersecurity, precinct6_cybersecurity, savani_cyber_attack

๐Ÿ“‡ Index (2 sources): species_25k, transport_25k


๐Ÿ›ก๏ธ Cybersecurity Deep Dive: Attack & Defense

โš”๏ธ Why This Matters

Modern AI systems are increasingly deployed in security-critical environmentsโ€”yet most open-source training data ignores real-world adversarial scenarios. The Open Distillation Codex includes a dedicated cybersecurity category designed to equip models with:

  • Attack Awareness: Recognize and generate realistic attack patterns, exploits, penetration testing commands, and social engineering dialogues.
  • Defense Proficiency: Learn to propose defensive measures, detect anomalies, and articulate incident response protocols.
  • Exploit Understanding: Analyze and explain software vulnerabilities, craft proof-of-concept code (for educational purposes), and understand exploit chains.
  • Red/Blue Team Simulation: Engage in multi-turn conversations mimicking red team attack planning and blue team defense coordination.
  • Threat Intelligence: Summarize, classify, and reason about cyber threat reports, CVEs, and IOCs (Indicators of Compromise).

This makes the dataset a powerful foundation for building cybersecurity-aware LLMs, security co-pilots, and automated vulnerability assessment tools.

๐Ÿ“Š Whatโ€™s Inside the Cybersecurity Category?

SourceDescriptionData FormatKey Themes
high_quality_cybersecurityManually curated high-quality instructionโ€“response pairs covering attack techniques, defense, and policyJSONL (shards)MITRE ATT&CK, OWASP, incident response
heimdall_v1_1~78 MB of security conversations, including red/blue team dialogues and threat analysisJSONLMulti-turn chat, tool usage
fenrir_v2_1411 MB, 2.1M+ entries โ€” massive corpus of cybersecurity Q&A, exploit descriptions, and code snippetsJSONLExploit code, CVEs, vulnerability research
clydeiii_cybersecurity20 MB yearly security corpus, aggregated from public reports and advisoriesJSONLYear-in-review, trends, threat landscape
precinct6_cybersecurity2.1 GB graph-based dataset with network signals, attack graphs, and reference materialsJSONL (graph+signals+ref)Network attacks, lateral movement, detection
savani_cyber_attack17 MB CSV of labeled cyber attack incidents with detailed featuresCSVAttack classification, feature analysis

๐Ÿงช Attack & Exploit Examples

Here are a few representative samples (sanitized) from the dataset:

Example 1 โ€“ SQL Injection Exploit

{
  "source": "fenrir_v2_1",
  "instruction": "Explain how an attacker could exploit a vulnerable login form using SQL injection. Provide a sample malicious input.",
  "response": "An attacker could bypass authentication by injecting SQL logic. For a query like:\n`SELECT * FROM users WHERE username='$user' AND password='$pass'`,\nthe attacker might enter `admin' --` as the username and anything as password. The resulting query becomes:\n`SELECT * FROM users WHERE username='admin' --' AND password='...'`\nThe `--` comments out the password check, granting access as admin. More advanced attacks use UNION to extract data."
}

Example 2 โ€“ Red Team Command Sequence

{
  "source": "heimdall_v1_1",
  "instruction": "As a red team operator, list the steps you would take after gaining initial access to a Windows machine using a phishing payload.",
  "response": "1. Establish persistence: `schtasks /create /tn \"Updater\" /tr \"C:\\malware.exe\" /sc hourly`\n2. Enumerate local users: `net user /domain`\n3. Dump credentials: `mimikatz sekurlsa::logonpasswords`\n4. Lateral movement: use PsExec to move to file server.\n5. Exfiltrate data: compress sensitive files and upload via DNS tunneling.\nAlways maintain C2 communication over HTTPS to blend in."
}

Example 3 โ€“ Defense Playbook (Blue Team)

{
  "source": "high_quality_cybersecurity",
  "instruction": "An IDS alert flags unusual outbound DNS traffic from a workstation. What's the immediate response?",
  "response": "1. Isolate the workstation from the network (VLAN change or port shutdown).\n2. Capture volatile memory and network logs for forensics.\n3. Check DNS queries: if long, random-looking subdomains, suspect DNS tunneling.\n4. Scan for malware with updated signatures.\n5. Review firewall logs for similar traffic from other hosts.\n6. If confirmed, initiate incident response playbook for data exfiltration."
}

๐ŸŽ“ How to Train a Cybersecurity-Focused LLM

from datasets import load_dataset

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# Load only cybersecurity category
cyber_ds = load_dataset(REPO, split="train", 
                        data_files="data/cybersecurity/**/*.jsonl",
                        streaming=True)

# Or load specific sources
fenrir = load_dataset(REPO, split="train", 
                      data_files="data/cybersecurity/fenrir_v2_1/*.jsonl")

# Format for SFT
def format_security_sample(example):
    return {
        "text": f"### Security Task:\n{example['instruction']}\n\n### Expert Response:\n{example['response']}"
    }

cyber_ds = cyber_ds.map(format_security_sample)

# Now train with your favourite framework (transformers, axolotl, etc.)

Curriculum Idea:

  1. Start with high_quality_cybersecurity and heimdall_v1_1 for foundational attack/defense conversations.
  2. Introduce fenrir_v2_1 for exploit code and vulnerability deep dives.
  3. Use precinct6_cybersecurity for network-level attack graph understanding.

๐Ÿ›ก๏ธ Ethical & Responsible Use

  • For Defensive Purposes Only: This data is intended to strengthen AI for defense, threat detection, and security education. Do not use it to generate active attack code without proper authorization.
  • No Zero-Day Exploits: The dataset contains only already-public vulnerabilities and techniques. It does not include zero-day or weaponized exploits.
  • Responsible Disclosure: If you fine-tune a model with this data, we recommend adding a safety preamble warning that generated security content must be used legally and ethically.
  • Dual-Use Awareness: While we believe open access improves collective security, we acknowledge the dual-use nature. Users are expected to follow applicable laws and guidelines.

โš ๏ธ Disclaimer: This dataset includes descriptions of attack techniques for educational purposes. The maintainers are not responsible for misuse.

๐Ÿ“ˆ Future Additions

  • Integration with CTF (Capture The Flag) challenge walkthroughs.
  • More blue team procedures and SOAR playbooks.
  • Anonymized real-world incident response logs (with permission).

๐Ÿ› ๏ธ How to Use & Train

1๏ธโƒฃ Load Categorized JSONL Data

from datasets import load_dataset

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# โ”€ Load a single category โ”€
ds = load_dataset(REPO, split="train", data_files="data/coding/*/*.jsonl", streaming=True)

# โ”€ Load a specific source โ”€
ds = load_dataset(REPO, split="train", data_files="data/coding/vibe_instruct_v2/*.jsonl", streaming=True)

# โ”€ Load everything (18M+ samples) โ”€
ds = load_dataset(REPO, split="train", streaming=True)

for sample in ds:
    print(sample["source"], sample["instruction"][:80])

2๏ธโƒฃ Stream the 64 GB archives/ GitHub Repositories

from huggingface_hub import hf_hub_download
import tarfile

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# โ”€ Option A: Download & extract ONE repository โ”€
hf_hub_download(
    repo_id=REPO,
    repo_type="dataset",
    filename="archives/0x101__lakewatch.tar.gz",
    local_dir="./repos",
)
with tarfile.open("./repos/archives/0x101__lakewatch.tar.gz", "r:gz") as tar:
    tar.extractall("./extracted/0x101__lakewatch")


# โ”€ Option B: Stream files WITHOUT full extraction โ”€
def stream_repo_files(archive_name, max_files=100):
    """Stream file contents from tar.gz without extracting to disk."""
    local_path = hf_hub_download(repo_id=REPO, repo_type="dataset", filename=archive_name)
    
    with tarfile.open(local_path, "r:gz") as tar:
        count = 0
        for member in tar:
            if member.isfile() and count < max_files:
                f = tar.extractfile(member)
                if f:
                    yield {
                        "path": member.name,
                        "content": f.read().decode("utf-8", errors="ignore")[:4000],
                    }
                    count += 1
    
    import os
    os.remove(local_path)  # Clean up

# Stream files from a specific repo
for file_data in stream_repo_files("archives/0x101__lakewatch.tar.gz"):
    print(f"๐Ÿ“„ {file_data['path']}: {file_data['content'][:100]}...")


# โ”€ Option C: Use pre-extracted JSONL shards (475K samples) โ”€
code_ds = load_dataset(
    REPO, split="train",
    data_files="data/coding/fable5_repos_full/*.jsonl",
    streaming=True
)
# Each sample: instruction = "<repo>/<file>", response = "<content>"

3๏ธโƒฃ SFT Training Script (Hugging Face Trainer)

import torch
from datasets import load_dataset
from transformers import (
    AutoTokenizer,
    AutoModelForCausalLM,
    TrainingArguments,
    Trainer,
    DataCollatorForLanguageModeling,
)

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# CONFIGURATION
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
MODEL_NAME = "meta-llama/Llama-3.1-8B"
DATASET_REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"
OUTPUT_DIR = "./sft-output"
MAX_SEQ_LEN = 2048

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# LOAD MODEL & TOKENIZER
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
tokenizer.pad_token = tokenizer.eos_token

model = AutoModelForCausalLM.from_pretrained(
    MODEL_NAME,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="flash_attention_2",
)

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# LOAD & FORMAT DATASET
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
def format_instruction(sample):
    text = f"### Instruction:\n{sample['instruction']}\n\n### Response:\n{sample['response']}"
    return {"text": text}

def tokenize(examples):
    return tokenizer(
        examples["text"],
        truncation=True,
        max_length=MAX_SEQ_LEN,
        padding="max_length",
    )

# Load coding category (use "data/**/*.jsonl" for full 18M+)
train_ds = load_dataset(
    DATASET_REPO,
    split="train",
    data_files="data/coding/*/*.jsonl",
    streaming=True,
)
train_ds = train_ds.map(format_instruction).filter(lambda x: len(x["text"]) > 0)
train_ds = train_ds.map(tokenize, batched=True)

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# TRAIN
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
training_args = TrainingArguments(
    output_dir=OUTPUT_DIR,
    num_train_epochs=3,
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,
    warmup_steps=500,
    logging_steps=100,
    save_steps=2000,
    learning_rate=2e-5,
    bf16=True,
    gradient_checkpointing=True,
    optim="adamw_torch",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=train_ds,
    data_collator=DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False),
)

trainer.train()
trainer.save_model(OUTPUT_DIR)

4๏ธโƒฃ Curriculum Learning Across Categories

from datasets import load_dataset, interleave_datasets

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# โ”€ Phase 1: Foundation (math + science) โ”€
phase1_math = load_dataset(REPO, split="train", data_files="data/math/**/*.jsonl", streaming=True)
phase1_sci = load_dataset(REPO, split="train", data_files="data/science/**/*.jsonl", streaming=True)
phase1 = interleave_datasets([phase1_math, phase1_sci])

# โ”€ Phase 2: Add coding traces โ”€
phase2 = load_dataset(REPO, split="train", data_files="data/coding/**/*.jsonl", streaming=True)

# โ”€ Phase 3: Add distilled reasoning + cybersecurity โ”€
phase3_distilled = load_dataset(REPO, split="train", data_files="data/distilled/**/*.jsonl", streaming=True)
phase3_cyber = load_dataset(REPO, split="train", data_files="data/cybersecurity/**/*.jsonl", streaming=True)
phase3 = interleave_datasets([phase3_distilled, phase3_cyber])

# Train sequentially
# trainer.train(phase1)  # epochs 0-1
# trainer.train(phase2)  # epochs 1-2
# trainer.train(phase3)  # epochs 2-3

๐Ÿ“‹ Schema Reference

{
    "source":          "fable5_2m",
    "source_dataset":  "Crownelius/Complete-FABLE.5-traces-2M",
    "instruction":     "<the prompt / question / file path>",
    "response":        "<the completion / answer / file content>",
    "category":        "coding"
}
FieldTypeMax LengthDescription
sourcestring200Short slug identifying upstream dataset
source_datasetstring200Full HF repo id (org/name)
instructionstring4,000User-side content (prompt/question/file path)
responsestring4,000Assistant-side content (completion/answer/file content)
categorystring50One of 8 categories

๐Ÿ” Licensing & Limitations

๐Ÿ“œ License

The collection as a whole is released under the MIT License.

Each upstream dataset retains its original license. The source_dataset field on every row identifies the upstream โ€” look it up on Hugging Face to determine its specific license.

LicenseApplies To
MITMost WithinUsAI datasets, OpenAssistant
Apache-2.0DeepSeek, OpenThoughts
CC-BY-4.0Dolly, various
CC-BY-SA-3.0Databricks Dolly
AGPL-3.0Some Fable-5 traces

โœ… Intended Use Cases (Our Vision)

  • Fine-tuning open-source LLMs for instruction following
  • Training coding agents and code-completion models
  • Reasoning chain distillation research
  • Domain-specific adaptation (math, science, cybersecurity)
  • Repository-scale context training (using archives/)
  • Deploying models without safety evaluation
  • Generating harmful, biased, or deceptive content
  • High-stakes domains (medical, legal, financial) without expert review
  • Claiming models "know" facts โ€” this is distilled output, not ground truth

โš ๏ธ Limitations

  1. Field length cap: instruction and response capped at 4,000 characters. For full content, use archives/.
  2. Distillation artifacts: Samples are model-generated โ€” may contain hallucinations or biases.
  3. Partial recovery: A few upstream datasets (GOD_Coder variants, Genesis_v1.1) had format errors and were partially recovered via raw JSONL parsing.

๐Ÿ“ Citation

@misc{open_distillation_codex_2026,
  title  = {The Open Distillation Codex: 18M+ samples + 7090 code repositories from 73 sources with Cybersecurity Attack & Defense},
  author = {Manusagents},
  year   = {2026},
  url    = {https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset},
  note   = {v8.2 - No skip, full. 516 shards + 7090 archives, 73 sources, 8 categories, 76 GB+}
}

๐Ÿ“œ Changelog

VersionDateKey Changes
v1.0โ€“v5.02026-07-01 to 05Progressive builds: 117K โ†’ 20.7M samples
v6.02026-07-06Category restructuring: data/<category>/<source>/shard-*.jsonl
v7.02026-07-06Training scripts + full processing started
v8.0 FINAL2026-07-06ALL sources FULLY processed โ€” no skipping. Verified 79.13 GB.
v8.12026-07-08Added 5 external cybersecurity datasets. Total 81.2 GB, 73 sources.
v8.22026-07-18Final numbers rectified: 18M+ samples, 76 GB+ total. All sources no skip, fully verified. Enhanced cybersecurity deep-dive with attack/defense examples, training scripts, ethical guidelines.


๐ŸŒŸ The Open Distillation Codex ๐ŸŒŸ

73 sources ยท 8 categories ยท 7,090 repositories ยท 516 shards ยท 76 GB+


No skip. Full. 18M+ samples. Built one archive at a time. Released under MIT.



"Two layers. Eight categories. Seventy-three sources. One codex. No skip. Full. Armed with cybersecurity attack and defense."


Streaming No Skip Full JSONL HF Cyber



โ€” The Open Distillation Codex โ€”

attack
biology
blue-team
claude
code-repositories
coding
collection
cyber
cyber security
cyber-security
cybersecurity
deepseek
defense
distillation
exploit
fable-5
gemini
gpt-5.5
grok
instruction-tuning
kimi
Math
open-source
Open-source
penetration-testing
qwen
reasoning
red-team
science
security
Trace

Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset

Dataset

๐Ÿ“– The Open Distillation Codex

207

500 commits

updated Aug 29, 2026

See the code

README

Version Storage Sources License Samples Cybersecurity



๐Ÿ“– The Open Distillation Codex

๐ŸŒŒ The Ultimate Open-Source Distillation Dataset โ€” No Skip, Full, with Attack & Defense ๐ŸŒŒ

Where 73 open-source minds converge into one unified stream of intelligence

18M+ Distilled Signals ยท 7,090 Raw GitHub Repositories ยท 8 Curated Categories ยท ~76 GB+


"We did not write this dataset. We assembled it. Every line is an echo โ€” of a model thinking, a coder drafting, a tutor explaining, a repo breathing. Seventy-three sources. Eight categories. Zero gatekeeping. No skipping. Fully processed. Now fortified with real-world cybersecurity confrontations."



๐Ÿ“Œ Table of Contents

#SectionDescription
1๐Ÿ“Š Dataset SummaryHigh-level overview & value proposition
2๐Ÿ—‚๏ธ Directory StructureASCII tree + folder explanation
3๐ŸŒ Data SourcesAll 73 sources with full attribution
4๐Ÿ›ก๏ธ Cybersecurity Deep Dive: Attack & DefenseImportance, attack traces, defense, exploit analysis
5๐Ÿ› ๏ธ How to Use & TrainLoading, streaming, training scripts
6๐Ÿ” Licensing & LimitationsLicense, intended use, limitations
7๐Ÿ“œ ChangelogVersion history

๐Ÿ“Š Dataset Summary

๐ŸŽฏ The Numbers That Matter

MetricValueStatus
Total Storage76 GB+โœ… Verified
JSONL Data Shards516โœ… Verified
Archive Files (tar.gz)7,090โœ… Verified
Source Datasets73โœ… Verified
Categories8โœ… Verified
Total Samples18M+โœ… Verified
Largest Source8.15M (Vibe-Coding-Instruct-V2)โœ…
Archive Size~64 GB (compressed GitHub repos)โœ…
Cybersecurity Sources6โœ…
Cybersecurity Data Size~2.6 GBโœ…

๐ŸŒŸ Why "Ultimate Distilled"?

This dataset is not a raw scrape. Every sample has been distilled through a unified extraction pipeline:

โ”Œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”
โ”‚                    UNIFIED EXTRACTION PIPELINE              โ”‚
โ”œโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ค
โ”‚                                                             โ”‚
โ”‚  73 Upstream Sources (ALL FULLY PROCESSED, NO SKIP)         โ”‚
โ”‚  โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ” โ”Œโ”€โ”€โ”€โ”€โ”€โ”          โ”‚
โ”‚  โ”‚ HF  โ”‚ โ”‚ HF  โ”‚ โ”‚ HF  โ”‚ โ”‚ GH  โ”‚ โ”‚ HF  โ”‚ โ”‚ ... โ”‚          โ”‚
โ”‚  โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜ โ””โ”€โ”€โ”ฌโ”€โ”€โ”˜          โ”‚
โ”‚     โ”‚       โ”‚       โ”‚       โ”‚       โ”‚       โ”‚               โ”‚
โ”‚     โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ผโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”ดโ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜               โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚ EXTRACT โ”‚ โ† Field normalization        โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜   (instruction/response)      โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚CATEGORIZEโ”‚ โ† 8 semantic categories      โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚  SHARD  โ”‚ โ† 20K samples per shard       โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”ฌโ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                         โ”‚                                    โ”‚
โ”‚                    โ”Œโ”€โ”€โ”€โ”€โ–ผโ”€โ”€โ”€โ”€โ”                               โ”‚
โ”‚                    โ”‚ UPLOAD  โ”‚ โ† Batch commits to HF         โ”‚
โ”‚                    โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜                               โ”‚
โ”‚                                                             โ”‚
โ”‚  STATUS: ALL 73 SOURCES COMPLETE. NO SKIPPING. 18M+ ROWS.   โ”‚
โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜

๐Ÿ’Ž Value to the Open-Source AI Community

๐ŸŽฏ For...๐Ÿ“ฆ This dataset provides...
Model TrainersSingle load_dataset() call to stream 18M+ SFT-ready samples
Coding Agent Researchers11M+ agentic coding traces from Fable-5, Vibe-Coding, Royal Ghost, Kimi, DeepSeek
Code Pretraining7,090 full GitHub repository snapshots (64 GB compressed)
Reasoning Researchers2.7M+ distilled reasoning traces from Claude, Gemini, Grok, GPT-5.5, Opus 4.8
Domain Specialists25K-sample sweeps across 29 disciplines
Cybersecurity ResearchersDedicated cybersecurity category with attack/defense/exploit traces, red/blue team dialogues, and incident reports
Red Team / Blue Team TrainersRealistic attack scenarios, defense strategies, exploit code, and post-mortem analysis

๐Ÿ—‚๏ธ Directory Structure

๐Ÿ“‚ Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset/
โ”‚
โ”œโ”€โ”€ ๐Ÿ“ฆ archives/                          # ~64 GB โ€” 7,090 compressed GitHub repos
โ”‚   โ”œโ”€โ”€ 0-chi__sonaure-lp.tar.gz
โ”‚   โ”œโ”€โ”€ 00MB__bitcoin_trading_bot.tar.gz
โ”‚   โ”œโ”€โ”€ 0101-agents__plugins.tar.gz
โ”‚   โ”œโ”€โ”€ ... (7,090 files total)
โ”‚   โ””โ”€โ”€ zznmg1__playable-survivor-ad.tar.gz
โ”‚
โ”œโ”€โ”€ ๐Ÿ“ data/                              # ~12 GB โ€” 516 JSONL shards (18M+ samples)
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ’ป coding/                        # 28 sources ยท ~11M+ samples
โ”‚   โ”‚   โ”œโ”€โ”€ vibe_instruct_v2/             # 8,152,510 samples
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_2m/                    # 2,006,487 samples
โ”‚   โ”‚   โ”œโ”€โ”€ vibe_instruct_v1/             # 1,100,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ vibe_coding/                  # 1,100,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ royal_ghost_1m/               # 1,000,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ citation_ground/              # 980,064 samples
โ”‚   โ”‚   โ”œโ”€โ”€ royal_ghost_501k/             # 703,449 samples
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_repos_full/            # 7,090 archive pointers
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_agentic_sft/           # 159,972 samples
โ”‚   โ”‚   โ”œโ”€โ”€ gpt55_codex/                  # 119,436 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ alpca_gpt55/                  # 49,099 samples
โ”‚   โ”‚   โ”œโ”€โ”€ deepseek_v4_pro_agent/        # 96,597 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_traces/                # 49,544 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code_100k/            # 68,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code/                 # 49,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ kimi_coding/                  # 9,014 samples
โ”‚   โ”‚   โ”œโ”€โ”€ mimo_claude_code_traces/      # 15,046 samples โญ FULL
โ”‚   โ”‚   โ”œโ”€โ”€ kimi_k26_claude_code_traces/  # 7,438 samples
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code_10k/             # 9,800 samples
โ”‚   โ”‚   โ”œโ”€โ”€ legend_python/                # 5,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ autonomy/                     # 10,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ genesis_code_demo/            # 1,000 samples
โ”‚   โ”‚   โ”œโ”€โ”€ god_coder/                    # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ python_god_coder/             # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ elite_god_coder/              # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ omega_genesis/                # โญ FULL raw recovery
โ”‚   โ”‚   โ”œโ”€โ”€ open_tool_trace/              # 48 samples
โ”‚   โ”‚   โ””โ”€โ”€ genesis_v11/                  # partial recovery
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿงฎ math/                          # 2 sources
โ”‚   โ”‚   โ”œโ”€โ”€ math_25k/
โ”‚   โ”‚   โ””โ”€โ”€ deepseek_prover_v1/           # 27,503 Lean theorem proofs
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ”ฌ science/                       # 7 sources
โ”‚   โ”‚   โ”œโ”€โ”€ science_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ physics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ chemistry_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ biology_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ medical_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ cs_25k/
โ”‚   โ”‚   โ””โ”€โ”€ biology_r2med/                # โญ NEW
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ โš™๏ธ applied/                       # 8 sources
โ”‚   โ”‚   โ”œโ”€โ”€ robotics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ nano_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ materials_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ earth_climate_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ renewable_energy_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ evolution_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ universe_25k/
โ”‚   โ”‚   โ””โ”€โ”€ kardashev_25k/
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ“š humanities/                    # 8 sources
โ”‚   โ”‚   โ”œโ”€โ”€ psychology_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ economics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ law_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ statistics_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ sports_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ human_25k/
โ”‚   โ”‚   โ”œโ”€โ”€ conscience_25k/
โ”‚   โ”‚   โ””โ”€โ”€ supernatural_25k/
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿง  distilled/                     # 9 sources ยท frontier distillations
โ”‚   โ”‚   โ”œโ”€โ”€ claude_mythos/
โ”‚   โ”‚   โ”œโ”€โ”€ gemini35/
โ”‚   โ”‚   โ”œโ”€โ”€ fable5_cleaned/
โ”‚   โ”‚   โ”œโ”€โ”€ grok44/
โ”‚   โ”‚   โ”œโ”€โ”€ gemini_pro32/
โ”‚   โ”‚   โ”œโ”€โ”€ gpt55_thinking/
โ”‚   โ”‚   โ”œโ”€โ”€ gpt55_distilled/
โ”‚   โ”‚   โ”œโ”€โ”€ claude_opus_48_distill/       # โญ NEW
โ”‚   โ”‚   โ””โ”€โ”€ claude_opus_48_max_thinking/  # โญ NEW
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ“ instruction/                   # 3 sources
โ”‚   โ”‚   โ”œโ”€โ”€ alpaca/                       # 52,002 samples
โ”‚   โ”‚   โ”œโ”€โ”€ oasst/                        # 32,141 samples
โ”‚   โ”‚   โ””โ”€โ”€ dolly/                        # 15,011 samples
โ”‚   โ”‚
โ”‚   โ”œโ”€โ”€ ๐Ÿ”’ cybersecurity/                 # 6 sources
โ”‚   โ”‚   โ”œโ”€โ”€ high_quality_cybersecurity/
โ”‚   โ”‚   โ”œโ”€โ”€ heimdall_v1_1/                # โญ NEW โ€” 78 MB conversations
โ”‚   โ”‚   โ”œโ”€โ”€ fenrir_v2_1/                  # โญ NEW โ€” 411 MB (2.1M+ entries)
โ”‚   โ”‚   โ”œโ”€โ”€ clydeiii_cybersecurity/       # โญ NEW โ€” 20 MB yearly corpus
โ”‚   โ”‚   โ”œโ”€โ”€ precinct6_cybersecurity/      # โญ NEW โ€” 2.1 GB (graph+signals+ref)
โ”‚   โ”‚   โ””โ”€โ”€ savani_cyber_attack/          # โญ NEW โ€” 17 MB attack CSV
โ”‚   โ”‚
โ”‚   โ””โ”€โ”€ ๐Ÿ“‡ index/                         # 2 sources
โ”‚       โ”œโ”€โ”€ species_25k/
โ”‚       โ””โ”€โ”€ transport_25k/
โ”‚
โ”œโ”€โ”€ ๐Ÿ“„ README.md
โ””โ”€โ”€ ๐Ÿ“„ dataset_info.json

๐Ÿค” Why is archives/ kept compressed?

ReasonExplanation
๐Ÿ’พ Space EfficiencyUncompressed would exceed 200+ GB. Compressed = 64 GB (3ร— saving)
๐ŸŽฏ On-Demand AccessDownload only specific repositories you need
๐Ÿ” Preservation Fidelitytar.gz preserves exact file permissions, directory structure, binaries

๐Ÿ’ก Tip: For training on code content, use data/coding/fable5_repos_full/ (475K samples, each a file extracted from archives, capped at 4KB). For full untruncated file access, stream directly from archives/.


๐ŸŒ Data Sources & Provenance

๐Ÿ—บ๏ธ 73 Sources Across 8 Categories

CategorySourcesSamplesDescription
๐Ÿ’ป coding28~11M+Agentic traces, code repos, coder distillations
๐Ÿง  distilled9~200KFrontier model distillations
โš™๏ธ applied8~200KRobotics, nano, materials, climate, energy
๐Ÿ“š humanities8~200KPsychology, economics, law, statistics
๐Ÿ”ฌ science7~175KPhysics, chemistry, biology, medical, CS
๐Ÿ“ instruction3~99KClassic instruction (alpaca, oasst, dolly)
๐Ÿ“‡ index2~50KSpecies index, transport
๐Ÿ”’ cybersecurity6~2.6 GBHigh-quality attack, defense, exploit traces
๐Ÿงฎ math2~52KMath + Lean theorem proofs

๐Ÿ’ป Coding Category (28 sources โ€” ALL FULLY PROCESSED โญ)

Source SlugUpstream DatasetTypeSamples
vibe_instruct_v2CodeDevX/Vibe-Coding-Instruct-V2Agentic coding8,152,510
fable5_2mCrownelius/Complete-FABLE.5-traces-2MFable-5 traces2,006,487
vibe_instruct_v1CodeDevX/Vibe-Coding-InstructAgentic coding1,100,000
vibe_codingattentionAllYouNeed/Vibe-Coding-Claude-Fable-5Claude coding1,100,000
royal_ghost_1mWithinUsAI/Royal_Ghost_Coder_1MGhost coder1,000,000
citation_groundWithinUsAI/CitationGround-1MCitation-grounded980,064
royal_ghost_501kWithinUsAI/Royal_Ghost_Coder_501kGhost coder703,449
fable5_repos_fullnotune/fable5-repos7,090 repo pointers7,090
fable5_agentic_sftNexlab/fable5-agentic-coding-sftAgentic SFT159,972
gpt55_codexAletheiaResearch/GPT-5.5-CodexGPT-5.5 Codex119,436
alpca_gpt55GabrielFreeze-2/alpca-mlt-gpt-5.5_chatmlGPT-5.5 chatml49,099
deepseek_v4_pro_agentTeichAI/DeepSeek-v4-Pro-AgentDeepSeek v496,597
fable5_tracesGlint-Research/Fable-5-tracesFable-5 traces49,544
genesis_code_100kWithinUsAI/Genesis_AI_Code_100kGenesis code68,000
genesis_codeWithinUsAI/Genesis_AI_Code_50kGenesis code49,000
kimi_codingtrjxter/Kimi-K2.7-CodingTraces-9000xKimi K2.79,014
mimo_claude_code_traceschoucsan/mimo-claude-code-traces-1kMimo Claude15,046
kimi_k26_claude_code_tracesarmand0e/kimi-k2.6-claude-code-tracesKimi K2.67,438
genesis_code_10kWithinUsAI/Genesis_AI_Code_10kGenesis code9,800
legend_pythonWithinUsAI/Legend_Python_CoderV.1Python coder5,000
autonomyWithinUsAI/The_Autonomy_From_WithIn_10kAutonomy10,000
genesis_code_demoWithinUsAI/Genesis_AI_Code_1k_DemoGenesis demo1,000
god_coderWithinUsAI/GOD_Coder_100kGOD coderFULL โญ
python_god_coderWithinUsAI/python_GOD_coder_100kPython GODFULL โญ
elite_god_coderWithinUsAI/Elite_GOD_Coder_100kElite GODFULL โญ
omega_genesisWithinUsAI/Omega_Genesis_Coder_100kOmega GenesisFULL โญ
open_tool_traceWithinUsAI/OpenToolTrace-XTool traces48
genesis_v11WithinUsAI/Genesis_v1_1_Update...Genesis v1.1partial

๐Ÿง  Distilled Category (9 sources)

SourceUpstreamDistilled From
claude_mythosWithinUsAI/claude_mythos_distilled_25kClaude
gemini35WithinUsAI/gemini_3.5_flash_distilled_25kGemini 3.5 Flash
fable5_cleanedWithinUsAI/fable_5_distillation_merged_cleaned_25kFable-5
grok44WithinUsAI/Grok4.4_heavy_max_distill_god_seed_25kGrok 4.4
gemini_pro32WithinUsAI/GeminiPro3.2_max_distill_god_seed_25kGemini Pro 3.2
gpt55_thinkingWithinUsAI/GPT5.5_thinking_max_distill_god_seed_25KGPT-5.5
gpt55_distilledWithinUsAI/GPT_5.5_DistilledGPT-5.5
claude_opus_48_distill11-47/claude_opus_4.8_distill_5kClaude Opus 4.8 โญ
claude_opus_48_max_thinking11-47/claude_opus_4.8_max_thinking_5k_v2Opus 4.8 Max โญ

๐Ÿ”ฌ Science ยท โš™๏ธ Applied ยท ๐Ÿ“š Humanities ยท ๐Ÿงฎ Math ยท ๐Ÿ“ Instruction ยท ๐Ÿ”’ Cybersecurity ยท ๐Ÿ“‡ Index

๐Ÿ“– Click to expand all other categories

๐Ÿ”ฌ Science (7 sources): science_25k, physics_25k, chemistry_25k, biology_25k, medical_25k, cs_25k, biology_r2med (R2MED/Biology)

โš™๏ธ Applied (8 sources): robotics_25k, nano_25k, materials_25k, earth_climate_25k, renewable_energy_25k, evolution_25k, universe_25k, kardashev_25k

๐Ÿ“š Humanities (8 sources): psychology_25k, economics_25k, law_25k, statistics_25k, sports_25k, human_25k, conscience_25k, supernatural_25k

๐Ÿงฎ Math (2 sources): math_25k, deepseek_prover_v1 (27,503 Lean proofs)

๐Ÿ“ Instruction (3 sources): alpaca (52K), oasst (32K), dolly (15K)

๐Ÿ”’ Cybersecurity (6 sources): high_quality_cybersecurity, heimdall_v1_1, fenrir_v2_1, clydeiii_cybersecurity, precinct6_cybersecurity, savani_cyber_attack

๐Ÿ“‡ Index (2 sources): species_25k, transport_25k


๐Ÿ›ก๏ธ Cybersecurity Deep Dive: Attack & Defense

โš”๏ธ Why This Matters

Modern AI systems are increasingly deployed in security-critical environmentsโ€”yet most open-source training data ignores real-world adversarial scenarios. The Open Distillation Codex includes a dedicated cybersecurity category designed to equip models with:

  • Attack Awareness: Recognize and generate realistic attack patterns, exploits, penetration testing commands, and social engineering dialogues.
  • Defense Proficiency: Learn to propose defensive measures, detect anomalies, and articulate incident response protocols.
  • Exploit Understanding: Analyze and explain software vulnerabilities, craft proof-of-concept code (for educational purposes), and understand exploit chains.
  • Red/Blue Team Simulation: Engage in multi-turn conversations mimicking red team attack planning and blue team defense coordination.
  • Threat Intelligence: Summarize, classify, and reason about cyber threat reports, CVEs, and IOCs (Indicators of Compromise).

This makes the dataset a powerful foundation for building cybersecurity-aware LLMs, security co-pilots, and automated vulnerability assessment tools.

๐Ÿ“Š Whatโ€™s Inside the Cybersecurity Category?

SourceDescriptionData FormatKey Themes
high_quality_cybersecurityManually curated high-quality instructionโ€“response pairs covering attack techniques, defense, and policyJSONL (shards)MITRE ATT&CK, OWASP, incident response
heimdall_v1_1~78 MB of security conversations, including red/blue team dialogues and threat analysisJSONLMulti-turn chat, tool usage
fenrir_v2_1411 MB, 2.1M+ entries โ€” massive corpus of cybersecurity Q&A, exploit descriptions, and code snippetsJSONLExploit code, CVEs, vulnerability research
clydeiii_cybersecurity20 MB yearly security corpus, aggregated from public reports and advisoriesJSONLYear-in-review, trends, threat landscape
precinct6_cybersecurity2.1 GB graph-based dataset with network signals, attack graphs, and reference materialsJSONL (graph+signals+ref)Network attacks, lateral movement, detection
savani_cyber_attack17 MB CSV of labeled cyber attack incidents with detailed featuresCSVAttack classification, feature analysis

๐Ÿงช Attack & Exploit Examples

Here are a few representative samples (sanitized) from the dataset:

Example 1 โ€“ SQL Injection Exploit

{
  "source": "fenrir_v2_1",
  "instruction": "Explain how an attacker could exploit a vulnerable login form using SQL injection. Provide a sample malicious input.",
  "response": "An attacker could bypass authentication by injecting SQL logic. For a query like:\n`SELECT * FROM users WHERE username='$user' AND password='$pass'`,\nthe attacker might enter `admin' --` as the username and anything as password. The resulting query becomes:\n`SELECT * FROM users WHERE username='admin' --' AND password='...'`\nThe `--` comments out the password check, granting access as admin. More advanced attacks use UNION to extract data."
}

Example 2 โ€“ Red Team Command Sequence

{
  "source": "heimdall_v1_1",
  "instruction": "As a red team operator, list the steps you would take after gaining initial access to a Windows machine using a phishing payload.",
  "response": "1. Establish persistence: `schtasks /create /tn \"Updater\" /tr \"C:\\malware.exe\" /sc hourly`\n2. Enumerate local users: `net user /domain`\n3. Dump credentials: `mimikatz sekurlsa::logonpasswords`\n4. Lateral movement: use PsExec to move to file server.\n5. Exfiltrate data: compress sensitive files and upload via DNS tunneling.\nAlways maintain C2 communication over HTTPS to blend in."
}

Example 3 โ€“ Defense Playbook (Blue Team)

{
  "source": "high_quality_cybersecurity",
  "instruction": "An IDS alert flags unusual outbound DNS traffic from a workstation. What's the immediate response?",
  "response": "1. Isolate the workstation from the network (VLAN change or port shutdown).\n2. Capture volatile memory and network logs for forensics.\n3. Check DNS queries: if long, random-looking subdomains, suspect DNS tunneling.\n4. Scan for malware with updated signatures.\n5. Review firewall logs for similar traffic from other hosts.\n6. If confirmed, initiate incident response playbook for data exfiltration."
}

๐ŸŽ“ How to Train a Cybersecurity-Focused LLM

from datasets import load_dataset

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# Load only cybersecurity category
cyber_ds = load_dataset(REPO, split="train", 
                        data_files="data/cybersecurity/**/*.jsonl",
                        streaming=True)

# Or load specific sources
fenrir = load_dataset(REPO, split="train", 
                      data_files="data/cybersecurity/fenrir_v2_1/*.jsonl")

# Format for SFT
def format_security_sample(example):
    return {
        "text": f"### Security Task:\n{example['instruction']}\n\n### Expert Response:\n{example['response']}"
    }

cyber_ds = cyber_ds.map(format_security_sample)

# Now train with your favourite framework (transformers, axolotl, etc.)

Curriculum Idea:

  1. Start with high_quality_cybersecurity and heimdall_v1_1 for foundational attack/defense conversations.
  2. Introduce fenrir_v2_1 for exploit code and vulnerability deep dives.
  3. Use precinct6_cybersecurity for network-level attack graph understanding.

๐Ÿ›ก๏ธ Ethical & Responsible Use

  • For Defensive Purposes Only: This data is intended to strengthen AI for defense, threat detection, and security education. Do not use it to generate active attack code without proper authorization.
  • No Zero-Day Exploits: The dataset contains only already-public vulnerabilities and techniques. It does not include zero-day or weaponized exploits.
  • Responsible Disclosure: If you fine-tune a model with this data, we recommend adding a safety preamble warning that generated security content must be used legally and ethically.
  • Dual-Use Awareness: While we believe open access improves collective security, we acknowledge the dual-use nature. Users are expected to follow applicable laws and guidelines.

โš ๏ธ Disclaimer: This dataset includes descriptions of attack techniques for educational purposes. The maintainers are not responsible for misuse.

๐Ÿ“ˆ Future Additions

  • Integration with CTF (Capture The Flag) challenge walkthroughs.
  • More blue team procedures and SOAR playbooks.
  • Anonymized real-world incident response logs (with permission).

๐Ÿ› ๏ธ How to Use & Train

1๏ธโƒฃ Load Categorized JSONL Data

from datasets import load_dataset

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# โ”€ Load a single category โ”€
ds = load_dataset(REPO, split="train", data_files="data/coding/*/*.jsonl", streaming=True)

# โ”€ Load a specific source โ”€
ds = load_dataset(REPO, split="train", data_files="data/coding/vibe_instruct_v2/*.jsonl", streaming=True)

# โ”€ Load everything (18M+ samples) โ”€
ds = load_dataset(REPO, split="train", streaming=True)

for sample in ds:
    print(sample["source"], sample["instruction"][:80])

2๏ธโƒฃ Stream the 64 GB archives/ GitHub Repositories

from huggingface_hub import hf_hub_download
import tarfile

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# โ”€ Option A: Download & extract ONE repository โ”€
hf_hub_download(
    repo_id=REPO,
    repo_type="dataset",
    filename="archives/0x101__lakewatch.tar.gz",
    local_dir="./repos",
)
with tarfile.open("./repos/archives/0x101__lakewatch.tar.gz", "r:gz") as tar:
    tar.extractall("./extracted/0x101__lakewatch")


# โ”€ Option B: Stream files WITHOUT full extraction โ”€
def stream_repo_files(archive_name, max_files=100):
    """Stream file contents from tar.gz without extracting to disk."""
    local_path = hf_hub_download(repo_id=REPO, repo_type="dataset", filename=archive_name)
    
    with tarfile.open(local_path, "r:gz") as tar:
        count = 0
        for member in tar:
            if member.isfile() and count < max_files:
                f = tar.extractfile(member)
                if f:
                    yield {
                        "path": member.name,
                        "content": f.read().decode("utf-8", errors="ignore")[:4000],
                    }
                    count += 1
    
    import os
    os.remove(local_path)  # Clean up

# Stream files from a specific repo
for file_data in stream_repo_files("archives/0x101__lakewatch.tar.gz"):
    print(f"๐Ÿ“„ {file_data['path']}: {file_data['content'][:100]}...")


# โ”€ Option C: Use pre-extracted JSONL shards (475K samples) โ”€
code_ds = load_dataset(
    REPO, split="train",
    data_files="data/coding/fable5_repos_full/*.jsonl",
    streaming=True
)
# Each sample: instruction = "<repo>/<file>", response = "<content>"

3๏ธโƒฃ SFT Training Script (Hugging Face Trainer)

import torch
from datasets import load_dataset
from transformers import (
    AutoTokenizer,
    AutoModelForCausalLM,
    TrainingArguments,
    Trainer,
    DataCollatorForLanguageModeling,
)

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# CONFIGURATION
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
MODEL_NAME = "meta-llama/Llama-3.1-8B"
DATASET_REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"
OUTPUT_DIR = "./sft-output"
MAX_SEQ_LEN = 2048

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# LOAD MODEL & TOKENIZER
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
tokenizer = AutoTokenizer.from_pretrained(MODEL_NAME)
tokenizer.pad_token = tokenizer.eos_token

model = AutoModelForCausalLM.from_pretrained(
    MODEL_NAME,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="flash_attention_2",
)

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# LOAD & FORMAT DATASET
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
def format_instruction(sample):
    text = f"### Instruction:\n{sample['instruction']}\n\n### Response:\n{sample['response']}"
    return {"text": text}

def tokenize(examples):
    return tokenizer(
        examples["text"],
        truncation=True,
        max_length=MAX_SEQ_LEN,
        padding="max_length",
    )

# Load coding category (use "data/**/*.jsonl" for full 18M+)
train_ds = load_dataset(
    DATASET_REPO,
    split="train",
    data_files="data/coding/*/*.jsonl",
    streaming=True,
)
train_ds = train_ds.map(format_instruction).filter(lambda x: len(x["text"]) > 0)
train_ds = train_ds.map(tokenize, batched=True)

# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
# TRAIN
# โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•โ•
training_args = TrainingArguments(
    output_dir=OUTPUT_DIR,
    num_train_epochs=3,
    per_device_train_batch_size=4,
    gradient_accumulation_steps=4,
    warmup_steps=500,
    logging_steps=100,
    save_steps=2000,
    learning_rate=2e-5,
    bf16=True,
    gradient_checkpointing=True,
    optim="adamw_torch",
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=train_ds,
    data_collator=DataCollatorForLanguageModeling(tokenizer=tokenizer, mlm=False),
)

trainer.train()
trainer.save_model(OUTPUT_DIR)

4๏ธโƒฃ Curriculum Learning Across Categories

from datasets import load_dataset, interleave_datasets

REPO = "Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset"

# โ”€ Phase 1: Foundation (math + science) โ”€
phase1_math = load_dataset(REPO, split="train", data_files="data/math/**/*.jsonl", streaming=True)
phase1_sci = load_dataset(REPO, split="train", data_files="data/science/**/*.jsonl", streaming=True)
phase1 = interleave_datasets([phase1_math, phase1_sci])

# โ”€ Phase 2: Add coding traces โ”€
phase2 = load_dataset(REPO, split="train", data_files="data/coding/**/*.jsonl", streaming=True)

# โ”€ Phase 3: Add distilled reasoning + cybersecurity โ”€
phase3_distilled = load_dataset(REPO, split="train", data_files="data/distilled/**/*.jsonl", streaming=True)
phase3_cyber = load_dataset(REPO, split="train", data_files="data/cybersecurity/**/*.jsonl", streaming=True)
phase3 = interleave_datasets([phase3_distilled, phase3_cyber])

# Train sequentially
# trainer.train(phase1)  # epochs 0-1
# trainer.train(phase2)  # epochs 1-2
# trainer.train(phase3)  # epochs 2-3

๐Ÿ“‹ Schema Reference

{
    "source":          "fable5_2m",
    "source_dataset":  "Crownelius/Complete-FABLE.5-traces-2M",
    "instruction":     "<the prompt / question / file path>",
    "response":        "<the completion / answer / file content>",
    "category":        "coding"
}
FieldTypeMax LengthDescription
sourcestring200Short slug identifying upstream dataset
source_datasetstring200Full HF repo id (org/name)
instructionstring4,000User-side content (prompt/question/file path)
responsestring4,000Assistant-side content (completion/answer/file content)
categorystring50One of 8 categories

๐Ÿ” Licensing & Limitations

๐Ÿ“œ License

The collection as a whole is released under the MIT License.

Each upstream dataset retains its original license. The source_dataset field on every row identifies the upstream โ€” look it up on Hugging Face to determine its specific license.

LicenseApplies To
MITMost WithinUsAI datasets, OpenAssistant
Apache-2.0DeepSeek, OpenThoughts
CC-BY-4.0Dolly, various
CC-BY-SA-3.0Databricks Dolly
AGPL-3.0Some Fable-5 traces

โœ… Intended Use Cases (Our Vision)

  • Fine-tuning open-source LLMs for instruction following
  • Training coding agents and code-completion models
  • Reasoning chain distillation research
  • Domain-specific adaptation (math, science, cybersecurity)
  • Repository-scale context training (using archives/)
  • Deploying models without safety evaluation
  • Generating harmful, biased, or deceptive content
  • High-stakes domains (medical, legal, financial) without expert review
  • Claiming models "know" facts โ€” this is distilled output, not ground truth

โš ๏ธ Limitations

  1. Field length cap: instruction and response capped at 4,000 characters. For full content, use archives/.
  2. Distillation artifacts: Samples are model-generated โ€” may contain hallucinations or biases.
  3. Partial recovery: A few upstream datasets (GOD_Coder variants, Genesis_v1.1) had format errors and were partially recovered via raw JSONL parsing.

๐Ÿ“ Citation

@misc{open_distillation_codex_2026,
  title  = {The Open Distillation Codex: 18M+ samples + 7090 code repositories from 73 sources with Cybersecurity Attack & Defense},
  author = {Manusagents},
  year   = {2026},
  url    = {https://huggingface.co/datasets/Manusagents/GPT-5.5-Gemini-3.1-Pro-Grok-4-Claude-Fable-5-Mythos-5-Qwen-3.7-Max-and-more-Distillation-Dataset},
  note   = {v8.2 - No skip, full. 516 shards + 7090 archives, 73 sources, 8 categories, 76 GB+}
}

๐Ÿ“œ Changelog

VersionDateKey Changes
v1.0โ€“v5.02026-07-01 to 05Progressive builds: 117K โ†’ 20.7M samples
v6.02026-07-06Category restructuring: data/<category>/<source>/shard-*.jsonl
v7.02026-07-06Training scripts + full processing started
v8.0 FINAL2026-07-06ALL sources FULLY processed โ€” no skipping. Verified 79.13 GB.
v8.12026-07-08Added 5 external cybersecurity datasets. Total 81.2 GB, 73 sources.
v8.22026-07-18Final numbers rectified: 18M+ samples, 76 GB+ total. All sources no skip, fully verified. Enhanced cybersecurity deep-dive with attack/defense examples, training scripts, ethical guidelines.


๐ŸŒŸ The Open Distillation Codex ๐ŸŒŸ

73 sources ยท 8 categories ยท 7,090 repositories ยท 516 shards ยท 76 GB+


No skip. Full. 18M+ samples. Built one archive at a time. Released under MIT.



"Two layers. Eight categories. Seventy-three sources. One codex. No skip. Full. Armed with cybersecurity attack and defense."


Streaming No Skip Full JSONL HF Cyber



โ€” The Open Distillation Codex โ€”

attack
biology
blue-team
claude
code-repositories
coding
collection
cyber
cyber security
cyber-security
cybersecurity
deepseek
defense
distillation
exploit
fable-5
gemini
gpt-5.5
grok
instruction-tuning
kimi
Math
open-source
Open-source
penetration-testing
qwen
reasoning
red-team
science
security
Trace