SEC 10-K Filings for Evaluating Chunking Algorithms
Ficha is a dataset of SEC 10-K financial filings designed to evaluate how well chunking algorithms handle formal business documents with complex financial terminology, tables, and structured sections.
This dataset tests chunking algorithms on:
| Field | Description |
|---|---|
ticker | Stock ticker symbol |
company | Company name |
filing_type | Type of SEC filing |
filing_date | Date of filing |
text | Full text of the filing |
| Field | Description |
|---|---|
ticker | Stock ticker symbol |
company | Company name |
question | Question about the filing |
answer | Answer to the question |
chunk-must-contain | Text passage that must be in the retrieved chunk |
from datasets import load_dataset
# Load corpus
corpus = load_dataset("chonkie-ai/ficha", "corpus", split="train")
# Load questions
questions = load_dataset("chonkie-ai/ficha", "questions", split="train")
Ficha is part of the Massive Text Chunking Benchmark (MTCB), a comprehensive benchmark for evaluating RAG chunking strategies.
CC-BY-4.0
10 commits
SEC 10-K Filings for Evaluating Chunking Algorithms
Ficha is a dataset of SEC 10-K financial filings designed to evaluate how well chunking algorithms handle formal business documents with complex financial terminology, tables, and structured sections.
This dataset tests chunking algorithms on:
| Field | Description |
|---|---|
ticker | Stock ticker symbol |
company | Company name |
filing_type | Type of SEC filing |
filing_date | Date of filing |
text | Full text of the filing |
| Field | Description |
|---|---|
ticker | Stock ticker symbol |
company | Company name |
question | Question about the filing |
answer | Answer to the question |
chunk-must-contain | Text passage that must be in the retrieved chunk |
from datasets import load_dataset
# Load corpus
corpus = load_dataset("chonkie-ai/ficha", "corpus", split="train")
# Load questions
questions = load_dataset("chonkie-ai/ficha", "questions", split="train")
Ficha is part of the Massive Text Chunking Benchmark (MTCB), a comprehensive benchmark for evaluating RAG chunking strategies.
CC-BY-4.0
10 commits