FAVA datasets include: annotation data and training data.
The annotation dataset includes 460 annotated passages identifying and editing errors using our hallucination taxonomy. This dataset was used for the fine-grained error detection task, using the annotated passages as the gold passages.
The training data includes 35k training instances of erroneous input and corrected output pairs using our synthetic data generation pipeline.
FAVA datasets include: annotation data and training data.
The annotation dataset includes 460 annotated passages identifying and editing errors using our hallucination taxonomy. This dataset was used for the fine-grained error detection task, using the annotated passages as the gold passages.
The training data includes 35k training instances of erroneous input and corrected output pairs using our synthetic data generation pipeline.