omkarthawakar/FannOrFlop

Dataset

Fann or Flop is the first comprehensive benchmark designed to evaluate large language models (LLMs) on their ability to understand Arabic poetry. It contains nearly 7,000 poem-explanation pairs covering 12 poetic eras, 21 genres, and multiple meters, providing a culturally rich and linguistically challenging testbed for Arabic NLP.

11

5 commits

1 linked in READMEs

updated Aug 21, 2025

See the code

README

Fann Or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding

Fann or Flop is the first comprehensive benchmark designed to evaluate large language models (LLMs) on their ability to understand Arabic poetry. It contains nearly 7,000 poem-explanation pairs covering 12 poetic eras, 21 genres, and multiple meters, providing a culturally rich and linguistically challenging testbed for Arabic NLP.


Latest Updates

🔥🔥 [20 Aug 2025] 🔥🔥 Fann or Flop accepted to EMNLP 2025 main track.
🔥 [26 May 2025] Fann or Flop the 1st benchmark for assessing the LLM's ability to comprehend and analyze Arabic poetry is released.
🤗 [19 Feb 2025] Fann or Flop dataset available on Hugging Face.


Key Features

  • Expert-Annotated Explanations: Verse-level commentary verified by native Arabic scholars.
  • 12 Historical Eras: From Pre-Islamic and Umayyad to Modern poetry.
  • Multi-Dimensional Evaluation: Faithfulness, fluency, metaphor, historical context, and rhetorical awareness.
  • Structured Taxonomy: Each poem tagged with meter, genre, and era.
  • QA-Style Format: Ideal for generative and comprehension-based evaluation in LLMs.

Dataset Summary

  • Name: Fann or Flop
  • Language: Arabic
  • Samples: 6,984 poem–explanation pairs
  • Task: Explanation generation, comprehension, QA-style evaluation
  • Annotation Level: Verse-level and poem-level explanations
  • Genres: مدح, هجاء, رثاء, غزل, etc.
  • Eras Covered: Pre-Islamic to Modern (e.g., Jahiliyyah, Abbasid, Ottoman, Contemporary)
  • Poetic Meters: الكامل, الطويل, البسيط, free verse, etc.

Dataset Structure

Each entry in the dataset contains:

FieldTypeDescription
idstringUnique poem identifier
titlestringTitle of the poem
authorstringName of the poet
sourcestringURL to original poem
tagslist[str]Meter, genre, and era (e.g., "الكامل", "مدح", "العصر الحديث")
meterstringPoetic meter (e.g., الكامل, الطويل)
genrestringPoetic genre (e.g., مدح, هجاء)
erastringHistorical era of the poem
verse_countintNumber of verses
poem_versesstringFull poem text (formatted with verse numbers)
explanationlist[dict]List of dictionaries, each containing a verse and its detailed explanation
raw_explanationstringFull poem explanation in paragraph format

Tasks and Use Cases

Fann or Flop can be used for a wide range of tasks, including:

  • Poetic Explanation Generation (LLM text generation)
  • Cultural and Historical QA (question answering from classical content)
  • Verse-Level Comprehension
  • Metrical & Stylistic Classification
  • Cultural Understanding Evaluation

Evaluation & Metrics

  • BLEU / chrF(++): Lexical overlap
  • BERTScore: Semantic similarity (AraBERT, etc.)
  • Textual Entailment: Consistency (mDeBERTa)
  • Human Evaluation: 0–10 scale scoring:
  • Literal understanding
  • Thematic/emotional depth
  • Cultural grounding
  • Stylistic sensitivity
  • Coherence and clarity

Model Benchmark Comparison on Fann or Flop

ModelBLEUchrF(++)BERTScoreTextual EntailmentFaithfulness / ConsistencyFluency / GrammaticalityInterpretive Depth
Closed Models
GPT-4o-2024-08-06 (OpenAI, 2024)0.03950.28820.64100.67753.92 (± 0.99)4.96 (± 0.20)7.52
GPT-4o-mini-2024-07-18 (OpenAI, 2024)0.03950.25420.61240.43832.91 (± 0.75)4.28 (± 0.57)7.50
Gemini-2.5-Flash (AI, 2025b)0.01530.26180.63190.74754.25 (± 1.00)4.98 (± 0.16)7.22
Gemini-2.0-Flash (AI, 2025a)0.03950.26180.63930.71543.99 (± 1.04)4.95 (± 0.22)6.50
Gemini-1.5-Pro (Reid et al., 2024)0.03950.26180.63330.61803.59 (± 1.00)4.80 (± 0.41)5.38
Fanar-Star (Team et al., 2025)0.01380.15380.56770.64682.16 (± 0.92)3.40 (± 0.76)2.88
Open Models
Deepseek-V3 (Liu et al., 2024)0.03950.27710.63350.51173.36 (± 0.91)4.98 (± 0.16)4.75
Deepseek-R1 (Guo et al., 2025)0.03950.27710.63350.51173.38 (± 0.92)4.98 (± 0.16)4.25
Llama-3.3-70B (Meta AI, 2024)0.01530.26180.63930.53642.51 (± 0.90)3.37 (± 0.73)7.20
Qwen-3 (Team, 2025)0.02960.28370.61580.64683.98 (± 0.90)4.73 (± 0.45)6.50
Aya-Expanse (Dang et al., 2024)0.03290.27710.63280.64683.76 (± 0.90)4.68 (± 0.47)5.88
Jais (Sengupta et al., 2023)0.03120.26980.62450.60233.21 (± 0.88)4.35 (± 0.52)5.35
ALLaM-7B (Bari et al., 2024)0.01190.04630.53750.59971.32 (± 0.62)2.11 (± 0.89)3.12
AceGPT-v2-70B-Chat (Huang et al., 2023)0.04020.04120.57590.60612.52 (± 0.91)3.46 (± 0.95)4.12

Citation

If you use Fann Or Flop dataset in your research, please consider citing:

@misc{alghallabi2025fannflopmultigenremultiera,
      title={Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs}, 
      author={Wafa Alghallabi and Ritesh Thawkar and Sara Ghaboura and Ketan More and Omkar Thawakar and Hisham Cholakkal and Salman Khan and Rao Muhammad Anwer},
      year={2025},
      eprint={2505.18152},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.18152}, 
}
Arabic
ArabicPoemUnderstanding
ArabicReasoning

omkarthawakar/FannOrFlop

Dataset

Fann or Flop is the first comprehensive benchmark designed to evaluate large language models (LLMs) on their ability to understand Arabic poetry. It contains nearly 7,000 poem-explanation pairs covering 12 poetic eras, 21 genres, and multiple meters, providing a culturally rich and linguistically challenging testbed for Arabic NLP.

11

5 commits

1 linked in READMEs

updated Aug 21, 2025

See the code

README

Fann Or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding

Fann or Flop is the first comprehensive benchmark designed to evaluate large language models (LLMs) on their ability to understand Arabic poetry. It contains nearly 7,000 poem-explanation pairs covering 12 poetic eras, 21 genres, and multiple meters, providing a culturally rich and linguistically challenging testbed for Arabic NLP.


Latest Updates

🔥🔥 [20 Aug 2025] 🔥🔥 Fann or Flop accepted to EMNLP 2025 main track.
🔥 [26 May 2025] Fann or Flop the 1st benchmark for assessing the LLM's ability to comprehend and analyze Arabic poetry is released.
🤗 [19 Feb 2025] Fann or Flop dataset available on Hugging Face.


Key Features

  • Expert-Annotated Explanations: Verse-level commentary verified by native Arabic scholars.
  • 12 Historical Eras: From Pre-Islamic and Umayyad to Modern poetry.
  • Multi-Dimensional Evaluation: Faithfulness, fluency, metaphor, historical context, and rhetorical awareness.
  • Structured Taxonomy: Each poem tagged with meter, genre, and era.
  • QA-Style Format: Ideal for generative and comprehension-based evaluation in LLMs.

Dataset Summary

  • Name: Fann or Flop
  • Language: Arabic
  • Samples: 6,984 poem–explanation pairs
  • Task: Explanation generation, comprehension, QA-style evaluation
  • Annotation Level: Verse-level and poem-level explanations
  • Genres: مدح, هجاء, رثاء, غزل, etc.
  • Eras Covered: Pre-Islamic to Modern (e.g., Jahiliyyah, Abbasid, Ottoman, Contemporary)
  • Poetic Meters: الكامل, الطويل, البسيط, free verse, etc.

Dataset Structure

Each entry in the dataset contains:

FieldTypeDescription
idstringUnique poem identifier
titlestringTitle of the poem
authorstringName of the poet
sourcestringURL to original poem
tagslist[str]Meter, genre, and era (e.g., "الكامل", "مدح", "العصر الحديث")
meterstringPoetic meter (e.g., الكامل, الطويل)
genrestringPoetic genre (e.g., مدح, هجاء)
erastringHistorical era of the poem
verse_countintNumber of verses
poem_versesstringFull poem text (formatted with verse numbers)
explanationlist[dict]List of dictionaries, each containing a verse and its detailed explanation
raw_explanationstringFull poem explanation in paragraph format

Tasks and Use Cases

Fann or Flop can be used for a wide range of tasks, including:

  • Poetic Explanation Generation (LLM text generation)
  • Cultural and Historical QA (question answering from classical content)
  • Verse-Level Comprehension
  • Metrical & Stylistic Classification
  • Cultural Understanding Evaluation

Evaluation & Metrics

  • BLEU / chrF(++): Lexical overlap
  • BERTScore: Semantic similarity (AraBERT, etc.)
  • Textual Entailment: Consistency (mDeBERTa)
  • Human Evaluation: 0–10 scale scoring:
  • Literal understanding
  • Thematic/emotional depth
  • Cultural grounding
  • Stylistic sensitivity
  • Coherence and clarity

Model Benchmark Comparison on Fann or Flop

ModelBLEUchrF(++)BERTScoreTextual EntailmentFaithfulness / ConsistencyFluency / GrammaticalityInterpretive Depth
Closed Models
GPT-4o-2024-08-06 (OpenAI, 2024)0.03950.28820.64100.67753.92 (± 0.99)4.96 (± 0.20)7.52
GPT-4o-mini-2024-07-18 (OpenAI, 2024)0.03950.25420.61240.43832.91 (± 0.75)4.28 (± 0.57)7.50
Gemini-2.5-Flash (AI, 2025b)0.01530.26180.63190.74754.25 (± 1.00)4.98 (± 0.16)7.22
Gemini-2.0-Flash (AI, 2025a)0.03950.26180.63930.71543.99 (± 1.04)4.95 (± 0.22)6.50
Gemini-1.5-Pro (Reid et al., 2024)0.03950.26180.63330.61803.59 (± 1.00)4.80 (± 0.41)5.38
Fanar-Star (Team et al., 2025)0.01380.15380.56770.64682.16 (± 0.92)3.40 (± 0.76)2.88
Open Models
Deepseek-V3 (Liu et al., 2024)0.03950.27710.63350.51173.36 (± 0.91)4.98 (± 0.16)4.75
Deepseek-R1 (Guo et al., 2025)0.03950.27710.63350.51173.38 (± 0.92)4.98 (± 0.16)4.25
Llama-3.3-70B (Meta AI, 2024)0.01530.26180.63930.53642.51 (± 0.90)3.37 (± 0.73)7.20
Qwen-3 (Team, 2025)0.02960.28370.61580.64683.98 (± 0.90)4.73 (± 0.45)6.50
Aya-Expanse (Dang et al., 2024)0.03290.27710.63280.64683.76 (± 0.90)4.68 (± 0.47)5.88
Jais (Sengupta et al., 2023)0.03120.26980.62450.60233.21 (± 0.88)4.35 (± 0.52)5.35
ALLaM-7B (Bari et al., 2024)0.01190.04630.53750.59971.32 (± 0.62)2.11 (± 0.89)3.12
AceGPT-v2-70B-Chat (Huang et al., 2023)0.04020.04120.57590.60612.52 (± 0.91)3.46 (± 0.95)4.12

Citation

If you use Fann Or Flop dataset in your research, please consider citing:

@misc{alghallabi2025fannflopmultigenremultiera,
      title={Fann or Flop: A Multigenre, Multiera Benchmark for Arabic Poetry Understanding in LLMs}, 
      author={Wafa Alghallabi and Ritesh Thawkar and Sara Ghaboura and Ketan More and Omkar Thawakar and Hisham Cholakkal and Salman Khan and Rao Muhammad Anwer},
      year={2025},
      eprint={2505.18152},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.18152}, 
}
Arabic
ArabicPoemUnderstanding
ArabicReasoning