MMLottieBench is a comprehensive evaluation protocol for multi-modal vector animation generation. The lack of mature and standardized benchmarks and metrics for vector animation generation poses significant challenges in evaluating (1) the quality of generated vector animations and (2) the extent to which generators faithfully follow multi-modal instructions.
Our benchmark addresses these challenges by providing:
MMLottieBench aims to construct a benchmark that:
| Feature | Type | Description |
|---|---|---|
| id | string | MD5 hash uniquely identifying each sample |
| text | string | Text description or prompt (may be None for Video-to-Lottie) |
| image | image | Reference image (may be None for Text-to-Lottie and Video-to-Lottie) |
| video | video | Reference video (may be None for Text-to-Lottie and Text-Image-to-Lottie) |
| task_type | string | Task type: "Text-to-Lottie", "Text-Image-to-Lottie", or "Video-to-Lottie" |
| subset | string | Subset type: "Real" or "Synthetic" |
| url | string | Source URL for synthetic samples (None for real samples) |
The Real Subset consists of samples curated from artist-designed Lottie animations collected from professional designers. All evaluation samples are strictly disjoint from the training data, ensuring assessment on genuinely unseen, real-world content.
Key Features:
To ensure the fairness and long-term robustness of our benchmark—particularly to mitigate potential contamination from future models trained on overly similar data—we construct a complementary Synthetic Subset via instruction-based synthesis using state-of-the-art generative models.
We synthesize 150 textual prompts using GPT-4o with carefully designed meta-prompts. The generation instruction ensures high-quality, diverse, and challenging animation prompts suitable for evaluating Lottie generation models.
Key Requirements:
Motion Complexity Distribution:
Object Type Coverage:
Allowed Motion Primitives:
Example Prompts:
Why Synthetic Subset?
The synthetic nature of MMLottieBench's Synthetic Subset provides several key advantages:
| Advantage | Description |
|---|---|
| True Generalization Test | Models cannot have seen these exact samples during training |
| Controlled Diversity | Systematic coverage of styles, complexities, and animation patterns |
| Reproducibility | The entire synthesis process is documented and released |
| Fairness | No model has an unfair advantage from training data overlap |
| Long-term Robustness | Reduces risk of benchmark contamination in future models |
MMLottieBench evaluates models across multiple dimensions to comprehensively assess both visual quality and semantic alignment:
Both alignment metrics use Claude-3.5-Sonnet as an LLM judge. Invalid generations are omitted from evaluation, and blank outputs receive a score of 0.
We provide comprehensive quantitative comparisons between state-of-the-art baseline methods across both Real Subset and Synthetic Subset. Bold numbers and underlined numbers represent the best and second-best performance respectively.
| Methods | Time(s) | # Tokens | Success Rate | FVD↓ | CLIP↑ | Obj.↑ | Motion↑ |
|---|---|---|---|---|---|---|---|
| DeepSeekV3 | 43.40 | 2.3k | 9.3% | 671.80 | 0.2677 | 1.51 | 2.09 |
| Qwen2.5-VL(3B) | 27.97 | 0.5k | 0.0% | - | - | - | - |
| GPT-5 | 43.40 | 1.4k | 12.7% | 715.73 | 0.2600 | 0.73 | 0.71 |
| Recraft | - | 54.1k | 77.3% | 300.70 | 0.2950 | 4.70 | 4.68 |
| Ours | 33.71 | 21.2k | 88.3% | 202.14 | 0.2748 | 4.44 | 5.94 |
| Methods | Time(s) | # Tokens | Success Rate | FVD↓ | CLIP↑ | Obj.↑ | Motion↑ |
|---|---|---|---|---|---|---|---|
| Qwen2.5-VL(3B) | 33.60 | 0.4k | 0.0% | - | - | - | - |
| GPT-5 | 31.18 | 1.5k | 28.0% | 546.65 | 0.2557 | 1.18 | 0.95 |
| AniClipart | 1212.34 | - | 87.3% | 266.46 | 0.2935 | 4.51 | 3.47 |
| Livesketch | 723.23 | - | 91.3% | 868.18 | 0.2309 | 2.84 | 2.42 |
| Ours | 88.57 | 23.4k | 93.3% | 180.27 | 0.2666 | 5.10 | 4.44 |
| Methods | Time(s) | # Tokens | Success Rate | FVD↓ | PSNR↑ | SSIM↑ | DINO↑ |
|---|---|---|---|---|---|---|---|
| Qwen2.5-VL(3B) | 49.31 | 1.0k | 0.0% | - | - | - | - |
| GPT-5 | 45.61 | 1.1k | 9.2% | 639.13 | 13.34 | 0.81 | 0.80 |
| Gemini3.1-Pro | 16.19 | 1.0k | 0.0% | 1076.22 | 14.54 | 0.79 | 0.88 |
| Ours | 110.77 | 36.8k | 88.1% | 227.11 | 16.08 | 0.82 | 0.92 |
| Methods | Time(s) | # Tokens | Success Rate | FVD↓ | CLIP↑ | Obj.↑ | Motion↑ |
|---|---|---|---|---|---|---|---|
| DeepSeekV3 | 56.71 | 2.3k | 7.4% | 483.11 | 0.2677 | 1.43 | 1.98 |
| Qwen2.5-VL(3B) | 94.36 | 0.4k | 0.0% | - | - | - | - |
| GPT-5 | 57.59 | 0.9k | 8.8% | 637.29 | 0.2600 | 0.45 | 0.66 |
| Recraft | - | 50.8k | 77.3% | 438.97 | 0.2950 | 4.33 | 3.12 |
| Ours | 37.93 | 13.4k | 82.1% | 206.35 | 0.2748 | 4.31 | 5.63 |
| Methods | Time(s) | # Tokens | Success Rate | FVD↓ | CLIP↑ | Obj.↑ | Motion↑ |
|---|---|---|---|---|---|---|---|
| Qwen2.5-VL(3B) | 31.06 | 0.3k | 0.0% | - | - | - | - |
| GPT-5 | 37.80 | 1.2k | 22.0% | 560.11 | 0.2557 | 1.02 | 0.66 |
| AniClipart | 1123.24 | - | 88.7% | 308.54 | 0.2935 | 4.11 | 2.79 |
| Livesketch | 742.23 | - | 91.9% | 1058.32 | 0.2309 | 2.01 | 1.91 |
| Ours | 84.80 | 16.3k | 92.9% | 225.45 | 0.2666 | 4.44 | 3.98 |
| Methods | Time(s) | # Tokens | Success Rate | FVD↓ | PSNR↑ | SSIM↑ | DINO↑ |
|---|---|---|---|---|---|---|---|
| Qwen2.5-VL(3B) | 42.19 | 1.1k | 0.0% | - | - | - | - |
| GPT-5 | 26.26 | 0.9k | 7.4% | 576.52 | 13.33 | 0.71 | 0.78 |
| Gemini3.1-Pro | 13.77 | 1.3k | 0.0% | 1550.65 | 13.89 | 0.75 | 0.83 |
| Ours | 109.53 | 41.4k | 80.7% | 342.65 | 15.76 | 0.79 | 0.88 |
Superior Visual Quality: Our method consistently achieves the best FVD scores across all tasks and subsets, demonstrating superior visual quality in generated Lottie animations.
Strong Motion Alignment: Our model excels in motion alignment, significantly outperforming baselines in capturing and reproducing complex animation patterns (5.94 vs 4.68 on Real Text-to-Lottie).
High Success Rate: With success rates of 80-93%, our method reliably generates valid Lottie animations, far exceeding general-purpose VLMs like GPT-4o (7-28%) and Qwen2.5-VL (0%).
Balanced Performance: While some baselines excel in specific metrics (e.g., Recraft in CLIP score), our method achieves the best overall balance across visual quality, semantic alignment, and generation reliability.
Token Efficiency Trade-off: Our method uses more tokens (13-42k) compared to VLM baselines (0.3-2.3k) but significantly fewer than optimization-based methods like Recraft (50-54k), striking a balance between expressiveness and efficiency.
from datasets import load_dataset
# Load the entire dataset
dataset = load_dataset("OmniLottie/MMLottieBench")
# Access specific subsets
real_subset = dataset["real"]
synthetic_subset = dataset["synthetic"]
# Filter by task type
text2lottie = real_subset.filter(lambda x: x["task_type"] == "Text-to-Lottie")
image2lottie = real_subset.filter(lambda x: x["task_type"] == "Text-Image-to-Lottie")
video2lottie = real_subset.filter(lambda x: x["task_type"] == "Video-to-Lottie")
# Example: iterate over text-to-lottie samples
for sample in text2lottie:
print(f"ID: {sample['id']}")
print(f"Text: {sample['text']}")
print(f"Subset: {sample['subset']}")
# Generate Lottie animation based on the prompt
| Subset | Text-to-Lottie | Text-Image-to-Lottie | Video-to-Lottie | Total |
|---|---|---|---|---|
| Real | 150 | 150 | 150 | 450 |
| Synthetic | 150 | 150 | 150 | 450 |
| Total | 300 | 300 | 300 | 900 |
If you use MMLottieBench in your research, please cite:
@article{yang2026omnilottie,
title={OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens},
author={Yiying Yang and Wei Cheng and Sijin Chen and Honghao Fu and Xianfang Zeng and Yujun Cai and Gang Yu and Xinjun Ma},
journal={arXiv preprint arxiv:2603.02138},
year={2026}
}
This dataset is released under the Apache 2.0 License.
For questions, issues, or contributions, please open an issue on our GitHub repository or contact us at [25113050158@m.fudan.edu.cn].
Note: This benchmark is designed for research purposes to advance the field of vector animation generation. All synthetic data generation processes are fully documented to ensure transparency and reproducibility.
15 commits
MMLottieBench is a comprehensive evaluation protocol for multi-modal vector animation generation. The lack of mature and standardized benchmarks and metrics for vector animation generation poses significant challenges in evaluating (1) the quality of generated vector animations and (2) the extent to which generators faithfully follow multi-modal instructions.
Our benchmark addresses these challenges by providing:
MMLottieBench aims to construct a benchmark that:
| Feature | Type | Description |
|---|---|---|
| id | string | MD5 hash uniquely identifying each sample |
| text | string | Text description or prompt (may be None for Video-to-Lottie) |
| image | image | Reference image (may be None for Text-to-Lottie and Video-to-Lottie) |
| video | video | Reference video (may be None for Text-to-Lottie and Text-Image-to-Lottie) |
| task_type | string | Task type: "Text-to-Lottie", "Text-Image-to-Lottie", or "Video-to-Lottie" |
| subset | string | Subset type: "Real" or "Synthetic" |
| url | string | Source URL for synthetic samples (None for real samples) |
The Real Subset consists of samples curated from artist-designed Lottie animations collected from professional designers. All evaluation samples are strictly disjoint from the training data, ensuring assessment on genuinely unseen, real-world content.
Key Features:
To ensure the fairness and long-term robustness of our benchmark—particularly to mitigate potential contamination from future models trained on overly similar data—we construct a complementary Synthetic Subset via instruction-based synthesis using state-of-the-art generative models.
We synthesize 150 textual prompts using GPT-4o with carefully designed meta-prompts. The generation instruction ensures high-quality, diverse, and challenging animation prompts suitable for evaluating Lottie generation models.
Key Requirements:
Motion Complexity Distribution:
Object Type Coverage:
Allowed Motion Primitives:
Example Prompts:
Why Synthetic Subset?
The synthetic nature of MMLottieBench's Synthetic Subset provides several key advantages:
| Advantage | Description |
|---|---|
| True Generalization Test | Models cannot have seen these exact samples during training |
| Controlled Diversity | Systematic coverage of styles, complexities, and animation patterns |
| Reproducibility | The entire synthesis process is documented and released |
| Fairness | No model has an unfair advantage from training data overlap |
| Long-term Robustness | Reduces risk of benchmark contamination in future models |
MMLottieBench evaluates models across multiple dimensions to comprehensively assess both visual quality and semantic alignment:
Both alignment metrics use Claude-3.5-Sonnet as an LLM judge. Invalid generations are omitted from evaluation, and blank outputs receive a score of 0.
We provide comprehensive quantitative comparisons between state-of-the-art baseline methods across both Real Subset and Synthetic Subset. Bold numbers and underlined numbers represent the best and second-best performance respectively.
| Methods | Time(s) | # Tokens | Success Rate | FVD↓ | CLIP↑ | Obj.↑ | Motion↑ |
|---|---|---|---|---|---|---|---|
| DeepSeekV3 | 43.40 | 2.3k | 9.3% | 671.80 | 0.2677 | 1.51 | 2.09 |
| Qwen2.5-VL(3B) | 27.97 | 0.5k | 0.0% | - | - | - | - |
| GPT-5 | 43.40 | 1.4k | 12.7% | 715.73 | 0.2600 | 0.73 | 0.71 |
| Recraft | - | 54.1k | 77.3% | 300.70 | 0.2950 | 4.70 | 4.68 |
| Ours | 33.71 | 21.2k | 88.3% | 202.14 | 0.2748 | 4.44 | 5.94 |
| Methods | Time(s) | # Tokens | Success Rate | FVD↓ | CLIP↑ | Obj.↑ | Motion↑ |
|---|---|---|---|---|---|---|---|
| Qwen2.5-VL(3B) | 33.60 | 0.4k | 0.0% | - | - | - | - |
| GPT-5 | 31.18 | 1.5k | 28.0% | 546.65 | 0.2557 | 1.18 | 0.95 |
| AniClipart | 1212.34 | - | 87.3% | 266.46 | 0.2935 | 4.51 | 3.47 |
| Livesketch | 723.23 | - | 91.3% | 868.18 | 0.2309 | 2.84 | 2.42 |
| Ours | 88.57 | 23.4k | 93.3% | 180.27 | 0.2666 | 5.10 | 4.44 |
| Methods | Time(s) | # Tokens | Success Rate | FVD↓ | PSNR↑ | SSIM↑ | DINO↑ |
|---|---|---|---|---|---|---|---|
| Qwen2.5-VL(3B) | 49.31 | 1.0k | 0.0% | - | - | - | - |
| GPT-5 | 45.61 | 1.1k | 9.2% | 639.13 | 13.34 | 0.81 | 0.80 |
| Gemini3.1-Pro | 16.19 | 1.0k | 0.0% | 1076.22 | 14.54 | 0.79 | 0.88 |
| Ours | 110.77 | 36.8k | 88.1% | 227.11 | 16.08 | 0.82 | 0.92 |
| Methods | Time(s) | # Tokens | Success Rate | FVD↓ | CLIP↑ | Obj.↑ | Motion↑ |
|---|---|---|---|---|---|---|---|
| DeepSeekV3 | 56.71 | 2.3k | 7.4% | 483.11 | 0.2677 | 1.43 | 1.98 |
| Qwen2.5-VL(3B) | 94.36 | 0.4k | 0.0% | - | - | - | - |
| GPT-5 | 57.59 | 0.9k | 8.8% | 637.29 | 0.2600 | 0.45 | 0.66 |
| Recraft | - | 50.8k | 77.3% | 438.97 | 0.2950 | 4.33 | 3.12 |
| Ours | 37.93 | 13.4k | 82.1% | 206.35 | 0.2748 | 4.31 | 5.63 |
| Methods | Time(s) | # Tokens | Success Rate | FVD↓ | CLIP↑ | Obj.↑ | Motion↑ |
|---|---|---|---|---|---|---|---|
| Qwen2.5-VL(3B) | 31.06 | 0.3k | 0.0% | - | - | - | - |
| GPT-5 | 37.80 | 1.2k | 22.0% | 560.11 | 0.2557 | 1.02 | 0.66 |
| AniClipart | 1123.24 | - | 88.7% | 308.54 | 0.2935 | 4.11 | 2.79 |
| Livesketch | 742.23 | - | 91.9% | 1058.32 | 0.2309 | 2.01 | 1.91 |
| Ours | 84.80 | 16.3k | 92.9% | 225.45 | 0.2666 | 4.44 | 3.98 |
| Methods | Time(s) | # Tokens | Success Rate | FVD↓ | PSNR↑ | SSIM↑ | DINO↑ |
|---|---|---|---|---|---|---|---|
| Qwen2.5-VL(3B) | 42.19 | 1.1k | 0.0% | - | - | - | - |
| GPT-5 | 26.26 | 0.9k | 7.4% | 576.52 | 13.33 | 0.71 | 0.78 |
| Gemini3.1-Pro | 13.77 | 1.3k | 0.0% | 1550.65 | 13.89 | 0.75 | 0.83 |
| Ours | 109.53 | 41.4k | 80.7% | 342.65 | 15.76 | 0.79 | 0.88 |
Superior Visual Quality: Our method consistently achieves the best FVD scores across all tasks and subsets, demonstrating superior visual quality in generated Lottie animations.
Strong Motion Alignment: Our model excels in motion alignment, significantly outperforming baselines in capturing and reproducing complex animation patterns (5.94 vs 4.68 on Real Text-to-Lottie).
High Success Rate: With success rates of 80-93%, our method reliably generates valid Lottie animations, far exceeding general-purpose VLMs like GPT-4o (7-28%) and Qwen2.5-VL (0%).
Balanced Performance: While some baselines excel in specific metrics (e.g., Recraft in CLIP score), our method achieves the best overall balance across visual quality, semantic alignment, and generation reliability.
Token Efficiency Trade-off: Our method uses more tokens (13-42k) compared to VLM baselines (0.3-2.3k) but significantly fewer than optimization-based methods like Recraft (50-54k), striking a balance between expressiveness and efficiency.
from datasets import load_dataset
# Load the entire dataset
dataset = load_dataset("OmniLottie/MMLottieBench")
# Access specific subsets
real_subset = dataset["real"]
synthetic_subset = dataset["synthetic"]
# Filter by task type
text2lottie = real_subset.filter(lambda x: x["task_type"] == "Text-to-Lottie")
image2lottie = real_subset.filter(lambda x: x["task_type"] == "Text-Image-to-Lottie")
video2lottie = real_subset.filter(lambda x: x["task_type"] == "Video-to-Lottie")
# Example: iterate over text-to-lottie samples
for sample in text2lottie:
print(f"ID: {sample['id']}")
print(f"Text: {sample['text']}")
print(f"Subset: {sample['subset']}")
# Generate Lottie animation based on the prompt
| Subset | Text-to-Lottie | Text-Image-to-Lottie | Video-to-Lottie | Total |
|---|---|---|---|---|
| Real | 150 | 150 | 150 | 450 |
| Synthetic | 150 | 150 | 150 | 450 |
| Total | 300 | 300 | 300 | 900 |
If you use MMLottieBench in your research, please cite:
@article{yang2026omnilottie,
title={OmniLottie: Generating Vector Animations via Parameterized Lottie Tokens},
author={Yiying Yang and Wei Cheng and Sijin Chen and Honghao Fu and Xianfang Zeng and Yujun Cai and Gang Yu and Xinjun Ma},
journal={arXiv preprint arxiv:2603.02138},
year={2026}
}
This dataset is released under the Apache 2.0 License.
For questions, issues, or contributions, please open an issue on our GitHub repository or contact us at [25113050158@m.fudan.edu.cn].
Note: This benchmark is designed for research purposes to advance the field of vector animation generation. All synthetic data generation processes are fully documented to ensure transparency and reproducibility.
15 commits