VidScribe is a diagnostic benchmark for visual text in video generation. It has four tasks: T2V (render text from a prompt), R2V (transfer text identity from a reference image), I2V (keep text intact under motion from a first frame), and V2V (edit localized text in an existing video). Every sample is labeled on 12 factor axes (F1–F12).
Release status. This repository hosts the public half of the benchmark: 403 of 803 human-verified samples. The remaining samples will be released after the paper is accepted.
AI-generated content. Everything in
generated_videos/is AI-generated, as are some of the reference images and source videos.
| Task | Released / full | Input media |
|---|---|---|
| T2V | 125 / 250 | none |
| R2V | 70 / 139 | 70 reference images |
| I2V | 115 / 229 | 115 first frames |
| V2V | 93 / 185 | 93 source videos |
The repository also contains results for 7 commercial video generators on these samples: per-sample scores, a leaderboard, and the 2,821 scored videos (about 10 GB).
Each task was halved by stratified sampling with a fixed seed. For every value of F1–F12, the coverage attributes, the data source, and the difficulty quartile, the released share is within 2.1 percentage points of the full task. All samples shown in the paper are included.
benchmark/<task>/metadata_*.json ground truth; a task is the union of its files
benchmark/<task>/{images,videos}/ model inputs
results/leaderboard.csv model-level scores (%)
results/scores_per_sample.csv one row per (model, sample), scores in [0, 1]
results/metric_keys.json raw metric key -> ID, name, group
results/raw/<model>.jsonl per-metric value, validity, N/A reason
generated_videos/<model>/<id>.mp4 scored model outputs
viewer/*.parquet flat copies for the Dataset Viewer
All file paths stored in the data are relative to the repository root. Model ids: minimax-h3, wan-3.0, seedance-2.5, gemini-omni-1.1-flash, happyhorse, pixverse-v6, kling-3.0-omni.
sample_id, task, promptF1_text_amount … F12_content_dynamicslanguage, content_domain, visual_style, text_colorfont_name, text_color_desc, motion_pattern, occluder_type, deform_source, text_position, temporal_evidence, glyph_profile, v2v_edit_opgt_text (for V2V, the source text), gt_text_after (V2V edit target), gt_text_switched (after a content switch), gt_text_visible (visible part under off-screen cropping)reference_image (R2V, I2V), source_video (V2V)Fields that do not apply are null. For multi-level text, per-text fields are lists with one entry per level.
| Group | Metrics |
|---|---|
| Text Fidelity | A1 Content Accuracy, A2 Glyph Correctness |
| Temporal Stability | B1 Content Stability, B2 Appearance Stability |
| Instruction Compliance | C1 Color, C2 Stroke Weight, C3 Typographic Style, C4 Carrier Attachment, C5 Spatial Placement, C6 Layout Hierarchy |
| Non-Text Preservation | D1 Background Consistency |
| V2V probes (not in Overall) | E1 Editing Residue (lower is better), E2 Source Text Preservation |
Overall averages each metric over the samples where it is valid, then averages metrics within each group, then averages the 4 groups with equal weight. Inapplicable metrics are left empty and excluded. Failed generations are still scored (low) and flagged with model_output_failure.
These numbers are close to, but not identical with, the full-benchmark results in the paper.
| Model | Overall | A1 | A2 | B1 | B2 | C1 | C2 | C3 | C4 | C5 | C6 | D1 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MiniMax H3 | 79.0 | 83.0 | 56.6 | 75.4 | 82.6 | 82.2 | 73.4 | 73.9 | 74.2 | 71.8 | 81.8 | 90.8 |
| Gemini Omni 1.1 Flash | 78.3 | 79.4 | 51.9 | 74.6 | 84.5 | 82.3 | 73.0 | 75.4 | 73.3 | 76.8 | 84.0 | 90.4 |
| Wan 3.0 | 78.0 | 82.0 | 45.5 | 75.9 | 85.2 | 83.0 | 75.1 | 77.8 | 78.9 | 76.4 | 82.5 | 88.8 |
| Seedance 2.5 | 77.2 | 75.2 | 57.8 | 71.3 | 82.6 | 84.8 | 71.4 | 69.7 | 74.7 | 68.8 | 78.9 | 90.8 |
| HappyHorse | 74.5 | 72.6 | 48.9 | 66.2 | 80.3 | 82.6 | 71.2 | 69.7 | 72.2 | 69.8 | 81.4 | 89.4 |
| PixVerse V6 | 73.4 | 72.6 | 47.9 | 64.1 | 78.2 | 83.0 | 73.7 | 73.5 | 68.2 | 67.4 | 82.5 | 87.5 |
| Kling3.0-Omni | 70.9 | 58.5 | 45.1 | 56.0 | 80.2 | 79.6 | 72.7 | 67.6 | 71.8 | 68.1 | 83.2 | 89.8 |
| Model | T2V | R2V | I2V | V2V | E1 ↓ | E2 ↑ |
|---|---|---|---|---|---|---|
| MiniMax H3 | 80.4 | 82.1 | 83.5 | 67.8 | 31.6 | 50.0 |
| Gemini Omni 1.1 Flash | 81.0 | 76.6 | 81.7 | 70.8 | 26.3 | 72.3 |
| Wan 3.0 | 78.7 | 79.7 | 82.2 | 69.8 | 22.8 | 65.9 |
| Seedance 2.5 | 79.8 | 80.0 | 82.6 | 63.8 | 43.9 | 62.1 |
| HappyHorse | 75.4 | 78.2 | 80.2 | 61.9 | 31.6 | 39.1 |
| PixVerse V6 | 71.6 | 74.9 | 80.3 | 65.2 | 15.8 | 31.6 |
| Kling3.0-Omni | 66.8 | 72.3 | 84.0 | 57.4 | 36.8 | 50.2 |
from huggingface_hub import hf_hub_download, snapshot_download
from pathlib import Path
import json, pandas as pd
repo = "Vicky0720/VidScribe"
root = Path(snapshot_download(repo, repo_type="dataset", ignore_patterns=["generated_videos/*"]))
gt = [r for f in sorted(root.glob("benchmark/*/metadata_*.json"))
for r in json.loads(f.read_text(encoding="utf-8"))]
scores = pd.read_csv(root / "results/scores_per_sample.csv")
video = hf_hub_download(repo, scores.generated_video[0], repo_type="dataset")
real_textocr_ or real_lsvt_ are derived from TextOCR and LSVT and remain subject to their terms. The generated videos and scores concern third-party commercial services, reflect the model versions available at evaluation time, and remain subject to each provider's terms, including any restriction on training competing models.sample_id, and the material will be removed.@article{vidscribe2026,
title = {Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation},
author = {TODO},
journal = {arXiv preprint arXiv:TODO},
year = {2026}
}
VidScribe is a diagnostic benchmark for visual text in video generation. It has four tasks: T2V (render text from a prompt), R2V (transfer text identity from a reference image), I2V (keep text intact under motion from a first frame), and V2V (edit localized text in an existing video). Every sample is labeled on 12 factor axes (F1–F12).
Release status. This repository hosts the public half of the benchmark: 403 of 803 human-verified samples. The remaining samples will be released after the paper is accepted.
AI-generated content. Everything in
generated_videos/is AI-generated, as are some of the reference images and source videos.
| Task | Released / full | Input media |
|---|---|---|
| T2V | 125 / 250 | none |
| R2V | 70 / 139 | 70 reference images |
| I2V | 115 / 229 | 115 first frames |
| V2V | 93 / 185 | 93 source videos |
The repository also contains results for 7 commercial video generators on these samples: per-sample scores, a leaderboard, and the 2,821 scored videos (about 10 GB).
Each task was halved by stratified sampling with a fixed seed. For every value of F1–F12, the coverage attributes, the data source, and the difficulty quartile, the released share is within 2.1 percentage points of the full task. All samples shown in the paper are included.
benchmark/<task>/metadata_*.json ground truth; a task is the union of its files
benchmark/<task>/{images,videos}/ model inputs
results/leaderboard.csv model-level scores (%)
results/scores_per_sample.csv one row per (model, sample), scores in [0, 1]
results/metric_keys.json raw metric key -> ID, name, group
results/raw/<model>.jsonl per-metric value, validity, N/A reason
generated_videos/<model>/<id>.mp4 scored model outputs
viewer/*.parquet flat copies for the Dataset Viewer
All file paths stored in the data are relative to the repository root. Model ids: minimax-h3, wan-3.0, seedance-2.5, gemini-omni-1.1-flash, happyhorse, pixverse-v6, kling-3.0-omni.
sample_id, task, promptF1_text_amount … F12_content_dynamicslanguage, content_domain, visual_style, text_colorfont_name, text_color_desc, motion_pattern, occluder_type, deform_source, text_position, temporal_evidence, glyph_profile, v2v_edit_opgt_text (for V2V, the source text), gt_text_after (V2V edit target), gt_text_switched (after a content switch), gt_text_visible (visible part under off-screen cropping)reference_image (R2V, I2V), source_video (V2V)Fields that do not apply are null. For multi-level text, per-text fields are lists with one entry per level.
| Group | Metrics |
|---|---|
| Text Fidelity | A1 Content Accuracy, A2 Glyph Correctness |
| Temporal Stability | B1 Content Stability, B2 Appearance Stability |
| Instruction Compliance | C1 Color, C2 Stroke Weight, C3 Typographic Style, C4 Carrier Attachment, C5 Spatial Placement, C6 Layout Hierarchy |
| Non-Text Preservation | D1 Background Consistency |
| V2V probes (not in Overall) | E1 Editing Residue (lower is better), E2 Source Text Preservation |
Overall averages each metric over the samples where it is valid, then averages metrics within each group, then averages the 4 groups with equal weight. Inapplicable metrics are left empty and excluded. Failed generations are still scored (low) and flagged with model_output_failure.
These numbers are close to, but not identical with, the full-benchmark results in the paper.
| Model | Overall | A1 | A2 | B1 | B2 | C1 | C2 | C3 | C4 | C5 | C6 | D1 |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| MiniMax H3 | 79.0 | 83.0 | 56.6 | 75.4 | 82.6 | 82.2 | 73.4 | 73.9 | 74.2 | 71.8 | 81.8 | 90.8 |
| Gemini Omni 1.1 Flash | 78.3 | 79.4 | 51.9 | 74.6 | 84.5 | 82.3 | 73.0 | 75.4 | 73.3 | 76.8 | 84.0 | 90.4 |
| Wan 3.0 | 78.0 | 82.0 | 45.5 | 75.9 | 85.2 | 83.0 | 75.1 | 77.8 | 78.9 | 76.4 | 82.5 | 88.8 |
| Seedance 2.5 | 77.2 | 75.2 | 57.8 | 71.3 | 82.6 | 84.8 | 71.4 | 69.7 | 74.7 | 68.8 | 78.9 | 90.8 |
| HappyHorse | 74.5 | 72.6 | 48.9 | 66.2 | 80.3 | 82.6 | 71.2 | 69.7 | 72.2 | 69.8 | 81.4 | 89.4 |
| PixVerse V6 | 73.4 | 72.6 | 47.9 | 64.1 | 78.2 | 83.0 | 73.7 | 73.5 | 68.2 | 67.4 | 82.5 | 87.5 |
| Kling3.0-Omni | 70.9 | 58.5 | 45.1 | 56.0 | 80.2 | 79.6 | 72.7 | 67.6 | 71.8 | 68.1 | 83.2 | 89.8 |
| Model | T2V | R2V | I2V | V2V | E1 ↓ | E2 ↑ |
|---|---|---|---|---|---|---|
| MiniMax H3 | 80.4 | 82.1 | 83.5 | 67.8 | 31.6 | 50.0 |
| Gemini Omni 1.1 Flash | 81.0 | 76.6 | 81.7 | 70.8 | 26.3 | 72.3 |
| Wan 3.0 | 78.7 | 79.7 | 82.2 | 69.8 | 22.8 | 65.9 |
| Seedance 2.5 | 79.8 | 80.0 | 82.6 | 63.8 | 43.9 | 62.1 |
| HappyHorse | 75.4 | 78.2 | 80.2 | 61.9 | 31.6 | 39.1 |
| PixVerse V6 | 71.6 | 74.9 | 80.3 | 65.2 | 15.8 | 31.6 |
| Kling3.0-Omni | 66.8 | 72.3 | 84.0 | 57.4 | 36.8 | 50.2 |
from huggingface_hub import hf_hub_download, snapshot_download
from pathlib import Path
import json, pandas as pd
repo = "Vicky0720/VidScribe"
root = Path(snapshot_download(repo, repo_type="dataset", ignore_patterns=["generated_videos/*"]))
gt = [r for f in sorted(root.glob("benchmark/*/metadata_*.json"))
for r in json.loads(f.read_text(encoding="utf-8"))]
scores = pd.read_csv(root / "results/scores_per_sample.csv")
video = hf_hub_download(repo, scores.generated_video[0], repo_type="dataset")
real_textocr_ or real_lsvt_ are derived from TextOCR and LSVT and remain subject to their terms. The generated videos and scores concern third-party commercial services, reflect the model versions available at evaluation time, and remain subject to each provider's terms, including any restriction on training competing models.sample_id, and the material will be removed.@article{vidscribe2026,
title = {Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation},
author = {TODO},
journal = {arXiv preprint arXiv:TODO},
year = {2026}
}