Vicky0720/VidScribe

Dataset

VidScribe

1

15 commits

1 linked in READMEs

updated Sep 28, 2026

See the code

README

VidScribe

VidScribe is a diagnostic benchmark for visual text in video generation. It has four tasks: T2V (render text from a prompt), R2V (transfer text identity from a reference image), I2V (keep text intact under motion from a first frame), and V2V (edit localized text in an existing video). Every sample is labeled on 12 factor axes (F1–F12).

Release status. This repository hosts the public half of the benchmark: 403 of 803 human-verified samples. The remaining samples will be released after the paper is accepted.

AI-generated content. Everything in generated_videos/ is AI-generated, as are some of the reference images and source videos.

Contents

TaskReleased / fullInput media
T2V125 / 250none
R2V70 / 13970 reference images
I2V115 / 229115 first frames
V2V93 / 18593 source videos

The repository also contains results for 7 commercial video generators on these samples: per-sample scores, a leaderboard, and the 2,821 scored videos (about 10 GB).

Each task was halved by stratified sampling with a fixed seed. For every value of F1–F12, the coverage attributes, the data source, and the difficulty quartile, the released share is within 2.1 percentage points of the full task. All samples shown in the paper are included.

benchmark/<task>/metadata_*.json       ground truth; a task is the union of its files
benchmark/<task>/{images,videos}/      model inputs
results/leaderboard.csv                model-level scores (%)
results/scores_per_sample.csv          one row per (model, sample), scores in [0, 1]
results/metric_keys.json               raw metric key -> ID, name, group
results/raw/<model>.jsonl              per-metric value, validity, N/A reason
generated_videos/<model>/<id>.mp4      scored model outputs
viewer/*.parquet                       flat copies for the Dataset Viewer

All file paths stored in the data are relative to the repository root. Model ids: minimax-h3, wan-3.0, seedance-2.5, gemini-omni-1.1-flash, happyhorse, pixverse-v6, kling-3.0-omni.

Metadata

  • sample_id, task, prompt
  • Factor axes: F1_text_amount … F12_content_dynamics
  • Coverage attributes: language, content_domain, visual_style, text_color
  • Conditional controls: font_name, text_color_desc, motion_pattern, occluder_type, deform_source, text_position, temporal_evidence, glyph_profile, v2v_edit_op
  • Target text: gt_text (for V2V, the source text), gt_text_after (V2V edit target), gt_text_switched (after a content switch), gt_text_visible (visible part under off-screen cropping)
  • Inputs: reference_image (R2V, I2V), source_video (V2V)

Fields that do not apply are null. For multi-level text, per-text fields are lists with one entry per level.

Metrics

GroupMetrics
Text FidelityA1 Content Accuracy, A2 Glyph Correctness
Temporal StabilityB1 Content Stability, B2 Appearance Stability
Instruction ComplianceC1 Color, C2 Stroke Weight, C3 Typographic Style, C4 Carrier Attachment, C5 Spatial Placement, C6 Layout Hierarchy
Non-Text PreservationD1 Background Consistency
V2V probes (not in Overall)E1 Editing Residue (lower is better), E2 Source Text Preservation

Overall averages each metric over the samples where it is valid, then averages metrics within each group, then averages the 4 groups with equal weight. Inapplicable metrics are left empty and excluded. Failed generations are still scored (low) and flagged with model_output_failure.

Leaderboard on the released half (%)

These numbers are close to, but not identical with, the full-benchmark results in the paper.

ModelOverallA1A2B1B2C1C2C3C4C5C6D1
MiniMax H379.083.056.675.482.682.273.473.974.271.881.890.8
Gemini Omni 1.1 Flash78.379.451.974.684.582.373.075.473.376.884.090.4
Wan 3.078.082.045.575.985.283.075.177.878.976.482.588.8
Seedance 2.577.275.257.871.382.684.871.469.774.768.878.990.8
HappyHorse74.572.648.966.280.382.671.269.772.269.881.489.4
PixVerse V673.472.647.964.178.283.073.773.568.267.482.587.5
Kling3.0-Omni70.958.545.156.080.279.672.767.671.868.183.289.8
ModelT2VR2VI2VV2VE1 ↓E2 ↑
MiniMax H380.482.183.567.831.650.0
Gemini Omni 1.1 Flash81.076.681.770.826.372.3
Wan 3.078.779.782.269.822.865.9
Seedance 2.579.880.082.663.843.962.1
HappyHorse75.478.280.261.931.639.1
PixVerse V671.674.980.365.215.831.6
Kling3.0-Omni66.872.384.057.436.850.2

Usage

from huggingface_hub import hf_hub_download, snapshot_download
from pathlib import Path
import json, pandas as pd

repo = "Vicky0720/VidScribe"
root = Path(snapshot_download(repo, repo_type="dataset", ignore_patterns=["generated_videos/*"]))
gt = [r for f in sorted(root.glob("benchmark/*/metadata_*.json"))
      for r in json.loads(f.read_text(encoding="utf-8"))]
scores = pd.read_csv(root / "results/scores_per_sample.csv")
video = hf_hub_download(repo, scores.generated_video[0], repo_type="dataset")

Data notes and terms

  • Processing. Embedded metadata was removed from all media files; videos were remuxed without re-encoding. The generated videos are the exact bitstreams that were scored, and their visible content (including any service watermarks) is unchanged.
  • Privacy. Faces are masked in the I2V first frames and the V2V source videos. Some media contain incidental scene text. Do not use this dataset to identify or profile people.
  • Third-party material. Files prefixed real_textocr_ or real_lsvt_ are derived from TextOCR and LSVT and remain subject to their terms. The generated videos and scores concern third-party commercial services, reflect the model versions available at evaluation time, and remain subject to each provider's terms, including any restriction on training competing models.
  • Scores. All scores come from an automated evaluator and may contain errors.
  • Use. Non-commercial research only.
  • Takedown. Rights holders can open a discussion with the sample_id, and the material will be removed.

Citation

@article{vidscribe2026,
  title   = {Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation},
  author  = {TODO},
  journal = {arXiv preprint arXiv:TODO},
  year    = {2026}
}
ai-generated
benchmark
text-rendering
video-editing
video-generation

Vicky0720/VidScribe

Dataset

VidScribe

1

15 commits

1 linked in READMEs

updated Sep 28, 2026

See the code

README

VidScribe

VidScribe is a diagnostic benchmark for visual text in video generation. It has four tasks: T2V (render text from a prompt), R2V (transfer text identity from a reference image), I2V (keep text intact under motion from a first frame), and V2V (edit localized text in an existing video). Every sample is labeled on 12 factor axes (F1–F12).

Release status. This repository hosts the public half of the benchmark: 403 of 803 human-verified samples. The remaining samples will be released after the paper is accepted.

AI-generated content. Everything in generated_videos/ is AI-generated, as are some of the reference images and source videos.

Contents

TaskReleased / fullInput media
T2V125 / 250none
R2V70 / 13970 reference images
I2V115 / 229115 first frames
V2V93 / 18593 source videos

The repository also contains results for 7 commercial video generators on these samples: per-sample scores, a leaderboard, and the 2,821 scored videos (about 10 GB).

Each task was halved by stratified sampling with a fixed seed. For every value of F1–F12, the coverage attributes, the data source, and the difficulty quartile, the released share is within 2.1 percentage points of the full task. All samples shown in the paper are included.

benchmark/<task>/metadata_*.json       ground truth; a task is the union of its files
benchmark/<task>/{images,videos}/      model inputs
results/leaderboard.csv                model-level scores (%)
results/scores_per_sample.csv          one row per (model, sample), scores in [0, 1]
results/metric_keys.json               raw metric key -> ID, name, group
results/raw/<model>.jsonl              per-metric value, validity, N/A reason
generated_videos/<model>/<id>.mp4      scored model outputs
viewer/*.parquet                       flat copies for the Dataset Viewer

All file paths stored in the data are relative to the repository root. Model ids: minimax-h3, wan-3.0, seedance-2.5, gemini-omni-1.1-flash, happyhorse, pixverse-v6, kling-3.0-omni.

Metadata

  • sample_id, task, prompt
  • Factor axes: F1_text_amount … F12_content_dynamics
  • Coverage attributes: language, content_domain, visual_style, text_color
  • Conditional controls: font_name, text_color_desc, motion_pattern, occluder_type, deform_source, text_position, temporal_evidence, glyph_profile, v2v_edit_op
  • Target text: gt_text (for V2V, the source text), gt_text_after (V2V edit target), gt_text_switched (after a content switch), gt_text_visible (visible part under off-screen cropping)
  • Inputs: reference_image (R2V, I2V), source_video (V2V)

Fields that do not apply are null. For multi-level text, per-text fields are lists with one entry per level.

Metrics

GroupMetrics
Text FidelityA1 Content Accuracy, A2 Glyph Correctness
Temporal StabilityB1 Content Stability, B2 Appearance Stability
Instruction ComplianceC1 Color, C2 Stroke Weight, C3 Typographic Style, C4 Carrier Attachment, C5 Spatial Placement, C6 Layout Hierarchy
Non-Text PreservationD1 Background Consistency
V2V probes (not in Overall)E1 Editing Residue (lower is better), E2 Source Text Preservation

Overall averages each metric over the samples where it is valid, then averages metrics within each group, then averages the 4 groups with equal weight. Inapplicable metrics are left empty and excluded. Failed generations are still scored (low) and flagged with model_output_failure.

Leaderboard on the released half (%)

These numbers are close to, but not identical with, the full-benchmark results in the paper.

ModelOverallA1A2B1B2C1C2C3C4C5C6D1
MiniMax H379.083.056.675.482.682.273.473.974.271.881.890.8
Gemini Omni 1.1 Flash78.379.451.974.684.582.373.075.473.376.884.090.4
Wan 3.078.082.045.575.985.283.075.177.878.976.482.588.8
Seedance 2.577.275.257.871.382.684.871.469.774.768.878.990.8
HappyHorse74.572.648.966.280.382.671.269.772.269.881.489.4
PixVerse V673.472.647.964.178.283.073.773.568.267.482.587.5
Kling3.0-Omni70.958.545.156.080.279.672.767.671.868.183.289.8
ModelT2VR2VI2VV2VE1 ↓E2 ↑
MiniMax H380.482.183.567.831.650.0
Gemini Omni 1.1 Flash81.076.681.770.826.372.3
Wan 3.078.779.782.269.822.865.9
Seedance 2.579.880.082.663.843.962.1
HappyHorse75.478.280.261.931.639.1
PixVerse V671.674.980.365.215.831.6
Kling3.0-Omni66.872.384.057.436.850.2

Usage

from huggingface_hub import hf_hub_download, snapshot_download
from pathlib import Path
import json, pandas as pd

repo = "Vicky0720/VidScribe"
root = Path(snapshot_download(repo, repo_type="dataset", ignore_patterns=["generated_videos/*"]))
gt = [r for f in sorted(root.glob("benchmark/*/metadata_*.json"))
      for r in json.loads(f.read_text(encoding="utf-8"))]
scores = pd.read_csv(root / "results/scores_per_sample.csv")
video = hf_hub_download(repo, scores.generated_video[0], repo_type="dataset")

Data notes and terms

  • Processing. Embedded metadata was removed from all media files; videos were remuxed without re-encoding. The generated videos are the exact bitstreams that were scored, and their visible content (including any service watermarks) is unchanged.
  • Privacy. Faces are masked in the I2V first frames and the V2V source videos. Some media contain incidental scene text. Do not use this dataset to identify or profile people.
  • Third-party material. Files prefixed real_textocr_ or real_lsvt_ are derived from TextOCR and LSVT and remain subject to their terms. The generated videos and scores concern third-party commercial services, reflect the model versions available at evaluation time, and remain subject to each provider's terms, including any restriction on training competing models.
  • Scores. All scores come from an automated evaluator and may contain errors.
  • Use. Non-commercial research only.
  • Takedown. Rights holders can open a discussion with the sample_id, and the material will be removed.

Citation

@article{vidscribe2026,
  title   = {Beyond Legibility: Benchmarking Visual Text Rendering and In-Place Editing in Unified Video Generation},
  author  = {TODO},
  journal = {arXiv preprint arXiv:TODO},
  year    = {2026}
}
ai-generated
benchmark
text-rendering
video-editing
video-generation