Messimm/Kandinsky-WM-1.0-Physics-IQ-Verified

Dataset

Kandinsky-WM 1.0 Physics-IQ Verified Artifacts

0

15 commits

2 linked in READMEs

updated Jul 29, 2026

See the code

README

Kandinsky-WM 1.0 Physics-IQ Verified Artifacts

This dataset contains the artifacts for the Kandinsky-WM 1.0 General Physics result on Physics-IQ Verified.

The dataset contains four generated runs with seeds 32768, 1109, 1137, and 16384. Each run contains 198 Physics-IQ take-1 views.

Generated artifacts

  • videos/seed_/ contains the generated videos for one seed.
  • masks/seed_/24FPS/ contains the generated binary masks used to score those videos.
  • eval_results/seed_/ contains the detailed per-scene score CSV, metrics JSON, and score plots for one seed.
  • logs/seed_/ contains the generation and evaluation logs for one seed.

Prompt pipeline

The prompt files in common/ show every stage of the prompt pipeline:

  • descriptions_base.csv is the starting Physics-IQ best-practice prompt (BPP) CSV.
  • descriptions_bpp_static_camera.csv is the BPP CSV after adding this prefix to every prompt: "The camera is strictly static, locked off, and does not move."
  • enhance_physics_iq_bpp_temporal_v2.py and enhance_physics_iq_bpp_prompts_qwen.py are the exact temporal enhancer and its helper used in this experiment. The same method is released in the Kandinsky-WM repository under prompt_enhancers/physics_iq/.
  • descriptions_bpp_static_enchancedV2.csv is the final enhanced prompt CSV used for video generation.
  • descriptions_bpp_static_enchancedV2_take1.csv is the take-1 subset used by Physics-IQ Verified evaluation.

The temporal enhancer uses each benchmark first frame and a Qwen3-VL endpoint to turn the static-camera BPP prompt into a grounded temporal description. The final CSV is included so the released video runs can be checked with exactly the prompts that were used, without regenerating VLM outputs.

The real GT videos and GT masks are part of the official Physics-IQ distribution and are not duplicated here.

image-to-video
physics-iq
reproducibility
video-generation

Messimm/Kandinsky-WM-1.0-Physics-IQ-Verified

Dataset

Kandinsky-WM 1.0 Physics-IQ Verified Artifacts

0

15 commits

2 linked in READMEs

updated Jul 29, 2026

See the code

README

Kandinsky-WM 1.0 Physics-IQ Verified Artifacts

This dataset contains the artifacts for the Kandinsky-WM 1.0 General Physics result on Physics-IQ Verified.

The dataset contains four generated runs with seeds 32768, 1109, 1137, and 16384. Each run contains 198 Physics-IQ take-1 views.

Generated artifacts

  • videos/seed_/ contains the generated videos for one seed.
  • masks/seed_/24FPS/ contains the generated binary masks used to score those videos.
  • eval_results/seed_/ contains the detailed per-scene score CSV, metrics JSON, and score plots for one seed.
  • logs/seed_/ contains the generation and evaluation logs for one seed.

Prompt pipeline

The prompt files in common/ show every stage of the prompt pipeline:

  • descriptions_base.csv is the starting Physics-IQ best-practice prompt (BPP) CSV.
  • descriptions_bpp_static_camera.csv is the BPP CSV after adding this prefix to every prompt: "The camera is strictly static, locked off, and does not move."
  • enhance_physics_iq_bpp_temporal_v2.py and enhance_physics_iq_bpp_prompts_qwen.py are the exact temporal enhancer and its helper used in this experiment. The same method is released in the Kandinsky-WM repository under prompt_enhancers/physics_iq/.
  • descriptions_bpp_static_enchancedV2.csv is the final enhanced prompt CSV used for video generation.
  • descriptions_bpp_static_enchancedV2_take1.csv is the take-1 subset used by Physics-IQ Verified evaluation.

The temporal enhancer uses each benchmark first frame and a Qwen3-VL endpoint to turn the static-camera BPP prompt into a grounded temporal description. The final CSV is included so the released video runs can be checked with exactly the prompts that were used, without regenerating VLM outputs.

The real GT videos and GT masks are part of the official Physics-IQ distribution and are not duplicated here.

image-to-video
physics-iq
reproducibility
video-generation