tiedong/Molmo2-SGCoT-Demo

Space

1

stars

8

commits

1

linked in READMEs

Apr 2, 2026

updated

chain-of-thought
gradio
spatial-grounding
video-understanding
visual-tracking

README

Molmo2-SGCoT: Visual Entity Tracking Demo

Can Vision-Language Models Solve the Shell Game?

This demo runs Molmo2-SGCoT, a fine-tuned Molmo2-8B that generates explicit object tracking trajectories (Spatiotemporal Grounded Chain-of-Thought) before answering visual tracking questions.

  • Upload a cup-shuffling or card-tracking video
  • The model outputs spatial coordinates at 0.5s intervals, tracing the target object
  • View the predicted trajectory visualization overlaid on the video
  • See the final answer: Left, Middle, or Right

Key Results

Trained with only 300 synthetic samples, Molmo2-SGCoT achieves 91% accuracy on VET-Bench — a benchmark where state-of-the-art VLMs (Gemini, Qwen, etc.) score at random chance (~33%).

Contributors

tiedong

8 commits

tiedong/Molmo2-SGCoT-Demo

Space

1

stars

8

commits

1

linked in READMEs

Apr 2, 2026

updated

chain-of-thought
gradio
spatial-grounding
video-understanding
visual-tracking

README

Molmo2-SGCoT: Visual Entity Tracking Demo

Can Vision-Language Models Solve the Shell Game?

This demo runs Molmo2-SGCoT, a fine-tuned Molmo2-8B that generates explicit object tracking trajectories (Spatiotemporal Grounded Chain-of-Thought) before answering visual tracking questions.

  • Upload a cup-shuffling or card-tracking video
  • The model outputs spatial coordinates at 0.5s intervals, tracing the target object
  • View the predicted trajectory visualization overlaid on the video
  • See the final answer: Left, Middle, or Right

Key Results

Trained with only 300 synthetic samples, Molmo2-SGCoT achieves 91% accuracy on VET-Bench — a benchmark where state-of-the-art VLMs (Gemini, Qwen, etc.) score at random chance (~33%).

Contributors

tiedong

8 commits