Can Vision-Language Models Solve the Shell Game?
This demo runs Molmo2-SGCoT, a fine-tuned Molmo2-8B that generates explicit object tracking trajectories (Spatiotemporal Grounded Chain-of-Thought) before answering visual tracking questions.
Trained with only 300 synthetic samples, Molmo2-SGCoT achieves 91% accuracy on VET-Bench — a benchmark where state-of-the-art VLMs (Gemini, Qwen, etc.) score at random chance (~33%).
8 commits
Can Vision-Language Models Solve the Shell Game?
This demo runs Molmo2-SGCoT, a fine-tuned Molmo2-8B that generates explicit object tracking trajectories (Spatiotemporal Grounded Chain-of-Thought) before answering visual tracking questions.
Trained with only 300 synthetic samples, Molmo2-SGCoT achieves 91% accuracy on VET-Bench — a benchmark where state-of-the-art VLMs (Gemini, Qwen, etc.) score at random chance (~33%).
8 commits