anyeZHY/tesseract

Model

We propose TesserAct, the 4D Embodied World Model, which takes input images and text instruction to generate RGB, depth,

7

6 commits

1 linked in READMEs

updated May 1, 2025

See the code

README

TesserAct: Learning 4D Embodied World Models

Haoyu Zhen*, Qiao Sun*, Hongxin Zhang, Junyan Li, Siyuan Zhou, Yilun Du, Chuang Gan

Paper PDF  |  Project Page  |  Model on Hugging Face  |  Code

We propose TesserAct, the 4D Embodied World Model, which takes input images and text instruction to generate RGB, depth, and normal videos, reconstructing a 4D scene and predicting actions.

diffusers
image-to-video
safetensors

Contributors

anyeZHY

5 commits

nielsr

1 commits

anyeZHY/tesseract

Model

We propose TesserAct, the 4D Embodied World Model, which takes input images and text instruction to generate RGB, depth,

7

6 commits

1 linked in READMEs

updated May 1, 2025

See the code

README

TesserAct: Learning 4D Embodied World Models

Haoyu Zhen*, Qiao Sun*, Hongxin Zhang, Junyan Li, Siyuan Zhou, Yilun Du, Chuang Gan

Paper PDF  |  Project Page  |  Model on Hugging Face  |  Code

We propose TesserAct, the 4D Embodied World Model, which takes input images and text instruction to generate RGB, depth, and normal videos, reconstructing a 4D scene and predicting actions.

diffusers
image-to-video
safetensors

Contributors

anyeZHY

5 commits

nielsr

1 commits