Checkpoint of https://huggingface.co/papers/2511.19418.
This CoVT checkpoint is aligned with 8 Segmentation tokens, 4 Depth tokens, and 4 DINO tokens.
These task-specific tokens are integrated into the model’s embedding space to enhance 2D-awareness, 3D-awareness, and patch-level feature representations.
6 commits
Checkpoint of https://huggingface.co/papers/2511.19418.
This CoVT checkpoint is aligned with 8 Segmentation tokens, 4 Depth tokens, and 4 DINO tokens.
These task-specific tokens are integrated into the model’s embedding space to enhance 2D-awareness, 3D-awareness, and patch-level feature representations.
6 commits