33
8 commits
updated Apr 30, 2025
UniScene: Unified Occupancy-centric Driving Scene Generation
GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Prediction
Don't Shake the Wheel: Momentum-AwarePlanning in End-to-End Autonomous Driving
DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
V2X - R: Cooperative LiDAR - 4D Radar Fusion with Denoising Diffusion for 3D Object Detection
OmniDrive: A Holistic Vision - Language Dataset for Autonomous Driving with counter Factual Reasoning
DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene
ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation
Rethinking Lanes and Points in complex Scenarios for Monocular 3D Lane Detection
ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling
DexHandDiff: Interaction - aware Diffusion Planning for Adaptive Dexterous Manipulation
G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulation
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins
Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability
--paper: https://arxiv.org/abs/2503.08481
Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
DynRefer: Delving into Region-level Multi-modality Tasks via Dynamic Resolution
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
VTON360: High-Fidelity Virtual Try-On from Any Viewing Direction
Cross-modal Causal Relation Alignment for Video Question Grounding
-- paper: https://arxiv.org/abs/2503.07635
-- code: https://github.com/WissingChen/CRA-GQA
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
-- paper: https://arxiv.org/abs/2503.03190
-- code: https://github.com/LZ-CH/DSPNet
Reproducible Vision-Language Models Meet Concepts Out of Pre-Training
LLM-driven Multimodal and Multi-Identity Listening Head Generation
DAGSM: Disentangled Avatar Generation with GS-enhanced Mesh
-- paper: https://arxiv.org/abs/2411.15205
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
-- paper: https://arxiv.org/abs/2407.08706
FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model
-- paper: https://arxiv.org/abs/2503.19839
-- code: https://zjgans.github.io/fireedit.github.io/
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
-- paper: https://arxiv.org/abs/2409.18042
-- code: https://emova-anonymous.github.io/
PS-Diffusion: Photorealistic Subject-Driven Image Editing with Disentangled Control and Attention
Boosting the Dual-Stream Architecture in Ultra-High Resolution Segmentation with Resolution-Biased Uncertainty Estimation
No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image Recognition
Empowering Large Language Models with 3D Situation Awareness
-- paper: https://arxiv.org/abs/2503.23024
Rethinking Query-based Transformer for Continual Image Segmentation
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
code: https://eval3d.git
Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator
SSL4Eco: A Global Seasonal Dataset for Geospatial Foundation Models in Ecology
BiasBench: A reproducible benchmark for tuning the biases of event cameras
Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models
Dynamic Camera Poses and Where to Find Them
PICO: Reconstructing 3D People In Contact with Objects
VEU-Bench: Towards Comprehensive Understanding of Video Editing
Dual Prompting Image Restoration with Diffusion Transformers
SemanticSugarBeets: A Multi-Task Framework and Dataset for Inspecting Harvest and Storage Characteristics of Sugar Beets
PRaDA: Projective Radial Distortion Averaging
CLOC: Contrastive Learning for Ordinal Classification with Multi-Margin N-pair Loss
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
code: https://showlab.git
PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning
RepNet-VSR: Reparameterizable Architecture for High-Fidelity Video Super-Resolution
MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World
Plug-and-Play Versatile Compressed Video Enhancement
code: https://huimin-zeng.git
Acquire and then Adapt: Squeezing out Text-to-Image Model for Image Restoration
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: KwaiSR Dataset and Study
code: https://lixinustc.git
DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
An LLM-enabled Multi-Agent Autonomous Mechatronics Design Framework
CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning
EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins
Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping
CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates
Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization
SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging
Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding
Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects
MESA: Text-Driven Terrain Generation Using Latent Diffusion and Global Copernicus Data
Neural Motion Simulator: Pushing the Limit of World Models in Reinforcement Learning
Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
From Broadcast to Minimap: Achieving State-of-the-Art SoccerNet Game State Reconstruction
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
Data Scaling Laws for End-to-End Autonomous Driving
Decision SpikeFormer: Spike-Driven Transformer for Decision Making
Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions
ZFusion: An Effective Fuser of Camera and 4D Radar for 3D Object Perception in Autonomous Driving
AI Hiring with LLMs: A Context-Aware and Explainable Multi-Agent Framework for Resume Screening
Exploration-Driven Generative Interactive Environments
Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation
MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
Scenario Dreamer: Vectorized Latent Diffusion for Generating Driving Simulation Environments
Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations
CoMapGS: Covisibility Map-based Gaussian Splatting for Sparse Novel View Synthesis
BioX-CPath: Biologically-driven Explainable Diagnostics for Multistain IHC Computational Pathology
CoLLM: A Large Language Model for Composed Image Retrieval
Attention IoU: Examining Biases in CelebA using Attention Maps
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation
HiLoTs: High-Low Temporal Sensitive Representation Learning for Semi-Supervised LiDAR Segmentation in Autonomous Driving
CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model
Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval
Vision-Language Gradient Descent-driven All-in-One Deep Unfolding Networks
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
Generating Multimodal Driving Scenes via Next-Scene Prediction
When Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven Approach
Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning
MP-GUI: Modality Perception with MLLMs for GUI Understanding
DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation
SALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
Efficient Motion-Aware Video MLLM
V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
Minding Fuzzy Regions: A Data-driven Alternating Learning Paradigm for Stable Lesion Segmentation
DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation
Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation
GaussHDR: High Dynamic Range Gaussian Splatting via Learning Unified 3D and 2D Local Tone Mapping
SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
Out-of-Distribution Segmentation in Autonomous Driving: Problems and State of the Art
DexGrasp Anything: Towards Universal Robotic Dexterous Grasping with Physics Awareness
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
Denoising Functional Maps: Diffusion Models for Shape Correspondence
KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation
Order-Robust Class Incremental Learning: Graph-Driven Dynamic Similarity Grouping
CarPlanner: Consistent Auto-regressive Trajectory Planning for Large-scale Reinforcement Learning in Autonomous Driving
Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving Scene
Universal Actions for Enhanced Embodied Foundation Models
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
MotionMap: Representing Multimodality in Human Pose Forecasting
Adapter Merging with Centroid Prototype Mapping for Scalable Class-Incremental Learning
GME: Improving Universal Multimodal Retrieval by Multimodal LLMs
Empowering LLMs to Understand and Generate Complex Vector Graphics
StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
UniScene: Unified Occupancy-centric Driving Scene Generation
DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruction
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
Navigation World Models
Task-driven Image Fusion with Learnable Fusion Loss
UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Grasping
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
Voxel-Aggregated Feature Synthesis: Efficient Dense Mapping for Simulated 3D Reasoning
Decoupling Fine Detail and Global Geometry for Compressed Depth Map Super-Resolution
Driving by the Rules: A Benchmark for Integrating Traffic Sign Regulations into Vectorized HD Map
Style-Editor: Text-driven object-centric style editing
3D-MVP: 3D Multiview Pretraining for Robotic Manipulation
Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
8 commits
33
8 commits
updated Apr 30, 2025
UniScene: Unified Occupancy-centric Driving Scene Generation
GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Prediction
Don't Shake the Wheel: Momentum-AwarePlanning in End-to-End Autonomous Driving
DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
V2X - R: Cooperative LiDAR - 4D Radar Fusion with Denoising Diffusion for 3D Object Detection
OmniDrive: A Holistic Vision - Language Dataset for Autonomous Driving with counter Factual Reasoning
DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene
ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration
DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation
Rethinking Lanes and Points in complex Scenarios for Monocular 3D Lane Detection
ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling
DexHandDiff: Interaction - aware Diffusion Planning for Adaptive Dexterous Manipulation
G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulation
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins
Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability
--paper: https://arxiv.org/abs/2503.08481
Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models
Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model
DynRefer: Delving into Region-level Multi-modality Tasks via Dynamic Resolution
OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations
VTON360: High-Fidelity Virtual Try-On from Any Viewing Direction
Cross-modal Causal Relation Alignment for Video Question Grounding
-- paper: https://arxiv.org/abs/2503.07635
-- code: https://github.com/WissingChen/CRA-GQA
DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering
-- paper: https://arxiv.org/abs/2503.03190
-- code: https://github.com/LZ-CH/DSPNet
Reproducible Vision-Language Models Meet Concepts Out of Pre-Training
LLM-driven Multimodal and Multi-Identity Listening Head Generation
DAGSM: Disentangled Avatar Generation with GS-enhanced Mesh
-- paper: https://arxiv.org/abs/2411.15205
HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models
-- paper: https://arxiv.org/abs/2407.08706
FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model
-- paper: https://arxiv.org/abs/2503.19839
-- code: https://zjgans.github.io/fireedit.github.io/
EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions
-- paper: https://arxiv.org/abs/2409.18042
-- code: https://emova-anonymous.github.io/
PS-Diffusion: Photorealistic Subject-Driven Image Editing with Disentangled Control and Attention
Boosting the Dual-Stream Architecture in Ultra-High Resolution Segmentation with Resolution-Biased Uncertainty Estimation
No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image Recognition
Empowering Large Language Models with 3D Situation Awareness
-- paper: https://arxiv.org/abs/2503.23024
Rethinking Query-based Transformer for Continual Image Segmentation
Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation
code: https://eval3d.git
Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator
SSL4Eco: A Global Seasonal Dataset for Geospatial Foundation Models in Ecology
BiasBench: A reproducible benchmark for tuning the biases of event cameras
Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models
Dynamic Camera Poses and Where to Find Them
PICO: Reconstructing 3D People In Contact with Objects
VEU-Bench: Towards Comprehensive Understanding of Video Editing
Dual Prompting Image Restoration with Diffusion Transformers
SemanticSugarBeets: A Multi-Task Framework and Dataset for Inspecting Harvest and Storage Characteristics of Sugar Beets
PRaDA: Projective Radial Distortion Averaging
CLOC: Contrastive Learning for Ordinal Classification with Multi-Margin N-pair Loss
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
code: https://showlab.git
PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning
RepNet-VSR: Reparameterizable Architecture for High-Fidelity Video Super-Resolution
MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World
Plug-and-Play Versatile Compressed Video Enhancement
code: https://huimin-zeng.git
Acquire and then Adapt: Squeezing out Text-to-Image Model for Image Restoration
Improving Sound Source Localization with Joint Slot Attention on Image and Audio
NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: KwaiSR Dataset and Study
code: https://lixinustc.git
DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding
LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale
An LLM-enabled Multi-Agent Autonomous Mechatronics Design Framework
CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning
EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance
RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins
Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction
ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping
CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates
Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization
SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging
Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding
Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects
MESA: Text-Driven Terrain Generation Using Latent Diffusion and Global Copernicus Data
Neural Motion Simulator: Pushing the Limit of World Models in Reinforcement Learning
Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation
From Broadcast to Minimap: Achieving State-of-the-Art SoccerNet Game State Reconstruction
OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance
Data Scaling Laws for End-to-End Autonomous Driving
Decision SpikeFormer: Spike-Driven Transformer for Decision Making
Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions
ZFusion: An Effective Fuser of Camera and 4D Radar for 3D Object Perception in Autonomous Driving
AI Hiring with LLMs: A Context-Aware and Explainable Multi-Agent Framework for Resume Screening
Exploration-Driven Generative Interactive Environments
Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation
MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving
Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs
Scenario Dreamer: Vectorized Latent Diffusion for Generating Driving Simulation Environments
Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations
CoMapGS: Covisibility Map-based Gaussian Splatting for Sparse Novel View Synthesis
BioX-CPath: Biologically-driven Explainable Diagnostics for Multistain IHC Computational Pathology
CoLLM: A Large Language Model for Composed Image Retrieval
Attention IoU: Examining Biases in CelebA using Attention Maps
AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers
Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation
HiLoTs: High-Low Temporal Sensitive Representation Learning for Semi-Supervised LiDAR Segmentation in Autonomous Driving
CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model
Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval
Vision-Language Gradient Descent-driven All-in-One Deep Unfolding Networks
MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations
Generating Multimodal Driving Scenes via Next-Scene Prediction
When Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven Approach
Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning
MP-GUI: Modality Perception with MLLMs for GUI Understanding
DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation
Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation
SALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos
Efficient Motion-Aware Video MLLM
V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents
Minding Fuzzy Regions: A Data-driven Alternating Learning Paradigm for Stable Lesion Segmentation
DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation
Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation
GaussHDR: High Dynamic Range Gaussian Splatting via Learning Unified 3D and 2D Local Tone Mapping
SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment
Out-of-Distribution Segmentation in Autonomous Driving: Problems and State of the Art
DexGrasp Anything: Towards Universal Robotic Dexterous Grasping with Physics Awareness
VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation
A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning
Denoising Functional Maps: Diffusion Models for Shape Correspondence
KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation
Order-Robust Class Incremental Learning: Graph-Driven Dynamic Similarity Grouping
CarPlanner: Consistent Auto-regressive Trajectory Planning for Large-scale Reinforcement Learning in Autonomous Driving
Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving Scene
Universal Actions for Enhanced Embodied Foundation Models
OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
MotionMap: Representing Multimodality in Human Pose Forecasting
Adapter Merging with Centroid Prototype Mapping for Scalable Class-Incremental Learning
GME: Improving Universal Multimodal Retrieval by Multimodal LLMs
Empowering LLMs to Understand and Generate Complex Vector Graphics
StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements
DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation
UniScene: Unified Occupancy-centric Driving Scene Generation
DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruction
Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection
Navigation World Models
Task-driven Image Fusion with Learnable Fusion Loss
UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Grasping
Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding
Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos
SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation
Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering
RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics
DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving
Voxel-Aggregated Feature Synthesis: Efficient Dense Mapping for Simulated 3D Reasoning
Decoupling Fine Detail and Global Geometry for Compressed Depth Map Super-Resolution
Driving by the Rules: A Benchmark for Integrating Traffic Sign Regulations into Vectorized HD Map
Style-Editor: Text-driven object-centric style editing
3D-MVP: 3D Multiview Pretraining for Robotic Manipulation
Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation
3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination
VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos
8 commits