autodriving-heart/CVPR2025-Papers-about-Autonomous-Driving-and-Embodied-AI

33

8 commits

updated Apr 30, 2025

See the code

README

CVPR 2025-Papers-about-Autonomous-Driving-and-Embodied-AI

Autonomous-Driving

(1) Occupancy

UniScene: Unified Occupancy-centric Driving Scene Generation

GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Prediction

(2) End-to-End Autonomous-Driving

Don't Shake the Wheel: Momentum-AwarePlanning in End-to-End Autonomous Driving

DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

(3) Multi-modal Fusion

V2X - R: Cooperative LiDAR - 4D Radar Fusion with Denoising Diffusion for 3D Object Detection

(4) Vision Language Model

OmniDrive: A Holistic Vision - Language Dataset for Autonomous Driving with counter Factual Reasoning

(5) World Model

DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene

ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration

(6) Closed Loop

DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation

(7) Lane Detection

Rethinking Lanes and Points in complex Scenarios for Monocular 3D Lane Detection

(8) Motion Prediction

ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling

Embodied-AI

(1) Diffusion Model

DexHandDiff: Interaction - aware Diffusion Planning for Adaptive Dexterous Manipulation

(2) 3D Semantic Flow

G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulation

(3) Dual-Arm Robot

RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins

Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method

(4) Visual Language Model

Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection

PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability

--paper: https://arxiv.org/abs/2503.08481

Others

Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models

Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model

DynRefer: Delving into Region-level Multi-modality Tasks via Dynamic Resolution

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

VTON360: High-Fidelity Virtual Try-On from Any Viewing Direction

Cross-modal Causal Relation Alignment for Video Question Grounding

-- paper: https://arxiv.org/abs/2503.07635

-- code: https://github.com/WissingChen/CRA-GQA

DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

-- paper: https://arxiv.org/abs/2503.03190

-- code: https://github.com/LZ-CH/DSPNet

Reproducible Vision-Language Models Meet Concepts Out of Pre-Training

LLM-driven Multimodal and Multi-Identity Listening Head Generation

DAGSM: Disentangled Avatar Generation with GS-enhanced Mesh

-- paper: https://arxiv.org/abs/2411.15205

HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models

-- paper: https://arxiv.org/abs/2407.08706

FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model

-- paper: https://arxiv.org/abs/2503.19839

-- code: https://zjgans.github.io/fireedit.github.io/

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

-- paper: https://arxiv.org/abs/2409.18042

-- code: https://emova-anonymous.github.io/

PS-Diffusion: Photorealistic Subject-Driven Image Editing with Disentangled Control and Attention

Boosting the Dual-Stream Architecture in Ultra-High Resolution Segmentation with Resolution-Biased Uncertainty Estimation

No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image Recognition

Empowering Large Language Models with 3D Situation Awareness

-- paper: https://arxiv.org/abs/2503.23024

Rethinking Query-based Transformer for Continual Image Segmentation

Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation

Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator

SSL4Eco: A Global Seasonal Dataset for Geospatial Foundation Models in Ecology

BiasBench: A reproducible benchmark for tuning the biases of event cameras

Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models

Dynamic Camera Poses and Where to Find Them

PICO: Reconstructing 3D People In Contact with Objects

VEU-Bench: Towards Comprehensive Understanding of Video Editing

Dual Prompting Image Restoration with Diffusion Transformers

SemanticSugarBeets: A Multi-Task Framework and Dataset for Inspecting Harvest and Storage Characteristics of Sugar Beets

PRaDA: Projective Radial Distortion Averaging

CLOC: Contrastive Learning for Ordinal Classification with Multi-Margin N-pair Loss

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning

RepNet-VSR: Reparameterizable Architecture for High-Fidelity Video Super-Resolution

MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World

Plug-and-Play Versatile Compressed Video Enhancement

Acquire and then Adapt: Squeezing out Text-to-Image Model for Image Restoration

Improving Sound Source Localization with Joint Slot Attention on Image and Audio

NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: KwaiSR Dataset and Study

DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding

Other Unfiled papers

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

An LLM-enabled Multi-Agent Autonomous Mechatronics Design Framework

CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning

EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance

RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins

Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction

ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping

CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates

Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization

SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging

Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding

Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects

MESA: Text-Driven Terrain Generation Using Latent Diffusion and Global Copernicus Data

Neural Motion Simulator: Pushing the Limit of World Models in Reinforcement Learning

Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation

From Broadcast to Minimap: Achieving State-of-the-Art SoccerNet Game State Reconstruction

OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance

Data Scaling Laws for End-to-End Autonomous Driving

Decision SpikeFormer: Spike-Driven Transformer for Decision Making

Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions

ZFusion: An Effective Fuser of Camera and 4D Radar for 3D Object Perception in Autonomous Driving

AI Hiring with LLMs: A Context-Aware and Explainable Multi-Agent Framework for Resume Screening

Exploration-Driven Generative Interactive Environments

Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation

MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving

Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs

Scenario Dreamer: Vectorized Latent Diffusion for Generating Driving Simulation Environments

Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations

CoMapGS: Covisibility Map-based Gaussian Splatting for Sparse Novel View Synthesis

BioX-CPath: Biologically-driven Explainable Diagnostics for Multistain IHC Computational Pathology

CoLLM: A Large Language Model for Composed Image Retrieval

Attention IoU: Examining Biases in CelebA using Attention Maps

AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers

Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation

HiLoTs: High-Low Temporal Sensitive Representation Learning for Semi-Supervised LiDAR Segmentation in Autonomous Driving

CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model

Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval

Vision-Language Gradient Descent-driven All-in-One Deep Unfolding Networks

MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations

Generating Multimodal Driving Scenes via Next-Scene Prediction

When Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven Approach

Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning

MP-GUI: Modality Perception with MLLMs for GUI Understanding

DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation

Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation

SALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing

Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Efficient Motion-Aware Video MLLM

V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents

Minding Fuzzy Regions: A Data-driven Alternating Learning Paradigm for Stable Lesion Segmentation

DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation

Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation

GaussHDR: High Dynamic Range Gaussian Splatting via Learning Unified 3D and 2D Local Tone Mapping

SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment

Out-of-Distribution Segmentation in Autonomous Driving: Problems and State of the Art

DexGrasp Anything: Towards Universal Robotic Dexterous Grasping with Physics Awareness

VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning

Denoising Functional Maps: Diffusion Models for Shape Correspondence

KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation

Order-Robust Class Incremental Learning: Graph-Driven Dynamic Similarity Grouping

CarPlanner: Consistent Auto-regressive Trajectory Planning for Large-scale Reinforcement Learning in Autonomous Driving

Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving Scene

Universal Actions for Enhanced Embodied Foundation Models

OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

MotionMap: Representing Multimodality in Human Pose Forecasting

Adapter Merging with Centroid Prototype Mapping for Scalable Class-Incremental Learning

GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Empowering LLMs to Understand and Generate Complex Vector Graphics

StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation

UniScene: Unified Occupancy-centric Driving Scene Generation

DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruction

Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection

Navigation World Models

Task-driven Image Fusion with Learnable Fusion Loss

UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Grasping

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment

CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

Voxel-Aggregated Feature Synthesis: Efficient Dense Mapping for Simulated 3D Reasoning

Decoupling Fine Detail and Global Geometry for Compressed Depth Map Super-Resolution

Driving by the Rules: A Benchmark for Integrating Traffic Sign Regulations into Vectorized HD Map

Style-Editor: Text-driven object-centric style editing

3D-MVP: 3D Multiview Pretraining for Robotic Manipulation

Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation

3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination

VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Contributors

Core9724

8 commits

autodriving-heart/CVPR2025-Papers-about-Autonomous-Driving-and-Embodied-AI

33

8 commits

updated Apr 30, 2025

See the code

README

CVPR 2025-Papers-about-Autonomous-Driving-and-Embodied-AI

Autonomous-Driving

(1) Occupancy

UniScene: Unified Occupancy-centric Driving Scene Generation

GaussianFormer-2: Probabilistic Gaussian Superposition for Efficient 3D Occupancy Prediction

(2) End-to-End Autonomous-Driving

Don't Shake the Wheel: Momentum-AwarePlanning in End-to-End Autonomous Driving

DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

(3) Multi-modal Fusion

V2X - R: Cooperative LiDAR - 4D Radar Fusion with Denoising Diffusion for 3D Object Detection

(4) Vision Language Model

OmniDrive: A Holistic Vision - Language Dataset for Autonomous Driving with counter Factual Reasoning

(5) World Model

DriveDreamer4D: World Models Are Effective Data Machines for 4D Driving Scene

ReconDreamer: Crafting World Models for Driving Scene Reconstruction via Online Restoration

(6) Closed Loop

DrivingSphere: Building a High-fidelity 4D World for Closed-loop Simulation

(7) Lane Detection

Rethinking Lanes and Points in complex Scenarios for Monocular 3D Lane Detection

(8) Motion Prediction

ModeSeq: Taming Sparse Multimodal Motion Prediction with Sequential Mode Modeling

Embodied-AI

(1) Diffusion Model

DexHandDiff: Interaction - aware Diffusion Planning for Adaptive Dexterous Manipulation

(2) 3D Semantic Flow

G3Flow: Generative 3D Semantic Flow for Pose-aware and Generalizable Object Manipulation

(3) Dual-Arm Robot

RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins

Towards Long-Horizon Vision-Language Navigation: Platform, Benchmark and Method

(4) Visual Language Model

Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection

PhysVLM: Enabling Visual Language Models to Understand Robotic Physical Reachability

--paper: https://arxiv.org/abs/2503.08481

Others

Mamba4D: Efficient 4D Point Cloud Video Understanding with Disentangled Spatial-Temporal State Space Models

Timestep Embedding Tells: It's Time to Cache for Video Diffusion Model

DynRefer: Delving into Region-level Multi-modality Tasks via Dynamic Resolution

OmniDocBench: Benchmarking Diverse PDF Document Parsing with Comprehensive Annotations

VTON360: High-Fidelity Virtual Try-On from Any Viewing Direction

Cross-modal Causal Relation Alignment for Video Question Grounding

-- paper: https://arxiv.org/abs/2503.07635

-- code: https://github.com/WissingChen/CRA-GQA

DSPNet: Dual-vision Scene Perception for Robust 3D Question Answering

-- paper: https://arxiv.org/abs/2503.03190

-- code: https://github.com/LZ-CH/DSPNet

Reproducible Vision-Language Models Meet Concepts Out of Pre-Training

LLM-driven Multimodal and Multi-Identity Listening Head Generation

DAGSM: Disentangled Avatar Generation with GS-enhanced Mesh

-- paper: https://arxiv.org/abs/2411.15205

HiRes-LLaVA: Restoring Fragmentation Input in High-Resolution Large Vision-Language Models

-- paper: https://arxiv.org/abs/2407.08706

FireEdit: Fine-grained Instruction-based Image Editing via Region-aware Vision Language Model

-- paper: https://arxiv.org/abs/2503.19839

-- code: https://zjgans.github.io/fireedit.github.io/

EMOVA: Empowering Language Models to See, Hear and Speak with Vivid Emotions

-- paper: https://arxiv.org/abs/2409.18042

-- code: https://emova-anonymous.github.io/

PS-Diffusion: Photorealistic Subject-Driven Image Editing with Disentangled Control and Attention

Boosting the Dual-Stream Architecture in Ultra-High Resolution Segmentation with Resolution-Biased Uncertainty Estimation

No Pains, More Gains: Recycling Sub-Salient Patches for Efficient High-Resolution Image Recognition

Empowering Large Language Models with 3D Situation Awareness

-- paper: https://arxiv.org/abs/2503.23024

Rethinking Query-based Transformer for Continual Image Segmentation

Eval3D: Interpretable and Fine-grained Evaluation for 3D Generation

Seeing Soundscapes: Audio-Visual Generation and Separation from Soundscapes Using Audio-Visual Separator

SSL4Eco: A Global Seasonal Dataset for Geospatial Foundation Models in Ecology

BiasBench: A reproducible benchmark for tuning the biases of event cameras

Enhancing Privacy-Utility Trade-offs to Mitigate Memorization in Diffusion Models

Dynamic Camera Poses and Where to Find Them

PICO: Reconstructing 3D People In Contact with Objects

VEU-Bench: Towards Comprehensive Understanding of Video Editing

Dual Prompting Image Restoration with Diffusion Transformers

SemanticSugarBeets: A Multi-Task Framework and Dataset for Inspecting Harvest and Storage Characteristics of Sugar Beets

PRaDA: Projective Radial Distortion Averaging

CLOC: Contrastive Learning for Ordinal Classification with Multi-Margin N-pair Loss

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

PointLoRA: Low-Rank Adaptation with Token Selection for Point Cloud Learning

RepNet-VSR: Reparameterizable Architecture for High-Fidelity Video Super-Resolution

MirrorVerse: Pushing Diffusion Models to Realistically Reflect the World

Plug-and-Play Versatile Compressed Video Enhancement

Acquire and then Adapt: Squeezing out Text-to-Image Model for Image Restoration

Improving Sound Source Localization with Joint Slot Attention on Image and Audio

NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement: KwaiSR Dataset and Study

DyFo: A Training-Free Dynamic Focus Visual Search for Enhancing LMMs in Fine-Grained Visual Understanding

Other Unfiled papers

LiveCC: Learning Video LLM with Streaming Speech Transcription at Scale

An LLM-enabled Multi-Agent Autonomous Mechatronics Design Framework

CheXWorld: Exploring Image World Modeling for Radiograph Representation Learning

EchoWorld: Learning Motion-Aware World Models for Echocardiography Probe Guidance

RoboTwin: Dual-Arm Robot Benchmark with Generative Digital Twins

Rethinking Temporal Fusion with a Unified Gradient Descent View for 3D Semantic Occupancy Prediction

ZeroGrasp: Zero-Shot Shape Reconstruction Enabled Robotic Grasping

CleanMAP: Distilling Multimodal LLMs for Confidence-Driven Crowdsourced HD Map Updates

Foundation Models for Remote Sensing: An Analysis of MLLMs for Object Localization

SilVar-Med: A Speech-Driven Visual Language Model for Explainable Abnormality Detection in Medical Imaging

Ges3ViG: Incorporating Pointing Gestures into Language-Based 3D Visual Grounding for Embodied Reference Understanding

Gen3DEval: Using vLLMs for Automatic Evaluation of Generated 3D Objects

MESA: Text-Driven Terrain Generation Using Latent Diffusion and Global Copernicus Data

Neural Motion Simulator: Pushing the Limit of World Models in Reinforcement Learning

Two by Two: Learning Multi-Task Pairwise Objects Assembly for Generalizable Robot Manipulation

From Broadcast to Minimap: Achieving State-of-the-Art SoccerNet Game State Reconstruction

OCC-MLLM-CoT-Alpha: Towards Multi-stage Occlusion Recognition Based on Large Language Models via 3D-Aware Supervision and Chain-of-Thoughts Guidance

Data Scaling Laws for End-to-End Autonomous Driving

Decision SpikeFormer: Spike-Driven Transformer for Decision Making

Shape My Moves: Text-Driven Shape-Aware Synthesis of Human Motions

ZFusion: An Effective Fuser of Camera and 4D Radar for 3D Object Perception in Autonomous Driving

AI Hiring with LLMs: A Context-Aware and Explainable Multi-Agent Framework for Resume Screening

Exploration-Driven Generative Interactive Environments

Think Small, Act Big: Primitive Prompt Learning for Lifelong Robot Manipulation

MPDrive: Improving Spatial Understanding with Marker-Based Prompt Learning for Autonomous Driving

Chapter-Llama: Efficient Chaptering in Hour-Long Videos with LLMs

Scenario Dreamer: Vectorized Latent Diffusion for Generating Driving Simulation Environments

Protecting Your Video Content: Disrupting Automated Video-based LLM Annotations

CoMapGS: Covisibility Map-based Gaussian Splatting for Sparse Novel View Synthesis

BioX-CPath: Biologically-driven Explainable Diagnostics for Multistain IHC Computational Pathology

CoLLM: A Large Language Model for Composed Image Retrieval

Attention IoU: Examining Biases in CelebA using Attention Maps

AudCast: Audio-Driven Human Video Generation by Cascaded Diffusion Transformers

Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation

HiLoTs: High-Low Temporal Sensitive Representation Learning for Semi-Supervised LiDAR Segmentation in Autonomous Driving

CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model

Missing Target-Relevant Information Prediction with World Model for Accurate Zero-Shot Composed Image Retrieval

Vision-Language Gradient Descent-driven All-in-One Deep Unfolding Networks

MASH-VLM: Mitigating Action-Scene Hallucination in Video-LLMs through Disentangled Spatial-Temporal Representations

Generating Multimodal Driving Scenes via Next-Scene Prediction

When Domain Generalization meets Generalized Category Discovery: An Adaptive Task-Arithmetic Driven Approach

Bridging Past and Future: End-to-End Autonomous Driving with Historical Prediction and Planning

MP-GUI: Modality Perception with MLLMs for GUI Understanding

DIFFVSGG: Diffusion-Driven Online Video Scene Graph Generation

Conformal Prediction and MLLM aided Uncertainty Quantification in Scene Graph Generation

SALAD: Skeleton-aware Latent Diffusion for Text-driven Motion Generation and Editing

Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos

Efficient Motion-Aware Video MLLM

V-Stylist: Video Stylization via Collaboration and Reflection of MLLM Agents

Minding Fuzzy Regions: A Data-driven Alternating Learning Paradigm for Stable Lesion Segmentation

DriveGEN: Generalized and Robust 3D Detection in Driving via Controllable Text-to-Image Diffusion Generation

Spatial-Temporal Graph Diffusion Policy with Kinematic Modeling for Bimanual Robotic Manipulation

GaussHDR: High Dynamic Range Gaussian Splatting via Learning Unified 3D and 2D Local Tone Mapping

SimLingo: Vision-Only Closed-Loop Autonomous Driving with Language-Action Alignment

Out-of-Distribution Segmentation in Autonomous Driving: Problems and State of the Art

DexGrasp Anything: Towards Universal Robotic Dexterous Grasping with Physics Awareness

VidBot: Learning Generalizable 3D Actions from In-the-Wild 2D Human Videos for Zero-Shot Robotic Manipulation

A Data-Centric Revisit of Pre-Trained Vision Models for Robot Learning

Denoising Functional Maps: Diffusion Models for Shape Correspondence

KeyFace: Expressive Audio-Driven Facial Animation for Long Sequences via KeyFrame Interpolation

Order-Robust Class Incremental Learning: Graph-Driven Dynamic Similarity Grouping

CarPlanner: Consistent Auto-regressive Trajectory Planning for Large-scale Reinforcement Learning in Autonomous Driving

Transfer Your Perspective: Controllable 3D Generation from Any Viewpoint in a Driving Scene

Universal Actions for Enhanced Embodied Foundation Models

OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?

MotionMap: Representing Multimodality in Human Pose Forecasting

Adapter Merging with Centroid Prototype Mapping for Scalable Class-Incremental Learning

GME: Improving Universal Multimodal Retrieval by Multimodal LLMs

Empowering LLMs to Understand and Generate Complex Vector Graphics

StyleStudio: Text-Driven Style Transfer with Selective Control of Style Elements

DiffSensei: Bridging Multi-Modal LLMs and Diffusion Models for Customized Manga Generation

UniScene: Unified Occupancy-centric Driving Scene Generation

DualPM: Dual Posed-Canonical Point Maps for 3D Shape and Pose Reconstruction

Code-as-Monitor: Constraint-aware Visual Programming for Reactive and Proactive Robotic Failure Detection

Navigation World Models

Task-driven Image Fusion with Learnable Fusion Loss

UniGraspTransformer: Simplified Policy Distillation for Scalable Dexterous Robotic Grasping

Video-3D LLM: Learning Position-Aware Video Representation for 3D Scene Understanding

Immune: Improving Safety Against Jailbreaks in Multi-modal LLMs via Inference-Time Alignment

CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos

SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation

Augmenting Multimodal LLMs with Self-Reflective Tokens for Knowledge-based Visual Question Answering

RoboSpatial: Teaching Spatial Understanding to 2D and 3D Vision-Language Models for Robotics

DiffusionDrive: Truncated Diffusion Model for End-to-End Autonomous Driving

Voxel-Aggregated Feature Synthesis: Efficient Dense Mapping for Simulated 3D Reasoning

Decoupling Fine Detail and Global Geometry for Compressed Depth Map Super-Resolution

Driving by the Rules: A Benchmark for Integrating Traffic Sign Regulations into Vectorized HD Map

Style-Editor: Text-driven object-centric style editing

3D-MVP: 3D Multiview Pretraining for Robotic Manipulation

Mitigating the Human-Robot Domain Discrepancy in Visual Pre-training for Robotic Manipulation

3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination

VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Contributors

Core9724

8 commits