jmwang0117/Video4Robot

List of papers on video-centric robot learning

23

7 commits

updated Nov 16, 2024

See the code

README

🤖 Video-Centric Approaches to General Robotic Manipulation Learning

Junming Wang


Methods:

YearVenuePaper TitleLink
2024.11arXivVidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation----
2024.11arXivGrounding Video Models to Actions through Goal Conditioned ExplorationProject Page
2024.11MicrosoftIGOR: Image-GOal RepresentationsProject Page
2024.11Physical Intelligenceπ0: A Vision-Language-Action Flow Model for General Robot ControlProject Page
2024.10arXivVideoAgent: Self-Improving Video GenerationProject Page
2024.10arXivRobots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot DatasetProject Page
2024.10CoRLOKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video ImitationProject Page
2024.10CoRLDifferentiable Robot RenderingProject Page
2024.10arXivLatent Action Pretraining from VideosProject Page
2024.10arXivTowards Synergistic, Generalized, and Efficient Dual-System for Robotic ManipulationProject Page
2024.10arXivVLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model----
2024.10arXivGR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot ManipulationProject Page
2024.09arXivDynaMo: In-Domain Dynamics Pretraining for Visuo-Motor ControlProject Page
2024.09arXivGen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot ManipulationProject Page
2024.09NeurIPSClosed-Loop Visuomotor Control with Generative Expectation for Robotic ManipulationProject Page
2024.07CoRLFlow as the Cross-Domain Manipulation InterfaceProject Page
2024.07arXivThis&That: Language-Gesture Controlled Video Generation for Robot PlanningProject Page
2024.06arXivARDuP: Active Region Video Diffusion for Universal Policies----
2024.06CoRLDreamitate: Real-World Visuomotor Policy Learning via Video GenerationProject Page
2024.05ECCVTrack2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot ManipulationProject Page
2023.12RSSAny-point Trajectory Modeling for Policy LearningProject Page
2023.10arXivLearning to Act from Actionless Videos through Dense CorrespondencesProject Page
2023.10arXivVideo Language PlanningProject Page

jmwang0117/Video4Robot

List of papers on video-centric robot learning

23

7 commits

updated Nov 16, 2024

See the code

README

🤖 Video-Centric Approaches to General Robotic Manipulation Learning

Junming Wang


Methods:

YearVenuePaper TitleLink
2024.11arXivVidMan: Exploiting Implicit Dynamics from Video Diffusion Model for Effective Robot Manipulation----
2024.11arXivGrounding Video Models to Actions through Goal Conditioned ExplorationProject Page
2024.11MicrosoftIGOR: Image-GOal RepresentationsProject Page
2024.11Physical Intelligenceπ0: A Vision-Language-Action Flow Model for General Robot ControlProject Page
2024.10arXivVideoAgent: Self-Improving Video GenerationProject Page
2024.10arXivRobots Pre-train Robots: Manipulation-Centric Robotic Representation from Large-Scale Robot DatasetProject Page
2024.10CoRLOKAMI: Teaching Humanoid Robots Manipulation Skills through Single Video ImitationProject Page
2024.10CoRLDifferentiable Robot RenderingProject Page
2024.10arXivLatent Action Pretraining from VideosProject Page
2024.10arXivTowards Synergistic, Generalized, and Efficient Dual-System for Robotic ManipulationProject Page
2024.10arXivVLM See, Robot Do: Human Demo Video to Robot Action Plan via Vision Language Model----
2024.10arXivGR-2: A Generative Video-Language-Action Model with Web-Scale Knowledge for Robot ManipulationProject Page
2024.09arXivDynaMo: In-Domain Dynamics Pretraining for Visuo-Motor ControlProject Page
2024.09arXivGen2Act: Human Video Generation in Novel Scenarios enables Generalizable Robot ManipulationProject Page
2024.09NeurIPSClosed-Loop Visuomotor Control with Generative Expectation for Robotic ManipulationProject Page
2024.07CoRLFlow as the Cross-Domain Manipulation InterfaceProject Page
2024.07arXivThis&That: Language-Gesture Controlled Video Generation for Robot PlanningProject Page
2024.06arXivARDuP: Active Region Video Diffusion for Universal Policies----
2024.06CoRLDreamitate: Real-World Visuomotor Policy Learning via Video GenerationProject Page
2024.05ECCVTrack2Act: Predicting Point Tracks from Internet Videos enables Generalizable Robot ManipulationProject Page
2023.12RSSAny-point Trajectory Modeling for Policy LearningProject Page
2023.10arXivLearning to Act from Actionless Videos through Dense CorrespondencesProject Page
2023.10arXivVideo Language PlanningProject Page