Gen Li1,#,‡, Hanqing Wang2,#, Jingliang Li1,#, Yifan Han3,#, Jindou Jia1, Tao Lin3, Yuanzhe Liu4, Yutong Wang5, Bo Zhao3, Fangqiang Ding2, Anh Nguyen6, Laura Sevilla-Lara7, Huazhe Xu8, Gregory S. Chirikjian9,
Marc Pollefeys10, Oier Mees10, Hui Xiong2,†, Jianfei Yang1,†
1Nanyang Technological University 2HKUST(GZ) 3SJTU 4UIUC 5The University of Sydney
6University of Liverpool 7The University of Edinburgh
8MBZUAI 9Tsinghua University 10ETH Zurich
#Equal Contribution ‡Project Lead †Corresponding Author
A curated collection of papers on affordance learning for embodied AI.
🧭 Exploring Embodied AI and Embodied perception? We hope this collection proves useful in your journey. If you'd like to support the project, feel free to ⭐️ the repo and share it with your peers. Contributions are warmly welcome!
📢 This list is actively maintained, and community contributions are always appreciated!
Feel free to open a pull request if you find any relevant papers.
This repository accompanies the survey From Passive Perception to Active Interaction: A Survey of Affordance Learning for Embodied AI and maintains a curated collection of papers, datasets, and benchmarks.
As robots and embodied agents move into real-world applications, they must understand not only what objects are, but also where, why, and how they can interact with them. Following the survey, the list is organized around three complementary questions:
A paper may be cross-referenced when the survey discusses it in more than one role (for example, both perception and reasoning). A separate section collects datasets and benchmarks. Venue labels use the formally published version whenever one is available; otherwise the first public preprint is marked arXiv.
The taxonomy follows the survey's functional pipeline rather than only model architecture. It expands each primary category into second- and third-level categories. Cross-listing is intentional when a method contributes to multiple stages or perspectives.
| Primary Category | Secondary Category | Third-Level Categories | Main Distinction |
|---|---|---|---|
| Affordance Perception | Visual affordance perception | Object-centric grounding; scene-level grounding; weakly supervised perception | Grounds masks, heatmaps, or keypoints on the 2D image plane. |
| Affordance Perception | Spatial affordance perception | 3D object grounding; 3D scene grounding | Grounds functional regions or geometric structures directly in metric 3D space. |
| Affordance Perception | Interaction-driven affordance perception | Demonstration/HOI video; HOI image; interaction-conditioned 3D; interaction-grounded scene perception | Derives supervision or context from observed human-object interactions. |
| Affordance Perception | Generalizable affordance perception | Example-based transfer; open-set grounding | Transfers to novel objects, labels, queries, or interaction contexts. |
| Affordance Reasoning | Relation-based reasoning | Probabilistic/semantic relations; object-pair and scene context; manipulation graphs | Infers affordances through explicit relations among objects, actions, agents, and scene context. |
| Affordance Reasoning | Language-centric reasoning | Grounded skill selection; integrated LLM/MLLM prediction; modular semantic-to-spatial grounding; sequential reasoning | Uses language to interpret intent and select or ground task-relevant affordances. |
| Affordance Reasoning | Agentic reasoning | Predefined workflows; adaptive workflows | Uses planning, memory, verification, or iterative tool/model calls over multiple steps. |
| Affordance-Guided Action | Hierarchical affordance-to-action | Primitive-based execution; retrieval-based transfer; planner-based execution | Predicts affordances first and then converts them into actions through a separate execution module. |
| Affordance-Guided Action | Affordance-integrated policy learning | Affordance as explicit input; implicit representation; optimization signal | Integrates affordance information directly into policy representation, input, or optimization. |
The Venue/Date column prioritizes the formal venue and publication year. Papers without a confirmed venue are labeled arXiv. The Name column records the method or system name explicitly introduced by the authors, including names stated only in the abstract or main text rather than in the title; - means that no explicit method name has been verified and no acronym is inferred from the title.
Survey groups: object-centric · scene-level · interaction-driven · language- and reasoning-oriented · action-oriented.
⭐ Help us grow this repository! If you know any valuable works we’ve missed, don’t hesitate to contribute — every suggestion makes a difference!
We welcome and appreciate all contributions! Here’s how you can help:
📄 Add or Update a Paper
Contribute by adding a new paper or improving details of an existing one. Please consider the most appropriate category for the work.
✍️ Use Consistent Formatting
Follow the format of the existing entries to maintain clarity and consistency across the list.
🔗 Include Abstract Link
If the paper is from arXiv, use the /abs/ link format for the abstract (e.g., https://arxiv.org/abs/xxxx.xxxxx).
💡 Explain Your Edit (Optional but Helpful)
A short note on why you think the paper deserves to be added or updated is appreciated and helps maintainers process your PR faster.
✅ Don't worry about getting everything perfect!
Minor mistakes are totally fine — we’ll help fix them. What matters most is your contribution. Let's highlight your awesome work together!
Thanks for the wonderful researchers focusing on affordance learning and embodied AI
This project is licensed under the MIT License.
Gen Li1,#,‡, Hanqing Wang2,#, Jingliang Li1,#, Yifan Han3,#, Jindou Jia1, Tao Lin3, Yuanzhe Liu4, Yutong Wang5, Bo Zhao3, Fangqiang Ding2, Anh Nguyen6, Laura Sevilla-Lara7, Huazhe Xu8, Gregory S. Chirikjian9,
Marc Pollefeys10, Oier Mees10, Hui Xiong2,†, Jianfei Yang1,†
1Nanyang Technological University 2HKUST(GZ) 3SJTU 4UIUC 5The University of Sydney
6University of Liverpool 7The University of Edinburgh
8MBZUAI 9Tsinghua University 10ETH Zurich
#Equal Contribution ‡Project Lead †Corresponding Author
A curated collection of papers on affordance learning for embodied AI.
🧭 Exploring Embodied AI and Embodied perception? We hope this collection proves useful in your journey. If you'd like to support the project, feel free to ⭐️ the repo and share it with your peers. Contributions are warmly welcome!
📢 This list is actively maintained, and community contributions are always appreciated!
Feel free to open a pull request if you find any relevant papers.
This repository accompanies the survey From Passive Perception to Active Interaction: A Survey of Affordance Learning for Embodied AI and maintains a curated collection of papers, datasets, and benchmarks.
As robots and embodied agents move into real-world applications, they must understand not only what objects are, but also where, why, and how they can interact with them. Following the survey, the list is organized around three complementary questions:
A paper may be cross-referenced when the survey discusses it in more than one role (for example, both perception and reasoning). A separate section collects datasets and benchmarks. Venue labels use the formally published version whenever one is available; otherwise the first public preprint is marked arXiv.
The taxonomy follows the survey's functional pipeline rather than only model architecture. It expands each primary category into second- and third-level categories. Cross-listing is intentional when a method contributes to multiple stages or perspectives.
| Primary Category | Secondary Category | Third-Level Categories | Main Distinction |
|---|---|---|---|
| Affordance Perception | Visual affordance perception | Object-centric grounding; scene-level grounding; weakly supervised perception | Grounds masks, heatmaps, or keypoints on the 2D image plane. |
| Affordance Perception | Spatial affordance perception | 3D object grounding; 3D scene grounding | Grounds functional regions or geometric structures directly in metric 3D space. |
| Affordance Perception | Interaction-driven affordance perception | Demonstration/HOI video; HOI image; interaction-conditioned 3D; interaction-grounded scene perception | Derives supervision or context from observed human-object interactions. |
| Affordance Perception | Generalizable affordance perception | Example-based transfer; open-set grounding | Transfers to novel objects, labels, queries, or interaction contexts. |
| Affordance Reasoning | Relation-based reasoning | Probabilistic/semantic relations; object-pair and scene context; manipulation graphs | Infers affordances through explicit relations among objects, actions, agents, and scene context. |
| Affordance Reasoning | Language-centric reasoning | Grounded skill selection; integrated LLM/MLLM prediction; modular semantic-to-spatial grounding; sequential reasoning | Uses language to interpret intent and select or ground task-relevant affordances. |
| Affordance Reasoning | Agentic reasoning | Predefined workflows; adaptive workflows | Uses planning, memory, verification, or iterative tool/model calls over multiple steps. |
| Affordance-Guided Action | Hierarchical affordance-to-action | Primitive-based execution; retrieval-based transfer; planner-based execution | Predicts affordances first and then converts them into actions through a separate execution module. |
| Affordance-Guided Action | Affordance-integrated policy learning | Affordance as explicit input; implicit representation; optimization signal | Integrates affordance information directly into policy representation, input, or optimization. |
The Venue/Date column prioritizes the formal venue and publication year. Papers without a confirmed venue are labeled arXiv. The Name column records the method or system name explicitly introduced by the authors, including names stated only in the abstract or main text rather than in the title; - means that no explicit method name has been verified and no acronym is inferred from the title.
Survey groups: object-centric · scene-level · interaction-driven · language- and reasoning-oriented · action-oriented.
⭐ Help us grow this repository! If you know any valuable works we’ve missed, don’t hesitate to contribute — every suggestion makes a difference!
We welcome and appreciate all contributions! Here’s how you can help:
📄 Add or Update a Paper
Contribute by adding a new paper or improving details of an existing one. Please consider the most appropriate category for the work.
✍️ Use Consistent Formatting
Follow the format of the existing entries to maintain clarity and consistency across the list.
🔗 Include Abstract Link
If the paper is from arXiv, use the /abs/ link format for the abstract (e.g., https://arxiv.org/abs/xxxx.xxxxx).
💡 Explain Your Edit (Optional but Helpful)
A short note on why you think the paper deserves to be added or updated is appreciated and helps maintainers process your PR faster.
✅ Don't worry about getting everything perfect!
Minor mistakes are totally fine — we’ll help fix them. What matters most is your contribution. Let's highlight your awesome work together!
Thanks for the wonderful researchers focusing on affordance learning and embodied AI
This project is licensed under the MIT License.