The full data pool of AFUN (arXiv:2606.02551):
183,657 data points for affordance segmentation and 3D interaction-motion prediction,
i.e. the fitted-curve pool of Table 7 from which the curated
AFUN training set (44,749) was
selected. AFUN β AFUN_pool. Same folder format as AFUN, same trajectory.json schemas.
It comes in two parts:
| part | data points | what each folder contains |
|---|---|---|
| full | 77,432 | obs_frame.png, obs_frame_depth.npy, sam_mask.png, trajectory.json |
| ego4d (annotations only) | 106,225 | sam_mask.png, trajectory.json, provenance.json β RGB/depth are rebuilt from your own Ego4D download, see below |
The 106,225 ego4d data points are human videos from Ego4D,
whose license allows redistributing annotations but not the video frames themselves.
Everything we produced for them is here; the frame is one script call away.
Of the 223,334 fitted curves in Table 7, 39,677 are not in this release: their affordance
mask was not tracked at the observation frame, or their trajectory.json lacks motion_2d
(HOI4D).
pip install -U huggingface_hub
hf download AFUN-dataset/AFUN_pool --repo-type dataset --local-dir afun_pool
cd afun_pool
for f in data/*.tar.zst; do tar --zstd -xf "$f"; done
Download β 363 GiB (of which the ego4d part is 0.26 GiB), extracted β 748 GiB. After extraction:
afun_pool/
βββ manifest.json # index β one entry per data point, with `part`
βββ reconstruct_ego4d_frames.py # rebuilds obs_frame.png for the ego4d part
βββ <source>/<episode>/<interval>/<cam>/ # 77,432 folders (part = "full")
β βββ obs_frame.png # RGB frame
β βββ obs_frame_depth.npy # float32 HΓW depth, millimeters
β βββ sam_mask.png # affordance mask (non-zero = actionable region)
β βββ trajectory.json # 3D motion + camera intrinsics
βββ ego4d/<episode>/<interval>/<cam>/ # 106,225 folders (part = "ego4d")
βββ sam_mask.png
βββ trajectory.json
βββ provenance.json # which Ego4D frame this is (+ pixel hash)
Loading a data point is identical to AFUN β manifest.json entries carry path,
dataset / episode_id / interval / cam, language, shard, plus part.
trajectory.json follows the two AFUN schemas (top-level fields for robot / human sources,
nested trajectories[] for scenefun3d); see the
AFUN README.
Each ego4d folder's provenance.json records the Ego4D video and frame:
{
"ego4d_video_uid": "0031d268-818c-4ec4-a804-935be610a61a",
"ego4d_frame_index": 56979, // 0-based frame in the full_scale video
"fps": 30.0,
"image_hw": [1440, 1920],
"rgb_sha256": "b22fda36β¦", // sha256 of the raw HΓWΓ3 uint8 pixels
"language": "Spread adhesive with the trowel over the underlayment.",
"vitra_episode_file": "ego4d_other/episodic_annotations/Ego4D_0031d268-β¦_ep_000859.npy",
"episode_frame_index": 0, // index within that VITRA-1M episode
"depth": { "model": "depth-anything/DA3NESTED-GIANT-LARGE-1.1", "...": "..." }
}
Get Ego4D access at https://ego4d-data.org and download only the videos you need:
pip install ego4d av Pillow
python reconstruct_ego4d_frames.py --pool-root ego4d --list-uids > uids.txt
ego4d --output_directory ~/ego4d --datasets full_scale --version v2 --video_uid_file uids.txt
Decode the frames into place and verify them against the recorded pixel hashes:
python reconstruct_ego4d_frames.py --pool-root ego4d --ego4d-root ~/ego4d --verify
This writes obs_frame.png into every ego4d/... folder using the same decoder and
frame indexing the annotations were made with (PyAV, native resolution, no resizing);
--verify confirms each decoded frame matches rgb_sha256.
The intervals were taken from VITRA-1M (MIT), and
vitra_episode_file / episode_frame_index locate the same frame in its episode files.
Depth maps are not shipped for this part; provenance.json β depth gives the settings we
used (Depth Anything 3 streaming video depth, every 4th frame, metric millimeters) if you
want to regenerate them.
| key | dataset | part | data points |
|---|---|---|---|
| ego4d | Ego4D via VITRA-1M (human videos) | ego4d | 106,225 |
| scenefun3d | SceneFun3D | full | 49,706 |
| vitra_epic | EPIC-KITCHENS via VITRA-1M (human videos) | full | 9,155 |
| robomind | RoboMIND | full | 7,077 |
| calvin | CALVIN | full | 3,105 |
| droid | DROID | full | 2,751 |
| robomind2 | RoboMIND 2 | full | 2,356 |
| rh20t_human | RH20T human demos | full | 1,422 |
| rh20t | RH20T | full | 1,044 |
| agibot | AgiBot World | full | 783 |
| rlbench | RLBench | full | 33 |
AFUN's vitra source corresponds to vitra_epic βͺ ego4d here. EPIC-KITCHENS frames are
redistributed under CC BY-NC 4.0 (Damen et al.); Ego4D frames are not redistributed.
The evaluation sets (AFUN_eval)
are disjoint from this pool at the sample level.
@article{wang2026afun,
title = {AFUN: Towards an Affordance Foundation Model for Functionality Understanding},
author = {Wang, Zhaoning and Zhong, Yi and Fu, Jiawei and Christensen, Henrik I. and Gao, Jun},
journal = {arXiv preprint arXiv:2606.02551},
year = {2026}
}
The full data pool of AFUN (arXiv:2606.02551):
183,657 data points for affordance segmentation and 3D interaction-motion prediction,
i.e. the fitted-curve pool of Table 7 from which the curated
AFUN training set (44,749) was
selected. AFUN β AFUN_pool. Same folder format as AFUN, same trajectory.json schemas.
It comes in two parts:
| part | data points | what each folder contains |
|---|---|---|
| full | 77,432 | obs_frame.png, obs_frame_depth.npy, sam_mask.png, trajectory.json |
| ego4d (annotations only) | 106,225 | sam_mask.png, trajectory.json, provenance.json β RGB/depth are rebuilt from your own Ego4D download, see below |
The 106,225 ego4d data points are human videos from Ego4D,
whose license allows redistributing annotations but not the video frames themselves.
Everything we produced for them is here; the frame is one script call away.
Of the 223,334 fitted curves in Table 7, 39,677 are not in this release: their affordance
mask was not tracked at the observation frame, or their trajectory.json lacks motion_2d
(HOI4D).
pip install -U huggingface_hub
hf download AFUN-dataset/AFUN_pool --repo-type dataset --local-dir afun_pool
cd afun_pool
for f in data/*.tar.zst; do tar --zstd -xf "$f"; done
Download β 363 GiB (of which the ego4d part is 0.26 GiB), extracted β 748 GiB. After extraction:
afun_pool/
βββ manifest.json # index β one entry per data point, with `part`
βββ reconstruct_ego4d_frames.py # rebuilds obs_frame.png for the ego4d part
βββ <source>/<episode>/<interval>/<cam>/ # 77,432 folders (part = "full")
β βββ obs_frame.png # RGB frame
β βββ obs_frame_depth.npy # float32 HΓW depth, millimeters
β βββ sam_mask.png # affordance mask (non-zero = actionable region)
β βββ trajectory.json # 3D motion + camera intrinsics
βββ ego4d/<episode>/<interval>/<cam>/ # 106,225 folders (part = "ego4d")
βββ sam_mask.png
βββ trajectory.json
βββ provenance.json # which Ego4D frame this is (+ pixel hash)
Loading a data point is identical to AFUN β manifest.json entries carry path,
dataset / episode_id / interval / cam, language, shard, plus part.
trajectory.json follows the two AFUN schemas (top-level fields for robot / human sources,
nested trajectories[] for scenefun3d); see the
AFUN README.
Each ego4d folder's provenance.json records the Ego4D video and frame:
{
"ego4d_video_uid": "0031d268-818c-4ec4-a804-935be610a61a",
"ego4d_frame_index": 56979, // 0-based frame in the full_scale video
"fps": 30.0,
"image_hw": [1440, 1920],
"rgb_sha256": "b22fda36β¦", // sha256 of the raw HΓWΓ3 uint8 pixels
"language": "Spread adhesive with the trowel over the underlayment.",
"vitra_episode_file": "ego4d_other/episodic_annotations/Ego4D_0031d268-β¦_ep_000859.npy",
"episode_frame_index": 0, // index within that VITRA-1M episode
"depth": { "model": "depth-anything/DA3NESTED-GIANT-LARGE-1.1", "...": "..." }
}
Get Ego4D access at https://ego4d-data.org and download only the videos you need:
pip install ego4d av Pillow
python reconstruct_ego4d_frames.py --pool-root ego4d --list-uids > uids.txt
ego4d --output_directory ~/ego4d --datasets full_scale --version v2 --video_uid_file uids.txt
Decode the frames into place and verify them against the recorded pixel hashes:
python reconstruct_ego4d_frames.py --pool-root ego4d --ego4d-root ~/ego4d --verify
This writes obs_frame.png into every ego4d/... folder using the same decoder and
frame indexing the annotations were made with (PyAV, native resolution, no resizing);
--verify confirms each decoded frame matches rgb_sha256.
The intervals were taken from VITRA-1M (MIT), and
vitra_episode_file / episode_frame_index locate the same frame in its episode files.
Depth maps are not shipped for this part; provenance.json β depth gives the settings we
used (Depth Anything 3 streaming video depth, every 4th frame, metric millimeters) if you
want to regenerate them.
| key | dataset | part | data points |
|---|---|---|---|
| ego4d | Ego4D via VITRA-1M (human videos) | ego4d | 106,225 |
| scenefun3d | SceneFun3D | full | 49,706 |
| vitra_epic | EPIC-KITCHENS via VITRA-1M (human videos) | full | 9,155 |
| robomind | RoboMIND | full | 7,077 |
| calvin | CALVIN | full | 3,105 |
| droid | DROID | full | 2,751 |
| robomind2 | RoboMIND 2 | full | 2,356 |
| rh20t_human | RH20T human demos | full | 1,422 |
| rh20t | RH20T | full | 1,044 |
| agibot | AgiBot World | full | 783 |
| rlbench | RLBench | full | 33 |
AFUN's vitra source corresponds to vitra_epic βͺ ego4d here. EPIC-KITCHENS frames are
redistributed under CC BY-NC 4.0 (Damen et al.); Ego4D frames are not redistributed.
The evaluation sets (AFUN_eval)
are disjoint from this pool at the sample level.
@article{wang2026afun,
title = {AFUN: Towards an Affordance Foundation Model for Functionality Understanding},
author = {Wang, Zhaoning and Zhong, Yi and Fu, Jiawei and Christensen, Henrik I. and Gao, Jun},
journal = {arXiv preprint arXiv:2606.02551},
year = {2026}
}