AFUN-dataset/AFUN_pool

Dataset

AFUN_pool

0

50 commits

2 linked in READMEs

updated Sep 6, 2026

See the code

README

AFUN_pool

The full data pool of AFUN (arXiv:2606.02551): 183,657 data points for affordance segmentation and 3D interaction-motion prediction, i.e. the fitted-curve pool of Table 7 from which the curated AFUN training set (44,749) was selected. AFUN βŠ‚ AFUN_pool. Same folder format as AFUN, same trajectory.json schemas.

It comes in two parts:

partdata pointswhat each folder contains
full77,432obs_frame.png, obs_frame_depth.npy, sam_mask.png, trajectory.json
ego4d (annotations only)106,225sam_mask.png, trajectory.json, provenance.json β€” RGB/depth are rebuilt from your own Ego4D download, see below

The 106,225 ego4d data points are human videos from Ego4D, whose license allows redistributing annotations but not the video frames themselves. Everything we produced for them is here; the frame is one script call away.

Of the 223,334 fitted curves in Table 7, 39,677 are not in this release: their affordance mask was not tracked at the observation frame, or their trajectory.json lacks motion_2d (HOI4D).

Download & extract

pip install -U huggingface_hub
hf download AFUN-dataset/AFUN_pool --repo-type dataset --local-dir afun_pool
cd afun_pool
for f in data/*.tar.zst; do tar --zstd -xf "$f"; done

Download β‰ˆ 363 GiB (of which the ego4d part is 0.26 GiB), extracted β‰ˆ 748 GiB. After extraction:

afun_pool/
β”œβ”€β”€ manifest.json                          # index β€” one entry per data point, with `part`
β”œβ”€β”€ reconstruct_ego4d_frames.py            # rebuilds obs_frame.png for the ego4d part
β”œβ”€β”€ <source>/<episode>/<interval>/<cam>/   # 77,432 folders  (part = "full")
β”‚   β”œβ”€β”€ obs_frame.png                      # RGB frame
β”‚   β”œβ”€β”€ obs_frame_depth.npy                # float32 HΓ—W depth, millimeters
β”‚   β”œβ”€β”€ sam_mask.png                       # affordance mask (non-zero = actionable region)
β”‚   └── trajectory.json                    # 3D motion + camera intrinsics
└── ego4d/<episode>/<interval>/<cam>/      # 106,225 folders (part = "ego4d")
    β”œβ”€β”€ sam_mask.png
    β”œβ”€β”€ trajectory.json
    └── provenance.json                    # which Ego4D frame this is (+ pixel hash)

Loading a data point is identical to AFUN β€” manifest.json entries carry path, dataset / episode_id / interval / cam, language, shard, plus part. trajectory.json follows the two AFUN schemas (top-level fields for robot / human sources, nested trajectories[] for scenefun3d); see the AFUN README.

Rebuilding the Ego4D frames

Each ego4d folder's provenance.json records the Ego4D video and frame:

{
  "ego4d_video_uid": "0031d268-818c-4ec4-a804-935be610a61a",
  "ego4d_frame_index": 56979,           // 0-based frame in the full_scale video
  "fps": 30.0,
  "image_hw": [1440, 1920],
  "rgb_sha256": "b22fda36…",            // sha256 of the raw HΓ—WΓ—3 uint8 pixels
  "language": "Spread adhesive with the trowel over the underlayment.",
  "vitra_episode_file": "ego4d_other/episodic_annotations/Ego4D_0031d268-…_ep_000859.npy",
  "episode_frame_index": 0,             // index within that VITRA-1M episode
  "depth": { "model": "depth-anything/DA3NESTED-GIANT-LARGE-1.1", "...": "..." }
}
  1. Get Ego4D access at https://ego4d-data.org and download only the videos you need:

    pip install ego4d av Pillow
    python reconstruct_ego4d_frames.py --pool-root ego4d --list-uids > uids.txt
    ego4d --output_directory ~/ego4d --datasets full_scale --version v2 --video_uid_file uids.txt
    
  2. Decode the frames into place and verify them against the recorded pixel hashes:

    python reconstruct_ego4d_frames.py --pool-root ego4d --ego4d-root ~/ego4d --verify
    

    This writes obs_frame.png into every ego4d/... folder using the same decoder and frame indexing the annotations were made with (PyAV, native resolution, no resizing); --verify confirms each decoded frame matches rgb_sha256.

The intervals were taken from VITRA-1M (MIT), and vitra_episode_file / episode_frame_index locate the same frame in its episode files. Depth maps are not shipped for this part; provenance.json β†’ depth gives the settings we used (Depth Anything 3 streaming video depth, every 4th frame, metric millimeters) if you want to regenerate them.

Sources

keydatasetpartdata points
ego4dEgo4D via VITRA-1M (human videos)ego4d106,225
scenefun3dSceneFun3Dfull49,706
vitra_epicEPIC-KITCHENS via VITRA-1M (human videos)full9,155
robomindRoboMINDfull7,077
calvinCALVINfull3,105
droidDROIDfull2,751
robomind2RoboMIND 2full2,356
rh20t_humanRH20T human demosfull1,422
rh20tRH20Tfull1,044
agibotAgiBot Worldfull783
rlbenchRLBenchfull33

AFUN's vitra source corresponds to vitra_epic βˆͺ ego4d here. EPIC-KITCHENS frames are redistributed under CC BY-NC 4.0 (Damen et al.); Ego4D frames are not redistributed. The evaluation sets (AFUN_eval) are disjoint from this pool at the sample level.

Citation

@article{wang2026afun,
  title   = {AFUN: Towards an Affordance Foundation Model for Functionality Understanding},
  author  = {Wang, Zhaoning and Zhong, Yi and Fu, Jiawei and Christensen, Henrik I. and Gao, Jun},
  journal = {arXiv preprint arXiv:2606.02551},
  year    = {2026}
}
3d-motion
affordance
manipulation
robotics
segmentation

AFUN-dataset/AFUN_pool

Dataset

AFUN_pool

0

50 commits

2 linked in READMEs

updated Sep 6, 2026

See the code

README

AFUN_pool

The full data pool of AFUN (arXiv:2606.02551): 183,657 data points for affordance segmentation and 3D interaction-motion prediction, i.e. the fitted-curve pool of Table 7 from which the curated AFUN training set (44,749) was selected. AFUN βŠ‚ AFUN_pool. Same folder format as AFUN, same trajectory.json schemas.

It comes in two parts:

partdata pointswhat each folder contains
full77,432obs_frame.png, obs_frame_depth.npy, sam_mask.png, trajectory.json
ego4d (annotations only)106,225sam_mask.png, trajectory.json, provenance.json β€” RGB/depth are rebuilt from your own Ego4D download, see below

The 106,225 ego4d data points are human videos from Ego4D, whose license allows redistributing annotations but not the video frames themselves. Everything we produced for them is here; the frame is one script call away.

Of the 223,334 fitted curves in Table 7, 39,677 are not in this release: their affordance mask was not tracked at the observation frame, or their trajectory.json lacks motion_2d (HOI4D).

Download & extract

pip install -U huggingface_hub
hf download AFUN-dataset/AFUN_pool --repo-type dataset --local-dir afun_pool
cd afun_pool
for f in data/*.tar.zst; do tar --zstd -xf "$f"; done

Download β‰ˆ 363 GiB (of which the ego4d part is 0.26 GiB), extracted β‰ˆ 748 GiB. After extraction:

afun_pool/
β”œβ”€β”€ manifest.json                          # index β€” one entry per data point, with `part`
β”œβ”€β”€ reconstruct_ego4d_frames.py            # rebuilds obs_frame.png for the ego4d part
β”œβ”€β”€ <source>/<episode>/<interval>/<cam>/   # 77,432 folders  (part = "full")
β”‚   β”œβ”€β”€ obs_frame.png                      # RGB frame
β”‚   β”œβ”€β”€ obs_frame_depth.npy                # float32 HΓ—W depth, millimeters
β”‚   β”œβ”€β”€ sam_mask.png                       # affordance mask (non-zero = actionable region)
β”‚   └── trajectory.json                    # 3D motion + camera intrinsics
└── ego4d/<episode>/<interval>/<cam>/      # 106,225 folders (part = "ego4d")
    β”œβ”€β”€ sam_mask.png
    β”œβ”€β”€ trajectory.json
    └── provenance.json                    # which Ego4D frame this is (+ pixel hash)

Loading a data point is identical to AFUN β€” manifest.json entries carry path, dataset / episode_id / interval / cam, language, shard, plus part. trajectory.json follows the two AFUN schemas (top-level fields for robot / human sources, nested trajectories[] for scenefun3d); see the AFUN README.

Rebuilding the Ego4D frames

Each ego4d folder's provenance.json records the Ego4D video and frame:

{
  "ego4d_video_uid": "0031d268-818c-4ec4-a804-935be610a61a",
  "ego4d_frame_index": 56979,           // 0-based frame in the full_scale video
  "fps": 30.0,
  "image_hw": [1440, 1920],
  "rgb_sha256": "b22fda36…",            // sha256 of the raw HΓ—WΓ—3 uint8 pixels
  "language": "Spread adhesive with the trowel over the underlayment.",
  "vitra_episode_file": "ego4d_other/episodic_annotations/Ego4D_0031d268-…_ep_000859.npy",
  "episode_frame_index": 0,             // index within that VITRA-1M episode
  "depth": { "model": "depth-anything/DA3NESTED-GIANT-LARGE-1.1", "...": "..." }
}
  1. Get Ego4D access at https://ego4d-data.org and download only the videos you need:

    pip install ego4d av Pillow
    python reconstruct_ego4d_frames.py --pool-root ego4d --list-uids > uids.txt
    ego4d --output_directory ~/ego4d --datasets full_scale --version v2 --video_uid_file uids.txt
    
  2. Decode the frames into place and verify them against the recorded pixel hashes:

    python reconstruct_ego4d_frames.py --pool-root ego4d --ego4d-root ~/ego4d --verify
    

    This writes obs_frame.png into every ego4d/... folder using the same decoder and frame indexing the annotations were made with (PyAV, native resolution, no resizing); --verify confirms each decoded frame matches rgb_sha256.

The intervals were taken from VITRA-1M (MIT), and vitra_episode_file / episode_frame_index locate the same frame in its episode files. Depth maps are not shipped for this part; provenance.json β†’ depth gives the settings we used (Depth Anything 3 streaming video depth, every 4th frame, metric millimeters) if you want to regenerate them.

Sources

keydatasetpartdata points
ego4dEgo4D via VITRA-1M (human videos)ego4d106,225
scenefun3dSceneFun3Dfull49,706
vitra_epicEPIC-KITCHENS via VITRA-1M (human videos)full9,155
robomindRoboMINDfull7,077
calvinCALVINfull3,105
droidDROIDfull2,751
robomind2RoboMIND 2full2,356
rh20t_humanRH20T human demosfull1,422
rh20tRH20Tfull1,044
agibotAgiBot Worldfull783
rlbenchRLBenchfull33

AFUN's vitra source corresponds to vitra_epic βˆͺ ego4d here. EPIC-KITCHENS frames are redistributed under CC BY-NC 4.0 (Damen et al.); Ego4D frames are not redistributed. The evaluation sets (AFUN_eval) are disjoint from this pool at the sample level.

Citation

@article{wang2026afun,
  title   = {AFUN: Towards an Affordance Foundation Model for Functionality Understanding},
  author  = {Wang, Zhaoning and Zhong, Yi and Fu, Jiawei and Christensen, Henrik I. and Gao, Jun},
  journal = {arXiv preprint arXiv:2606.02551},
  year    = {2026}
}
3d-motion
affordance
manipulation
robotics
segmentation