AFUN-dataset/AFUN

Dataset

AFUN

0

74 commits

1 linked in READMEs

updated Sep 3, 2026

See the code

README

AFUN

Training data of AFUN (arXiv:2606.02551): 44,749 data points for affordance segmentation and 3D interaction-motion prediction. Each data point is one folder: an RGB frame, its depth map, the ground-truth affordance mask, the ground-truth 3D motion, and a language instruction.

Download & extract

pip install -U huggingface_hub
hf download AFUN-dataset/AFUN --repo-type dataset --local-dir afun_train
cd afun_train
for f in data/*.tar.zst; do tar --zstd -xf "$f"; done

Download β‰ˆ 231 GiB, extracted β‰ˆ 522 GiB. After extraction:

afun_train/
β”œβ”€β”€ manifest.json                        # index β€” one entry per data point
└── <source>/<episode>/<interval>/<cam>/ # 44,749 folders
    β”œβ”€β”€ obs_frame.png                    # RGB frame
    β”œβ”€β”€ obs_frame_depth.npy              # float32 HΓ—W depth, millimeters
    β”œβ”€β”€ sam_mask.png                     # affordance mask (non-zero = actionable region)
    └── trajectory.json                  # 3D motion + camera intrinsics

Load a data point

import json, numpy as np
from PIL import Image

m = json.load(open("manifest.json"))
s = m["samples"][0]
rgb   = np.array(Image.open(f"{s['path']}/obs_frame.png"))
depth = np.load(f"{s['path']}/obs_frame_depth.npy")          # millimeters
mask  = np.array(Image.open(f"{s['path']}/sam_mask.png")) > 0
traj  = json.load(open(f"{s['path']}/trajectory.json"))
print(s["language"], traj["camera_info"]["intrinsics"])

Each manifest entry has path (the folder), dataset / episode_id / interval / cam, the instruction (language, with variants in queries), and shard (which archive contains it).

trajectory.json

3D positions are in the camera frame, in meters. camera_info holds the intrinsics (fx, fy, cx, cy), the distortion model, and T_base_to_cam.

There are two schemas, because SceneFun3D scenes are annotated differently from robot and human videos:

A. Robot / human sources (droid, robomind, agibot, rh20t, rh20t_human, calvin, rlbench, vitra) β€” one interaction per file, fields at the top level:

fieldmeaning
trajectory_3dGT motion of the interaction point, [{frame_idx, position_3d}, ...] β€” the curve-fitted (denoised) track
motion_2dstart / end pixel of the motion
spline_params.ctrlcontrol points of the fitted 3D curve (the training target is sampled from this curve)

B. scenefun3d β€” SceneFun3D is a set of annotated 3D scans, not videos. A scene can have several annotated functional parts (a drawer, a window, a tap), so the motions live in a list called trajectories, one entry per annotation. Each entry has the same fields as schema A (trajectory_3d, motion_2d, spline_params) plus:

fieldmeaning
annot_idthe original SceneFun3D annotation id
interval_languagethe instruction for this annotation
motion_typerot (hinged: door, window) or trans (sliding: drawer)
scenefun3d_motion_paramsthe analytic motion (see below)

scenefun3d_motion_params is what makes this source distinctive β€” the motion is given in closed form, not just as samples:

  • motion_type: "rot" β†’ motion_dir_cam (rotation axis), origin_cam (a point on the axis, i.e. the hinge), ref_cam (reference point), angle_rad (e.g. 1.5708 = 90Β°)
  • motion_type: "trans" β†’ motion_dir_cam (slide direction), origin_cam (start point), distance_m (e.g. 0.3), orient (inwards / outwards)

trajectory_3d is 15 points sampled from those parameters. In practice trajectories has length 1 (~95% of files; the rest have 2).

Reading either schema:

traj = json.load(open(f"{s['path']}/trajectory.json"))
if "trajectories" in traj:          # scenefun3d
    motion = traj["trajectories"][0]     # [0] is enough for almost every file
else:                               # robot / human sources
    motion = traj
points = [p["position_3d"] for p in motion["trajectory_3d"]]   # camera frame, meters

Sources

keydatasetdata points
scenefun3dSceneFun3D39,772
robomindRoboMIND2,197
vitraVITRA (human videos)1,205
droidDROID816
rh20t_humanRH20T human demos315
agibotAgiBot World299
rh20tRH20T93
calvinCALVIN46
rlbenchRLBench6

The evaluation sets are released separately as AFUN_eval and are disjoint from this set at the sample level.

Citation

@article{wang2026afun,
  title   = {AFUN: Towards an Affordance Foundation Model for Functionality Understanding},
  author  = {Wang, Zhaoning and Zhong, Yi and Fu, Jiawei and Christensen, Henrik I. and Gao, Jun},
  journal = {arXiv preprint arXiv:2606.02551},
  year    = {2026}
}
3d-motion
affordance
manipulation
robotics
segmentation

AFUN-dataset/AFUN

Dataset

AFUN

0

74 commits

1 linked in READMEs

updated Sep 3, 2026

See the code

README

AFUN

Training data of AFUN (arXiv:2606.02551): 44,749 data points for affordance segmentation and 3D interaction-motion prediction. Each data point is one folder: an RGB frame, its depth map, the ground-truth affordance mask, the ground-truth 3D motion, and a language instruction.

Download & extract

pip install -U huggingface_hub
hf download AFUN-dataset/AFUN --repo-type dataset --local-dir afun_train
cd afun_train
for f in data/*.tar.zst; do tar --zstd -xf "$f"; done

Download β‰ˆ 231 GiB, extracted β‰ˆ 522 GiB. After extraction:

afun_train/
β”œβ”€β”€ manifest.json                        # index β€” one entry per data point
└── <source>/<episode>/<interval>/<cam>/ # 44,749 folders
    β”œβ”€β”€ obs_frame.png                    # RGB frame
    β”œβ”€β”€ obs_frame_depth.npy              # float32 HΓ—W depth, millimeters
    β”œβ”€β”€ sam_mask.png                     # affordance mask (non-zero = actionable region)
    └── trajectory.json                  # 3D motion + camera intrinsics

Load a data point

import json, numpy as np
from PIL import Image

m = json.load(open("manifest.json"))
s = m["samples"][0]
rgb   = np.array(Image.open(f"{s['path']}/obs_frame.png"))
depth = np.load(f"{s['path']}/obs_frame_depth.npy")          # millimeters
mask  = np.array(Image.open(f"{s['path']}/sam_mask.png")) > 0
traj  = json.load(open(f"{s['path']}/trajectory.json"))
print(s["language"], traj["camera_info"]["intrinsics"])

Each manifest entry has path (the folder), dataset / episode_id / interval / cam, the instruction (language, with variants in queries), and shard (which archive contains it).

trajectory.json

3D positions are in the camera frame, in meters. camera_info holds the intrinsics (fx, fy, cx, cy), the distortion model, and T_base_to_cam.

There are two schemas, because SceneFun3D scenes are annotated differently from robot and human videos:

A. Robot / human sources (droid, robomind, agibot, rh20t, rh20t_human, calvin, rlbench, vitra) β€” one interaction per file, fields at the top level:

fieldmeaning
trajectory_3dGT motion of the interaction point, [{frame_idx, position_3d}, ...] β€” the curve-fitted (denoised) track
motion_2dstart / end pixel of the motion
spline_params.ctrlcontrol points of the fitted 3D curve (the training target is sampled from this curve)

B. scenefun3d β€” SceneFun3D is a set of annotated 3D scans, not videos. A scene can have several annotated functional parts (a drawer, a window, a tap), so the motions live in a list called trajectories, one entry per annotation. Each entry has the same fields as schema A (trajectory_3d, motion_2d, spline_params) plus:

fieldmeaning
annot_idthe original SceneFun3D annotation id
interval_languagethe instruction for this annotation
motion_typerot (hinged: door, window) or trans (sliding: drawer)
scenefun3d_motion_paramsthe analytic motion (see below)

scenefun3d_motion_params is what makes this source distinctive β€” the motion is given in closed form, not just as samples:

  • motion_type: "rot" β†’ motion_dir_cam (rotation axis), origin_cam (a point on the axis, i.e. the hinge), ref_cam (reference point), angle_rad (e.g. 1.5708 = 90Β°)
  • motion_type: "trans" β†’ motion_dir_cam (slide direction), origin_cam (start point), distance_m (e.g. 0.3), orient (inwards / outwards)

trajectory_3d is 15 points sampled from those parameters. In practice trajectories has length 1 (~95% of files; the rest have 2).

Reading either schema:

traj = json.load(open(f"{s['path']}/trajectory.json"))
if "trajectories" in traj:          # scenefun3d
    motion = traj["trajectories"][0]     # [0] is enough for almost every file
else:                               # robot / human sources
    motion = traj
points = [p["position_3d"] for p in motion["trajectory_3d"]]   # camera frame, meters

Sources

keydatasetdata points
scenefun3dSceneFun3D39,772
robomindRoboMIND2,197
vitraVITRA (human videos)1,205
droidDROID816
rh20t_humanRH20T human demos315
agibotAgiBot World299
rh20tRH20T93
calvinCALVIN46
rlbenchRLBench6

The evaluation sets are released separately as AFUN_eval and are disjoint from this set at the sample level.

Citation

@article{wang2026afun,
  title   = {AFUN: Towards an Affordance Foundation Model for Functionality Understanding},
  author  = {Wang, Zhaoning and Zhong, Yi and Fu, Jiawei and Christensen, Henrik I. and Gao, Jun},
  journal = {arXiv preprint arXiv:2606.02551},
  year    = {2026}
}
3d-motion
affordance
manipulation
robotics
segmentation