AFUN-dataset/AFUN_eval

Dataset

AFUN_eval

0

49 commits

1 linked in READMEs

updated Sep 3, 2026

See the code

README

AFUN_eval

The three 3D-motion evaluation sets of AFUN (arXiv:2606.02551, Table 3).

subsetfoldersamplesdescription
AFUN testafun_test/121held-out split across six robot/human sources
SceneFun3D testscenefun3d_test/721from the original SceneFun3D validation visits
RoboMIND2 testrobomind2_test/156out-of-domain robot data

The training data is released separately as AFUN; these evaluation sets are disjoint from it at the sample level.

Download

pip install -U huggingface_hub
hf download AFUN-dataset/AFUN_eval --repo-type dataset --local-dir afun_eval

Layout

Each subset has a manifest.json (one entry per sample: path, source fields, and language β€” the task instruction) plus one folder per sample:

<subset>/<source>/<episode>/<interval>/<cam>/
β”œβ”€β”€ obs_frame.png                    # RGB frame
β”œβ”€β”€ obs_frame_depth.npy              # float32 HΓ—W depth, millimeters
β”œβ”€β”€ sam_mask.png                     # GT affordance mask (non-zero = functional region)
└── trajectory.json                  # GT 3D motion + camera intrinsics

trajectory.json

3D positions are in the camera frame, in meters. Every file has the two fields the evaluation reads:

  • camera_info β€” intrinsics (fx, fy, cx, cy), distortion, T_base_to_cam
  • trajectory_3d β€” ground-truth motion of the interaction point: [{frame_idx, position_3d}, ...]
traj = json.load(open(f"{s['path']}/trajectory.json"))
gt = [p["position_3d"] for p in traj["trajectory_3d"]]   # camera frame, meters
fx = traj["camera_info"]["intrinsics"]["fx"]

afun_test / robomind2_test files also carry motion_2d (start/end pixel) and the fitted spline_params; the evaluation protocol does not read them.

Evaluation protocol (paper Β§5.2.3)

The method receives obs_frame.png, obs_frame_depth.npy, the camera intrinsics, and the instruction (from manifest.json), and predicts a 3D motion curve in the camera frame.

  • Prediction and ground-truth trajectory_3d are linearly resampled to 50 points each.
  • ADE / FDE β€” mean / final-point L2 distance in meters, computed in absolute scale (subscript a) and in relative scale (subscript r: both curves shifted so their first point is the origin).
  • CIM (contact-in-mask) β€” the first point of the predicted curve, pinhole-projected to pixels with the intrinsics, must land inside the GT mask (sam_mask.png > 127); projecting outside the image is a miss. Reported as the hit rate over all samples; samples where the method produced no prediction count as misses.
  • #fail β€” number of samples the method could not produce a prediction for.

Reproducing the paper numbers

The instruction lives in manifest.json, so a loader needs to join it with the sample folder. Minimal example:

import json, pathlib
sub = pathlib.Path("afun_eval/afun_test")
for s in json.load(open(sub / "manifest.json"))["samples"]:
    cam = sub / s["path"]
    instruction = s["language"]
    # cam/obs_frame.png, cam/obs_frame_depth.npy, cam/sam_mask.png, cam/trajectory.json

Evaluated with the paper's checkpoint, a fresh download of this release reproduces Table 3:

subsetADE_aFDE_aADE_rFDE_rCIM %#fail
AFUN test (121)0.0980.1390.0800.13581.00
SceneFun3D test (721)0.3510.4410.1350.26067.31
RoboMIND2 test (156)0.2540.3230.1770.27662.20

Measured on a fresh hf download of this dataset, every metric lands within 0.005 m (and CIM within 0.7 points) of the table above; the residual is bfloat16 inference non-determinism.

Citation

@article{wang2026afun,
  title   = {AFUN: Towards an Affordance Foundation Model for Functionality Understanding},
  author  = {Wang, Zhaoning and Zhong, Yi and Fu, Jiawei and Christensen, Henrik I. and Gao, Jun},
  journal = {arXiv preprint arXiv:2606.02551},
  year    = {2026}
}
3d-motion
affordance
manipulation
robotics
segmentation

AFUN-dataset/AFUN_eval

Dataset

AFUN_eval

0

49 commits

1 linked in READMEs

updated Sep 3, 2026

See the code

README

AFUN_eval

The three 3D-motion evaluation sets of AFUN (arXiv:2606.02551, Table 3).

subsetfoldersamplesdescription
AFUN testafun_test/121held-out split across six robot/human sources
SceneFun3D testscenefun3d_test/721from the original SceneFun3D validation visits
RoboMIND2 testrobomind2_test/156out-of-domain robot data

The training data is released separately as AFUN; these evaluation sets are disjoint from it at the sample level.

Download

pip install -U huggingface_hub
hf download AFUN-dataset/AFUN_eval --repo-type dataset --local-dir afun_eval

Layout

Each subset has a manifest.json (one entry per sample: path, source fields, and language β€” the task instruction) plus one folder per sample:

<subset>/<source>/<episode>/<interval>/<cam>/
β”œβ”€β”€ obs_frame.png                    # RGB frame
β”œβ”€β”€ obs_frame_depth.npy              # float32 HΓ—W depth, millimeters
β”œβ”€β”€ sam_mask.png                     # GT affordance mask (non-zero = functional region)
└── trajectory.json                  # GT 3D motion + camera intrinsics

trajectory.json

3D positions are in the camera frame, in meters. Every file has the two fields the evaluation reads:

  • camera_info β€” intrinsics (fx, fy, cx, cy), distortion, T_base_to_cam
  • trajectory_3d β€” ground-truth motion of the interaction point: [{frame_idx, position_3d}, ...]
traj = json.load(open(f"{s['path']}/trajectory.json"))
gt = [p["position_3d"] for p in traj["trajectory_3d"]]   # camera frame, meters
fx = traj["camera_info"]["intrinsics"]["fx"]

afun_test / robomind2_test files also carry motion_2d (start/end pixel) and the fitted spline_params; the evaluation protocol does not read them.

Evaluation protocol (paper Β§5.2.3)

The method receives obs_frame.png, obs_frame_depth.npy, the camera intrinsics, and the instruction (from manifest.json), and predicts a 3D motion curve in the camera frame.

  • Prediction and ground-truth trajectory_3d are linearly resampled to 50 points each.
  • ADE / FDE β€” mean / final-point L2 distance in meters, computed in absolute scale (subscript a) and in relative scale (subscript r: both curves shifted so their first point is the origin).
  • CIM (contact-in-mask) β€” the first point of the predicted curve, pinhole-projected to pixels with the intrinsics, must land inside the GT mask (sam_mask.png > 127); projecting outside the image is a miss. Reported as the hit rate over all samples; samples where the method produced no prediction count as misses.
  • #fail β€” number of samples the method could not produce a prediction for.

Reproducing the paper numbers

The instruction lives in manifest.json, so a loader needs to join it with the sample folder. Minimal example:

import json, pathlib
sub = pathlib.Path("afun_eval/afun_test")
for s in json.load(open(sub / "manifest.json"))["samples"]:
    cam = sub / s["path"]
    instruction = s["language"]
    # cam/obs_frame.png, cam/obs_frame_depth.npy, cam/sam_mask.png, cam/trajectory.json

Evaluated with the paper's checkpoint, a fresh download of this release reproduces Table 3:

subsetADE_aFDE_aADE_rFDE_rCIM %#fail
AFUN test (121)0.0980.1390.0800.13581.00
SceneFun3D test (721)0.3510.4410.1350.26067.31
RoboMIND2 test (156)0.2540.3230.1770.27662.20

Measured on a fresh hf download of this dataset, every metric lands within 0.005 m (and CIM within 0.7 points) of the table above; the residual is bfloat16 inference non-determinism.

Citation

@article{wang2026afun,
  title   = {AFUN: Towards an Affordance Foundation Model for Functionality Understanding},
  author  = {Wang, Zhaoning and Zhong, Yi and Fu, Jiawei and Christensen, Henrik I. and Gao, Jun},
  journal = {arXiv preprint arXiv:2606.02551},
  year    = {2026}
}
3d-motion
affordance
manipulation
robotics
segmentation