Demo for recognizing the main stages of hand-object manipulation in everyday videos.
Python
0
0 commits
updated Sep 28, 2026
Hand action recognition (hand-object manipulation stages) in everyday videos with Hiera-Hand (ECCVW 2024). It tracks people, localizes their hands, and labels every hand in every frame as grasp, hold, operate, or release.
⚠️ This is a standalone demo based on the original research work; for training, evaluation, and the dataset, see idiap/childplay_hand.

./setup_env.sh # needs uv; Python 3.11 venv in .venv
source .venv/bin/activate
Fetch the checkpoints from Zenodo into checkpoints/:
python src/download_checkpoints.py # manipulation (~0.4 GB download)
python src/download_checkpoints.py --task all # + object (~0.8 GB download)
python src/demo.py input.mp4 --output result.mp4
| Option | Default | |
|---|---|---|
--stride N | 1 | Run Hiera every N frames; 2 ≈ 2× faster |
--task | manipulation | or object (object in hand) |
--device | auto | CUDA → MPS → CPU |
--batch-size | 16 / 4 / 2 | CUDA / MPS / CPU |
--smoothing-window | 5 | frames; 0 = off |
--save-intermediates | off | writes pose.pkl, pred.pkl |
Works best on clips where people are fully visible, without camera cuts.
Note: The paper used HRNet-W32 for pose; this demo uses YOLO26m-pose for speed, so predictions may differ slightly from the reported results.
src/hiera/): Apache-2.0, © Meta.@inproceedings{Farkhondeh_ECCVW_2024,
author = {Farkhondeh*, Arya and Tafasca*, Samy and Odobez, Jean-Marc},
title = {ChildPlay-Hand: A Dataset of Hand Manipulations in the Wild},
booktitle = {Proceedings of the European Conference on Computer Vision (ECCV) Workshops},
year = {2024},
note = {* Equal contribution}
}
Python
93.6%
Jupyter Notebook
5.3%
Shell
1.1%
Demo for recognizing the main stages of hand-object manipulation in everyday videos.
Python
0
0 commits
updated Sep 28, 2026
Hand action recognition (hand-object manipulation stages) in everyday videos with Hiera-Hand (ECCVW 2024). It tracks people, localizes their hands, and labels every hand in every frame as grasp, hold, operate, or release.
⚠️ This is a standalone demo based on the original research work; for training, evaluation, and the dataset, see idiap/childplay_hand.

./setup_env.sh # needs uv; Python 3.11 venv in .venv
source .venv/bin/activate
Fetch the checkpoints from Zenodo into checkpoints/:
python src/download_checkpoints.py # manipulation (~0.4 GB download)
python src/download_checkpoints.py --task all # + object (~0.8 GB download)
python src/demo.py input.mp4 --output result.mp4
| Option | Default | |
|---|---|---|
--stride N | 1 | Run Hiera every N frames; 2 ≈ 2× faster |
--task | manipulation | or object (object in hand) |
--device | auto | CUDA → MPS → CPU |
--batch-size | 16 / 4 / 2 | CUDA / MPS / CPU |
--smoothing-window | 5 | frames; 0 = off |
--save-intermediates | off | writes pose.pkl, pred.pkl |
Works best on clips where people are fully visible, without camera cuts.
Note: The paper used HRNet-W32 for pose; this demo uses YOLO26m-pose for speed, so predictions may differ slightly from the reported results.
src/hiera/): Apache-2.0, © Meta.@inproceedings{Farkhondeh_ECCVW_2024,
author = {Farkhondeh*, Arya and Tafasca*, Samy and Odobez, Jean-Marc},
title = {ChildPlay-Hand: A Dataset of Hand Manipulations in the Wild},
booktitle = {Proceedings of the European Conference on Computer Vision (ECCV) Workshops},
year = {2024},
note = {* Equal contribution}
}
Python
93.6%
Jupyter Notebook
5.3%
Shell
1.1%