Linzhan/UniML3D

Dataset

UniML3D

11

184 commits

1 linked in READMEs

updated Oct 1, 2026

See the code

README

UniML3D

arXiv Hugging Face paper page Project page GitHub

UniML3D: text-paired motion clips for humanoids, dragons, crabs, snakes, octopuses, cheetahs and hands

UniML3D is the text-paired, topology-annotated motion dataset behind UniMate (SIGGRAPH Asia 2026): motion clips from three sources with very different skeletons β€” Mixamo humanoids, Truebones ZOO animals and rigged Objaverse-XL objects β€” brought into one canonical layout, captioned, and annotated with cleaned joint names, a body-plan category and a facing-direction joint pair per skeleton. Every annotation in it was generated by this project's own data pipeline, not inherited from a third-party annotation set.

TruebonesMixamoObjaverse-XLtotal
clips1,0972,31710,35513,769
skeletons (object types)7417,3557,430
frames (30 fps)111,743260,0901,629,4792,001,312
duration62.1 min144.5 min905.3 min18.5 h
joints per skeleton9 – 140653 – 117
clip length (frames)median 81 Β· max 556median 70 Β· max 1,401median 68 Β· max 22,353

This dataset is actively maintained and will keep being updated. The annotations are re-reviewed in passes, so captions, cleaned joint names, facing pairs, body-plan categories and the per-clip quality filters can all change between revisions β€” and with them the clip and skeleton counts above. Pin a revision if you need a snapshot that does not move under you:

hf download Linzhan/UniML3D --type dataset --revision <commit-sha> --local-dir dataset

Distribution of clip lengths across the three sources

Distribution of joint counts across the three sources

Clip lengths (left: share of clips per log-spaced bin; right: share of clips at least N frames long, with the 60 / 90 / 120 / 180-frame training windows marked) and joint counts per skeleton (Mixamo is a single 65-joint rig, hence the step). Both figures are regenerated from clip_frames.json / joint_count.json by data_process/tools/vis_clip_frames.py and vis_joint_count.py; the per-source versions sit next to each JSON as export/<dataset>/*_distribution.png.

Getting the data

hf download Linzhan/UniML3D --type dataset --local-dir dataset                          # everything, β‰ˆ 51 GB
hf download Linzhan/UniML3D --type dataset --local-dir dataset --include "export/truebones/*"   # one source
hf download Linzhan/UniML3D --type dataset --local-dir dataset --include "export/*/*.json" \
    --include "export/*/*.txt" --include "export/*/*.png"                                  # annotations and figures only

export/ is the processed dataset: one NPZ per clip in a uniform format, plus every caption and annotation. It is exactly what data_process/feature_extraction consumes to build the canonicalized training tensors, so the training layer is derived locally rather than shipped here.

Browse it first

Two tables at the repository root carry the whole index, so the dataset viewer above and the datasets loader both work without fetching a single motion file:

tablerowsone row is
clips.csv13,769a clip: dataset, clip, object_type, frames, fps, duration_sec, num_joints, category, caption, video, tpose
skeletons.csv7,430a skeleton: dataset, object_type, num_joints, num_clips, total_frames, duration_sec, category, face_right_joint, face_left_joint, face_source, body_axis, tpose

video and tpose hold repo-relative paths, so any row resolves to a real file:

from datasets import load_dataset
from huggingface_hub import hf_hub_download

clips = load_dataset("Linzhan/UniML3D", "clips", split="train")
skels = load_dataset("Linzhan/UniML3D", "skeletons", split="train")

birds = skels.filter(lambda r: r["category"] == "avian")            # 143 skeletons
long_ = clips.filter(lambda r: r["frames"] >= 60)                   # usable at a 60-frame window
mp4 = hf_hub_download("Linzhan/UniML3D", clips[0]["video"], repo_type="dataset")

Layout

One directory per source, all of it the processed export layer:

The repository root mirrors the pipeline's dataset/ directory, so a download lands directly in the layout UniMate reads:

export/<dataset>/                     stage 1-3.  <dataset> ∈ {truebones, mixamo, objaverse}
  motions/<clip>.npz                  one export NPZ per clip, in the source rig
  videos/<clip>.mp4                   skeleton preview of every clip
                                      (objaverse: both are hash-sharded, see below)
  videos_captioned/<clip>.mp4         skeleton preview with the clip's caption burned in
                                      (objaverse: only clips outside filtered_clips.txt / filtered_objects.txt)
  tpose/<object_type>.png             rest-pose skeleton render per object type
  summary.json                        clip / frame totals and fps
  clip_frames.json                    {clip: n_frames}
  joint_count.json                    {object_type: n_joints}
  joint_names.json                    {object_type: [raw joint names]}            (stage 1)
  clean_joint_names.json              {object_type: [canonical anatomical names]} (stage 3)
  face_joint_names.json               {object_type: facing-direction joint pair}  (stage 3)
  motion_captions.json                {clip: caption}                             (stage 2)
  category_groups.json                {body-plan category: [object types]}        (stage 2)
  filtered_clips.txt                  clips stage 4 skips
  clip_trims.txt                      clips whose first N frames stage 4 drops
  activity_keep.txt                   (objaverse) small real motions stage 4 keeps
  filtered_objects.txt                (objaverse) whole rigs stage 4 skips
  rig_flags.json                      (objaverse) every flagged rig, with category and reason
  clip_frames_distribution.png        frames-per-clip distribution figure
  joint_count_distribution.png        joints-per-skeleton distribution figure

assets/                               figures used by this card

Sizes: export/motions β‰ˆ 48 GB (46 GB of it Objaverse-XL), export/videos β‰ˆ 2.5 GB, export/videos_captioned β‰ˆ 4.2 GB, export/tpose β‰ˆ 0.7 GB, sidecars β‰ˆ 26 MB.

Clip naming. Clips are named <object_type>-<action>, and the object type never contains -, so clip.split('-', 1)[0] recovers it. Mixamo is a single shared rig, so its clips carry no object-type prefix and its object type is simply mixamo.

Hash-sharded folders. The Hub allows at most 10,000 files per directory, and Objaverse-XL has 10,355 clips, so export/objaverse/motions/ and export/objaverse/videos/ are stored as 64 buckets β€” <sub>/<xx>/<file>, where xx is sha1(filename) mod 64 in hex. Every other folder is flat. Either read both levels,

glob.glob(f'{export_dir}/motions/*.npz') + glob.glob(f'{export_dir}/motions/*/*.npz')

or flatten the buckets once after downloading (hardlinks, so it costs no disk):

for d in dataset/export/objaverse/{motions,videos}; do find "$d" -mindepth 2 -type f -exec ln -f -t "$d" -- {} +; done

Data formats

Export NPZ β€” export/<dataset>/motions/<clip>.npz

One file per source clip, float64, written by stage 1 from the source FBX/GLB with Blender: control and helper bones that carry no skin weight are pruned (jointly across all clips of a skeleton, so every clip of an object type shares one topology), the rig is converted to Y-up, frames identical to the rest pose are dropped, and clips shorter than 5 frames are skipped.

fieldshapemeaning
rest_local_pos / rest_local_rot(J, 3) / (J, 4)rest-pose local translations / rotations (quaternions, w-first)
anim_local_pos / anim_local_rot(T, J, 3) / (T, J, 4)per-frame local translations / rotations
offsets(J, 3)bind-pose bone offsets
parents(J,)parent indices, -1 for the root
names(J,)bone name strings
skin_matrix(V, J)skinning weights (empty (0, J) for Mixamo, whose source clips carry no mesh)
fps, action_namescalarframe rate (30) and source action

Captions β€” motion_captions.json

One sentence per clip, object-relative directions, no appearance. Every caption is produced by this dataset's own pipeline: a vision-language model (Qwen3.5-9B) reads a four-view render of the clip, and the source asset's own action name is used only as a disambiguating hint where the asset carries one.

Body-plan categories β€” category_groups.json

Every skeleton is assigned one of bipedal, quadrupedal, avian, insectoid, marine, serpentine, articulated_rigid (plus uncertain) by a vision-language model (Qwen3.5-9B, data_process/vlm_caption/classify_category.py) looking at its T-pose grid, the first frame of a clip, its joint-label histogram and its captions. Classifier mistakes found by hand were corrected in the shipped category_groups.json.

All 7,430 skeletons carry a category. The classifier leaves a rig uncertain when its votes disagree (158 Objaverse-XL rigs, none in Truebones or Mixamo); those still have full motion data and annotations β€” the category only drives optional category-balanced sampling.

Joint annotations

clean_joint_names.json maps every raw bone name to a canonical anatomical vocabulary (mixamorig:LeftUpLeg β†’ Left Thigh, R_momo β†’ Right Thigh). face_joint_names.json names, per skeleton, the bilaterally symmetric pair whose left-right vector defines the facing direction:

"Alligator": {"r_hip": {"raw": "R_momo", "clean": "Right Thigh"},
              "l_hip": {"raw": "L_momo", "clean": "Left Thigh"}, "source": "thigh"}

Serpentine rigs with no bilateral symmetry use a head/tail body-axis pair instead and carry "body_axis": true; a skeleton with no usable pair has "source": "empty" and is canonicalized with identity facing.

Skip lists

Optional lists sit next to the data and are read automatically by stage 4; delete one and that step is skipped.

  • filtered_clips.txt β€” individual clips to skip, each with its reason. Objaverse-XL has 453:

    • clips whose motion is discontinuous inside the frames stage 4 uses (several actions stitched together, a loop seam, a root teleport, garbage frames);
    • held poses;
    • exact duplicates of another clip on the same rig;
    • roots that fly or glide in a way the body motion does not explain;
    • visible per-joint jitter;
    • renders where the rig drives only part of the mesh;
    • clips excluded on hand review, including the reference pipeline's own hand deletions.

    Truebones has 5: four whose export disagrees with the object's reference (a skeleton with missing joints, a mirrored rig, a mis-posed limb), and one whose root moves implausibly fast.

    Mixamo has 7: bodies held in the air with nothing under them (a fall into an absent pool, a carried person without the carrier) and a root that glides with no steps.

  • clip_trims.txt (objaverse) β€” 54 clips kept after dropping a few opening frames. Each line is <clip> <N>: the first N frames of motions/<clip>.npz are a bind pose or an unrelated pose that snaps into the motion, and stage 4 drops them when it loads the clip. If you read the NPZs yourself, use anim_local_rot[N:] and anim_local_pos[N:] for these clips.

  • activity_keep.txt (objaverse) β€” 22 clips exempt from stage 4's low-activity filter: real motions confined to a few joints (a head turn, a nod, a clap, a wave) that the filter would otherwise read as a held pose.

  • filtered_objects.txt (objaverse) β€” 899 whole rigs:

    • 686 whose exported rest pose is unusable as the canonical T-pose: it lies flat, is rotated, reclined or upside-down, or the skeleton does not match the mesh;
    • 213 whose source asset is held out of training.
  • rig_flags.json (objaverse) records the full picture behind those lists: every flagged rig with a category (tpose_wrong, not_in_legacy_raw, empty_pair for a skeleton with no bilateral joint pair, bone_pair, facing_wrong) and the reason, plus the rigs whose automatic flag was cleared by hand (verified_ok). Only tpose_wrong and not_in_legacy_raw reach filtered_objects.txt; the other categories are informational and those rigs are kept.

Stage 4 additionally drops clips that fail its own quality checks (too short, too little motion). The table is the training set of the UniMate paper: the three filters took the 13,769 exported clips down to 11,850. The lists above have been extended since, so rerunning stage 4 on this export yields fewer Objaverse-XL clips.

exportedβˆ’ skip listsβˆ’ stage-4 quality= training clipsskeletonsframesduration
Truebones1,0971,0941,0231,0237495,49153.1 min
Mixamo2,3172,3172,1692,1691189,973105.5 min
Objaverse-XL10,3559,0038,6588,6586,479784,571435.9 min
total13,76912,41411,85011,8506,5541,070,0359.9 h

Previews

export/<ds>/videos/<clip>.mp4 is a matplotlib stick-figure render of the exported skeleton motion in its source rig (not the textured asset). tpose/<object_type>.png shows the pruned rest pose. videos_captioned/<clip>.mp4 renders the same export clip on a checkerboard floor, with its motion_captions.json caption as the title, so a caption can be checked against the motion it describes. For Objaverse-XL it covers only the clips that the skip lists keep.

Reproducing it

Every layer here is built by the open data pipeline at UniMate/data_process, which also documents each stage. Derive the training clips from this repository:

hf download Linzhan/UniML3D --type dataset --local-dir dataset --include "export/*"
bash data_process/scripts/run_extract_features.sh objaverse

Or rebuild the whole thing from the raw assets, which live in their own mirrors on the Hub β€” Mixamo-Animations-Characters, Objaverse-XL-Rigged-Animated with its renders, and Truebones-ZOO-Annotations:

bash data_process/scripts/run_download.sh objaverse          # raw assets from the Hub
bash data_process/scripts/run_export.sh objaverse            # stage 1: assets  -> NPZ + previews
bash data_process/scripts/run_caption_motion.sh objaverse    # stage 2: captions
bash data_process/scripts/run_joints_names_clean_rule.sh objaverse   # stage 3: annotations
bash data_process/scripts/run_extract_features.sh objaverse   # stage 4: training clips

Every stage is deterministic, so the same inputs regenerate the same files.

Licensing

The three sources keep their own terms; this repository does not relicense any of them.

  • Mixamo β€” Adobe's Mixamo terms apply to the motions (export/mixamo/motions) and their previews. Source mirror: Mixamo-Animations-Characters.
  • Objaverse-XL β€” every clip inherits the upstream licence of the Objaverse-XL object it was exported from; there is no blanket licence. The clip stem is the object id, so the licence resolves through the Objaverse-XL annotations. Source mirror and full notice: Objaverse-XL-Rigged-Animated.
  • Truebones β€” Truebones ZOO is a commercial library that may not be redistributed. export/truebones/motions/ is therefore not part of the public release: purchase the pack from Truebones, rebuild the per-clip FBXs with the Truebones-ZOO-Annotations build scripts, and re-run the data pipeline β€” it is deterministic and regenerates the NPZs byte-for-byte. The sidecars, previews and annotations for Truebones are included, and like every other annotation here they were generated by that pipeline rather than taken from a third-party annotation set.

Everything this repository adds on top of those sources β€” the captions for all three, the cleaned joint labels, the facing pairs, the body-plan categories, the skip lists, the index tables and the figures β€” was produced by the data pipeline in this project and is offered under ODC-BY 1.0, matching the derived-metadata licence of the source mirrors.

Citation

@article{mou2026unimate,
  title   = {UniMate: One Unified Model to Animate Diverse Skeletons},
  author  = {Mou, Linzhan and Lei, Jiahui and Dou, Zhiyang and Cai, Chenyue and Song, Chaoyue and Finkelstein, Adam and Rusinkiewicz, Szymon},
  journal = {arXiv preprint arXiv:2609.05415},
  year    = {2026}
}
3d
animation
heterogeneous-skeletons
mixamo
mocap
motion
objaverse
skeletal-animation
text-to-motion
truebones

Linzhan/UniML3D

Dataset

UniML3D

11

184 commits

1 linked in READMEs

updated Oct 1, 2026

See the code

README

UniML3D

arXiv Hugging Face paper page Project page GitHub

UniML3D: text-paired motion clips for humanoids, dragons, crabs, snakes, octopuses, cheetahs and hands

UniML3D is the text-paired, topology-annotated motion dataset behind UniMate (SIGGRAPH Asia 2026): motion clips from three sources with very different skeletons β€” Mixamo humanoids, Truebones ZOO animals and rigged Objaverse-XL objects β€” brought into one canonical layout, captioned, and annotated with cleaned joint names, a body-plan category and a facing-direction joint pair per skeleton. Every annotation in it was generated by this project's own data pipeline, not inherited from a third-party annotation set.

TruebonesMixamoObjaverse-XLtotal
clips1,0972,31710,35513,769
skeletons (object types)7417,3557,430
frames (30 fps)111,743260,0901,629,4792,001,312
duration62.1 min144.5 min905.3 min18.5 h
joints per skeleton9 – 140653 – 117
clip length (frames)median 81 Β· max 556median 70 Β· max 1,401median 68 Β· max 22,353

This dataset is actively maintained and will keep being updated. The annotations are re-reviewed in passes, so captions, cleaned joint names, facing pairs, body-plan categories and the per-clip quality filters can all change between revisions β€” and with them the clip and skeleton counts above. Pin a revision if you need a snapshot that does not move under you:

hf download Linzhan/UniML3D --type dataset --revision <commit-sha> --local-dir dataset

Distribution of clip lengths across the three sources

Distribution of joint counts across the three sources

Clip lengths (left: share of clips per log-spaced bin; right: share of clips at least N frames long, with the 60 / 90 / 120 / 180-frame training windows marked) and joint counts per skeleton (Mixamo is a single 65-joint rig, hence the step). Both figures are regenerated from clip_frames.json / joint_count.json by data_process/tools/vis_clip_frames.py and vis_joint_count.py; the per-source versions sit next to each JSON as export/<dataset>/*_distribution.png.

Getting the data

hf download Linzhan/UniML3D --type dataset --local-dir dataset                          # everything, β‰ˆ 51 GB
hf download Linzhan/UniML3D --type dataset --local-dir dataset --include "export/truebones/*"   # one source
hf download Linzhan/UniML3D --type dataset --local-dir dataset --include "export/*/*.json" \
    --include "export/*/*.txt" --include "export/*/*.png"                                  # annotations and figures only

export/ is the processed dataset: one NPZ per clip in a uniform format, plus every caption and annotation. It is exactly what data_process/feature_extraction consumes to build the canonicalized training tensors, so the training layer is derived locally rather than shipped here.

Browse it first

Two tables at the repository root carry the whole index, so the dataset viewer above and the datasets loader both work without fetching a single motion file:

tablerowsone row is
clips.csv13,769a clip: dataset, clip, object_type, frames, fps, duration_sec, num_joints, category, caption, video, tpose
skeletons.csv7,430a skeleton: dataset, object_type, num_joints, num_clips, total_frames, duration_sec, category, face_right_joint, face_left_joint, face_source, body_axis, tpose

video and tpose hold repo-relative paths, so any row resolves to a real file:

from datasets import load_dataset
from huggingface_hub import hf_hub_download

clips = load_dataset("Linzhan/UniML3D", "clips", split="train")
skels = load_dataset("Linzhan/UniML3D", "skeletons", split="train")

birds = skels.filter(lambda r: r["category"] == "avian")            # 143 skeletons
long_ = clips.filter(lambda r: r["frames"] >= 60)                   # usable at a 60-frame window
mp4 = hf_hub_download("Linzhan/UniML3D", clips[0]["video"], repo_type="dataset")

Layout

One directory per source, all of it the processed export layer:

The repository root mirrors the pipeline's dataset/ directory, so a download lands directly in the layout UniMate reads:

export/<dataset>/                     stage 1-3.  <dataset> ∈ {truebones, mixamo, objaverse}
  motions/<clip>.npz                  one export NPZ per clip, in the source rig
  videos/<clip>.mp4                   skeleton preview of every clip
                                      (objaverse: both are hash-sharded, see below)
  videos_captioned/<clip>.mp4         skeleton preview with the clip's caption burned in
                                      (objaverse: only clips outside filtered_clips.txt / filtered_objects.txt)
  tpose/<object_type>.png             rest-pose skeleton render per object type
  summary.json                        clip / frame totals and fps
  clip_frames.json                    {clip: n_frames}
  joint_count.json                    {object_type: n_joints}
  joint_names.json                    {object_type: [raw joint names]}            (stage 1)
  clean_joint_names.json              {object_type: [canonical anatomical names]} (stage 3)
  face_joint_names.json               {object_type: facing-direction joint pair}  (stage 3)
  motion_captions.json                {clip: caption}                             (stage 2)
  category_groups.json                {body-plan category: [object types]}        (stage 2)
  filtered_clips.txt                  clips stage 4 skips
  clip_trims.txt                      clips whose first N frames stage 4 drops
  activity_keep.txt                   (objaverse) small real motions stage 4 keeps
  filtered_objects.txt                (objaverse) whole rigs stage 4 skips
  rig_flags.json                      (objaverse) every flagged rig, with category and reason
  clip_frames_distribution.png        frames-per-clip distribution figure
  joint_count_distribution.png        joints-per-skeleton distribution figure

assets/                               figures used by this card

Sizes: export/motions β‰ˆ 48 GB (46 GB of it Objaverse-XL), export/videos β‰ˆ 2.5 GB, export/videos_captioned β‰ˆ 4.2 GB, export/tpose β‰ˆ 0.7 GB, sidecars β‰ˆ 26 MB.

Clip naming. Clips are named <object_type>-<action>, and the object type never contains -, so clip.split('-', 1)[0] recovers it. Mixamo is a single shared rig, so its clips carry no object-type prefix and its object type is simply mixamo.

Hash-sharded folders. The Hub allows at most 10,000 files per directory, and Objaverse-XL has 10,355 clips, so export/objaverse/motions/ and export/objaverse/videos/ are stored as 64 buckets β€” <sub>/<xx>/<file>, where xx is sha1(filename) mod 64 in hex. Every other folder is flat. Either read both levels,

glob.glob(f'{export_dir}/motions/*.npz') + glob.glob(f'{export_dir}/motions/*/*.npz')

or flatten the buckets once after downloading (hardlinks, so it costs no disk):

for d in dataset/export/objaverse/{motions,videos}; do find "$d" -mindepth 2 -type f -exec ln -f -t "$d" -- {} +; done

Data formats

Export NPZ β€” export/<dataset>/motions/<clip>.npz

One file per source clip, float64, written by stage 1 from the source FBX/GLB with Blender: control and helper bones that carry no skin weight are pruned (jointly across all clips of a skeleton, so every clip of an object type shares one topology), the rig is converted to Y-up, frames identical to the rest pose are dropped, and clips shorter than 5 frames are skipped.

fieldshapemeaning
rest_local_pos / rest_local_rot(J, 3) / (J, 4)rest-pose local translations / rotations (quaternions, w-first)
anim_local_pos / anim_local_rot(T, J, 3) / (T, J, 4)per-frame local translations / rotations
offsets(J, 3)bind-pose bone offsets
parents(J,)parent indices, -1 for the root
names(J,)bone name strings
skin_matrix(V, J)skinning weights (empty (0, J) for Mixamo, whose source clips carry no mesh)
fps, action_namescalarframe rate (30) and source action

Captions β€” motion_captions.json

One sentence per clip, object-relative directions, no appearance. Every caption is produced by this dataset's own pipeline: a vision-language model (Qwen3.5-9B) reads a four-view render of the clip, and the source asset's own action name is used only as a disambiguating hint where the asset carries one.

Body-plan categories β€” category_groups.json

Every skeleton is assigned one of bipedal, quadrupedal, avian, insectoid, marine, serpentine, articulated_rigid (plus uncertain) by a vision-language model (Qwen3.5-9B, data_process/vlm_caption/classify_category.py) looking at its T-pose grid, the first frame of a clip, its joint-label histogram and its captions. Classifier mistakes found by hand were corrected in the shipped category_groups.json.

All 7,430 skeletons carry a category. The classifier leaves a rig uncertain when its votes disagree (158 Objaverse-XL rigs, none in Truebones or Mixamo); those still have full motion data and annotations β€” the category only drives optional category-balanced sampling.

Joint annotations

clean_joint_names.json maps every raw bone name to a canonical anatomical vocabulary (mixamorig:LeftUpLeg β†’ Left Thigh, R_momo β†’ Right Thigh). face_joint_names.json names, per skeleton, the bilaterally symmetric pair whose left-right vector defines the facing direction:

"Alligator": {"r_hip": {"raw": "R_momo", "clean": "Right Thigh"},
              "l_hip": {"raw": "L_momo", "clean": "Left Thigh"}, "source": "thigh"}

Serpentine rigs with no bilateral symmetry use a head/tail body-axis pair instead and carry "body_axis": true; a skeleton with no usable pair has "source": "empty" and is canonicalized with identity facing.

Skip lists

Optional lists sit next to the data and are read automatically by stage 4; delete one and that step is skipped.

  • filtered_clips.txt β€” individual clips to skip, each with its reason. Objaverse-XL has 453:

    • clips whose motion is discontinuous inside the frames stage 4 uses (several actions stitched together, a loop seam, a root teleport, garbage frames);
    • held poses;
    • exact duplicates of another clip on the same rig;
    • roots that fly or glide in a way the body motion does not explain;
    • visible per-joint jitter;
    • renders where the rig drives only part of the mesh;
    • clips excluded on hand review, including the reference pipeline's own hand deletions.

    Truebones has 5: four whose export disagrees with the object's reference (a skeleton with missing joints, a mirrored rig, a mis-posed limb), and one whose root moves implausibly fast.

    Mixamo has 7: bodies held in the air with nothing under them (a fall into an absent pool, a carried person without the carrier) and a root that glides with no steps.

  • clip_trims.txt (objaverse) β€” 54 clips kept after dropping a few opening frames. Each line is <clip> <N>: the first N frames of motions/<clip>.npz are a bind pose or an unrelated pose that snaps into the motion, and stage 4 drops them when it loads the clip. If you read the NPZs yourself, use anim_local_rot[N:] and anim_local_pos[N:] for these clips.

  • activity_keep.txt (objaverse) β€” 22 clips exempt from stage 4's low-activity filter: real motions confined to a few joints (a head turn, a nod, a clap, a wave) that the filter would otherwise read as a held pose.

  • filtered_objects.txt (objaverse) β€” 899 whole rigs:

    • 686 whose exported rest pose is unusable as the canonical T-pose: it lies flat, is rotated, reclined or upside-down, or the skeleton does not match the mesh;
    • 213 whose source asset is held out of training.
  • rig_flags.json (objaverse) records the full picture behind those lists: every flagged rig with a category (tpose_wrong, not_in_legacy_raw, empty_pair for a skeleton with no bilateral joint pair, bone_pair, facing_wrong) and the reason, plus the rigs whose automatic flag was cleared by hand (verified_ok). Only tpose_wrong and not_in_legacy_raw reach filtered_objects.txt; the other categories are informational and those rigs are kept.

Stage 4 additionally drops clips that fail its own quality checks (too short, too little motion). The table is the training set of the UniMate paper: the three filters took the 13,769 exported clips down to 11,850. The lists above have been extended since, so rerunning stage 4 on this export yields fewer Objaverse-XL clips.

exportedβˆ’ skip listsβˆ’ stage-4 quality= training clipsskeletonsframesduration
Truebones1,0971,0941,0231,0237495,49153.1 min
Mixamo2,3172,3172,1692,1691189,973105.5 min
Objaverse-XL10,3559,0038,6588,6586,479784,571435.9 min
total13,76912,41411,85011,8506,5541,070,0359.9 h

Previews

export/<ds>/videos/<clip>.mp4 is a matplotlib stick-figure render of the exported skeleton motion in its source rig (not the textured asset). tpose/<object_type>.png shows the pruned rest pose. videos_captioned/<clip>.mp4 renders the same export clip on a checkerboard floor, with its motion_captions.json caption as the title, so a caption can be checked against the motion it describes. For Objaverse-XL it covers only the clips that the skip lists keep.

Reproducing it

Every layer here is built by the open data pipeline at UniMate/data_process, which also documents each stage. Derive the training clips from this repository:

hf download Linzhan/UniML3D --type dataset --local-dir dataset --include "export/*"
bash data_process/scripts/run_extract_features.sh objaverse

Or rebuild the whole thing from the raw assets, which live in their own mirrors on the Hub β€” Mixamo-Animations-Characters, Objaverse-XL-Rigged-Animated with its renders, and Truebones-ZOO-Annotations:

bash data_process/scripts/run_download.sh objaverse          # raw assets from the Hub
bash data_process/scripts/run_export.sh objaverse            # stage 1: assets  -> NPZ + previews
bash data_process/scripts/run_caption_motion.sh objaverse    # stage 2: captions
bash data_process/scripts/run_joints_names_clean_rule.sh objaverse   # stage 3: annotations
bash data_process/scripts/run_extract_features.sh objaverse   # stage 4: training clips

Every stage is deterministic, so the same inputs regenerate the same files.

Licensing

The three sources keep their own terms; this repository does not relicense any of them.

  • Mixamo β€” Adobe's Mixamo terms apply to the motions (export/mixamo/motions) and their previews. Source mirror: Mixamo-Animations-Characters.
  • Objaverse-XL β€” every clip inherits the upstream licence of the Objaverse-XL object it was exported from; there is no blanket licence. The clip stem is the object id, so the licence resolves through the Objaverse-XL annotations. Source mirror and full notice: Objaverse-XL-Rigged-Animated.
  • Truebones β€” Truebones ZOO is a commercial library that may not be redistributed. export/truebones/motions/ is therefore not part of the public release: purchase the pack from Truebones, rebuild the per-clip FBXs with the Truebones-ZOO-Annotations build scripts, and re-run the data pipeline β€” it is deterministic and regenerates the NPZs byte-for-byte. The sidecars, previews and annotations for Truebones are included, and like every other annotation here they were generated by that pipeline rather than taken from a third-party annotation set.

Everything this repository adds on top of those sources β€” the captions for all three, the cleaned joint labels, the facing pairs, the body-plan categories, the skip lists, the index tables and the figures β€” was produced by the data pipeline in this project and is offered under ODC-BY 1.0, matching the derived-metadata licence of the source mirrors.

Citation

@article{mou2026unimate,
  title   = {UniMate: One Unified Model to Animate Diverse Skeletons},
  author  = {Mou, Linzhan and Lei, Jiahui and Dou, Zhiyang and Cai, Chenyue and Song, Chaoyue and Finkelstein, Adam and Rusinkiewicz, Szymon},
  journal = {arXiv preprint arXiv:2609.05415},
  year    = {2026}
}
3d
animation
heterogeneous-skeletons
mixamo
mocap
motion
objaverse
skeletal-animation
text-to-motion
truebones