Linzhan/UniMate

Model

UniMate

20

11 commits

1 linked in READMEs

updated Sep 30, 2026

See the code

README

UniMate

One Unified Model to Animate Diverse Skeletons (SIGGRAPH Asia 2026)

Project Page · Paper · Video · Code · Dataset · Interactive Demo

UniMate teaser

Pretrained checkpoints for UniMate, a text-conditioned flow-matching model that generates motion for skeletons of arbitrary topology: animals, humanoids and rigged objects.

Models

ModelArchitectureTraining dataJointsStepsParams¹Status
unimate_uniml3d_f60_v2graph attention, AdaLN textUniML3D faaa817 (2026-09-27)5–70100k74.1MRecommended
unimate_uniml3d_f60_v2_full_cross_attnfull attention, cross-attention textUniML3D faaa817 (2026-09-27)5–70100k66.2MVariant
unimate_uniml3d_f60_previewgraph attention, AdaLN textUniML3D, pre-release build5–60120k74.1MSuperseded
unimate_mixamo_f60graph attention, AdaLN textUniML3D faaa817 (2026-09-27), Mixamo only22 (one rig)120k47.8MRecommended for Mixamo humanoids

¹ Denoiser only. The frozen google/flan-t5-base text encoder is downloaded from the Hub on first use.

Use each model's final checkpoint, checkpoints/checkpoint_step_<steps>.pt.

v2 versus preview. Raising the joint limit from 60 to 70 keeps 98.1% of the training clips instead of 86.9%, and 70 of the 74 Truebones species. Datasets are balanced before object types (sampler_dataset_alpha = 0.25), so Truebones, Mixamo and Objaverse-XL make up about 13%, 24% and 63% of the training samples; the single Mixamo rig previously got under 1%.

Mixamo humanoids. To test humanoid animation on the Mixamo rig, use unimate_mixamo_f60. It is trained on the 2,162 Mixamo clips alone, all on the one 22-joint rig, without topology augmentations. It animates that rig only; use a UniML3D model for any other skeleton.

Model details

ArchitectureTransformer denoiser, 10 layers (6 for unimate_mixamo_f60), width 512, 8 heads; flow matching with a linear path and velocity prediction
InputsEnglish motion description; the target skeleton's T-pose, joint hierarchy and joint names
Output60 frames at 30 fps, 12 features per joint: 3-D position, 6-D rotation relative to the T-pose, 3-D velocity. The root joint encodes height, facing and planar velocity instead
Text encodergoogle/flan-t5-base, frozen
GuidanceClassifier-free guidance on the caption: 10% caption dropout in training, default scale 3.0

Quick start

Run every command from the root of the code repository, with its unimate environment installed.

1. Build the dataset features. The sampler reads each target skeleton from dataset/features/<dataset>/. The Hub dataset ships the stage 1-3 export but not these stage-4 features, so build them from the revision the model was trained on:

hf download Linzhan/UniML3D --repo-type dataset \
    --revision faaa81773b315247b03f183548e3898dbec2ce80 --local-dir dataset
bash data_process/scripts/run_extract_features.sh truebones   # likewise mixamo, objaverse

unimate_mixamo_f60 needs only the mixamo features.

The Truebones motions come from a commercial pack and are not on the Hub; the dataset card explains how to rebuild them.

2. Download a model. This fetches the configuration, normalization statistics and final checkpoint of the recommended model:

hf download Linzhan/UniMate \
    --include "unimate_uniml3d_f60_v2/*.json" "unimate_uniml3d_f60_v2/*.npy" \
              "unimate_uniml3d_f60_v2/checkpoints/checkpoint_step_100000.pt" \
    --local-dir outputs

To test Mixamo humanoid animation, fetch unimate_mixamo_f60 instead:

hf download Linzhan/UniMate \
    --include "unimate_mixamo_f60/*.json" "unimate_mixamo_f60/*.npy" \
              "unimate_mixamo_f60/checkpoints/checkpoint_step_120000.pt" \
    --local-dir outputs

3. Generate motion from text. Write the prompts to test_cases.json with keys <object_type>-<case_id>. object_type is a skeleton in dataset/features/: a Truebones species, mixamo, or an Objaverse-XL object ID. case_id is a free-form tag that names the output files. With unimate_mixamo_f60, every key is mixamo-<case_id> and --exp_dir is outputs/unimate_mixamo_f60.

{
  "Horse-0": "An object rears up on its hind legs.",
  "mixamo-0": "An object jumps in place with both arms raised.",
  "144367de23534c28ad2e83fc8abbd9ea-0": "An object flaps its wings."
}
python -m unimate.inference.sample \
    --exp_dir outputs/unimate_uniml3d_f60_v2 \
    --test_cases_json test_cases.json \
    --num_repetitions 3

The sampler loads the EMA weights of the latest checkpoint in checkpoints/. It writes the motions (.npy, shape (60, J, 12) for a skeleton with J joints) and skeleton renders (.mp4) to outputs/unimate_uniml3d_f60_v2/samples/. In-betweening, joint-level editing, motion expansion and driving a rigged mesh are documented in the code README.

Training

OptimizerAdamW, learning rate 1e-4, betas (0.9, 0.99), weight decay 1e-5, gradient clipping at 1.0
ScheduleLinear warmup over the first 3% of steps, then cosine decay to 5% of the peak rate
LossMasked L2 flow-matching loss + 0.5 × geodesic rotation loss + 0.1 × velocity smoothness loss
EMADecay 0.9999; used at inference
Batch size16 per GPU (32 for unimate_mixamo_f60)
Data60-frame windows at 30 fps; joint addition, joint removal, pooling and perturbation augmentations (off for unimate_mixamo_f60)
ModelConfig (code repository)GPUsTraining time
unimate_uniml3d_f60_v2configs/uniml3d_60frames_graph_adaln_v2.json8× H100about 22 hours
unimate_uniml3d_f60_v2_full_cross_attnconfigs/uniml3d_60frames_full_cross_attn_v2.json8× H100about 36 hours²
unimate_uniml3d_f60_previewconfigs/uniml3d_60frames_graph_adaln.json6× H100about 23 hours
unimate_mixamo_f60configs/mixamo_60frames_graph_adaln.json4× H100about 12 hours

² Resumed once from the step-60k checkpoint.

Resume from a released checkpoint, or train from scratch:

bash scripts/run_train.sh configs/uniml3d_60frames_graph_adaln_v2.json -- \
    --resume outputs/unimate_uniml3d_f60_v2/checkpoints/checkpoint_step_100000.pt

accelerate launch --num_processes 8 -m unimate.training.train \
    --config configs/uniml3d_60frames_graph_adaln_v2.json

Repository layout

config.json                        index of the released models
LICENSE
<model>/
  config.json                      resolved training configuration; read by inference
  dataset_stats.npy                feature normalization statistics
  checkpoints/
    checkpoint_step_<N>.pt         model and EMA weights, optimizer and scheduler state; every 10k steps
  logs/                            TensorBoard curves, one event file per training job
  samples/
    step_<NNNNNN>/                 motions generated during training (step_000000: before training)
      <object_type>-<i>_fk.mp4     joints from forward kinematics of the predicted rotations
      <object_type>-<i>_ric.mp4    predicted joint positions

Limitations

  • Clip length: a sample is a fixed 60-frame window (2 seconds). Longer motion requires chaining samples with the expansion application.
  • Joint count: the v2 models are trained on skeletons with at most 70 joints, which excludes four Truebones species (Bear, Centipede, Monkey, Dragon); the preview model's limit is 60. Skeletons with more than 71 joints (61 for the preview model) cannot be sampled. unimate_mixamo_f60 has seen only the 22-joint Mixamo rig and cannot sample a larger skeleton.
  • New skeletons: there is no skeleton-only input. A new rig must be converted into a feature directory with data_process/scripts/run_preprocess_char.sh, and needs at least one animation clip.
  • Text: every training caption has the subject "An object"; the skeleton determines the character. Prompts should describe the motion only.

License

The checkpoints and the UniMate code are released under the MIT License. The training data remain under their source licenses: Adobe's Mixamo terms of use, the per-object licenses of Objaverse-XL, and the commercial Truebones ZOO license. Review these terms before using the models.

Citation

@article{mou2026unimate,
  title   = {UniMate: One Unified Model to Animate Diverse Skeletons},
  author  = {Mou, Linzhan and Lei, Jiahui and Dou, Zhiyang and Cai, Chenyue and Song, Chaoyue and Finkelstein, Adam and Rusinkiewicz, Szymon},
  journal = {arXiv preprint arXiv:2609.05415},
  year    = {2026}
}
character-animation
flow-matching
motion-generation
pytorch
skeletal-animation
tensorboard
text-to-motion
unimate

Linzhan/UniMate

Model

UniMate

20

11 commits

1 linked in READMEs

updated Sep 30, 2026

See the code

README

UniMate

One Unified Model to Animate Diverse Skeletons (SIGGRAPH Asia 2026)

Project Page · Paper · Video · Code · Dataset · Interactive Demo

UniMate teaser

Pretrained checkpoints for UniMate, a text-conditioned flow-matching model that generates motion for skeletons of arbitrary topology: animals, humanoids and rigged objects.

Models

ModelArchitectureTraining dataJointsStepsParams¹Status
unimate_uniml3d_f60_v2graph attention, AdaLN textUniML3D faaa817 (2026-09-27)5–70100k74.1MRecommended
unimate_uniml3d_f60_v2_full_cross_attnfull attention, cross-attention textUniML3D faaa817 (2026-09-27)5–70100k66.2MVariant
unimate_uniml3d_f60_previewgraph attention, AdaLN textUniML3D, pre-release build5–60120k74.1MSuperseded
unimate_mixamo_f60graph attention, AdaLN textUniML3D faaa817 (2026-09-27), Mixamo only22 (one rig)120k47.8MRecommended for Mixamo humanoids

¹ Denoiser only. The frozen google/flan-t5-base text encoder is downloaded from the Hub on first use.

Use each model's final checkpoint, checkpoints/checkpoint_step_<steps>.pt.

v2 versus preview. Raising the joint limit from 60 to 70 keeps 98.1% of the training clips instead of 86.9%, and 70 of the 74 Truebones species. Datasets are balanced before object types (sampler_dataset_alpha = 0.25), so Truebones, Mixamo and Objaverse-XL make up about 13%, 24% and 63% of the training samples; the single Mixamo rig previously got under 1%.

Mixamo humanoids. To test humanoid animation on the Mixamo rig, use unimate_mixamo_f60. It is trained on the 2,162 Mixamo clips alone, all on the one 22-joint rig, without topology augmentations. It animates that rig only; use a UniML3D model for any other skeleton.

Model details

ArchitectureTransformer denoiser, 10 layers (6 for unimate_mixamo_f60), width 512, 8 heads; flow matching with a linear path and velocity prediction
InputsEnglish motion description; the target skeleton's T-pose, joint hierarchy and joint names
Output60 frames at 30 fps, 12 features per joint: 3-D position, 6-D rotation relative to the T-pose, 3-D velocity. The root joint encodes height, facing and planar velocity instead
Text encodergoogle/flan-t5-base, frozen
GuidanceClassifier-free guidance on the caption: 10% caption dropout in training, default scale 3.0

Quick start

Run every command from the root of the code repository, with its unimate environment installed.

1. Build the dataset features. The sampler reads each target skeleton from dataset/features/<dataset>/. The Hub dataset ships the stage 1-3 export but not these stage-4 features, so build them from the revision the model was trained on:

hf download Linzhan/UniML3D --repo-type dataset \
    --revision faaa81773b315247b03f183548e3898dbec2ce80 --local-dir dataset
bash data_process/scripts/run_extract_features.sh truebones   # likewise mixamo, objaverse

unimate_mixamo_f60 needs only the mixamo features.

The Truebones motions come from a commercial pack and are not on the Hub; the dataset card explains how to rebuild them.

2. Download a model. This fetches the configuration, normalization statistics and final checkpoint of the recommended model:

hf download Linzhan/UniMate \
    --include "unimate_uniml3d_f60_v2/*.json" "unimate_uniml3d_f60_v2/*.npy" \
              "unimate_uniml3d_f60_v2/checkpoints/checkpoint_step_100000.pt" \
    --local-dir outputs

To test Mixamo humanoid animation, fetch unimate_mixamo_f60 instead:

hf download Linzhan/UniMate \
    --include "unimate_mixamo_f60/*.json" "unimate_mixamo_f60/*.npy" \
              "unimate_mixamo_f60/checkpoints/checkpoint_step_120000.pt" \
    --local-dir outputs

3. Generate motion from text. Write the prompts to test_cases.json with keys <object_type>-<case_id>. object_type is a skeleton in dataset/features/: a Truebones species, mixamo, or an Objaverse-XL object ID. case_id is a free-form tag that names the output files. With unimate_mixamo_f60, every key is mixamo-<case_id> and --exp_dir is outputs/unimate_mixamo_f60.

{
  "Horse-0": "An object rears up on its hind legs.",
  "mixamo-0": "An object jumps in place with both arms raised.",
  "144367de23534c28ad2e83fc8abbd9ea-0": "An object flaps its wings."
}
python -m unimate.inference.sample \
    --exp_dir outputs/unimate_uniml3d_f60_v2 \
    --test_cases_json test_cases.json \
    --num_repetitions 3

The sampler loads the EMA weights of the latest checkpoint in checkpoints/. It writes the motions (.npy, shape (60, J, 12) for a skeleton with J joints) and skeleton renders (.mp4) to outputs/unimate_uniml3d_f60_v2/samples/. In-betweening, joint-level editing, motion expansion and driving a rigged mesh are documented in the code README.

Training

OptimizerAdamW, learning rate 1e-4, betas (0.9, 0.99), weight decay 1e-5, gradient clipping at 1.0
ScheduleLinear warmup over the first 3% of steps, then cosine decay to 5% of the peak rate
LossMasked L2 flow-matching loss + 0.5 × geodesic rotation loss + 0.1 × velocity smoothness loss
EMADecay 0.9999; used at inference
Batch size16 per GPU (32 for unimate_mixamo_f60)
Data60-frame windows at 30 fps; joint addition, joint removal, pooling and perturbation augmentations (off for unimate_mixamo_f60)
ModelConfig (code repository)GPUsTraining time
unimate_uniml3d_f60_v2configs/uniml3d_60frames_graph_adaln_v2.json8× H100about 22 hours
unimate_uniml3d_f60_v2_full_cross_attnconfigs/uniml3d_60frames_full_cross_attn_v2.json8× H100about 36 hours²
unimate_uniml3d_f60_previewconfigs/uniml3d_60frames_graph_adaln.json6× H100about 23 hours
unimate_mixamo_f60configs/mixamo_60frames_graph_adaln.json4× H100about 12 hours

² Resumed once from the step-60k checkpoint.

Resume from a released checkpoint, or train from scratch:

bash scripts/run_train.sh configs/uniml3d_60frames_graph_adaln_v2.json -- \
    --resume outputs/unimate_uniml3d_f60_v2/checkpoints/checkpoint_step_100000.pt

accelerate launch --num_processes 8 -m unimate.training.train \
    --config configs/uniml3d_60frames_graph_adaln_v2.json

Repository layout

config.json                        index of the released models
LICENSE
<model>/
  config.json                      resolved training configuration; read by inference
  dataset_stats.npy                feature normalization statistics
  checkpoints/
    checkpoint_step_<N>.pt         model and EMA weights, optimizer and scheduler state; every 10k steps
  logs/                            TensorBoard curves, one event file per training job
  samples/
    step_<NNNNNN>/                 motions generated during training (step_000000: before training)
      <object_type>-<i>_fk.mp4     joints from forward kinematics of the predicted rotations
      <object_type>-<i>_ric.mp4    predicted joint positions

Limitations

  • Clip length: a sample is a fixed 60-frame window (2 seconds). Longer motion requires chaining samples with the expansion application.
  • Joint count: the v2 models are trained on skeletons with at most 70 joints, which excludes four Truebones species (Bear, Centipede, Monkey, Dragon); the preview model's limit is 60. Skeletons with more than 71 joints (61 for the preview model) cannot be sampled. unimate_mixamo_f60 has seen only the 22-joint Mixamo rig and cannot sample a larger skeleton.
  • New skeletons: there is no skeleton-only input. A new rig must be converted into a feature directory with data_process/scripts/run_preprocess_char.sh, and needs at least one animation clip.
  • Text: every training caption has the subject "An object"; the skeleton determines the character. Prompts should describe the motion only.

License

The checkpoints and the UniMate code are released under the MIT License. The training data remain under their source licenses: Adobe's Mixamo terms of use, the per-object licenses of Objaverse-XL, and the commercial Truebones ZOO license. Review these terms before using the models.

Citation

@article{mou2026unimate,
  title   = {UniMate: One Unified Model to Animate Diverse Skeletons},
  author  = {Mou, Linzhan and Lei, Jiahui and Dou, Zhiyang and Cai, Chenyue and Song, Chaoyue and Finkelstein, Adam and Rusinkiewicz, Szymon},
  journal = {arXiv preprint arXiv:2609.05415},
  year    = {2026}
}
character-animation
flow-matching
motion-generation
pytorch
skeletal-animation
tensorboard
text-to-motion
unimate