UniMate
One Unified Model to Animate Diverse Skeletons (SIGGRAPH Asia 2026)
Project Page · Paper · Video · Code · Dataset · Interactive Demo
Pretrained checkpoints of UniMate, a text-conditioned flow-matching model that generates motion for skeletons of any topology: animals, humanoids and rigged objects.
Models
| Model | Architecture | Training data | Joints | Steps | Params¹ | Status |
|---|---|---|---|---|---|---|
unimate_uniml3d_f60_v3 |
graph attention, AdaLN text | UniML3D d0f19d9 |
5–70 | 150k | 74.1M | Recommended |
unimate_uniml3d_f60_v2 |
graph attention, AdaLN text | UniML3D faaa817 |
5–70 | 100k | 74.1M | Previous version |
unimate_uniml3d_f60_v2_cross_attn |
full attention, cross-attention text | UniML3D faaa817 |
5–70 | 100k | 66.2M | Variant |
unimate_mixamo_f60_v2 |
graph attention, AdaLN text | UniML3D faaa817, Mixamo |
22 (one rig) | 120k | 47.8M | Recommended for Mixamo |
unimate_truebones_f60_v2 |
graph attention, AdaLN text | UniML3D faaa817, Truebones |
5–85 | 80k | 47.8M | Recommended for Truebones |
unimate_uniml3d_f60_v1_preview |
graph attention, AdaLN text | UniML3D pre-release build | 5–60 | 120k | 74.1M | Superseded |
¹ Denoiser only; the frozen google/flan-t5-base text encoder is downloaded on first use.
The Mixamo model animates the 22-joint Mixamo rig only. The Truebones model covers 73 of the 74 species (all but Dragon), including Bear, Centipede and Monkey, which exceed the UniML3D models' limit. Use each model's latest checkpoint.
Model details
| Architecture | Transformer denoiser, width 512, 8 heads, 10 layers (6 for the Mixamo and Truebones models); flow matching, linear path, velocity prediction |
| Inputs | English motion description; the skeleton's T-pose, hierarchy and cleaned joint labels |
| Output | 60 frames at 30 fps; per joint: position (3), rotation relative to the T-pose (6), velocity (3) |
| Text | Frozen google/flan-t5-base; classifier-free guidance (10% caption dropout, default scale 3.0) |
Quick start
From the root of the code repository, in its unimate environment:
# 1. download a model (for another model or checkpoint, change the folder and the step)
hf download Linzhan/UniMate --repo-type model --local-dir outputs \
--include "unimate_uniml3d_f60_v3/*.json" --include "unimate_uniml3d_f60_v3/*.npy" \
--include "unimate_uniml3d_f60_v3/checkpoints/checkpoint_step_150000.pt"
# 2. generate for an example asset shipped with the code (assets/examples/ also has an eagle and a shark)
python -m unimate.inference.sample --exp_dir outputs/unimate_uniml3d_f60_v3 --asset assets/examples/unitree_go2 \
--prompt "An object trots forward." "An object rears up on its hind legs." --num_repetitions 3 \
--output_dir outputs/samples/custom
# 3. drive each asset's mesh with its motions (one GLB / FBX per motion)
bash scripts/run_animate_motion.sh outputs/samples/custom
For your own rigged GLB / glTF / FBX, python -m data_process.rig_preprocess run --input <asset> --output_dir outputs/rig/<name> labels its joints, stops so you can check them, and builds an asset directory that --asset takes. Skeletons of the UniML3D dataset are named <dataset>:<object_type> (e.g. mixamo, objaverse:<uid>) once it is downloaded into dataset/. The code README covers both, along with prompting, in-betweening, joint editing and motion expansion.
Training
AdamW (lr 1e-4, betas 0.9 / 0.99, weight decay 1e-5, clipping 1.0), 3% linear warmup then cosine decay to 5%, EMA 0.9999 (used at inference). Loss: masked L2 flow matching + 0.5 × geodesic rotation + 0.1 × velocity smoothness. 60-frame windows at 30 fps, batch 16 per GPU, with joint addition / removal, pooling and perturbation augmentations (Mixamo model: batch 32, no augmentation).
| Model | GPUs | Time |
|---|---|---|
unimate_uniml3d_f60_v3 |
8× H100 | ≈ 35 h |
unimate_uniml3d_f60_v2 |
8× H100 | ≈ 22 h |
unimate_uniml3d_f60_v2_cross_attn |
8× H100 | ≈ 36 h |
unimate_mixamo_f60_v2 |
4× H100 | ≈ 12 h |
unimate_truebones_f60_v2 |
4× H100 | ≈ 15 h |
unimate_uniml3d_f60_v1_preview |
6× H100 | ≈ 23 h |
Each model's config.json holds the model and data configuration that inference reads. The code repository's configs/uniml3d_60frames_graph_adaln.json follows the v3 recipe (at 120k steps; v3 trained 150k), and configs/{mixamo,truebones}_60frames_graph_adaln.json those of the Mixamo and Truebones models:
accelerate launch --num_processes 8 -m unimate.training.train --config configs/uniml3d_60frames_graph_adaln.json
Repository layout
config.json index of the models
<model>/
config.json model and data configuration, read by inference
dataset_stats.npy normalization statistics
checkpoints/ checkpoint_step_<N>.pt every 10k steps (model, EMA, optimizer, scheduler)
logs/ TensorBoard event files
samples/step_<NNNNNN>/ skeleton renders of motions sampled during training
unimate_uniml3d_f60_v3 holds the checkpoints from 100k to 150k steps (150k recommended), without logs or samples.
Versions
- v3 (2026-10-05), over v2: newer data (UniML3D
d0f19d9: re-reviewed annotations; 480 Objaverse-XL rigs authored lying down or upside down stood upright instead of filtered); three captions per clip (generic 35%, detailed 25%, normal 40%); fixed per-dataset sampling weights; one normalization pool for all datasets; logit-normal flow time; 150k steps. - v2 (2026-09-28), over v1_preview: UniML3D
faaa817; joint limit 70 instead of 60 (98.1% of the clips instead of 86.9%); datasets balanced before object types (Mixamo from under 1% to about 24% of the samples); 100k steps. Includes the cross-attention variant and the Mixamo and Truebones models. - v1_preview (2026-09-27): first release, on a pre-release build of UniML3D.
Limitations
- Contact. No unified contact model: contact-rich motions can slide, drift, hover or penetrate the ground; foot locking or IK post-processing helps where contacts are well defined.
- Rare topologies and motions. The data is long-tailed toward humanoids and common locomotion; rare skeletons and unusual motions can come out static, jittery or off-prompt.
- Format. Fixed 60-frame (2 s) samples (longer motion via expansion); a skeleton must fit the model's joint limit; prompts should describe the motion only, as every training caption starts with "An object".
See the paper (Section 6, Appendix F) for details.
License
Checkpoints: CC BY-NC 4.0. Code: MIT. This covers only our rights in the weights; the training data keep their source terms:
| Model | Training data |
|---|---|
unimate_uniml3d_f60_v3, _v2, _v2_cross_attn, _v1_preview |
Truebones ZOO, Mixamo, Objaverse-XL |
unimate_mixamo_f60_v2 |
Mixamo |
unimate_truebones_f60_v2 |
Truebones ZOO |
- Truebones ZOO: a commercial asset pack by Truebones. Third parties allege that some of its animal assets come from commercial games (among them Resident Evil, Skyrim, Bless Online and Conan Exiles). We have not verified these claims; until they are resolved, treat the provenance of the Truebones-trained models as unclear.
- Mixamo: Adobe's Mixamo terms of use.
- Objaverse-XL: the license of each source object; some are non-commercial.
We make no warranty that the checkpoints, or motions generated with them, are free of third-party rights. Rights holders with a concern can open a discussion on this repository or contact the authors.
Citation
@article{mou2026unimate,
title = {UniMate: One Unified Model to Animate Diverse Skeletons},
author = {Mou, Linzhan and Lei, Jiahui and Dou, Zhiyang and Cai, Chenyue and Song, Chaoyue and Finkelstein, Adam and Rusinkiewicz, Szymon},
journal = {arXiv preprint arXiv:2609.05415},
year = {2026}
}
- Downloads last month
- 340
