cy0307/a-mdm-text-to-motion

Model

0

stars

3

commits

2

linked in READMEs

Jun 28, 2026

updated

advanced
embodied-ai
gpu
ropedia-academy
todo
track-ag
track-lm

README

MDM — text-to-motion 🚧 not trained yet

Generate 3D human motion from a text prompt with the Motion Diffusion Model.

Status — documented recipe (placeholder). A production-grade pipeline from Ropedia Academy for an advanced, GPU-heavy task. Everything below — base model, objective, dataset, config, the exact evaluation — is specified; the weights / metrics / figures land here automatically when you run the notebook on a GPU (one click below). Try the trained models live in the Ropedia demos Space.

At a glance

Base modelPretrained MDM checkpoint (or train from scratch)
Taskhuman-motion generation from text
Training objectiveDDPM denoising of motion sequences conditioned on text (CLIP) embeddings.
TrackA · Human modeling
Built onGuyTevet/motion-diffusion-model
NotebookOpen In Colab
Compute / storage / timeGPU required — see the Compute · storage · time table in the notebook

Dataset

  • Source: HumanML3D — text↔motion pairs (~14k motions).

Training config

GPU-scale — the notebook ships a demo profile (free Colab T4) and a full profile, with an exact Compute · storage · time table. Hyperparameters (optimizer, steps, batch, LoRA rank, …) are in the training cell.

Evaluation results

Pending — run the notebook on a GPU to fill this in. This lab reports FID · R-precision@1/2/3 · Diversity on a held-out split (see its Evaluate cell).

Inference example

No weights are published yet. After a GPU run, load the checkpoint/adapter the notebook saves (it also has a ready inference cell). Base model: Pretrained MDM checkpoint (or train from scratch).

How to fill this repo

  1. Open the notebook in ColabRuntime → GPU → Run all (runs the real pipeline).
  2. Run its Publish to the Hugging Face Hub step (or HfApi().upload_folder(...)) — the checkpoint + metrics.json + figures replace this placeholder.
  • Train / run on a GPU · [ ] upload weights · [ ] add metrics.json · [ ] add figures · [ ] swap in the real results card

Limitations

Not yet trained — no numbers to report. The pipeline is GPU-heavy (see the compute table); on free Colab use the demo-scale settings. This is an educational, reproducible recipe, not a tuned production release.

License

Code: MIT (this repository). The base model (GuyTevet/motion-diffusion-model) and dataset are each under their own licenses — check the upstream source before redistribution.

Citation

@misc{ropedia_academy,
  title  = {Ropedia Academy: an interactive course on embodied & spatial AI},
  author = {Ropedia Academy},
  year   = {2026},
  howpublished = {\url{https://chaoyue0307.github.io/ropedia-academy/}}
}

Method / original work: Tevet et al., MDM, ICLR 2023 (arXiv:2209.14916).


Documented placeholder in the Ropedia Academy collection — train it on a GPU to publish the real model. Contributions welcome on GitHub.

Contributors

cy0307

3 commits

cy0307/a-mdm-text-to-motion

Model

0

stars

3

commits

2

linked in READMEs

Jun 28, 2026

updated

advanced
embodied-ai
gpu
ropedia-academy
todo
track-ag
track-lm

README

MDM — text-to-motion 🚧 not trained yet

Generate 3D human motion from a text prompt with the Motion Diffusion Model.

Status — documented recipe (placeholder). A production-grade pipeline from Ropedia Academy for an advanced, GPU-heavy task. Everything below — base model, objective, dataset, config, the exact evaluation — is specified; the weights / metrics / figures land here automatically when you run the notebook on a GPU (one click below). Try the trained models live in the Ropedia demos Space.

At a glance

Base modelPretrained MDM checkpoint (or train from scratch)
Taskhuman-motion generation from text
Training objectiveDDPM denoising of motion sequences conditioned on text (CLIP) embeddings.
TrackA · Human modeling
Built onGuyTevet/motion-diffusion-model
NotebookOpen In Colab
Compute / storage / timeGPU required — see the Compute · storage · time table in the notebook

Dataset

  • Source: HumanML3D — text↔motion pairs (~14k motions).

Training config

GPU-scale — the notebook ships a demo profile (free Colab T4) and a full profile, with an exact Compute · storage · time table. Hyperparameters (optimizer, steps, batch, LoRA rank, …) are in the training cell.

Evaluation results

Pending — run the notebook on a GPU to fill this in. This lab reports FID · R-precision@1/2/3 · Diversity on a held-out split (see its Evaluate cell).

Inference example

No weights are published yet. After a GPU run, load the checkpoint/adapter the notebook saves (it also has a ready inference cell). Base model: Pretrained MDM checkpoint (or train from scratch).

How to fill this repo

  1. Open the notebook in ColabRuntime → GPU → Run all (runs the real pipeline).
  2. Run its Publish to the Hugging Face Hub step (or HfApi().upload_folder(...)) — the checkpoint + metrics.json + figures replace this placeholder.
  • Train / run on a GPU · [ ] upload weights · [ ] add metrics.json · [ ] add figures · [ ] swap in the real results card

Limitations

Not yet trained — no numbers to report. The pipeline is GPU-heavy (see the compute table); on free Colab use the demo-scale settings. This is an educational, reproducible recipe, not a tuned production release.

License

Code: MIT (this repository). The base model (GuyTevet/motion-diffusion-model) and dataset are each under their own licenses — check the upstream source before redistribution.

Citation

@misc{ropedia_academy,
  title  = {Ropedia Academy: an interactive course on embodied & spatial AI},
  author = {Ropedia Academy},
  year   = {2026},
  howpublished = {\url{https://chaoyue0307.github.io/ropedia-academy/}}
}

Method / original work: Tevet et al., MDM, ICLR 2023 (arXiv:2209.14916).


Documented placeholder in the Ropedia Academy collection — train it on a GPU to publish the real model. Contributions welcome on GitHub.

Contributors

cy0307

3 commits