0
stars
5
commits
1
linked in READMEs
Jun 28, 2026
updated
A policy trained by supervised imitation of expert demonstrations; 100% rollout success.
Trained from scratch in Ropedia Academy — an interactive, bilingual course on embodied & spatial AI. Educational model: small and quick to train; the value is the method and a reproducible pipeline, not a leaderboard score. Try it live in the Ropedia demos Space.
Adam (lr 3e-3), 800 steps; cross-entropy on ~2k expert (state→action) pairs; 6×6 grid.
| metric | value | meaning |
|---|---|---|
imitation_acc (final) | 1.0 | |
rollout_success | 1.0 | fraction of start states from which the policy reaches the goal |

import torch
state = torch.load("policy.pt", map_location="cpu") # this repo's checkpoint
# Rebuild the exact module from the lab notebook (see "Reproduce"), then:
# model.load_state_dict(state); model.eval()
Educational scale. Trained quickly on CPU on small or synthetic data, so absolute numbers are not competitive with production systems — the value is the method and a reproducible pipeline. No large-scale data, no hyperparameter sweep, and no multi-seed variance is reported. Not for production use.
Compounding error / distribution shift once it leaves the expert's states (no recovery) — needs DAgger to fix.
One click: open the notebook in Colab → Runtime → GPU → Run all, then run its Publish to the Hugging Face Hub cell.
From a shell:
git clone https://github.com/ChaoYue0307/ropedia-academy.git && cd ropedia-academy
pip install torch numpy matplotlib scikit-learn scikit-image gymnasium
jupyter nbconvert --to notebook --execute notebooks/training/AG_behavior_cloning.ipynb --output run.ipynb
# optional: override training length, e.g. STEPS=2000 (or EPISODES=600) before running
figure.pngmetrics.jsonpolicy.ptCode & weights: MIT (this repository) — educational use encouraged.
Data: generated procedurally in the notebook — no external dataset.
If you use this model or the course materials, please cite:
@misc{ropedia_academy,
title = {Ropedia Academy: an interactive course on embodied & spatial AI},
author = {Ropedia Academy},
year = {2026},
howpublished = {\url{https://chaoyue0307.github.io/ropedia-academy/}}
}
Method / original work: Pomerleau, ALVINN, 1988; Ross et al., DAgger, AISTATS 2011.
Part of the Ropedia Academy trained-model collection. Contributions & issues welcome on GitHub.
5 commits
0
stars
5
commits
1
linked in READMEs
Jun 28, 2026
updated
A policy trained by supervised imitation of expert demonstrations; 100% rollout success.
Trained from scratch in Ropedia Academy — an interactive, bilingual course on embodied & spatial AI. Educational model: small and quick to train; the value is the method and a reproducible pipeline, not a leaderboard score. Try it live in the Ropedia demos Space.
Adam (lr 3e-3), 800 steps; cross-entropy on ~2k expert (state→action) pairs; 6×6 grid.
| metric | value | meaning |
|---|---|---|
imitation_acc (final) | 1.0 | |
rollout_success | 1.0 | fraction of start states from which the policy reaches the goal |

import torch
state = torch.load("policy.pt", map_location="cpu") # this repo's checkpoint
# Rebuild the exact module from the lab notebook (see "Reproduce"), then:
# model.load_state_dict(state); model.eval()
Educational scale. Trained quickly on CPU on small or synthetic data, so absolute numbers are not competitive with production systems — the value is the method and a reproducible pipeline. No large-scale data, no hyperparameter sweep, and no multi-seed variance is reported. Not for production use.
Compounding error / distribution shift once it leaves the expert's states (no recovery) — needs DAgger to fix.
One click: open the notebook in Colab → Runtime → GPU → Run all, then run its Publish to the Hugging Face Hub cell.
From a shell:
git clone https://github.com/ChaoYue0307/ropedia-academy.git && cd ropedia-academy
pip install torch numpy matplotlib scikit-learn scikit-image gymnasium
jupyter nbconvert --to notebook --execute notebooks/training/AG_behavior_cloning.ipynb --output run.ipynb
# optional: override training length, e.g. STEPS=2000 (or EPISODES=600) before running
figure.pngmetrics.jsonpolicy.ptCode & weights: MIT (this repository) — educational use encouraged.
Data: generated procedurally in the notebook — no external dataset.
If you use this model or the course materials, please cite:
@misc{ropedia_academy,
title = {Ropedia Academy: an interactive course on embodied & spatial AI},
author = {Ropedia Academy},
year = {2026},
howpublished = {\url{https://chaoyue0307.github.io/ropedia-academy/}}
}
Method / original work: Pomerleau, ALVINN, 1988; Ross et al., DAgger, AISTATS 2011.
Part of the Ropedia Academy trained-model collection. Contributions & issues welcome on GitHub.
5 commits