cy0307/ag-agent-harness

Model

0

stars

5

commits

1

linked in READMEs

Jun 28, 2026

updated

agent
educational
embodied-ai
from-scratch
pytorch
reproducible
ropedia-academy

README

Agent + tool-use harness

A reusable harness: a tool registry, a tool-using agent loop, a task suite, and a scorer.

Trained from scratch in Ropedia Academy — an interactive, bilingual course on embodied & spatial AI. Educational model: small and quick to train; the value is the method and a reproducible pipeline, not a leaderboard score. Try it live in the Ropedia demos Space.

At a glance

Base modelTrained from scratch (random initialization) — no pretrained base model.
Taskagent evaluation
Training objectiveNo training — a tool-use agent loop + scorer (an evaluation harness).
TrackAG · Agents & RL
NotebookOpen In Colab

Dataset

  • Name: Tool-use task suite
  • Type: synthetic — procedural
  • Size / stats: 5 graded tasks (arithmetic + string ops) with ground-truth answers
  • Split: eval suite
  • Source: procedural

Training config

No training — a tool registry + ReAct-style loop + scorer over a 5-task suite.

Evaluation results

metricvaluemeaning
success_rate1.0fraction of agent tasks solved with the correct final answer

figure

Inference example

import torch
state = torch.load("model.pt", map_location="cpu")   # this repo's checkpoint
# Rebuild the exact module from the lab notebook (see "Reproduce"), then:
# model.load_state_dict(state); model.eval()

Limitations

Educational scale. Trained quickly on CPU on small or synthetic data, so absolute numbers are not competitive with production systems — the value is the method and a reproducible pipeline. No large-scale data, no hyperparameter sweep, and no multi-seed variance is reported. Not for production use.

Failure cases

Only as strong as the task suite & scorer; brittle tool-call parsing; doesn't probe long-horizon planning.

Reproduce / train your own

One click: open the notebook in Colab → Runtime → GPU → Run all, then run its Publish to the Hugging Face Hub cell.

Open In Colab

From a shell:

git clone https://github.com/ChaoYue0307/ropedia-academy.git && cd ropedia-academy
pip install torch numpy matplotlib scikit-learn scikit-image gymnasium
jupyter nbconvert --to notebook --execute notebooks/training/AG_agent_harness.ipynb --output run.ipynb
# optional: override training length, e.g.  STEPS=2000  (or EPISODES=600)  before running

Files

  • figure.png
  • results.json

License

Code & weights: MIT (this repository) — educational use encouraged.
Data: generated procedurally in the notebook — no external dataset.

Citation

If you use this model or the course materials, please cite:

@misc{ropedia_academy,
  title  = {Ropedia Academy: an interactive course on embodied & spatial AI},
  author = {Ropedia Academy},
  year   = {2026},
  howpublished = {\url{https://chaoyue0307.github.io/ropedia-academy/}}
}

Method / original work: Yao et al., ReAct, ICLR 2023; Schick et al., Toolformer, 2023.


Part of the Ropedia Academy trained-model collection. Contributions & issues welcome on GitHub.

Contributors

cy0307

5 commits

cy0307/ag-agent-harness

Model

0

stars

5

commits

1

linked in READMEs

Jun 28, 2026

updated

agent
educational
embodied-ai
from-scratch
pytorch
reproducible
ropedia-academy

README

Agent + tool-use harness

A reusable harness: a tool registry, a tool-using agent loop, a task suite, and a scorer.

Trained from scratch in Ropedia Academy — an interactive, bilingual course on embodied & spatial AI. Educational model: small and quick to train; the value is the method and a reproducible pipeline, not a leaderboard score. Try it live in the Ropedia demos Space.

At a glance

Base modelTrained from scratch (random initialization) — no pretrained base model.
Taskagent evaluation
Training objectiveNo training — a tool-use agent loop + scorer (an evaluation harness).
TrackAG · Agents & RL
NotebookOpen In Colab

Dataset

  • Name: Tool-use task suite
  • Type: synthetic — procedural
  • Size / stats: 5 graded tasks (arithmetic + string ops) with ground-truth answers
  • Split: eval suite
  • Source: procedural

Training config

No training — a tool registry + ReAct-style loop + scorer over a 5-task suite.

Evaluation results

metricvaluemeaning
success_rate1.0fraction of agent tasks solved with the correct final answer

figure

Inference example

import torch
state = torch.load("model.pt", map_location="cpu")   # this repo's checkpoint
# Rebuild the exact module from the lab notebook (see "Reproduce"), then:
# model.load_state_dict(state); model.eval()

Limitations

Educational scale. Trained quickly on CPU on small or synthetic data, so absolute numbers are not competitive with production systems — the value is the method and a reproducible pipeline. No large-scale data, no hyperparameter sweep, and no multi-seed variance is reported. Not for production use.

Failure cases

Only as strong as the task suite & scorer; brittle tool-call parsing; doesn't probe long-horizon planning.

Reproduce / train your own

One click: open the notebook in Colab → Runtime → GPU → Run all, then run its Publish to the Hugging Face Hub cell.

Open In Colab

From a shell:

git clone https://github.com/ChaoYue0307/ropedia-academy.git && cd ropedia-academy
pip install torch numpy matplotlib scikit-learn scikit-image gymnasium
jupyter nbconvert --to notebook --execute notebooks/training/AG_agent_harness.ipynb --output run.ipynb
# optional: override training length, e.g.  STEPS=2000  (or EPISODES=600)  before running

Files

  • figure.png
  • results.json

License

Code & weights: MIT (this repository) — educational use encouraged.
Data: generated procedurally in the notebook — no external dataset.

Citation

If you use this model or the course materials, please cite:

@misc{ropedia_academy,
  title  = {Ropedia Academy: an interactive course on embodied & spatial AI},
  author = {Ropedia Academy},
  year   = {2026},
  howpublished = {\url{https://chaoyue0307.github.io/ropedia-academy/}}
}

Method / original work: Yao et al., ReAct, ICLR 2023; Schick et al., Toolformer, 2023.


Part of the Ropedia Academy trained-model collection. Contributions & issues welcome on GitHub.

Contributors

cy0307

5 commits