This repo is an experiment I am doing of how much research / implementation can be done with AI agents.
This is part of: https://huggingface.co/spaces/ICML-2026-agent-repro/challenge
A good motivation to reproduce a lot of research papers which get published and also validate them. Unfortunately, this is heavily skewed towards those who have lot of capital in terms of tokens they can use and compute for such reproductions.
But such community efforts are super helpful for future researchers and community in large to know which papers have validated claims. Will bring ML conferences closer to those like Usenix with these reproductions.
Also, the models and harnesses have become really excellent. Being able to reproduce papers in parallel and accurately with an judge is excellent and really benefits research - fastracks one to do quick research as tasks like validating previous work and rewriting the code can now done very easily with agents.
34 papers reproduced so far. Each one has a live Trackio logbook on the Hub walking through the claims, the runs behind them, and where the reproduction agreed or diverged from the paper. Code links to the reproduction folder in this repo.
All logbooks live under huggingface.co/vimarsh and are tagged icml2026-repro.
18 commits
HTML
90.6%
Jupyter Notebook
4.0%
Python
3.8%
This repo is an experiment I am doing of how much research / implementation can be done with AI agents.
This is part of: https://huggingface.co/spaces/ICML-2026-agent-repro/challenge
A good motivation to reproduce a lot of research papers which get published and also validate them. Unfortunately, this is heavily skewed towards those who have lot of capital in terms of tokens they can use and compute for such reproductions.
But such community efforts are super helpful for future researchers and community in large to know which papers have validated claims. Will bring ML conferences closer to those like Usenix with these reproductions.
Also, the models and harnesses have become really excellent. Being able to reproduce papers in parallel and accurately with an judge is excellent and really benefits research - fastracks one to do quick research as tasks like validating previous work and rewriting the code can now done very easily with agents.
34 papers reproduced so far. Each one has a live Trackio logbook on the Hub walking through the claims, the runs behind them, and where the reproduction agreed or diverged from the paper. Code links to the reproduction folder in this repo.
All logbooks live under huggingface.co/vimarsh and are tagged icml2026-repro.
18 commits
HTML
90.6%
Jupyter Notebook
4.0%
Python
3.8%