Research Skills · Agent-led post-training · Executable evaluation
Overview · Paper · Framework · Components · Research Skills · Eval Harness · Model · Security
iCoder is a research project on agent-led model development for RTL design and GPU kernel optimization. This repository publishes the project's evaluation harness and Research Skills bundle. The full technical report is available here, and model artifacts are available as iCoder-27B.
The project targets industrial coding: RTL design and GPU kernel optimization, where acceptance is determined by specialized executable toolchains and deployment quality rather than surface similarity alone.
[!NOTE] iCoder is human-guided and agent-led, not a fully autonomous self-improvement system. Its workflows require user-provided objectives, task instructions, verifiers, resource limits, and an isolated execution environment.
Research Skills guide the agent loop; task-native execution feeds evidence back into experimentation.
The project framework separates instructions supplied before an experiment from evidence produced during execution:
| Element | Function |
|---|---|
| Research Skills | Provide versioned task instructions, workflow constraints, and verifier requirements. |
| Agent loop | Applies those instructions, runs experiments, and reacts to execution feedback. |
| Executable verification | Uses task-native compilers, simulators, testbenches, and numerical oracles to produce structured outcomes. |
| Model release | Makes the resulting iCoder-27B artifacts available through Hugging Face. |
Data construction first creates a shared executable task pool. Parameter updates then proceed through SFT → OPSD → RLVR:
| Stage | Role |
|---|---|
| Data construction | Prepares executable RTL and GPU-kernel tasks and checks them with domain toolchains. |
| SFT (Supervised Fine-Tuning) | Uses verified teacher solutions to establish task capability. |
| OPSD (On-Policy Self-Distillation) | Trains on feedback collected from the model's own execution attempts. |
| RLVR (Reinforcement Learning with Verifiable Rewards) | Uses domain-specific execution outcomes as training signals. |
The stages form a feedback loop: evaluation can send the process back to data construction, verifier work, or an earlier training decision. This repository focuses on the released evaluation and Research Skills components; it does not include the complete training infrastructure.
| Component | Role |
|---|---|
eval_harness/ | Runs code-model evaluations across RTL and GPU-kernel benchmarks using local vLLM or an OpenAI-compatible endpoint. |
research_skills/ | Provides the versioned task instructions and workflow constraints used by the project. |
The harness turns generated artifacts into task-native evidence:
model endpoint → candidate generation → isolated verification → structured artifacts
It separates generation from resource-intensive verification, records structured outcomes and provenance fields, and keeps benchmark-specific logic behind a shared orchestration interface. See the harness guide.
Clone the repository once:
git clone https://github.com/bingreeky/iCoder.git
cd iCoder
To use the agent workflow, install the Skill directories and start with the
auto-post-training controller. The Research Skills guide
explains installation, bootstrap inputs, the Human Prior boundary, and the role
of every Skill.
To run model evaluation, enter eval_harness/ and follow the
toolchain, dataset, and endpoint setup.
iCoder/
├── eval_harness/ evaluation and executable-verification infrastructure
├── research_skills/ versioned Research Skills and release manifest
└── README.md project overview
The evaluation harness compiles and executes model-generated code. Run it only in an isolated, disposable environment with strict filesystem, process, resource, and network controls. Never place personal data, model credentials, or unrelated secrets on an evaluation host.
See the harness security guidance for operational details.
@article{yang2026icoder,
title = {iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model},
author = {Cheng Yang and Jiayang Lyu and Shangyuan Liu and Guibin Zhang and
Jiong Lin and Xinlei Yu and Junchi Yan and Shuicheng Yan and
Weinan E and Linfeng Zhang and Linfeng Zhang and Qibing Ren},
journal = {arXiv preprint arXiv:2609.29626},
year = {2026}
}
Citation metadata for the evaluation harness is available in
eval_harness/CITATION.cff. Model artifacts are
available from the i-Coder organization on Hugging Face.
The evaluation harness is released under the MIT License. Vendored components and benchmark datasets retain their original licenses; review the third-party notices before redistribution.
SystemVerilog
54.5%
Python
40.1%
Shell
4.3%
Research Skills · Agent-led post-training · Executable evaluation
Overview · Paper · Framework · Components · Research Skills · Eval Harness · Model · Security
iCoder is a research project on agent-led model development for RTL design and GPU kernel optimization. This repository publishes the project's evaluation harness and Research Skills bundle. The full technical report is available here, and model artifacts are available as iCoder-27B.
The project targets industrial coding: RTL design and GPU kernel optimization, where acceptance is determined by specialized executable toolchains and deployment quality rather than surface similarity alone.
[!NOTE] iCoder is human-guided and agent-led, not a fully autonomous self-improvement system. Its workflows require user-provided objectives, task instructions, verifiers, resource limits, and an isolated execution environment.
Research Skills guide the agent loop; task-native execution feeds evidence back into experimentation.
The project framework separates instructions supplied before an experiment from evidence produced during execution:
| Element | Function |
|---|---|
| Research Skills | Provide versioned task instructions, workflow constraints, and verifier requirements. |
| Agent loop | Applies those instructions, runs experiments, and reacts to execution feedback. |
| Executable verification | Uses task-native compilers, simulators, testbenches, and numerical oracles to produce structured outcomes. |
| Model release | Makes the resulting iCoder-27B artifacts available through Hugging Face. |
Data construction first creates a shared executable task pool. Parameter updates then proceed through SFT → OPSD → RLVR:
| Stage | Role |
|---|---|
| Data construction | Prepares executable RTL and GPU-kernel tasks and checks them with domain toolchains. |
| SFT (Supervised Fine-Tuning) | Uses verified teacher solutions to establish task capability. |
| OPSD (On-Policy Self-Distillation) | Trains on feedback collected from the model's own execution attempts. |
| RLVR (Reinforcement Learning with Verifiable Rewards) | Uses domain-specific execution outcomes as training signals. |
The stages form a feedback loop: evaluation can send the process back to data construction, verifier work, or an earlier training decision. This repository focuses on the released evaluation and Research Skills components; it does not include the complete training infrastructure.
| Component | Role |
|---|---|
eval_harness/ | Runs code-model evaluations across RTL and GPU-kernel benchmarks using local vLLM or an OpenAI-compatible endpoint. |
research_skills/ | Provides the versioned task instructions and workflow constraints used by the project. |
The harness turns generated artifacts into task-native evidence:
model endpoint → candidate generation → isolated verification → structured artifacts
It separates generation from resource-intensive verification, records structured outcomes and provenance fields, and keeps benchmark-specific logic behind a shared orchestration interface. See the harness guide.
Clone the repository once:
git clone https://github.com/bingreeky/iCoder.git
cd iCoder
To use the agent workflow, install the Skill directories and start with the
auto-post-training controller. The Research Skills guide
explains installation, bootstrap inputs, the Human Prior boundary, and the role
of every Skill.
To run model evaluation, enter eval_harness/ and follow the
toolchain, dataset, and endpoint setup.
iCoder/
├── eval_harness/ evaluation and executable-verification infrastructure
├── research_skills/ versioned Research Skills and release manifest
└── README.md project overview
The evaluation harness compiles and executes model-generated code. Run it only in an isolated, disposable environment with strict filesystem, process, resource, and network controls. Never place personal data, model credentials, or unrelated secrets on an evaluation host.
See the harness security guidance for operational details.
@article{yang2026icoder,
title = {iCoder-27B: Recursive AI-Led Development of Frontier Industrial Coding Model},
author = {Cheng Yang and Jiayang Lyu and Shangyuan Liu and Guibin Zhang and
Jiong Lin and Xinlei Yu and Junchi Yan and Shuicheng Yan and
Weinan E and Linfeng Zhang and Linfeng Zhang and Qibing Ren},
journal = {arXiv preprint arXiv:2609.29626},
year = {2026}
}
Citation metadata for the evaluation harness is available in
eval_harness/CITATION.cff. Model artifacts are
available from the i-Coder organization on Hugging Face.
The evaluation harness is released under the MIT License. Vendored components and benchmark datasets retain their original licenses; review the third-party notices before redistribution.
SystemVerilog
54.5%
Python
40.1%
Shell
4.3%