ukisai/Qwen3.8-27B-multi-turn-agent-sft

Dataset

Qwen3.8-multi-turn-agent-sft

9

7 commits

1 linked in READMEs

updated Aug 28, 2026

See the code
agents
code
software-engineering
terminal

README

Qwen3.8-multi-turn-agent-sft

Hello everyone! We are UkisAI, a small research lab from Europe.

We created this dataset based on the OpenThoughts-Agent-v1-SFT dataset. The traces in this release were generated with Qwen3.8-27B in FP16 using the Terminus-2 agentic harness.

This dataset contains approximately 15,200 agent traces covering terminal, coding, and software-engineering tasks, including tasks from nl2bash and InferredBugs.

Please feel free to try it, share feedback, report issues, or suggest improvements. We would be happy to hear from you!

Intended use

This is a supervised fine-tuning (SFT) dataset for training models to improve on agentic tasks. It can be used to warm-start a model's ability to interact with a terminal, use tools, and solve coding and software-engineering tasks. It may also serve as a starting point for further post-training or reinforcement learning.

The original OpenThoughts project used SFT traces from strong teacher agents as the first stage of training, followed by reinforcement learning. See the original dataset card for more background and the OpenThoughts-Agent project for the broader pipeline.

Acknowledgements

Many thanks to the OpenThoughts team for releasing the original dataset and agent-training pipeline openly.

Contributors

ukisai

7 commits

ukisai/Qwen3.8-27B-multi-turn-agent-sft

Dataset

Qwen3.8-multi-turn-agent-sft

9

7 commits

1 linked in READMEs

updated Aug 28, 2026

See the code
agents
code
software-engineering
terminal

README

Qwen3.8-multi-turn-agent-sft

Hello everyone! We are UkisAI, a small research lab from Europe.

We created this dataset based on the OpenThoughts-Agent-v1-SFT dataset. The traces in this release were generated with Qwen3.8-27B in FP16 using the Terminus-2 agentic harness.

This dataset contains approximately 15,200 agent traces covering terminal, coding, and software-engineering tasks, including tasks from nl2bash and InferredBugs.

Please feel free to try it, share feedback, report issues, or suggest improvements. We would be happy to hear from you!

Intended use

This is a supervised fine-tuning (SFT) dataset for training models to improve on agentic tasks. It can be used to warm-start a model's ability to interact with a terminal, use tools, and solve coding and software-engineering tasks. It may also serve as a starting point for further post-training or reinforcement learning.

The original OpenThoughts project used SFT traces from strong teacher agents as the first stage of training, followed by reinforcement learning. See the original dataset card for more background and the OpenThoughts-Agent project for the broader pipeline.

Acknowledgements

Many thanks to the OpenThoughts team for releasing the original dataset and agent-training pipeline openly.

Contributors

ukisai

7 commits