Qwen3.8-multi-turn-agent-sft
9
7 commits
1 linked in READMEs
updated Aug 28, 2026
Hello everyone! We are UkisAI, a small research lab from Europe.
We created this dataset based on the OpenThoughts-Agent-v1-SFT dataset. The traces in this release were generated with Qwen3.8-27B in FP16 using the Terminus-2 agentic harness.
This dataset contains approximately 15,200 agent traces covering terminal, coding, and software-engineering tasks, including tasks from nl2bash and InferredBugs.
Please feel free to try it, share feedback, report issues, or suggest improvements. We would be happy to hear from you!
This is a supervised fine-tuning (SFT) dataset for training models to improve on agentic tasks. It can be used to warm-start a model's ability to interact with a terminal, use tools, and solve coding and software-engineering tasks. It may also serve as a starting point for further post-training or reinforcement learning.
The original OpenThoughts project used SFT traces from strong teacher agents as the first stage of training, followed by reinforcement learning. See the original dataset card for more background and the OpenThoughts-Agent project for the broader pipeline.
Many thanks to the OpenThoughts team for releasing the original dataset and agent-training pipeline openly.
7 commits
Qwen3.8-multi-turn-agent-sft
9
7 commits
1 linked in READMEs
updated Aug 28, 2026
Hello everyone! We are UkisAI, a small research lab from Europe.
We created this dataset based on the OpenThoughts-Agent-v1-SFT dataset. The traces in this release were generated with Qwen3.8-27B in FP16 using the Terminus-2 agentic harness.
This dataset contains approximately 15,200 agent traces covering terminal, coding, and software-engineering tasks, including tasks from nl2bash and InferredBugs.
Please feel free to try it, share feedback, report issues, or suggest improvements. We would be happy to hear from you!
This is a supervised fine-tuning (SFT) dataset for training models to improve on agentic tasks. It can be used to warm-start a model's ability to interact with a terminal, use tools, and solve coding and software-engineering tasks. It may also serve as a starting point for further post-training or reinforcement learning.
The original OpenThoughts project used SFT traces from strong teacher agents as the first stage of training, followed by reinforcement learning. See the original dataset card for more background and the OpenThoughts-Agent project for the broader pipeline.
Many thanks to the OpenThoughts team for releasing the original dataset and agent-training pipeline openly.
7 commits