π Playground | π Technical report | π» GitHub | π Sign up for the API
Selene Mini is a state-of-the-art small language model-as-a-judge (SLMJ). Selene Mini achieves comparable performance to models 10x its size, outperforming GPT-4o on RewardBench, EvalBiasBench, and AutoJ.
Post-trained from Llama-3.1-8B across a wide range of evaluation tasks and scoring criteria, Selene Mini outperforms prior small evaluation models overall across 11 benchmarks covering three different types of tasks:
It is also the #1 8B generative model on RewardBench.
This repo features prompt templates used during training and hands-on examples for using Selene Mini.
Get in touch if you have any queries not covered in this repo.
Jupyter Notebook
100.0%
π Playground | π Technical report | π» GitHub | π Sign up for the API
Selene Mini is a state-of-the-art small language model-as-a-judge (SLMJ). Selene Mini achieves comparable performance to models 10x its size, outperforming GPT-4o on RewardBench, EvalBiasBench, and AutoJ.
Post-trained from Llama-3.1-8B across a wide range of evaluation tasks and scoring criteria, Selene Mini outperforms prior small evaluation models overall across 11 benchmarks covering three different types of tasks:
It is also the #1 8B generative model on RewardBench.
This repo features prompt templates used during training and hands-on examples for using Selene Mini.
Get in touch if you have any queries not covered in this repo.
Jupyter Notebook
100.0%