kaist-ai/mistral-orpo-alpha

Model

Mistral-ORPO-⍺ (7B)

8

8 commits

4 linked in READMEs

updated Mar 17, 2024

See the code

README

Mistral-ORPO-⍺ (7B)

Mistral-ORPO is a fine-tuned version of mistralai/Mistral-7B-v0.1 using the odds ratio preference optimization (ORPO). With ORPO, the model directly learns the preference without the supervised fine-tuning warmup phase. Mistral-ORPO-⍺ is fine-tuned exclusively on HuggingFaceH4/ultrafeedback_binarized.

👍 Model Performance

1) AlpacaEval & MT-Bench

Model NameSizeAlignMT-BenchAlpacaEval 1.0AlpacaEval 2.0
Mistral-ORPO-⍺7BORPO7.2387.9211.33
Mistral-ORPO7BORPO7.3291.4112.20
Zephyr β7BDPO7.3490.6010.99
TULU-2-DPO13BDPO7.0089.510.12
Llama-2-Chat7BRLHF6.2771.374.96
Llama-2-Chat13BRLHF6.6581.097.70

2) IFEval

Model TypePrompt-StrictPrompt-LooseInst-StrictInst-Loose
Mistral-ORPO-⍺0.50090.50830.59950.6163
Mistral-ORPO-β0.52870.55640.63550.6619

🗺️ MT-Bench by Category

image/png

🖥️ Inference

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("kaist-ai/mistral-orpo-alpha")
tokenizer = AutoTokenizer.from_pretrained("kaist-ai/mistral-orpo-alpha")

# Apply chat template
query = [{'role': 'user', 'content': 'Hi! How are you doing?'}]
prompt = tokenizer.apply_chat_template(query, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors='pt')

# Generation with specific configurations
output = model.generate(
  **inputs,
  max_new_tokens=128,
  do_sample=True,
  temperature=0.7
)
response = tokenizer.batch_decode(output)

#<|user|>
#Hi! How are you doing?</s>
#<|assistant|>
#I'm doing well, thank you! How are you?</s>

📎 Citation

@misc{hong2024orpo,
      title={ORPO: Monolithic Preference Optimization without Reference Model}, 
      author={Jiwoo Hong and Noah Lee and James Thorne},
      year={2024},
      eprint={2403.07691},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}
conversational
endpoints_compatible
mistral
model-index
safetensors
text-generation
text-generation-inference
transformers

Contributors

JW17

8 commits

kaist-ai/mistral-orpo-alpha

Model

Mistral-ORPO-⍺ (7B)

8

8 commits

4 linked in READMEs

updated Mar 17, 2024

See the code

README

Mistral-ORPO-⍺ (7B)

Mistral-ORPO is a fine-tuned version of mistralai/Mistral-7B-v0.1 using the odds ratio preference optimization (ORPO). With ORPO, the model directly learns the preference without the supervised fine-tuning warmup phase. Mistral-ORPO-⍺ is fine-tuned exclusively on HuggingFaceH4/ultrafeedback_binarized.

👍 Model Performance

1) AlpacaEval & MT-Bench

Model NameSizeAlignMT-BenchAlpacaEval 1.0AlpacaEval 2.0
Mistral-ORPO-⍺7BORPO7.2387.9211.33
Mistral-ORPO7BORPO7.3291.4112.20
Zephyr β7BDPO7.3490.6010.99
TULU-2-DPO13BDPO7.0089.510.12
Llama-2-Chat7BRLHF6.2771.374.96
Llama-2-Chat13BRLHF6.6581.097.70

2) IFEval

Model TypePrompt-StrictPrompt-LooseInst-StrictInst-Loose
Mistral-ORPO-⍺0.50090.50830.59950.6163
Mistral-ORPO-β0.52870.55640.63550.6619

🗺️ MT-Bench by Category

image/png

🖥️ Inference

from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("kaist-ai/mistral-orpo-alpha")
tokenizer = AutoTokenizer.from_pretrained("kaist-ai/mistral-orpo-alpha")

# Apply chat template
query = [{'role': 'user', 'content': 'Hi! How are you doing?'}]
prompt = tokenizer.apply_chat_template(query, tokenize=False, add_generation_prompt=True)
inputs = tokenizer(prompt, return_tensors='pt')

# Generation with specific configurations
output = model.generate(
  **inputs,
  max_new_tokens=128,
  do_sample=True,
  temperature=0.7
)
response = tokenizer.batch_decode(output)

#<|user|>
#Hi! How are you doing?</s>
#<|assistant|>
#I'm doing well, thank you! How are you?</s>

📎 Citation

@misc{hong2024orpo,
      title={ORPO: Monolithic Preference Optimization without Reference Model}, 
      author={Jiwoo Hong and Noah Lee and James Thorne},
      year={2024},
      eprint={2403.07691},
      archivePrefix={arXiv},
      primaryClass={cs.CL}
}
conversational
endpoints_compatible
mistral
model-index
safetensors
text-generation
text-generation-inference
transformers

Contributors

JW17

8 commits