4 repos
tatsu-lab/alpaca_eval
An automatic evaluator for instruction-following language models. Human-validated, high-quality,…
2,013
585 commits
wschella/llm-reliability
Code for the paper "Larger and more instructable language models become less reliable"
34
17 commits
princeton-nlp/LLMBar
[ICLR 2024] Evaluating Large Language Models at Evaluating Instruction Following
141
10 commits
WeOpenML/PandaLM
No description
924
77 commits