This is the project to quantify the issues with LLM's self evaluation
9
stars
11
commits
Python
primary language
Apr 18, 2024
updated
9 commits
1 commits
ouhenio/llms-overstimate-profoundness
2
IS2150534702/is-project
0
VietHoang1512/PSR
Think Twice, Generate Once: Safeguarding by Progressive Self-Reflection (EMNLP 2025 Findings)
5
Higashi-Masafumi/llm-error-injection
inject error in llm computation and eval
1
tigerchen52/reconfidencing_llms
nikhil9111924/SolidSnakeDrivers-BiasAudit
IIITH | M25-ANLP Project Repo
aiden200/MLLMS-FAIRNESS
Fairness experiments in Multimodal Large Language Models
Phylliida/ModelPreferences
Learning Preferences of LLMs
93.4%
Shell
6.6%