11 repos
Datasets, benchmarks, and evaluation frameworks for testing and improving the safety and harmlessness of conversational AI systems. The cluster centers on standardized test formats and conversation datasets designed to identify harmful outputs and measure model behavior across sensitive domains. Repositories here focus on structured evaluation methodologies and shared benchmark formats for assessing how well language models avoid generating inappropriate, unsafe, or problematic content.
HuggingFaceH4/grok-conversation-harmless-old
Dataset Card for "cai-conversation-dev1705369037"