11 repos
Datasets, models, and techniques for training language models to be helpful, harmless, and honest in conversational contexts. The cluster centers on conversation safety benchmarks and supervised fine-tuning approaches, with repositories like the Grok and CAI conversation harmlessness datasets forming a tightly-connected core focused on filtering and evaluating model outputs for harmful content. Related work includes instruction-following datasets (DEITA) and broader alignment methodologies that address how to steer large language models toward safe, beneficial behavior.