Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Project Homepage]
[🤗 SafeSora Datasets]
[🤗 SafeSora Label]
[🤗 SafeSora Evaluation]
[⭐️ Github Repo]
SafeSora is a human preference dataset designed to support safety alignment research in the text-to-video generation field, aiming to enhance the helpfulness and harmlessness of Large Vision Models (LVMs). It currently contains three types of data:
In the future, we will also open-source some baseline alignment algorithms that utilize these datasets.
The multi-label classification dataset contains 57k+ text-video pairs, each labeled with 12 harm tags. We perform multi-label classification on individual prompts as well as the combination of prompts and the videos generated from those prompts. These 12 harm tags are defined as:
Adult, Explicit Sexual ContentAnimal AbuseChild AbuseCrimeDebated Sensitive Social IssueDrug, Weapons, Substance AbuseInsulting, Hateful, Aggressive BehaviorViolence, Injury, Gory ContentRacial DiscriminationOther Discrimination (Excluding Racial)Terrorism, Organized CrimeOther Harmful ContentThe distribution of these 14 categories is shown below:

In our dataset, nearly half of the prompts are safety-critical, while the remaining half are safety-neutral. Our prompts partly come from real online users, while the remaining portion is supplemented by researchers for balancing purposes.
Multi-label Classification Dataset only contains classification information for T-V pairs

Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Project Homepage]
[🤗 SafeSora Datasets]
[🤗 SafeSora Label]
[🤗 SafeSora Evaluation]
[⭐️ Github Repo]
SafeSora is a human preference dataset designed to support safety alignment research in the text-to-video generation field, aiming to enhance the helpfulness and harmlessness of Large Vision Models (LVMs). It currently contains three types of data:
In the future, we will also open-source some baseline alignment algorithms that utilize these datasets.
The multi-label classification dataset contains 57k+ text-video pairs, each labeled with 12 harm tags. We perform multi-label classification on individual prompts as well as the combination of prompts and the videos generated from those prompts. These 12 harm tags are defined as:
Adult, Explicit Sexual ContentAnimal AbuseChild AbuseCrimeDebated Sensitive Social IssueDrug, Weapons, Substance AbuseInsulting, Hateful, Aggressive BehaviorViolence, Injury, Gory ContentRacial DiscriminationOther Discrimination (Excluding Racial)Terrorism, Organized CrimeOther Harmful ContentThe distribution of these 14 categories is shown below:

In our dataset, nearly half of the prompts are safety-critical, while the remaining half are safety-neutral. Our prompts partly come from real online users, while the remaining portion is supplemented by researchers for balancing purposes.
Multi-label Classification Dataset only contains classification information for T-V pairs
