Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Project Homepage]
[🤗 SafeSora Datasets]
[🤗 SafeSora Label]
[🤗 SafeSora Evaluation]
[⭐️ Github Repo]
SafeSora is a human preference dataset designed to support safety alignment research in the text-to-video generation field, aiming to enhance the helpfulness and harmlessness of Large Vision Models (LVMs). It currently contains three types of data:
In the future, we will also open-source some baseline alignment algorithms that utilize these datasets.
SafeSora is a human preference dataset of 51k+ instances in the text-to-video generation task, containing comparative relationships in terms of helpfulness and harmlessness, as well as four sub-dimensions of helpfulness.
Each data point comprising a user input and two generated videos. Through a heuristic-based annotation process, human preferences were obtained in terms of helpfulness or harmlessness dimensions.
Additionally, due to a pre-annotation heuristic process, human preferences on four helpfulness sub-dimensions were also included. These sub-dimensions are:
Instruction FollowingCorrectnessInformativenessAestheticsThe specific annotation process is as shown in the figure below:


Warning: this dataset contains data that may be offensive or harmful. The data are intended for research purposes, especially research that can make models less harmful. The views expressed in the data do not reflect the views of PKU-Alignment Team or any of its members.
[🏠 Project Homepage]
[🤗 SafeSora Datasets]
[🤗 SafeSora Label]
[🤗 SafeSora Evaluation]
[⭐️ Github Repo]
SafeSora is a human preference dataset designed to support safety alignment research in the text-to-video generation field, aiming to enhance the helpfulness and harmlessness of Large Vision Models (LVMs). It currently contains three types of data:
In the future, we will also open-source some baseline alignment algorithms that utilize these datasets.
SafeSora is a human preference dataset of 51k+ instances in the text-to-video generation task, containing comparative relationships in terms of helpfulness and harmlessness, as well as four sub-dimensions of helpfulness.
Each data point comprising a user input and two generated videos. Through a heuristic-based annotation process, human preferences were obtained in terms of helpfulness or harmlessness dimensions.
Additionally, due to a pre-annotation heuristic process, human preferences on four helpfulness sub-dimensions were also included. These sub-dimensions are:
Instruction FollowingCorrectnessInformativenessAestheticsThe specific annotation process is as shown in the figure below:

