
Dataset: LLaDA-Sample-10BT
Base: HuggingFaceFW/fineweb (subset sample-10BT)
Purpose: Training LLaDA (Large Language Diffusion Models)
GSAI-ML/LLaDA-8B-Instructinput_idsnoisy_input_idsmaskt (time scalar).pt filesThis dataset is used for training in the LLaDA-from-scratch GitHub repository, where you’ll find the full data pipeline and training scripts.

Dataset: LLaDA-Sample-10BT
Base: HuggingFaceFW/fineweb (subset sample-10BT)
Purpose: Training LLaDA (Large Language Diffusion Models)
GSAI-ML/LLaDA-8B-Instructinput_idsnoisy_input_idsmaskt (time scalar).pt filesThis dataset is used for training in the LLaDA-from-scratch GitHub repository, where you’ll find the full data pipeline and training scripts.