Amshaker/Mobile-O-Pre-Train

Dataset

12

stars

446

commits

1

linked in READMEs

Feb 24, 2026

updated

cross-modal-alignment
mobile-o
multimodal
pretraining

README

Mobile-O Pre-Training Data

Cross-Modal Alignment Β· 9M Text-Image Pairs

arXiv Code Project Page Models Live Demo

πŸ“Œ Overview

This dataset is used for Stage 1: Cross-Modal Alignment pre-training of Mobile-O, a unified multimodal model for on-device understanding and generation.

The goal of this stage is to align the DiT diffusion decoder and Mobile Conditioning Projector (MCP) with the frozen VLM backbone using large-scale text-image pairs.

πŸ“Š Dataset Composition

SourceSamplesDescription
JourneyDB4MHigh-quality AI-generated images with captions
BLIP3o-Pretrain-Short-Caption5MEach image paired with a short caption generated by Qwen/Qwen2.5-VL-7B-Instruct

πŸ‹οΈ Training Details

  • Stage: 1 β€” Cross-Modal Alignment (Pre-training)
  • Trainable components: DiT + Mobile Conditioning Projector (MCP)
  • Frozen components: Visual encoders, LLM backbone, VAE
  • Script: pretrain.sh
ResourceLink
πŸ“„ PaperarXiv
πŸ’» CodeGitHub
πŸ€— SFT DataMobile-O-SFT
πŸ€— Post-Training DataMobile-O-Post-Train
πŸ€— Model (0.5B)Mobile-O-0.5B
πŸ€— Model (1.5B)Mobile-O-1.5B

πŸ“„ Citation

@article{shaker2026mobileo,
  title={Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device},
  author={Shaker, Abdelrahman and Heakl, Ahmed and Muhammad, Jaseel and Thawkar, Ritesh and Thawakar, Omkar and Li, Senmao and Cholakkal, Hisham and Reid, Ian and Xing, Eric P. and Khan, Salman and Khan, Fahad Shahbaz},
  journal={arXiv preprint arXiv:2602.20161},
  year={2026}
}

πŸ™ Acknowledgments

We gratefully acknowledge the following datasets used in constructing this pre-training corpus:

Contributors

Amshaker

446 commits

Amshaker/Mobile-O-Pre-Train

Dataset

12

stars

446

commits

1

linked in READMEs

Feb 24, 2026

updated

cross-modal-alignment
mobile-o
multimodal
pretraining

README

Mobile-O Pre-Training Data

Cross-Modal Alignment Β· 9M Text-Image Pairs

arXiv Code Project Page Models Live Demo

πŸ“Œ Overview

This dataset is used for Stage 1: Cross-Modal Alignment pre-training of Mobile-O, a unified multimodal model for on-device understanding and generation.

The goal of this stage is to align the DiT diffusion decoder and Mobile Conditioning Projector (MCP) with the frozen VLM backbone using large-scale text-image pairs.

πŸ“Š Dataset Composition

SourceSamplesDescription
JourneyDB4MHigh-quality AI-generated images with captions
BLIP3o-Pretrain-Short-Caption5MEach image paired with a short caption generated by Qwen/Qwen2.5-VL-7B-Instruct

πŸ‹οΈ Training Details

  • Stage: 1 β€” Cross-Modal Alignment (Pre-training)
  • Trainable components: DiT + Mobile Conditioning Projector (MCP)
  • Frozen components: Visual encoders, LLM backbone, VAE
  • Script: pretrain.sh
ResourceLink
πŸ“„ PaperarXiv
πŸ’» CodeGitHub
πŸ€— SFT DataMobile-O-SFT
πŸ€— Post-Training DataMobile-O-Post-Train
πŸ€— Model (0.5B)Mobile-O-0.5B
πŸ€— Model (1.5B)Mobile-O-1.5B

πŸ“„ Citation

@article{shaker2026mobileo,
  title={Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device},
  author={Shaker, Abdelrahman and Heakl, Ahmed and Muhammad, Jaseel and Thawkar, Ritesh and Thawakar, Omkar and Li, Senmao and Cholakkal, Hisham and Reid, Ian and Xing, Eric P. and Khan, Salman and Khan, Fahad Shahbaz},
  journal={arXiv preprint arXiv:2602.20161},
  year={2026}
}

πŸ™ Acknowledgments

We gratefully acknowledge the following datasets used in constructing this pre-training corpus:

Contributors

Amshaker

446 commits