Amshaker/Mobile-O-Post-Train

Dataset

13

stars

27

commits

1

linked in READMEs

Feb 24, 2026

updated

mobile-o
multimodal
post-training
unified-training

README

Mobile-O Post-Training Data

Unified Multimodal Post-Training Β· ~105K Quadruplet Samples

arXiv Code Project Page Models Live Demo

πŸ“Œ Overview

This dataset is used for Stage 3: Unified Multimodal Post-Training of Mobile-O, a unified multimodal model for on-device understanding and generation.

The goal of this stage is to jointly improve both image generation and visual understanding through a multi-task objective using quadruplet samples.

πŸ“Š Dataset Format

Each sample is a quadruplet consisting of:

FieldDescription
Generation PromptText prompt for image generation
ImageCorresponding image
QuestionVisual understanding question about the image
AnswerGround-truth answer

Total samples: ~105K

πŸ‹οΈ Training Details

  • Stage: 3 β€” Unified Multimodal Post-Training
  • Trainable components: DiT + MCP + LLM (via LoRA) + Visual Encoder
  • Frozen components: VAE only
ResourceLink
πŸ“„ PaperarXiv
πŸ’» CodeGitHub
πŸ€— Pre-Training DataMobile-O-Pre-Train
πŸ€— SFT DataMobile-O-SFT
πŸ€— Model (0.5B)Mobile-O-0.5B
πŸ€— Model (1.5B)Mobile-O-1.5B

πŸ“„ Citation

@article{shaker2026mobileo,
  title={Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device},
  author={Shaker, Abdelrahman and Heakl, Ahmed and Muhammad, Jaseel and Thawkar, Ritesh and Thawakar, Omkar and Li, Senmao and Cholakkal, Hisham and Reid, Ian and Xing, Eric P. and Khan, Salman and Khan, Fahad Shahbaz},
  journal={arXiv preprint arXiv:2602.20161},
  year={2026}
}

βš–οΈ License

This dataset is released under CC BY-NC 4.0. For research purposes only.

Contributors

Amshaker

27 commits

Amshaker/Mobile-O-Post-Train

Dataset

13

stars

27

commits

1

linked in READMEs

Feb 24, 2026

updated

mobile-o
multimodal
post-training
unified-training

README

Mobile-O Post-Training Data

Unified Multimodal Post-Training Β· ~105K Quadruplet Samples

arXiv Code Project Page Models Live Demo

πŸ“Œ Overview

This dataset is used for Stage 3: Unified Multimodal Post-Training of Mobile-O, a unified multimodal model for on-device understanding and generation.

The goal of this stage is to jointly improve both image generation and visual understanding through a multi-task objective using quadruplet samples.

πŸ“Š Dataset Format

Each sample is a quadruplet consisting of:

FieldDescription
Generation PromptText prompt for image generation
ImageCorresponding image
QuestionVisual understanding question about the image
AnswerGround-truth answer

Total samples: ~105K

πŸ‹οΈ Training Details

  • Stage: 3 β€” Unified Multimodal Post-Training
  • Trainable components: DiT + MCP + LLM (via LoRA) + Visual Encoder
  • Frozen components: VAE only
ResourceLink
πŸ“„ PaperarXiv
πŸ’» CodeGitHub
πŸ€— Pre-Training DataMobile-O-Pre-Train
πŸ€— SFT DataMobile-O-SFT
πŸ€— Model (0.5B)Mobile-O-0.5B
πŸ€— Model (1.5B)Mobile-O-1.5B

πŸ“„ Citation

@article{shaker2026mobileo,
  title={Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device},
  author={Shaker, Abdelrahman and Heakl, Ahmed and Muhammad, Jaseel and Thawkar, Ritesh and Thawakar, Omkar and Li, Senmao and Cholakkal, Hisham and Reid, Ian and Xing, Eric P. and Khan, Salman and Khan, Fahad Shahbaz},
  journal={arXiv preprint arXiv:2602.20161},
  year={2026}
}

βš–οΈ License

This dataset is released under CC BY-NC 4.0. For research purposes only.

Contributors

Amshaker

27 commits