This dataset is used for Stage 2: Supervised Fine-Tuning (SFT) of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to improve image generation quality by fine-tuning on high-quality curated prompt-image pairs.
| Source | Samples | Description |
|---|---|---|
| BLIP3o | 60K | High-quality prompt-image pairs |
| ShareGPT-4o-Image | 45K | Curated image generation pairs |
| Total | ~105K |
| Resource | Link |
|---|---|
| π Paper | arXiv |
| π» Code | GitHub |
| π€ Pre-Training Data | Mobile-O-Pre-Train |
| π€ Post-Training Data | Mobile-O-Post-Train |
| π€ Model (0.5B) | Mobile-O-0.5B |
| π€ Model (1.5B) | Mobile-O-1.5B |
@article{shaker2026mobileo,
title={Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device},
author={Shaker, Abdelrahman and Heakl, Ahmed and Muhammad, Jaseel and Thawkar, Ritesh and Thawakar, Omkar and Li, Senmao and Cholakkal, Hisham and Reid, Ian and Xing, Eric P. and Khan, Salman and Khan, Fahad Shahbaz},
journal={arXiv preprint arXiv:2602.20161},
year={2026}
}
This dataset is released under CC BY-NC 4.0. For research purposes only.
26 commits
This dataset is used for Stage 2: Supervised Fine-Tuning (SFT) of Mobile-O, a unified multimodal model for on-device understanding and generation.
The goal of this stage is to improve image generation quality by fine-tuning on high-quality curated prompt-image pairs.
| Source | Samples | Description |
|---|---|---|
| BLIP3o | 60K | High-quality prompt-image pairs |
| ShareGPT-4o-Image | 45K | Curated image generation pairs |
| Total | ~105K |
| Resource | Link |
|---|---|
| π Paper | arXiv |
| π» Code | GitHub |
| π€ Pre-Training Data | Mobile-O-Pre-Train |
| π€ Post-Training Data | Mobile-O-Post-Train |
| π€ Model (0.5B) | Mobile-O-0.5B |
| π€ Model (1.5B) | Mobile-O-1.5B |
@article{shaker2026mobileo,
title={Mobile-O: Unified Multimodal Understanding and Generation on Mobile Device},
author={Shaker, Abdelrahman and Heakl, Ahmed and Muhammad, Jaseel and Thawkar, Ritesh and Thawakar, Omkar and Li, Senmao and Cholakkal, Hisham and Reid, Ian and Xing, Eric P. and Khan, Salman and Khan, Fahad Shahbaz},
journal={arXiv preprint arXiv:2602.20161},
year={2026}
}
This dataset is released under CC BY-NC 4.0. For research purposes only.
26 commits