stevelaskaridis/awesome-mobile-llm

Awesome Mobile LLMs

403

150 commits

updated Sep 3, 2026

See the code

README

Awesome Mobile LLMs Awesome

A curated list of LLMs and related studies targeted at mobile and embedded hardware

Last update: 3rd September 2026

If your publication/work is not included - and you think it should - please open an issue or reach out directly to @stevelaskaridis.

Let's try to make this list as useful as possible to researchers, engineers and practitioners all around the world.

Contents

Mobile-First LLMs

The following Table shows sub-3B models designed for on-device deployments, sorted by year.

NameYearSizesPrimary Group/AffiliationPublicationCode RepositoryHF Repository
2026
MobileMoE20260.3B, 0.5B, 0.9B active (1.3B, 2.8B, 5.3B total)Meta AIpaper--
Gemma 42026E2B, E4B, 26BGoogle DeepMindwebsitecodehuggingface
LFM2.52026350M, 1.2B, 1.5B, 1.6B, 2.6BLiquid AIwebsite, LFM2.5-2.6B blog-huggingface
MobileLLM-Flash2026350M, 650M, 1.4BMetapaper--
Apertus Mini20260.5B, 1.5B, 4BSwiss AI (EPFL, ETH Zurich, CSCS)papercodehuggingface
Qwen-3.520260.8B, 2B, ...Qwen Teamblogcodehuggingface
2025
LFM22025350M, 700M, 1.2B, 2.6B, 8.3B (1.5B active)Liquid AIpaper, website-huggingface
MobileLLM-R1.52025140M, 360M, 950MMetapapercodehuggingface
Nemotron-Flash20251B, 3BNvidiapaper, NeurIPS'25-huggingface
MobileLLM-Pro20251BMetapaper-huggingface
MobileLLM-R12025140M, 360M, 950MMetapapercodehuggingface
SmolLM320253BHuggingFaceblogcodehuggingface
Gemma 320251B, 4B, ...Google DeepMindpapercodehuggingface
Qwen-320250.6B, 1.7B, ...Qwen Teampapercodehuggingface
Pareto-Q2025125M, 350M, 600M, 1B, 1.5B, 3BMetapapercodehuggingface
2024
BlueLM-V20242.7BCUHK, Vivo AI Labpapercode-
PhoneLM20240.5B, 1.5BBUPTpapercodehuggingface
AMD-Llama-135m2024135MAMDblogcodehuggingface
SmolLM22024135M, 360M, 1.7BHuggingface-codehuggingface
Ministral20243B, ...Mistralblog-huggingface
Llama 3.220241B, 3BMetablogcodehuggingface
OLMoE20247B (1B active)AllenAIpapercodehuggingface
Spectra202499M - 3.9BNolanoAIpapercodehuggingface
Gemma 220242B, ...Googlepaper blogcodehuggingface
Apple Intelligence Foundation LMs20243BApplepaper--
SmolLM2024135M, 360M, 1.7BHuggingfaceblog-huggingface
Fox20241.6BTensorOperablog-huggingface
Qwen22024500M, 1.5B, ...Qwen Teampapercodehuggingface
OpenELM2024270M, 450M, 1.08B, 3.04BApplepapercodehuggingface
DCLM2024400M, 1B, ...Univerisy of Washington, Apple, Toyota Research Institute, ...papercodehuggingface
Phi-320243.8BMicrosoftwhitepapercodehuggingface
BitNet-b1.5820241.3B, 3B, ...Microsoftpapercodehuggingface
OLMo20241B, ...AllenAIpapercodehuggingface
Mobile LLMs2024125M, 250MMetapaper, ICML'24code-
Gemma20242B, ...Googlepaper, websitecode, gemma.cpphuggingface
MobiLlama20240.5B, 1BMBZUAIpapercodehuggingface
Stable LM 2 (Zephyr)20241.6BStability.aipaper-huggingface
TinyLlama20241.1BSingapore University of Technology and Designpapercodehuggingface
Gemini-Nano20241.8B, 3.25BGooglepaper--
2023
Stable LM (Zephyr)20233BStabilityblogcodehuggingface
OpenLM202311M, 25M, 87M, 160M, 411M, 830M, 1B, 3B, ...OpenLM team-codehuggingface
Phi-220232.7BMicrosoftwebsite-huggingface
Phi-1.520231.3BMicrosoftpaper-huggingface
Phi-120231.3BMicrosoftpaper-huggingface
RWKV2023169M, 430M, 1.5B, 3B, ...EleutherAIpapercodehuggingface
Cerebras-GPT2023111M, 256M, 590M, 1.3B, 2.7B ...Cerebraspapercodehuggingface
OPT2022125M, 350M, 1.3B, 2.7B, ...Metapapercodehuggingface
LaMini-LM202361M, 77M, 111M, 124M, 223M, 248M, 256M, 590M, 774M, 738M, 783M, 1.3B, 1.5B, ...MBZUAIpapercodehuggingface
Pythia202370M, 160M, 410M, 1B, 1.4B, 2.8B, ...EleutherAIpapercodehuggingface
2022
Galactica2022125M, 1.3B, ...Metapapercodehuggingface
BLOOM2022560M, 1.1B, 1.7B, 3B, ...BigSciencepapercodehuggingface
2021
XGLM2021564M, 1.7B, 2.9B, ...Metapapercodehuggingface
GPT-Neo2021125M, 350M, 1.3B, 2.7BEleutherAI-code, gpt-neoxhuggingface
2020
MobileBERT202015.1M, 25.3MCMU, Googlepapercodehuggingface
2019
BART2019140M, 400MMetapapercodehuggingface
DistilBERT201966MHuggingFacepapercodehuggingface
T5201960M, 220M, 770M, 3B, ...Googlepapercodehuggingface
TinyBERT201914.5MHuaweipapercodehuggingface
Megatron-LM2019336M, 1.3B, ...Nvidiapapercode-

Infrastructure / Deployment of LLMs on Device

This section showcases frameworks and contributions for supporting LLM inference on mobile and edge devices.

Deployment Frameworks

On-Device Inference Frameworks

These frameworks are primarily used to run models directly on-device, inside mobile apps, edge deployments, or tightly integrated local runtimes.

  • llama.cpp: Inference of Meta's LLaMA model (and others) in pure C/C++. Supports various platforms and builds on top of ggml (now gguf format).
    • LLMFarm: iOS frontend for llama.cpp
    • LLM.swift: iOS frontend for llama.cpp
    • Sherpa: Android frontend for llama.cpp
    • iAkashPaul/Portal: Wraps the example android app with tweaked UI, configs & additional model support
    • dusty-nv's llama.cpp: Containers for Jetson deployment of llama.cpp
    • Off Grid: Open-source React Native app for on-device LLM chat, vision models (SmolVLM, LLaVA), and Stable Diffusion image generation on iOS & Android.
    • Airgap: Open-source React Native framework for on-device, offline-first customer support chatbots. Runs Gemma 4 E2B locally via llama.rn. Seven industry templates (telco, retail, healthcare, banking, education, insurance, airlines) ship in the repo.
  • MLC-LLM: MLC LLM is a machine learning compiler and high-performance deployment engine for large language models. Supports various platforms and build on top of TVM.
  • PyTorch ExecuTorch: Solution for enabling on-device inference capabilities across mobile and edge devices including wearables, embedded devices and microcontrollers.
    • TorchChat: Codebase showcasing the ability to run large language models (LLMs) seamlessly across iOS and Android
  • Google MediaPipe: A suite of libraries and tools for you to quickly apply artificial intelligence (AI) and machine learning (ML) techniques in your applications. Support Android, iOS, Python and Web.
    • GoogleAI-Edge Gallery: Experimental app that puts the power of cutting-edge Generative AI models directly into your hands, running entirely on your Android and iOS devices.
  • Apple MLX: MLX is an array framework for machine learning research on Apple silicon, brought to you by Apple machine learning research. Builds upon lazy evaluation and unified memory architecture.
  • Apple Foundation Models SDK: Python bindings for Apple's Foundation Models framework, providing access to the on-device foundation model at the core of Apple Intelligence on macOS.
  • HF Swift Transformers: Swift Package to implement a transformers-like API in Swift
  • Alibaba MNN: MNN supports inference and training of deep learning models and for inference and training on-device.
  • llama2.c (More educational, see here for android port)
  • tinygrad: Simple neural network framework from tinycorp and @geohot
  • TinyChatEngine: Targeted at Nvidia, Apple M1 and RPi, from Song Han's (MIT) group.
  • Llama Stack (swift, kotlin): These libraries are a set of SDKs that provide a simple and effective way to integrate AI capabilities into your iOS/Android app, whether it is local (on-device) or remote inference.
  • OLMoE.Swift: Ai2 OLMoE is an AI chatbot powered by the OLMoE model. Unlike cloud-based AI assistants, OLMoE runs entirely on your device, ensuring complete privacy and offline accessibility—even in Flight Mode.
  • HuggingSnap: HuggingSnap is an iOS app that lets users quickly learn more about the places and objects around them. HuggingSnap runs SmolVLM2, a compact open multimodal model that accepts arbitrary sequences of image, videos, and text inputs to produce text outputs.
  • Flower Intelligence: Flower Intelligence is a cross-platform inference library that lets users seamlessly interact with Large-Language Models both locally and remotely in a secure and private way. The library was created by the Flower Labs team. It supports TypeScript, JavaScript and Swift backends.
  • ONNX Runtime: Cross-platform inference and training engine for ONNX models, with a Mobile package and execution providers (NNAPI, Core ML, XNNPACK, QNN) for on-device deployment on Android and iOS. See the Mobile deployment guide.

Local Network Model Serving

These frameworks are primarily used to host models on a laptop, desktop, or workstation and expose them over a local API to other devices on the same LAN.

  • LM Studio: Desktop application and local inference server for hosting models on your machine, with an OpenAI-compatible local API.
  • Ollama: Local model runner and server for hosting and serving models through a simple CLI and HTTP API.
  • Lemonade: Open-source local AI server for text, image, and speech workloads, designed to run privately on local PCs and compatible with OpenAI-style APIs.
  • llama.cpp: Can also be used as a lightweight local inference server for hosting GGUF models via CLI and HTTP server modes.
  • LocalAI: Self-hosted local inference server and OpenAI-compatible REST API for running LLM, vision, image, and audio workloads on local or on-prem hardware.
  • Locally AI: Native Apple-platform app for running AI models fully offline on iPhone, iPad, and Mac, optimized for Apple Silicon and on-device privacy.
  • vLLM: High-throughput inference and serving engine that can expose OpenAI-compatible local APIs, better suited to stronger desktops and workstations.
  • SGLang: High-performance model serving framework for local and distributed deployments, designed for low-latency and high-throughput inference.

Papers

2026

  • [SenSys'26] An Efficient Context Management System for On-Device LLMaaS
    Wangsong Yin et al.
    DOI
  • FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
    Yinpeng Wu, Yitong Chen, Lixiang Wang, et al.
    arXiv

2025

  • Apple Intelligence Foundation Language Models: Tech Report 2025
    Ethan Li, Anders Boesen Lindbo Larsen, Chen Zhang, et al.
    arXiv
  • [ACM Queue] Generative AI at the Edge: Challenges and Opportunities: The next phase in AI deployment
    Vijay Janapa Reddi
    DOI

2024

  • PowerInfer-2: Fast Large Language Model Inference on a Smartphone
    Zhenliang Xue, Yixin Song, Zeyu Mi, et al.
    arXiv Code
  • [MobiCom'24] Mobile Foundation Model as Firmware
    Jinliang Yuan, Chen Yang, Dongqi Cai, et al.
    Paper DOI Code
  • Merino: Entropy-driven Design for Generative Language Models on IoT Devicess
    Youpeng Zhao, Ming Lin, Huadong Tang, et al.
    arXiv
  • LLM as a System Service on Mobile Devices
    Wangsong Yin, Mengwei Xu, Yuanchun Li, et al.
    arXiv

2023

  • LinguaLinked: A Distributed Large Language Model Inference System for Mobile Devices
    Junchen Zhao, Yurun Song, Simeng Liu, et al.
    arXiv
  • LLMCad: Fast and Scalable On-device Large Language Model Inference
    Daliang Xu, Wangsong Yin, Xin Jin, et al.
    arXiv
  • EdgeMoE: Fast On-Device Inference of MoE-based Large Language Models
    Rongjie Yi, Liwei Guo, Shiyun Wei, et al.
    arXiv

2022

  • [IEEE Pervasive Computing] The Future of Consumer Edge-AI Computing
    Stefanos Laskaridis, Stylianos I. Venieris, Alexandros Kouris, et al.
    arXiv Talk

Benchmarking LLMs on Device

This section focuses on measurements and benchmarking efforts for assessing LLM performance when deployed on device.

Papers

2026

  • LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
    Pranay Tummalapalli, Sahil Arayakandy, Ritam Pal, Kautuk Kundan
    arXiv

2025

  • Intelligence Per Watt: Measuring Intelligence Efficiency of Local AI
    Jon Saad-Falcon, Avanika Narayan, Hakki Orhun Akengin, et al.
    arXiv
  • P/D-Device: Disaggregated Large Language Model between Cloud and Devices
    Yibo Jin, Yixu Xu, Yue Chen, et al.
    arXiv
  • Sometimes Painful but Promising: Feasibility and Trade-Offs of On-Device Language Model Inference
    Maximilian Abstreiter, Sasu Tarkoma, Roberto Morabito
    arXiv DOI
  • [ICLR'25] PalmBench: A Comprehensive Benchmark of Compressed Large Language Models on Mobile Platforms
    Yilong Li, Jingyu Liu, Hao Zhang, et al.
    arXiv Publication
  • [SEC'25] lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
    Haoxin Wang, Xiaolong Tu, Hongyu Ke, et al.
    DOI

2024

  • Large Language Model Performance Benchmarking on Mobile Platforms: A Thorough Evaluation
    Jie Xiao, Qianyi Huang, Xu Chen, et al.
    arXiv Publication
  • [EdgeFM @ MobiSys'24] Large Language Models on Mobile Devices: Measurements, Analysis, and Insights
    Xiang Li, Zhenyan Lu, Dongqi Cai, et al.
    DOI
  • MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
    Rithesh Murthy, Liangwei Yang, Juntao Tan, et al.
    arXiv
  • [MobiCom'24] MELTing point: Mobile Evaluation of Language Transformers
    Stefanos Laskaridis, Kleomenis Katevas, Lorenzo Minto, et al.
    arXiv DOI Talk Code

Mobile-Specific Optimisations

This section focuses on techniques and optimisations that target mobile-specific deployment.

Papers

2026

  • MobileMoE: Scaling On-Device Mixture of Experts
    Yanbei Chen et al.
    arXiv
  • Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
    Jinghe Zhang et al.
    arXiv
  • [MobiSys'26] ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
    Wangsong Yin, Daliang Xu, Mengwei Xu, et al.
    arXiv DOI
  • Efficient On-Device Diffusion LLM Inference with Mobile NPU
    Tuowei Wang, Yanfan Sun, Ju Ren
    arXiv

2025

  • [NeurIPS'25] Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
    Yonggan Fu, Xin Dong, Shizhe Diao, et al.
    arXiv Publication
  • [MobiCom '25] Elastic On-Device LLM Service
    Wangsong Yin, Rongjie Yi, Daliang Xu, et al.
    arXiv DOI
  • [MobiCom '25] Confidant: Customizing Transformer-based LLMs via Collaborative Training on Mobile Devices
    Yuhao Chen, Yuxuan Yan, Shuowei Ge, et al.
    DOI
  • [MobiCom '25] D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
    Haodong Wang, Qihua Zhou, Zicong Hong, et al.
    arXiv DOI
  • [CVPR'25 EDGE Workshop] Scaling On-Device GPU Inference for Large Generative Models
    Jiuqiang Tang, Raman Sarokin, Ekaterina Ignasheva, et al.
    arXiv Publication
  • ROMA: a Read-Only-Memory-based Accelerator for QLoRA-based On-Device LLM
    Liang Li, Xingke Yang, Wen Wu, et al.
    arXiv
  • [ASPLOS'25] Fast On-device LLM Inference with NPUs
    Daliang Xu, Hao Zhang, Liming Yang, et al.
    arXiv DOI Code

2024

  • Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
    Andrii Skliar, Ties van Rozendaal, Romain Lepert, et al.
    arXiv
  • PhoneLM: An Efficient and Capable Small Language Model Family through Principled Pre-training
    Rongjie Yi, Xiang Li, Weikai Xie, et al.
    arXiv Code
  • MobileQuant: Mobile-friendly Quantization for On-device Language Models
    Fuwen Tan, Royson Lee, Łukasz Dudziak, et al.
    arXiv Code
  • Gemma 2: Improving Open Language Models at a Practical Size
    Gemma Team, Morgane Riviere, Shreya Pathak, et al.
    arXiv Code
  • Apple Intelligence Foundation Language Models
    Tom Gunter, Zirui Wang, Chong Wang, et al.
    arXiv
  • EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Layerwise Unified Compression and Adaptive Layer Tuning and Voting
    Zhongzhi Yu, Zheng Wang, Yuhan Li, et al.
    arXiv Code
  • Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
    Marah Abdin, Jyoti Aneja, Hany Awadalla, et al.
    arXiv Code
  • Transformer-Lite: High-efficiency Deployment of Large Language Models on Mobile Phone GPUs
    Luchang Li, Sheng Qian, Jie Lu, et al.
    arXiv
  • Gemma: Open Models Based on Gemini Research and Technology
    Gemma Team, Google DeepMind
    Paper Code
  • MobiLlama: Towards Accurate and Lightweight Fully Transparent GPT
    Omkar Thawakar, Ashmal Vayani, Salman Khan, et al.
    arXiv Code
  • [ICML'24] MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
    Zechun Liu, Changsheng Zhao, Forrest Iandola, et al.
    arXiv Publication Code
  • [ICML'24] Rethinking Optimization and Architecture for Tiny Language Models
    Yehui Tang, Kai Han, Fangcheng Liu, et al.
    arXiv Publication Code
  • TinyLlama: An Open-Source Small Language Model
    Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, et al.
    arXiv Code

Applications

Papers

2024

  • Octopus v3: Technical Report for On-device Sub-billion Multimodal AI Agent
    Wei Chen, Zhiyuan Li
    arXiv
  • Octopus v2: On-device language model for super agent
    Wei Chen, Zhiyuan Li
    arXiv
  • Octopus: On-device language model for function calling of software APIs
    Wei Chen, Zhiyuan Li, Mingyuan Ma
    arXiv Hugging Face

2023

  • Revolutionizing Mobile Interaction: Enabling a 3 Billion Parameter GPT LLM on Mobile
    Samuel Carreira, Tomas Marques, Jose Ribeiro, Carlos Grilo
    arXiv
  • Towards an On-device Agent for Text Rewriting
    Yun Zhu, Yinxiao Liu, Felix Stahlberg, et al.
    arXiv

Multimodal LLMs

This section refers to multimodal LLMs, which integrate vision or other modalities in their tasks.

Papers

2026

  • Small Vision-Language Models are Smart Compressors for Long Video Understanding
    Junjie Fei et al.
    arXiv

2024

  • [CVPR 2024] MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
    Vasu, Pavan Kumar Anasosalu, Pouransari, Hadi, Faghri, Fartash, et al.
    CVF
  • TinyLLaVA: A Framework of Small-scale Large Multimodal Models
    Baichuan Zhou, Ying Hu, Xi Weng, et al.
    arXiv Code
  • MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
    Xiangxiang Chu, Limeng Qiao, Xinyu Zhang, et al.
    arXiv Code

2023

  • MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
    Xiangxiang Chu, Limeng Qiao, Xinyang Lin, et al.
    arXiv Code

Surveys on Efficient LLMs

This section includes survey papers on LLM efficiency, a topic very much related to deploying in constrained devices.

Papers

2025

  • GenAI at the Edge: Comprehensive Survey on Empowering Edge Devices
    Mozhgan Navardi, Romina Aalishah, Yuzhe Fu, et al.
    arXiv Publication
  • Demystifying Small Language Models for Edge Deployment
    Zhenyan Lu, Xiang Li, Dongqi Cai, et al.
    ACL DOI
  • Small Language Models (SLMs) Can Still Pack a Punch: A survey
    Shreyas Subramanian, Vikram Elango, Mecit Gungor
    arXiv

2024

  • A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
    Fali Wang, Zhiwei Zhang, Xianren Zhang, et al.
    arXiv
  • Small Language Models: Survey, Measurements, and Insights
    Zhenyan Lu, Xiang Li, Dongqi Cai, et al.
    arXiv
  • On-Device Language Models: A Comprehensive Review
    Jiajun Xu, Zhiyuan Li, Wei Chen, et al.
    arXiv
  • A Survey of Resource-efficient LLM and Multimodal Foundation Models
    Mengwei Xu, Wangsong Yin, Dongqi Cai, et al.
    arXiv

2023

  • Efficient Large Language Models: A Survey
    Zhongwei Wan, Xin Wang, Che Liu, et al.
    arXiv Code
  • Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
    Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, et al.
    arXiv
  • A Survey on Model Compression for Large Language Models
    Xunyu Zhu, Jian Li, Yong Liu, et al.
    arXiv

Training LLMs on Device

This section refers to papers attempting to train/fine-tune LLMs on device, in a standalone or federated manner.

Papers

2026

  • [MobiSys'26] FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs
    Kahou Tam, Wei Niu, Yu Bao, et al.
    arXiv DOI
  • Online-SDFT: Fine-Tuning Small Language Models for Continual Learning On-Device with Self-Distillation
    I-Ju Lin, Zhang-Wei Hong
    Blog Post Code

2025

  • Computational Bottlenecks of Training Small-scale Large Language Models
    Saleh Ashkboos, Iman Mirzadeh, Keivan Alizadeh, et al.
    arXiv
  • [ICML'25] On-device collaborative language modeling via a mixture of generalists and specialists
    Dongyang Fan, Bettina Messmer, Nikita Doikov, et al.
    arXiv Publication
  • MobiLLM: Enabling LLM Fine-Tuning on the Mobile Device via Server Assisted Side Tuning
    Liang Li, Xingke Yang, Wen Wu, et al.
    arXiv

2024

  • [Privacy in Natural Language Processing @ ACL'24] PocketLLM: Enabling On-Device Fine-Tuning for Personalized LLMs
    Dan Peng, Zhihui Fu
    ACL

2023

  • [MobiCom'23] Federated Few-Shot Learning for Mobile NLP
    Dongqi Cai, Shangguang Wang, Yaozong Wu, et al.
    arXiv DOI Code
  • FwdLLM: Efficient FedLLM using Forward Gradient
    Mengwei Xu, Dongqi Cai, Yaozong Wu, et al.
    arXiv Code
  • [Electronics'24] Forward Learning of Large Language Models by Consumer Devices
    Danilo Pietro Pau, Fabrizio Maria Aymone
    Paper
  • Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
    Herbert Woisetschläger, Alexander Isenko, Shiqiang Wang, et al.
    arXiv
  • Federated Full-Parameter Tuning of Billion-Sized Language Models with Communication Cost under 18 Kilobytes
    Zhen Qin, Daoyuan Chen, Bingchen Qian, et al.
    arXiv Code

This section includes paper that are mobile-related, but not necessarily run on device.

Papers

2026

  • Xiaomi-GUI-0 Technical Report
    Wanxia Cao, Chengzhen Duan, Pei Fu, et al.
    arXiv
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management
    Guangyi Liu, Gao Wu, Congxiao Liu, et al.
    arXiv
  • Beyond the GUI Paradigm: Do Mobile Agents Need the Phone Screen?
    Li Gu, Zihuan Jiang, Linqiang Guo, et al.
    arXiv
  • CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents
    Siyu Shen, Fenghao Xu, Wenrui Diao, et al.
    arXiv
  • iOSWorld: A Benchmark for Personally Intelligent Phone Agents
    Lawrence Keunho Jang, Mareks Woodside, Geronimo Carom, et al.
    arXiv
  • MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
    Ziyun Zeng, Hang Hua, Bocheng Zou, et al.
    arXiv
  • How Mobile World Model Guides GUI Agents?
    Weikai Xu, Kun Huang, Yunren Feng, et al.
    arXiv
  • Mem-W: Latent Memory-Native GUI Agents
    Guibin Zhang, Yaohui Ling, Fanci Meng, et al.
    arXiv
  • ClawMobile: Rethinking Smartphone-Native Agentic Systems
    Lepeng Zhao, Zhenhua Zou, Shuo Li, et al.
    arXiv
  • Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
    Haiyang Xu, Xi Zhang, Haowei Liu, et al.
    arXiv
  • Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
    Hongchao Du, Shangyu Wu, Qiao Li, et al.
    arXiv
  • MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
    Guangyi Liu, Pengxiang Zhao, Yaozhen Liang, et al.
    arXiv

2025

  • Slm-mux: Orchestrating small language models for reasoning
    Chenyu Wang, Zishen Wan, Hao Kang, et al.
    arXiv
  • Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
    Zhen Yang, Zi-Yi Dou, Di Feng, et al.
    arXiv
  • [NeurIPS'25] OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
    Ramchalam Kinattinkara Ramakrishnan, Zhaocong Yuan, Shaojie Zhuo, et al.
    arXiv Publication
  • Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement
    Xuechen Zhang, Zijian Huang, Chenshun Ni, et al.
    arXiv
  • Small Language Models are the Future of Agentic AI
    Peter Belcak, Greg Heinrich, Shizhe Diao, et al.
    arXiv

2024

  • Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
    Junyang Wang, Haiyang Xu, Haitao Jia, et al.
    arXiv Code Demo
  • Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
    Keen You, Haotian Zhang, Eldon Schoop, et al.
    arXiv
  • Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
    Junyang Wang, Haiyang Xu, Jiabo Ye, et al.
    arXiv Code
  • [MobiCom'24] MobileGPT: Augmenting LLM with Human-like App Memory for Mobile Task Automation
    Sunjae Lee, Junyoung Choi, Jungjae Lee, et al.
    arXiv DOI
  • [MobiCom'24] AutoDroid: LLM-powered Task Automation in Android
    Hao Wen, Yuanchun Li, Guohong Liu, et al.
    arXiv DOI Code

2023

  • [NeurIPS'23] AndroidInTheWild: A Large-Scale Dataset For Android Device Control
    Christopher Rawles, Alice Li, Daniel Rodriguez, et al.
    arXiv Publication Code
  • GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
    An Yan, Zhengyuan Yang, Wanrong Zhu, et al.
    arXiv Code

Older

  • [ACL'20] Mapping Natural Language Instructions to Mobile UI Action Sequences
    Yang Li, Jiacong He, Xin Zhou, et al.
    arXiv Publication

Benchmarks

Leaderboards

Books and Courses

Industry Announcements

If you want to read more about related topics, here are some tangential awesome repositories to visit:

Contribute

Contributions welcome! Read the contribution guidelines first.

Contributors

stevelaskaridis

146 commits

iakashpaul

1 commits

OldCat-CN

1 commits

stevelaskaridis/awesome-mobile-llm

Awesome Mobile LLMs

403

150 commits

updated Sep 3, 2026

See the code

README

Awesome Mobile LLMs Awesome

A curated list of LLMs and related studies targeted at mobile and embedded hardware

Last update: 3rd September 2026

If your publication/work is not included - and you think it should - please open an issue or reach out directly to @stevelaskaridis.

Let's try to make this list as useful as possible to researchers, engineers and practitioners all around the world.

Contents

Mobile-First LLMs

The following Table shows sub-3B models designed for on-device deployments, sorted by year.

NameYearSizesPrimary Group/AffiliationPublicationCode RepositoryHF Repository
2026
MobileMoE20260.3B, 0.5B, 0.9B active (1.3B, 2.8B, 5.3B total)Meta AIpaper--
Gemma 42026E2B, E4B, 26BGoogle DeepMindwebsitecodehuggingface
LFM2.52026350M, 1.2B, 1.5B, 1.6B, 2.6BLiquid AIwebsite, LFM2.5-2.6B blog-huggingface
MobileLLM-Flash2026350M, 650M, 1.4BMetapaper--
Apertus Mini20260.5B, 1.5B, 4BSwiss AI (EPFL, ETH Zurich, CSCS)papercodehuggingface
Qwen-3.520260.8B, 2B, ...Qwen Teamblogcodehuggingface
2025
LFM22025350M, 700M, 1.2B, 2.6B, 8.3B (1.5B active)Liquid AIpaper, website-huggingface
MobileLLM-R1.52025140M, 360M, 950MMetapapercodehuggingface
Nemotron-Flash20251B, 3BNvidiapaper, NeurIPS'25-huggingface
MobileLLM-Pro20251BMetapaper-huggingface
MobileLLM-R12025140M, 360M, 950MMetapapercodehuggingface
SmolLM320253BHuggingFaceblogcodehuggingface
Gemma 320251B, 4B, ...Google DeepMindpapercodehuggingface
Qwen-320250.6B, 1.7B, ...Qwen Teampapercodehuggingface
Pareto-Q2025125M, 350M, 600M, 1B, 1.5B, 3BMetapapercodehuggingface
2024
BlueLM-V20242.7BCUHK, Vivo AI Labpapercode-
PhoneLM20240.5B, 1.5BBUPTpapercodehuggingface
AMD-Llama-135m2024135MAMDblogcodehuggingface
SmolLM22024135M, 360M, 1.7BHuggingface-codehuggingface
Ministral20243B, ...Mistralblog-huggingface
Llama 3.220241B, 3BMetablogcodehuggingface
OLMoE20247B (1B active)AllenAIpapercodehuggingface
Spectra202499M - 3.9BNolanoAIpapercodehuggingface
Gemma 220242B, ...Googlepaper blogcodehuggingface
Apple Intelligence Foundation LMs20243BApplepaper--
SmolLM2024135M, 360M, 1.7BHuggingfaceblog-huggingface
Fox20241.6BTensorOperablog-huggingface
Qwen22024500M, 1.5B, ...Qwen Teampapercodehuggingface
OpenELM2024270M, 450M, 1.08B, 3.04BApplepapercodehuggingface
DCLM2024400M, 1B, ...Univerisy of Washington, Apple, Toyota Research Institute, ...papercodehuggingface
Phi-320243.8BMicrosoftwhitepapercodehuggingface
BitNet-b1.5820241.3B, 3B, ...Microsoftpapercodehuggingface
OLMo20241B, ...AllenAIpapercodehuggingface
Mobile LLMs2024125M, 250MMetapaper, ICML'24code-
Gemma20242B, ...Googlepaper, websitecode, gemma.cpphuggingface
MobiLlama20240.5B, 1BMBZUAIpapercodehuggingface
Stable LM 2 (Zephyr)20241.6BStability.aipaper-huggingface
TinyLlama20241.1BSingapore University of Technology and Designpapercodehuggingface
Gemini-Nano20241.8B, 3.25BGooglepaper--
2023
Stable LM (Zephyr)20233BStabilityblogcodehuggingface
OpenLM202311M, 25M, 87M, 160M, 411M, 830M, 1B, 3B, ...OpenLM team-codehuggingface
Phi-220232.7BMicrosoftwebsite-huggingface
Phi-1.520231.3BMicrosoftpaper-huggingface
Phi-120231.3BMicrosoftpaper-huggingface
RWKV2023169M, 430M, 1.5B, 3B, ...EleutherAIpapercodehuggingface
Cerebras-GPT2023111M, 256M, 590M, 1.3B, 2.7B ...Cerebraspapercodehuggingface
OPT2022125M, 350M, 1.3B, 2.7B, ...Metapapercodehuggingface
LaMini-LM202361M, 77M, 111M, 124M, 223M, 248M, 256M, 590M, 774M, 738M, 783M, 1.3B, 1.5B, ...MBZUAIpapercodehuggingface
Pythia202370M, 160M, 410M, 1B, 1.4B, 2.8B, ...EleutherAIpapercodehuggingface
2022
Galactica2022125M, 1.3B, ...Metapapercodehuggingface
BLOOM2022560M, 1.1B, 1.7B, 3B, ...BigSciencepapercodehuggingface
2021
XGLM2021564M, 1.7B, 2.9B, ...Metapapercodehuggingface
GPT-Neo2021125M, 350M, 1.3B, 2.7BEleutherAI-code, gpt-neoxhuggingface
2020
MobileBERT202015.1M, 25.3MCMU, Googlepapercodehuggingface
2019
BART2019140M, 400MMetapapercodehuggingface
DistilBERT201966MHuggingFacepapercodehuggingface
T5201960M, 220M, 770M, 3B, ...Googlepapercodehuggingface
TinyBERT201914.5MHuaweipapercodehuggingface
Megatron-LM2019336M, 1.3B, ...Nvidiapapercode-

Infrastructure / Deployment of LLMs on Device

This section showcases frameworks and contributions for supporting LLM inference on mobile and edge devices.

Deployment Frameworks

On-Device Inference Frameworks

These frameworks are primarily used to run models directly on-device, inside mobile apps, edge deployments, or tightly integrated local runtimes.

  • llama.cpp: Inference of Meta's LLaMA model (and others) in pure C/C++. Supports various platforms and builds on top of ggml (now gguf format).
    • LLMFarm: iOS frontend for llama.cpp
    • LLM.swift: iOS frontend for llama.cpp
    • Sherpa: Android frontend for llama.cpp
    • iAkashPaul/Portal: Wraps the example android app with tweaked UI, configs & additional model support
    • dusty-nv's llama.cpp: Containers for Jetson deployment of llama.cpp
    • Off Grid: Open-source React Native app for on-device LLM chat, vision models (SmolVLM, LLaVA), and Stable Diffusion image generation on iOS & Android.
    • Airgap: Open-source React Native framework for on-device, offline-first customer support chatbots. Runs Gemma 4 E2B locally via llama.rn. Seven industry templates (telco, retail, healthcare, banking, education, insurance, airlines) ship in the repo.
  • MLC-LLM: MLC LLM is a machine learning compiler and high-performance deployment engine for large language models. Supports various platforms and build on top of TVM.
  • PyTorch ExecuTorch: Solution for enabling on-device inference capabilities across mobile and edge devices including wearables, embedded devices and microcontrollers.
    • TorchChat: Codebase showcasing the ability to run large language models (LLMs) seamlessly across iOS and Android
  • Google MediaPipe: A suite of libraries and tools for you to quickly apply artificial intelligence (AI) and machine learning (ML) techniques in your applications. Support Android, iOS, Python and Web.
    • GoogleAI-Edge Gallery: Experimental app that puts the power of cutting-edge Generative AI models directly into your hands, running entirely on your Android and iOS devices.
  • Apple MLX: MLX is an array framework for machine learning research on Apple silicon, brought to you by Apple machine learning research. Builds upon lazy evaluation and unified memory architecture.
  • Apple Foundation Models SDK: Python bindings for Apple's Foundation Models framework, providing access to the on-device foundation model at the core of Apple Intelligence on macOS.
  • HF Swift Transformers: Swift Package to implement a transformers-like API in Swift
  • Alibaba MNN: MNN supports inference and training of deep learning models and for inference and training on-device.
  • llama2.c (More educational, see here for android port)
  • tinygrad: Simple neural network framework from tinycorp and @geohot
  • TinyChatEngine: Targeted at Nvidia, Apple M1 and RPi, from Song Han's (MIT) group.
  • Llama Stack (swift, kotlin): These libraries are a set of SDKs that provide a simple and effective way to integrate AI capabilities into your iOS/Android app, whether it is local (on-device) or remote inference.
  • OLMoE.Swift: Ai2 OLMoE is an AI chatbot powered by the OLMoE model. Unlike cloud-based AI assistants, OLMoE runs entirely on your device, ensuring complete privacy and offline accessibility—even in Flight Mode.
  • HuggingSnap: HuggingSnap is an iOS app that lets users quickly learn more about the places and objects around them. HuggingSnap runs SmolVLM2, a compact open multimodal model that accepts arbitrary sequences of image, videos, and text inputs to produce text outputs.
  • Flower Intelligence: Flower Intelligence is a cross-platform inference library that lets users seamlessly interact with Large-Language Models both locally and remotely in a secure and private way. The library was created by the Flower Labs team. It supports TypeScript, JavaScript and Swift backends.
  • ONNX Runtime: Cross-platform inference and training engine for ONNX models, with a Mobile package and execution providers (NNAPI, Core ML, XNNPACK, QNN) for on-device deployment on Android and iOS. See the Mobile deployment guide.

Local Network Model Serving

These frameworks are primarily used to host models on a laptop, desktop, or workstation and expose them over a local API to other devices on the same LAN.

  • LM Studio: Desktop application and local inference server for hosting models on your machine, with an OpenAI-compatible local API.
  • Ollama: Local model runner and server for hosting and serving models through a simple CLI and HTTP API.
  • Lemonade: Open-source local AI server for text, image, and speech workloads, designed to run privately on local PCs and compatible with OpenAI-style APIs.
  • llama.cpp: Can also be used as a lightweight local inference server for hosting GGUF models via CLI and HTTP server modes.
  • LocalAI: Self-hosted local inference server and OpenAI-compatible REST API for running LLM, vision, image, and audio workloads on local or on-prem hardware.
  • Locally AI: Native Apple-platform app for running AI models fully offline on iPhone, iPad, and Mac, optimized for Apple Silicon and on-device privacy.
  • vLLM: High-throughput inference and serving engine that can expose OpenAI-compatible local APIs, better suited to stronger desktops and workstations.
  • SGLang: High-performance model serving framework for local and distributed deployments, designed for low-latency and high-throughput inference.

Papers

2026

  • [SenSys'26] An Efficient Context Management System for On-Device LLMaaS
    Wangsong Yin et al.
    DOI
  • FlexServe: A Fast and Secure LLM Serving System for Mobile Devices with Flexible Resource Isolation
    Yinpeng Wu, Yitong Chen, Lixiang Wang, et al.
    arXiv

2025

  • Apple Intelligence Foundation Language Models: Tech Report 2025
    Ethan Li, Anders Boesen Lindbo Larsen, Chen Zhang, et al.
    arXiv
  • [ACM Queue] Generative AI at the Edge: Challenges and Opportunities: The next phase in AI deployment
    Vijay Janapa Reddi
    DOI

2024

  • PowerInfer-2: Fast Large Language Model Inference on a Smartphone
    Zhenliang Xue, Yixin Song, Zeyu Mi, et al.
    arXiv Code
  • [MobiCom'24] Mobile Foundation Model as Firmware
    Jinliang Yuan, Chen Yang, Dongqi Cai, et al.
    Paper DOI Code
  • Merino: Entropy-driven Design for Generative Language Models on IoT Devicess
    Youpeng Zhao, Ming Lin, Huadong Tang, et al.
    arXiv
  • LLM as a System Service on Mobile Devices
    Wangsong Yin, Mengwei Xu, Yuanchun Li, et al.
    arXiv

2023

  • LinguaLinked: A Distributed Large Language Model Inference System for Mobile Devices
    Junchen Zhao, Yurun Song, Simeng Liu, et al.
    arXiv
  • LLMCad: Fast and Scalable On-device Large Language Model Inference
    Daliang Xu, Wangsong Yin, Xin Jin, et al.
    arXiv
  • EdgeMoE: Fast On-Device Inference of MoE-based Large Language Models
    Rongjie Yi, Liwei Guo, Shiyun Wei, et al.
    arXiv

2022

  • [IEEE Pervasive Computing] The Future of Consumer Edge-AI Computing
    Stefanos Laskaridis, Stylianos I. Venieris, Alexandros Kouris, et al.
    arXiv Talk

Benchmarking LLMs on Device

This section focuses on measurements and benchmarking efforts for assessing LLM performance when deployed on device.

Papers

2026

  • LLM Inference at the Edge: Mobile, NPU, and GPU Performance Efficiency Trade-offs Under Sustained Load
    Pranay Tummalapalli, Sahil Arayakandy, Ritam Pal, Kautuk Kundan
    arXiv

2025

  • Intelligence Per Watt: Measuring Intelligence Efficiency of Local AI
    Jon Saad-Falcon, Avanika Narayan, Hakki Orhun Akengin, et al.
    arXiv
  • P/D-Device: Disaggregated Large Language Model between Cloud and Devices
    Yibo Jin, Yixu Xu, Yue Chen, et al.
    arXiv
  • Sometimes Painful but Promising: Feasibility and Trade-Offs of On-Device Language Model Inference
    Maximilian Abstreiter, Sasu Tarkoma, Roberto Morabito
    arXiv DOI
  • [ICLR'25] PalmBench: A Comprehensive Benchmark of Compressed Large Language Models on Mobile Platforms
    Yilong Li, Jingyu Liu, Hao Zhang, et al.
    arXiv Publication
  • [SEC'25] lm-Meter: Unveiling Runtime Inference Latency for On-Device Language Models
    Haoxin Wang, Xiaolong Tu, Hongyu Ke, et al.
    DOI

2024

  • Large Language Model Performance Benchmarking on Mobile Platforms: A Thorough Evaluation
    Jie Xiao, Qianyi Huang, Xu Chen, et al.
    arXiv Publication
  • [EdgeFM @ MobiSys'24] Large Language Models on Mobile Devices: Measurements, Analysis, and Insights
    Xiang Li, Zhenyan Lu, Dongqi Cai, et al.
    DOI
  • MobileAIBench: Benchmarking LLMs and LMMs for On-Device Use Cases
    Rithesh Murthy, Liangwei Yang, Juntao Tan, et al.
    arXiv
  • [MobiCom'24] MELTing point: Mobile Evaluation of Language Transformers
    Stefanos Laskaridis, Kleomenis Katevas, Lorenzo Minto, et al.
    arXiv DOI Talk Code

Mobile-Specific Optimisations

This section focuses on techniques and optimisations that target mobile-specific deployment.

Papers

2026

  • MobileMoE: Scaling On-Device Mixture of Experts
    Yanbei Chen et al.
    arXiv
  • Quant.npu: Enabling Efficient Mobile NPU Inference for on-device LLMs via Fully Static Quantization
    Jinghe Zhang et al.
    arXiv
  • [MobiSys'26] ShadowNPU: System and Algorithm Co-design for NPU-Centric On-Device LLM Inference
    Wangsong Yin, Daliang Xu, Mengwei Xu, et al.
    arXiv DOI
  • Efficient On-Device Diffusion LLM Inference with Mobile NPU
    Tuowei Wang, Yanfan Sun, Ju Ren
    arXiv

2025

  • [NeurIPS'25] Nemotron-Flash: Towards Latency-Optimal Hybrid Small Language Models
    Yonggan Fu, Xin Dong, Shizhe Diao, et al.
    arXiv Publication
  • [MobiCom '25] Elastic On-Device LLM Service
    Wangsong Yin, Rongjie Yi, Daliang Xu, et al.
    arXiv DOI
  • [MobiCom '25] Confidant: Customizing Transformer-based LLMs via Collaborative Training on Mobile Devices
    Yuhao Chen, Yuxuan Yan, Shuowei Ge, et al.
    DOI
  • [MobiCom '25] D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
    Haodong Wang, Qihua Zhou, Zicong Hong, et al.
    arXiv DOI
  • [CVPR'25 EDGE Workshop] Scaling On-Device GPU Inference for Large Generative Models
    Jiuqiang Tang, Raman Sarokin, Ekaterina Ignasheva, et al.
    arXiv Publication
  • ROMA: a Read-Only-Memory-based Accelerator for QLoRA-based On-Device LLM
    Liang Li, Xingke Yang, Wen Wu, et al.
    arXiv
  • [ASPLOS'25] Fast On-device LLM Inference with NPUs
    Daliang Xu, Hao Zhang, Liming Yang, et al.
    arXiv DOI Code

2024

  • Mixture of Cache-Conditional Experts for Efficient Mobile Device Inference
    Andrii Skliar, Ties van Rozendaal, Romain Lepert, et al.
    arXiv
  • PhoneLM: An Efficient and Capable Small Language Model Family through Principled Pre-training
    Rongjie Yi, Xiang Li, Weikai Xie, et al.
    arXiv Code
  • MobileQuant: Mobile-friendly Quantization for On-device Language Models
    Fuwen Tan, Royson Lee, Łukasz Dudziak, et al.
    arXiv Code
  • Gemma 2: Improving Open Language Models at a Practical Size
    Gemma Team, Morgane Riviere, Shreya Pathak, et al.
    arXiv Code
  • Apple Intelligence Foundation Language Models
    Tom Gunter, Zirui Wang, Chong Wang, et al.
    arXiv
  • EDGE-LLM: Enabling Efficient Large Language Model Adaptation on Edge Devices via Layerwise Unified Compression and Adaptive Layer Tuning and Voting
    Zhongzhi Yu, Zheng Wang, Yuhan Li, et al.
    arXiv Code
  • Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
    Marah Abdin, Jyoti Aneja, Hany Awadalla, et al.
    arXiv Code
  • Transformer-Lite: High-efficiency Deployment of Large Language Models on Mobile Phone GPUs
    Luchang Li, Sheng Qian, Jie Lu, et al.
    arXiv
  • Gemma: Open Models Based on Gemini Research and Technology
    Gemma Team, Google DeepMind
    Paper Code
  • MobiLlama: Towards Accurate and Lightweight Fully Transparent GPT
    Omkar Thawakar, Ashmal Vayani, Salman Khan, et al.
    arXiv Code
  • [ICML'24] MobileLLM: Optimizing Sub-billion Parameter Language Models for On-Device Use Cases
    Zechun Liu, Changsheng Zhao, Forrest Iandola, et al.
    arXiv Publication Code
  • [ICML'24] Rethinking Optimization and Architecture for Tiny Language Models
    Yehui Tang, Kai Han, Fangcheng Liu, et al.
    arXiv Publication Code
  • TinyLlama: An Open-Source Small Language Model
    Peiyuan Zhang, Guangtao Zeng, Tianduo Wang, et al.
    arXiv Code

Applications

Papers

2024

  • Octopus v3: Technical Report for On-device Sub-billion Multimodal AI Agent
    Wei Chen, Zhiyuan Li
    arXiv
  • Octopus v2: On-device language model for super agent
    Wei Chen, Zhiyuan Li
    arXiv
  • Octopus: On-device language model for function calling of software APIs
    Wei Chen, Zhiyuan Li, Mingyuan Ma
    arXiv Hugging Face

2023

  • Revolutionizing Mobile Interaction: Enabling a 3 Billion Parameter GPT LLM on Mobile
    Samuel Carreira, Tomas Marques, Jose Ribeiro, Carlos Grilo
    arXiv
  • Towards an On-device Agent for Text Rewriting
    Yun Zhu, Yinxiao Liu, Felix Stahlberg, et al.
    arXiv

Multimodal LLMs

This section refers to multimodal LLMs, which integrate vision or other modalities in their tasks.

Papers

2026

  • Small Vision-Language Models are Smart Compressors for Long Video Understanding
    Junjie Fei et al.
    arXiv

2024

  • [CVPR 2024] MobileCLIP: Fast Image-Text Models through Multi-Modal Reinforced Training
    Vasu, Pavan Kumar Anasosalu, Pouransari, Hadi, Faghri, Fartash, et al.
    CVF
  • TinyLLaVA: A Framework of Small-scale Large Multimodal Models
    Baichuan Zhou, Ying Hu, Xi Weng, et al.
    arXiv Code
  • MobileVLM V2: Faster and Stronger Baseline for Vision Language Model
    Xiangxiang Chu, Limeng Qiao, Xinyu Zhang, et al.
    arXiv Code

2023

  • MobileVLM : A Fast, Strong and Open Vision Language Assistant for Mobile Devices
    Xiangxiang Chu, Limeng Qiao, Xinyang Lin, et al.
    arXiv Code

Surveys on Efficient LLMs

This section includes survey papers on LLM efficiency, a topic very much related to deploying in constrained devices.

Papers

2025

  • GenAI at the Edge: Comprehensive Survey on Empowering Edge Devices
    Mozhgan Navardi, Romina Aalishah, Yuzhe Fu, et al.
    arXiv Publication
  • Demystifying Small Language Models for Edge Deployment
    Zhenyan Lu, Xiang Li, Dongqi Cai, et al.
    ACL DOI
  • Small Language Models (SLMs) Can Still Pack a Punch: A survey
    Shreyas Subramanian, Vikram Elango, Mecit Gungor
    arXiv

2024

  • A Comprehensive Survey of Small Language Models in the Era of Large Language Models: Techniques, Enhancements, Applications, Collaboration with LLMs, and Trustworthiness
    Fali Wang, Zhiwei Zhang, Xianren Zhang, et al.
    arXiv
  • Small Language Models: Survey, Measurements, and Insights
    Zhenyan Lu, Xiang Li, Dongqi Cai, et al.
    arXiv
  • On-Device Language Models: A Comprehensive Review
    Jiajun Xu, Zhiyuan Li, Wei Chen, et al.
    arXiv
  • A Survey of Resource-efficient LLM and Multimodal Foundation Models
    Mengwei Xu, Wangsong Yin, Dongqi Cai, et al.
    arXiv

2023

  • Efficient Large Language Models: A Survey
    Zhongwei Wan, Xin Wang, Che Liu, et al.
    arXiv Code
  • Towards Efficient Generative Large Language Model Serving: A Survey from Algorithms to Systems
    Xupeng Miao, Gabriele Oliaro, Zhihao Zhang, et al.
    arXiv
  • A Survey on Model Compression for Large Language Models
    Xunyu Zhu, Jian Li, Yong Liu, et al.
    arXiv

Training LLMs on Device

This section refers to papers attempting to train/fine-tune LLMs on device, in a standalone or federated manner.

Papers

2026

  • [MobiSys'26] FBLayout: Optimizing Memory Layout for Efficient LLM Finetuning on Mobile GPUs
    Kahou Tam, Wei Niu, Yu Bao, et al.
    arXiv DOI
  • Online-SDFT: Fine-Tuning Small Language Models for Continual Learning On-Device with Self-Distillation
    I-Ju Lin, Zhang-Wei Hong
    Blog Post Code

2025

  • Computational Bottlenecks of Training Small-scale Large Language Models
    Saleh Ashkboos, Iman Mirzadeh, Keivan Alizadeh, et al.
    arXiv
  • [ICML'25] On-device collaborative language modeling via a mixture of generalists and specialists
    Dongyang Fan, Bettina Messmer, Nikita Doikov, et al.
    arXiv Publication
  • MobiLLM: Enabling LLM Fine-Tuning on the Mobile Device via Server Assisted Side Tuning
    Liang Li, Xingke Yang, Wen Wu, et al.
    arXiv

2024

  • [Privacy in Natural Language Processing @ ACL'24] PocketLLM: Enabling On-Device Fine-Tuning for Personalized LLMs
    Dan Peng, Zhihui Fu
    ACL

2023

  • [MobiCom'23] Federated Few-Shot Learning for Mobile NLP
    Dongqi Cai, Shangguang Wang, Yaozong Wu, et al.
    arXiv DOI Code
  • FwdLLM: Efficient FedLLM using Forward Gradient
    Mengwei Xu, Dongqi Cai, Yaozong Wu, et al.
    arXiv Code
  • [Electronics'24] Forward Learning of Large Language Models by Consumer Devices
    Danilo Pietro Pau, Fabrizio Maria Aymone
    Paper
  • Federated Fine-Tuning of LLMs on the Very Edge: The Good, the Bad, the Ugly
    Herbert Woisetschläger, Alexander Isenko, Shiqiang Wang, et al.
    arXiv
  • Federated Full-Parameter Tuning of Billion-Sized Language Models with Communication Cost under 18 Kilobytes
    Zhen Qin, Daoyuan Chen, Bingchen Qian, et al.
    arXiv Code

This section includes paper that are mobile-related, but not necessarily run on device.

Papers

2026

  • Xiaomi-GUI-0 Technical Report
    Wanxia Cao, Chengzhen Duan, Pei Fu, et al.
    arXiv
  • MemGUI-Agent: An End-to-End Long-Horizon Mobile GUI Agent with Proactive Context Management
    Guangyi Liu, Gao Wu, Congxiao Liu, et al.
    arXiv
  • Beyond the GUI Paradigm: Do Mobile Agents Need the Phone Screen?
    Li Gu, Zihuan Jiang, Linqiang Guo, et al.
    arXiv
  • CAPED: Context-Aware Privacy Exposure Defense for Mobile GUI Agents
    Siyu Shen, Fenghao Xu, Wenrui Diao, et al.
    arXiv
  • iOSWorld: A Benchmark for Personally Intelligent Phone Agents
    Lawrence Keunho Jang, Mareks Woodside, Geronimo Carom, et al.
    arXiv
  • MementoGUI: Learning Agentic Multimodal Memory Control for Long-Horizon GUI Agents
    Ziyun Zeng, Hang Hua, Bocheng Zou, et al.
    arXiv
  • How Mobile World Model Guides GUI Agents?
    Weikai Xu, Kun Huang, Yunren Feng, et al.
    arXiv
  • Mem-W: Latent Memory-Native GUI Agents
    Guibin Zhang, Yaohui Ling, Fanci Meng, et al.
    arXiv
  • ClawMobile: Rethinking Smartphone-Native Agentic Systems
    Lepeng Zhao, Zhenhua Zou, Shuo Li, et al.
    arXiv
  • Mobile-Agent-v3.5: Multi-platform Fundamental GUI Agents
    Haiyang Xu, Xi Zhang, Haowei Liu, et al.
    arXiv
  • Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
    Hongchao Du, Shangyu Wu, Qiao Li, et al.
    arXiv
  • MemGUI-Bench: Benchmarking Memory of Mobile GUI Agents in Dynamic Environments
    Guangyi Liu, Pengxiang Zhao, Yaozhen Liang, et al.
    arXiv

2025

  • Slm-mux: Orchestrating small language models for reasoning
    Chenyu Wang, Zishen Wan, Hao Kang, et al.
    arXiv
  • Ferret-UI Lite: Lessons from Building Small On-Device GUI Agents
    Zhen Yang, Zi-Yi Dou, Di Feng, et al.
    arXiv
  • [NeurIPS'25] OmniDraft: A Cross-vocabulary, Online Adaptive Drafter for On-device Speculative Decoding
    Ramchalam Kinattinkara Ramakrishnan, Zhaocong Yuan, Shaojie Zhuo, et al.
    arXiv Publication
  • Making Small Language Models Efficient Reasoners: Intervention, Supervision, Reinforcement
    Xuechen Zhang, Zijian Huang, Chenshun Ni, et al.
    arXiv
  • Small Language Models are the Future of Agentic AI
    Peter Belcak, Greg Heinrich, Shizhe Diao, et al.
    arXiv

2024

  • Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via Multi-Agent Collaboration
    Junyang Wang, Haiyang Xu, Haitao Jia, et al.
    arXiv Code Demo
  • Ferret-UI: Grounded Mobile UI Understanding with Multimodal LLMs
    Keen You, Haotian Zhang, Eldon Schoop, et al.
    arXiv
  • Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception
    Junyang Wang, Haiyang Xu, Jiabo Ye, et al.
    arXiv Code
  • [MobiCom'24] MobileGPT: Augmenting LLM with Human-like App Memory for Mobile Task Automation
    Sunjae Lee, Junyoung Choi, Jungjae Lee, et al.
    arXiv DOI
  • [MobiCom'24] AutoDroid: LLM-powered Task Automation in Android
    Hao Wen, Yuanchun Li, Guohong Liu, et al.
    arXiv DOI Code

2023

  • [NeurIPS'23] AndroidInTheWild: A Large-Scale Dataset For Android Device Control
    Christopher Rawles, Alice Li, Daniel Rodriguez, et al.
    arXiv Publication Code
  • GPT-4V in Wonderland: Large Multimodal Models for Zero-Shot Smartphone GUI Navigation
    An Yan, Zhengyuan Yang, Wanrong Zhu, et al.
    arXiv Code

Older

  • [ACL'20] Mapping Natural Language Instructions to Mobile UI Action Sequences
    Yang Li, Jiacong He, Xin Zhou, et al.
    arXiv Publication

Benchmarks

Leaderboards

Books and Courses

Industry Announcements

If you want to read more about related topics, here are some tangential awesome repositories to visit:

Contribute

Contributions welcome! Read the contribution guidelines first.

Contributors

stevelaskaridis

146 commits

iakashpaul

1 commits

OldCat-CN

1 commits