jeho-lee/Awesome-On-Device-AI-Systems

186

74 commits

updated Sep 14, 2026

See the code

README

Awesome On-Device AI Systems Awesome

A curated list of efficient on-device AI systems, including practical inference engines, benchmarks, and state-of-the-art research papers for mobile and edge devices.

This repository bridges the gap between Systems Research (academic papers) and Practical Deployment (engineering frameworks), focusing on optimizing ML models (e.g., LLM/VLMs, ViTs, etc.) on resource-constrained hardware.

📂 Table of Contents

🚀 Inference Engines

Frameworks and runtimes designed for deploying models on edge devices.

General ML Workloads

  • LiteRT (formerly TensorFlow Lite) - Google's framework for on-device inference.
  • ExecuTorch - PyTorch’s end-to-end solution for enabling on-device AI.
  • ONNX Runtime - Cross-platform inference engine for ONNX models.
  • MNN - Lightweight deep learning framework by Alibaba.
  • NCNN - High-performance NN inference framework by Tencent.

Vendor-Specific SDKs

  • Qualcomm QNN - Qualcomm AI Stack for Snapdragon NPUs/DSPs.
  • Apple Core ML - Framework to integrate ML models into iOS/macOS apps.
  • NVIDIA TensorRT - SDK for high-performance deep learning inference on NVIDIA GPUs (including Jetson).
  • Intel OpenVINO - Toolkit for optimizing and deploying AI inference on Intel hardware (CPU/GPU/NPU).
  • MediaTek NeuroPilot - AI ecosystem and SDK for MediaTek NPUs.

Task-specific SDKs

  • llama.cpp - LLM inference in C/C++ with minimal dependencies.
  • MLC LLM - Universal solution for deploying LLMs on any hardware (based on TVM).
  • TensorRT-LLM - NVIDIA GPU-optimized LLM inference library, relevant for Jetson-class edge devices.
  • mllm - A fast and lightweight LLM inference engine for mobile and edge devices.
  • MLX LM - LLM inference and fine-tuning toolkit built on MLX for Apple silicon.
  • FluidAudio - Local audio AI SDK for Apple platforms with ASR, speaker diarization, VAD, and TTS optimized for Apple Neural Engine.
  • RunAnywhere - Open-source SDK for running LLMs and multimodal models on-device across iOS, Android, and cross-platform apps.
  • Off Grid - Open-source iOS/Android app running LLMs (Llama, Qwen, Gemma, Phi, DeepSeek) entirely on-device via llama.cpp. Includes voice (whisper.cpp), vision, on-device image generation, and tool calling.

📝 Research Papers

Note: Some of the works are designed for inference acceleration on cloud/server infrastructure, which has much higher computational resources, but I also include them here if they can be potentially generalized to on-device inference use cases.

LLM Inference on Mobile SoCs

Hardware- and Compiler-aware On-device Inference

Attention Acceleration

Quantization/Sparsity

Application-centric On-device AI Systems

Multi-DNN / Heterogeneous Runtime Scheduling

On-device Training, Model Adaptation

Profilers

edge-computing
efficient-ai
machine-learning
mobile-systems
on-device-ai
resource-constrained-devices

Significant stargazers

Tommy Chiang

159 followers · starred Mar 2025

jeho-lee/Awesome-On-Device-AI-Systems

186

74 commits

updated Sep 14, 2026

See the code

README

Awesome On-Device AI Systems Awesome

A curated list of efficient on-device AI systems, including practical inference engines, benchmarks, and state-of-the-art research papers for mobile and edge devices.

This repository bridges the gap between Systems Research (academic papers) and Practical Deployment (engineering frameworks), focusing on optimizing ML models (e.g., LLM/VLMs, ViTs, etc.) on resource-constrained hardware.

📂 Table of Contents

🚀 Inference Engines

Frameworks and runtimes designed for deploying models on edge devices.

General ML Workloads

  • LiteRT (formerly TensorFlow Lite) - Google's framework for on-device inference.
  • ExecuTorch - PyTorch’s end-to-end solution for enabling on-device AI.
  • ONNX Runtime - Cross-platform inference engine for ONNX models.
  • MNN - Lightweight deep learning framework by Alibaba.
  • NCNN - High-performance NN inference framework by Tencent.

Vendor-Specific SDKs

  • Qualcomm QNN - Qualcomm AI Stack for Snapdragon NPUs/DSPs.
  • Apple Core ML - Framework to integrate ML models into iOS/macOS apps.
  • NVIDIA TensorRT - SDK for high-performance deep learning inference on NVIDIA GPUs (including Jetson).
  • Intel OpenVINO - Toolkit for optimizing and deploying AI inference on Intel hardware (CPU/GPU/NPU).
  • MediaTek NeuroPilot - AI ecosystem and SDK for MediaTek NPUs.

Task-specific SDKs

  • llama.cpp - LLM inference in C/C++ with minimal dependencies.
  • MLC LLM - Universal solution for deploying LLMs on any hardware (based on TVM).
  • TensorRT-LLM - NVIDIA GPU-optimized LLM inference library, relevant for Jetson-class edge devices.
  • mllm - A fast and lightweight LLM inference engine for mobile and edge devices.
  • MLX LM - LLM inference and fine-tuning toolkit built on MLX for Apple silicon.
  • FluidAudio - Local audio AI SDK for Apple platforms with ASR, speaker diarization, VAD, and TTS optimized for Apple Neural Engine.
  • RunAnywhere - Open-source SDK for running LLMs and multimodal models on-device across iOS, Android, and cross-platform apps.
  • Off Grid - Open-source iOS/Android app running LLMs (Llama, Qwen, Gemma, Phi, DeepSeek) entirely on-device via llama.cpp. Includes voice (whisper.cpp), vision, on-device image generation, and tool calling.

📝 Research Papers

Note: Some of the works are designed for inference acceleration on cloud/server infrastructure, which has much higher computational resources, but I also include them here if they can be potentially generalized to on-device inference use cases.

LLM Inference on Mobile SoCs

Hardware- and Compiler-aware On-device Inference

Attention Acceleration

Quantization/Sparsity

Application-centric On-device AI Systems

Multi-DNN / Heterogeneous Runtime Scheduling

On-device Training, Model Adaptation

Profilers

edge-computing
efficient-ai
machine-learning
mobile-systems
on-device-ai
resource-constrained-devices

Significant stargazers

Tommy Chiang

159 followers · starred Mar 2025