KAIWEILIUCC/Awesome-LLM-IoT-Papers

A collection of papers on LLM applications in the IoT field.

25

117 commits

updated Jan 21, 2026

See the code

README

Awesome-LLM-IoT-Papers

Awesome

Table of Contents

Surveys

A Survey of Foundation Models for IoT: Taxonomy and Criteria-Based Analysis.
arXiv

The role of Large Language Models in addressing IoT challenges: A systematic literature review.

Foundation Models for CPS-IoT: Opportunities and Challenges.
arXiv

LLMs and IoT: A Comprehensive Survey on Large Language Models and the Internet of Things.

A Review on Edge Large Language Models: Design, Execution, and Applications.

Large Language Models for Network Intrusion Detection Systems: Foundations, Implementations, and Future Directions.
arXiv

Toward Edge General Intelligence via Large Language Models: Opportunities and Challenges.

Large Language Models in Smart Grid: Applications and Risks.

Small Language Models: Survey, Measurements, and Insights.
arXiv

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities.
arXiv

LLM Agents

AutoDroid: LLM-powered Task Automation in Android (MobiCom 2024)
arXiv

MobileGPT: Augmenting LLM with Human-like App Memory for Mobile Task Automation (MobiCom 2024)
arXiv Star

Poster: Enabling Agent-centric Interaction on Smartphones with LLM-based UI Reassembling (MobiSys 2024)

LLMind: Orchestrating AI and IoT with LLM for Complex Task Execution.
arXiv

LLMind 2.0: Distributed IoT Automation with Natural Language M2M Communication and Lightweight LLM Agents.
arXiv

TaskSense: A Translation-like Approach for Tasking Heterogeneous Sensor Systems with LLMs. (SenSys 2025)

Toward Sensor-In-the-Loop LLM Agent: Benchmarks and Implications. (SenSys 2025)

ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions. (NeurIPS 2025)
arXiv

SensorMCP: A Model Context Protocol Server for Custom Sensor Tool Creation. (NetAISys 2025)

Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment.
arXiv

Empowering Agentic Video Analytics Systems with Video Language Models.
arXiv

AutoBridge: Automating Smart Device Integration with Centralized Platform.
arXiv

LLM-based Question-Answer Framework for Sensor-driven HVAC System Interaction.
arXiv

Towards Privacy-Preserving and Personalized Smart Homes via Tailored Small Language Models.
arXiv

UIShift: Enhancing VLM-based GUI Agents through Self-supervised Reinforcement Learning.
arXiv

MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents.
arXiv

MobileViews: A Large-Scale Mobile GUI Dataset.
arXiv

AutoDroid-V2: Boosting SLM-based GUI Agents via Code Generation.
arXiv

LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile Apps.
arXiv

Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents.
arXiv

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection.
arXiv

AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning.
arXiv

Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment.
arXiv

FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents.
arXiv

AgentSense: LLMs Empower Generalizable and Explainable Web-Based Participatory Urban Sensing.
arXiv

IndusGCC: A Data Benchmark and Evaluation Framework for GUI-Based General Computer Control in Industrial Automation.
arXiv

IoT-MCP: Bridging LLMs and IoT Systems Through Model Context Protocol.
arXiv

AsyncVoice Agent: Real-Time Explanation for LLM Planning and Reasoning.
arXiv

VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification.
arXiv

DroidCall: A Dataset for LLM-powered Android Intent Invocation.
arXiv

More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models.
arXiv

BAMAS: Structuring Budget-Aware Multi-Agent Systems.
arXiv

HiveMind: Contribution-Guided Online Prompt Optimization of LLM Multi-Agent Systems.
arXiv

DIMGen: Dynamic Intent Macro Generation for Efficient LLM-Driven Mobile Automation.

NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls.
arXiv

Edge FM

MELTing Point: Mobile Evaluation of Language Transformers (MobiCom 2024)
arXiv

Your Data, Your Model: A Framework for Training and Deploying Foundational Language Models for Embedded Devices. (MobiCom 2024)

Federated Black-box Prompt Tuning System for Large Language Models on the Edge. (MobiCom 2024)

A Framework for Training and Deploying Foundational Language Models for Embedded Sensing. (MobiCom 2024)

EdgeFM: Leveraging Foundation Model for Open-set Learning on the Edge (SenSys 2023)
arXiv

Confidant: Customizing Transformer-based LLMs via Collaborative Training on Mobile Devices (MobiCom 2025)

Dynamic Sparse Attention on Mobile SoCs
arXiv

MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
arXiv

PhoneLM: an Efficient and Capable Small Language Model Family through Principled Pre-training
arXiv

LLM as a System Service on Mobile Devices
arXiv

LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades
arXiv

ELMS: Elasticized Large Language Models On Mobile Devices
arXiv

Serving MoE Models on Resource-constrained Edge Devices via Dynamic Expert Swapping

HAPE: Hardware-Aware LLM Pruning For Efficient On-Device Inference Optimization

Demystifying Small Language Models for Edge Deployment

TinyLLM: A Framework for Training and Deploying Language Models at the Edge Computerss
arXiv

Confidant: Customizing Transformer-based LLMs via Collaborative Edge Training
arXiv

Modality Plug-and-Play: Runtime Modality Adaptation in LLM-Driven Autonomous Mobile Systems

D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
arXiv

Elastic On-Device LLM Service
arXiv

PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
arXiv

Sensor Data Understanding

Improving On-Device LLMs' Sensory Understanding with Embedding Interpolations (MobiCom 2024)

Penetrative AI: Making LLMs Comprehend the Physical World (ACL 2024)
arXiv

Exploring the Capabilities of LLMs for IMU-based Fine-grained Human Activity Understanding (ACM FMSys 2025)
arXiv

ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs. (HotMobile 2025)

SensorBench: Benchmarking LLMs in Coding-Based Sensor Processing (HotMobile 2025)

Making Sensing Interactive and Descriptive with LLMs: Context Reasoning from Multi-Sensor Data (HotMobile 2025)

Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality Alignment
arXiv

Empowering Agentic Video Analytics Systems with Video Language Models
arXiv

SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
arXiv

ChainStream: An LLM-based Framework for Unified Synthetic Sensing
arXiv

MASTER: A Multi-modal Foundation Model for Human Activity Recognition
arXiv

LLM-CoSen: Revisiting Collaborative Sensing With Large Language Models (LLMs)

Sensor Data Generation

High Resolution Millimeter Wave Imaging Based on FMCW Radar Systems at W-Band

SHADE-AD: An LLM-Based Framework for Synthesizing Activity Data of Alzheimer's Patients. (SenSys 2025)

DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs
arXiv

Spider: Any-to-Many Multimodal LLM
arXiv

CCC: cross-modal contrastive creator for end-to-end sign language generation

Code Generation

AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT Applications (MobiCom 2025)
arXiv Star

GPIoT: Tailoring Small Language Models for IoT Program Synthesis and Development (SenSys 2025)
arXiv Star

CheckMate: LLM-Powered Approximate Intermittent Computing (SenSys 2025)

Exploring Code Language Models for Automated HLS-based Hardware Generation: Benchmark, Infrastructure and Analysis

LLM for Complex Signal Processing in FPGA-based Software Defined Radios: A Case Study on FFT

WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning arXiv

DeepFeature: Iterative Context-aware Feature Generation for Wearable Biosignals arXiv

Interesting Applications

TransCompressor: LLM-Powered Multimodal Data Compression for Smart Transportation
arXiv

Llambda: An LLM-Empowered Low-Resolution Vision System for On-Device Human Behavior Understanding
arXiv

IoT-LLM: Enhancing Real-World IoT Task Reasoning with Large Language Models (Arxiv) Paper

LightLLM: A Versatile Large Language Model for Predictive Light Sensing (SenSys 2025)
arXiv

FlexiFly: Interfacing the Physical World with Foundation Models Empowered by Reconfigurable Drone Systems (SenSys 2025)
arXiv

Socialmind: LLM-based proactive ar social assistive system with human-like perception for in-situ live interactions (IMWUT 2025)
arXiv

TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable Platforms (IMWUT 2024)
arXiv

Exploring Foundation Models in Detecting Concerning Daily Functioning in Psychotherapeutic Context Based on Images from Smart Home Devices (ACM FMSys 2024)

Sensor2Scene: Foundation Model-Driven Interactive Realities (ACM FMSys 2024)

See Where You Read with Eye Gaze Tracking and Large Language Model
arXiv

Empower Vision Applications with LoRA LMM
arXiv

Can we make FCC Experts out of LLMs? (HotMobile 2025)

RouteLLM: A Large Language Model with Native Route Context Understanding to Enable Context-Aware Reasoning (IMWUT 2025)

Congestion Control System Optimization with Large Language Models arXiv

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments
arXiv

SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
arXiv

Toward Foundation Models for Online Complex Event Detection in CPS-IoT: A Case Study

RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
arXiv

Large Language Model-Guided Disentangled Belief Representation Learning on Polarized Social Graphs

Decoding the Silent Majority: Inducing Belief Augmented Social Graph with Large Language Model for Response Forecasting
arXiv

SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
arXiv

Leveraging AI Agents for Autonomous Networks: A Reference Architecture and Empirical Studies
arXiv

Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
arXiv

Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
arXiv

Large Language Models for Wireless Communications: From Adaptation to Autonomy
arXiv

ThermiKit: Edge-Optimized LWIR Analytics with Agent-Driven Interactions

All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles
arXiv

Edge-IoT and MLLMs for Education and Scene Understanding: Assisting Vision and Hearing-Impaired Individuals

SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
arXiv

LLM-Assisted IoT Testing: Finding Conformance Bugs in Matter SDKs

CBM-RAG: Demonstrating Enhanced Interpretability in Radiology Report Generation with Multi-Agent RAG and Concept Bottleneck Models arXiv

Multimodal LLM for Patient Activity Recognition: Integrating Video, Audio, and Text in Clinical Environments

MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models arXiv

Reasoning Visual Language Model for Chest X-Ray Analysis arXiv

Smart Health

DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge (IMWUT 2024)
arXiv

LLM-based Conversational AI Therapist for Daily Functioning Screening and Psychotherapeutic Intervention via Everyday Smart Devices
arXiv

AutoLife: Automatic Life Journaling with Smartphones and LLMs
arXiv

MDTeamGPT: A Self-Evolving LLM-based Multi-Agent Framework for Multi-Disciplinary Team Medical Consultation
arXiv

Introduction to the Special Issue on Large Language Models, Conversational Systems, and Generative AI in Healthβ€”Part 1

CataractBot: An LLM-powered Expert-in-the-Loop Chatbot for Cataract Patients

GLOSS: Group of LLMs for Open-ended Sensemaking of Passive Sensing Data for Health and Wellbeing.

Demo: Myotrainer: Muscle-Aware Motion Analysis and Feedback System for In-Home Resistance Training

A Disease-Centric Vision-Language Foundation Model for Precision Oncology in Kidney Cancer
arXiv

DermINO: Hybrid Pretraining for a Versatile Dermatology Foundation Model
arXiv

DynamiCare: A Dynamic Multi-Agent Framework for Interactive and Open-Ended Medical Decision-Making
arXiv

Demo Abstract: An LLM-Powered Multimodal Mobile Sensing System for Personalized and Interactive Health Behavior Analysis

DocCHA: Towards LLM-Augmented Interactive Online diagnosis System
arXiv

MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
arXiv

MedAgentSim: Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions
arXiv

MegaAgent: A Large-Scale Autonomous LLM-based Multi-Agent System Without Predefined SOPs
arXiv

Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
arXiv

Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
arXiv

SynLLM: A Comparative Analysis of Large Language Models for Medical Tabular Synthetic Data Generation via Prompt Engineering
arXiv

Myo-Trainer: A Vision-based Muscle-Aware Motion Feedback System for In-Home Resistance Training

MedAide: Towards an Omni Medical Aide via Specialized LLM-based Multi-Agent Collaboration
arXiv

MedPlan: A Two-Stage RAG-Based System for Personalized Medical Plan Generation
arXiv

Robotics

RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation
arXiv

NaVILA: Legged Robot Vision-Language-Action Model for Navigation
arXiv

Affordances-Oriented Planning using Foundation Models for Continuous Vision-Language Navigation
arXiv

FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph
arXiv

Multi-robot Rigid Formation Navigation via Synchronous Motion and Discrete-time Communication-Control Optimization
arXiv

Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
arXiv

See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation
arXiv

Guide-LLM: An Embodied LLM Agent and Text-Based Topological Map for Robotic Guidance of People with Visual Impairments
arXiv

Human-computer Interaction

Memoro: Using Large Language Models to Realize a Concise Interface for Real-Time Memory Augmentation
arXiv

Exploring Large Language Model as an Interactive Sports Coach: Lessons from a Single-Subject Half Marathon Preparation
arXiv

TransforMerger: Transformer-based Voice-Gesture Fusion for Robust Human-Robot Communication
arXiv

Resources

cps
cyber-physical-systems
internet-of-things
iot
large-language-models
llm
network
networking
system

Contributors

KAIWEILIUCC

113 commits

siyang-jiang

3 commits

genglinWang

1 commits

KAIWEILIUCC/Awesome-LLM-IoT-Papers

A collection of papers on LLM applications in the IoT field.

25

117 commits

updated Jan 21, 2026

See the code

README

Awesome-LLM-IoT-Papers

Awesome

Table of Contents

Surveys

A Survey of Foundation Models for IoT: Taxonomy and Criteria-Based Analysis.
arXiv

The role of Large Language Models in addressing IoT challenges: A systematic literature review.

Foundation Models for CPS-IoT: Opportunities and Challenges.
arXiv

LLMs and IoT: A Comprehensive Survey on Large Language Models and the Internet of Things.

A Review on Edge Large Language Models: Design, Execution, and Applications.

Large Language Models for Network Intrusion Detection Systems: Foundations, Implementations, and Future Directions.
arXiv

Toward Edge General Intelligence via Large Language Models: Opportunities and Challenges.

Large Language Models in Smart Grid: Applications and Risks.

Small Language Models: Survey, Measurements, and Insights.
arXiv

From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities.
arXiv

LLM Agents

AutoDroid: LLM-powered Task Automation in Android (MobiCom 2024)
arXiv

MobileGPT: Augmenting LLM with Human-like App Memory for Mobile Task Automation (MobiCom 2024)
arXiv Star

Poster: Enabling Agent-centric Interaction on Smartphones with LLM-based UI Reassembling (MobiSys 2024)

LLMind: Orchestrating AI and IoT with LLM for Complex Task Execution.
arXiv

LLMind 2.0: Distributed IoT Automation with Natural Language M2M Communication and Lightweight LLM Agents.
arXiv

TaskSense: A Translation-like Approach for Tasking Heterogeneous Sensor Systems with LLMs. (SenSys 2025)

Toward Sensor-In-the-Loop LLM Agent: Benchmarks and Implications. (SenSys 2025)

ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions. (NeurIPS 2025)
arXiv

SensorMCP: A Model Context Protocol Server for Custom Sensor Tool Creation. (NetAISys 2025)

Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment.
arXiv

Empowering Agentic Video Analytics Systems with Video Language Models.
arXiv

AutoBridge: Automating Smart Device Integration with Centralized Platform.
arXiv

LLM-based Question-Answer Framework for Sensor-driven HVAC System Interaction.
arXiv

Towards Privacy-Preserving and Personalized Smart Homes via Tailored Small Language Models.
arXiv

UIShift: Enhancing VLM-based GUI Agents through Self-supervised Reinforcement Learning.
arXiv

MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents.
arXiv

MobileViews: A Large-Scale Mobile GUI Dataset.
arXiv

AutoDroid-V2: Boosting SLM-based GUI Agents via Code Generation.
arXiv

LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile Apps.
arXiv

Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents.
arXiv

InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection.
arXiv

AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning.
arXiv

Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment.
arXiv

FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents.
arXiv

AgentSense: LLMs Empower Generalizable and Explainable Web-Based Participatory Urban Sensing.
arXiv

IndusGCC: A Data Benchmark and Evaluation Framework for GUI-Based General Computer Control in Industrial Automation.
arXiv

IoT-MCP: Bridging LLMs and IoT Systems Through Model Context Protocol.
arXiv

AsyncVoice Agent: Real-Time Explanation for LLM Planning and Reasoning.
arXiv

VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification.
arXiv

DroidCall: A Dataset for LLM-powered Android Intent Invocation.
arXiv

More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models.
arXiv

BAMAS: Structuring Budget-Aware Multi-Agent Systems.
arXiv

HiveMind: Contribution-Guided Online Prompt Optimization of LLM Multi-Agent Systems.
arXiv

DIMGen: Dynamic Intent Macro Generation for Efficient LLM-Driven Mobile Automation.

NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls.
arXiv

Edge FM

MELTing Point: Mobile Evaluation of Language Transformers (MobiCom 2024)
arXiv

Your Data, Your Model: A Framework for Training and Deploying Foundational Language Models for Embedded Devices. (MobiCom 2024)

Federated Black-box Prompt Tuning System for Large Language Models on the Edge. (MobiCom 2024)

A Framework for Training and Deploying Foundational Language Models for Embedded Sensing. (MobiCom 2024)

EdgeFM: Leveraging Foundation Model for Open-set Learning on the Edge (SenSys 2023)
arXiv

Confidant: Customizing Transformer-based LLMs via Collaborative Training on Mobile Devices (MobiCom 2025)

Dynamic Sparse Attention on Mobile SoCs
arXiv

MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
arXiv

PhoneLM: an Efficient and Capable Small Language Model Family through Principled Pre-training
arXiv

LLM as a System Service on Mobile Devices
arXiv

LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades
arXiv

ELMS: Elasticized Large Language Models On Mobile Devices
arXiv

Serving MoE Models on Resource-constrained Edge Devices via Dynamic Expert Swapping

HAPE: Hardware-Aware LLM Pruning For Efficient On-Device Inference Optimization

Demystifying Small Language Models for Edge Deployment

TinyLLM: A Framework for Training and Deploying Language Models at the Edge Computerss
arXiv

Confidant: Customizing Transformer-based LLMs via Collaborative Edge Training
arXiv

Modality Plug-and-Play: Runtime Modality Adaptation in LLM-Driven Autonomous Mobile Systems

D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
arXiv

Elastic On-Device LLM Service
arXiv

PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
arXiv

Sensor Data Understanding

Improving On-Device LLMs' Sensory Understanding with Embedding Interpolations (MobiCom 2024)

Penetrative AI: Making LLMs Comprehend the Physical World (ACL 2024)
arXiv

Exploring the Capabilities of LLMs for IMU-based Fine-grained Human Activity Understanding (ACM FMSys 2025)
arXiv

ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs. (HotMobile 2025)

SensorBench: Benchmarking LLMs in Coding-Based Sensor Processing (HotMobile 2025)

Making Sensing Interactive and Descriptive with LLMs: Context Reasoning from Multi-Sensor Data (HotMobile 2025)

Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality Alignment
arXiv

Empowering Agentic Video Analytics Systems with Video Language Models
arXiv

SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
arXiv

ChainStream: An LLM-based Framework for Unified Synthetic Sensing
arXiv

MASTER: A Multi-modal Foundation Model for Human Activity Recognition
arXiv

LLM-CoSen: Revisiting Collaborative Sensing With Large Language Models (LLMs)

Sensor Data Generation

High Resolution Millimeter Wave Imaging Based on FMCW Radar Systems at W-Band

SHADE-AD: An LLM-Based Framework for Synthesizing Activity Data of Alzheimer's Patients. (SenSys 2025)

DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs
arXiv

Spider: Any-to-Many Multimodal LLM
arXiv

CCC: cross-modal contrastive creator for end-to-end sign language generation

Code Generation

AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT Applications (MobiCom 2025)
arXiv Star

GPIoT: Tailoring Small Language Models for IoT Program Synthesis and Development (SenSys 2025)
arXiv Star

CheckMate: LLM-Powered Approximate Intermittent Computing (SenSys 2025)

Exploring Code Language Models for Automated HLS-based Hardware Generation: Benchmark, Infrastructure and Analysis

LLM for Complex Signal Processing in FPGA-based Software Defined Radios: A Case Study on FFT

WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning arXiv

DeepFeature: Iterative Context-aware Feature Generation for Wearable Biosignals arXiv

Interesting Applications

TransCompressor: LLM-Powered Multimodal Data Compression for Smart Transportation
arXiv

Llambda: An LLM-Empowered Low-Resolution Vision System for On-Device Human Behavior Understanding
arXiv

IoT-LLM: Enhancing Real-World IoT Task Reasoning with Large Language Models (Arxiv) Paper

LightLLM: A Versatile Large Language Model for Predictive Light Sensing (SenSys 2025)
arXiv

FlexiFly: Interfacing the Physical World with Foundation Models Empowered by Reconfigurable Drone Systems (SenSys 2025)
arXiv

Socialmind: LLM-based proactive ar social assistive system with human-like perception for in-situ live interactions (IMWUT 2025)
arXiv

TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable Platforms (IMWUT 2024)
arXiv

Exploring Foundation Models in Detecting Concerning Daily Functioning in Psychotherapeutic Context Based on Images from Smart Home Devices (ACM FMSys 2024)

Sensor2Scene: Foundation Model-Driven Interactive Realities (ACM FMSys 2024)

See Where You Read with Eye Gaze Tracking and Large Language Model
arXiv

Empower Vision Applications with LoRA LMM
arXiv

Can we make FCC Experts out of LLMs? (HotMobile 2025)

RouteLLM: A Large Language Model with Native Route Context Understanding to Enable Context-Aware Reasoning (IMWUT 2025)

Congestion Control System Optimization with Large Language Models arXiv

Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments
arXiv

SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
arXiv

Toward Foundation Models for Online Complex Event Detection in CPS-IoT: A Case Study

RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
arXiv

Large Language Model-Guided Disentangled Belief Representation Learning on Polarized Social Graphs

Decoding the Silent Majority: Inducing Belief Augmented Social Graph with Large Language Model for Response Forecasting
arXiv

SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
arXiv

Leveraging AI Agents for Autonomous Networks: A Reference Architecture and Empirical Studies
arXiv

Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
arXiv

Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
arXiv

Large Language Models for Wireless Communications: From Adaptation to Autonomy
arXiv

ThermiKit: Edge-Optimized LWIR Analytics with Agent-Driven Interactions

All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles
arXiv

Edge-IoT and MLLMs for Education and Scene Understanding: Assisting Vision and Hearing-Impaired Individuals

SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
arXiv

LLM-Assisted IoT Testing: Finding Conformance Bugs in Matter SDKs

CBM-RAG: Demonstrating Enhanced Interpretability in Radiology Report Generation with Multi-Agent RAG and Concept Bottleneck Models arXiv

Multimodal LLM for Patient Activity Recognition: Integrating Video, Audio, and Text in Clinical Environments

MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models arXiv

Reasoning Visual Language Model for Chest X-Ray Analysis arXiv

Smart Health

DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge (IMWUT 2024)
arXiv

LLM-based Conversational AI Therapist for Daily Functioning Screening and Psychotherapeutic Intervention via Everyday Smart Devices
arXiv

AutoLife: Automatic Life Journaling with Smartphones and LLMs
arXiv

MDTeamGPT: A Self-Evolving LLM-based Multi-Agent Framework for Multi-Disciplinary Team Medical Consultation
arXiv

Introduction to the Special Issue on Large Language Models, Conversational Systems, and Generative AI in Healthβ€”Part 1

CataractBot: An LLM-powered Expert-in-the-Loop Chatbot for Cataract Patients

GLOSS: Group of LLMs for Open-ended Sensemaking of Passive Sensing Data for Health and Wellbeing.

Demo: Myotrainer: Muscle-Aware Motion Analysis and Feedback System for In-Home Resistance Training

A Disease-Centric Vision-Language Foundation Model for Precision Oncology in Kidney Cancer
arXiv

DermINO: Hybrid Pretraining for a Versatile Dermatology Foundation Model
arXiv

DynamiCare: A Dynamic Multi-Agent Framework for Interactive and Open-Ended Medical Decision-Making
arXiv

Demo Abstract: An LLM-Powered Multimodal Mobile Sensing System for Personalized and Interactive Health Behavior Analysis

DocCHA: Towards LLM-Augmented Interactive Online diagnosis System
arXiv

MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
arXiv

MedAgentSim: Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions
arXiv

MegaAgent: A Large-Scale Autonomous LLM-based Multi-Agent System Without Predefined SOPs
arXiv

Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
arXiv

Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
arXiv

SynLLM: A Comparative Analysis of Large Language Models for Medical Tabular Synthetic Data Generation via Prompt Engineering
arXiv

Myo-Trainer: A Vision-based Muscle-Aware Motion Feedback System for In-Home Resistance Training

MedAide: Towards an Omni Medical Aide via Specialized LLM-based Multi-Agent Collaboration
arXiv

MedPlan: A Two-Stage RAG-Based System for Personalized Medical Plan Generation
arXiv

Robotics

RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation
arXiv

NaVILA: Legged Robot Vision-Language-Action Model for Navigation
arXiv

Affordances-Oriented Planning using Foundation Models for Continuous Vision-Language Navigation
arXiv

FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph
arXiv

Multi-robot Rigid Formation Navigation via Synchronous Motion and Discrete-time Communication-Control Optimization
arXiv

Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
arXiv

See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation
arXiv

Guide-LLM: An Embodied LLM Agent and Text-Based Topological Map for Robotic Guidance of People with Visual Impairments
arXiv

Human-computer Interaction

Memoro: Using Large Language Models to Realize a Concise Interface for Real-Time Memory Augmentation
arXiv

Exploring Large Language Model as an Interactive Sports Coach: Lessons from a Single-Subject Half Marathon Preparation
arXiv

TransforMerger: Transformer-based Voice-Gesture Fusion for Robust Human-Robot Communication
arXiv

Resources

cps
cyber-physical-systems
internet-of-things
iot
large-language-models
llm
network
networking
system

Contributors

KAIWEILIUCC

113 commits

siyang-jiang

3 commits

genglinWang

1 commits