A collection of papers on LLM applications in the IoT field.
25
117 commits
updated Jan 21, 2026
A Survey of Foundation Models for IoT: Taxonomy and Criteria-Based Analysis.
The role of Large Language Models in addressing IoT challenges: A systematic literature review.
Foundation Models for CPS-IoT: Opportunities and Challenges.
LLMs and IoT: A Comprehensive Survey on Large Language Models and the Internet of Things.
A Review on Edge Large Language Models: Design, Execution, and Applications.
Large Language Models for Network Intrusion Detection Systems: Foundations, Implementations, and Future Directions.
Toward Edge General Intelligence via Large Language Models: Opportunities and Challenges.
Large Language Models in Smart Grid: Applications and Risks.
Small Language Models: Survey, Measurements, and Insights.
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities.
AutoDroid: LLM-powered Task Automation in Android (MobiCom 2024)
MobileGPT: Augmenting LLM with Human-like App Memory for Mobile Task Automation (MobiCom 2024)
Poster: Enabling Agent-centric Interaction on Smartphones with LLM-based UI Reassembling (MobiSys 2024)
LLMind: Orchestrating AI and IoT with LLM for Complex Task Execution.
LLMind 2.0: Distributed IoT Automation with Natural Language M2M Communication and Lightweight LLM Agents.
TaskSense: A Translation-like Approach for Tasking Heterogeneous Sensor Systems with LLMs. (SenSys 2025)
Toward Sensor-In-the-Loop LLM Agent: Benchmarks and Implications. (SenSys 2025)
ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions. (NeurIPS 2025)
SensorMCP: A Model Context Protocol Server for Custom Sensor Tool Creation. (NetAISys 2025)
Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment.
Empowering Agentic Video Analytics Systems with Video Language Models.
AutoBridge: Automating Smart Device Integration with Centralized Platform.
LLM-based Question-Answer Framework for Sensor-driven HVAC System Interaction.
Towards Privacy-Preserving and Personalized Smart Homes via Tailored Small Language Models.
UIShift: Enhancing VLM-based GUI Agents through Self-supervised Reinforcement Learning.
MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents.
MobileViews: A Large-Scale Mobile GUI Dataset.
AutoDroid-V2: Boosting SLM-based GUI Agents via Code Generation.
LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile Apps.
Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents.
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection.
AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning.
Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment.
FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents.
AgentSense: LLMs Empower Generalizable and Explainable Web-Based Participatory Urban Sensing.
IndusGCC: A Data Benchmark and Evaluation Framework for GUI-Based General Computer Control in Industrial Automation.
IoT-MCP: Bridging LLMs and IoT Systems Through Model Context Protocol.
AsyncVoice Agent: Real-Time Explanation for LLM Planning and Reasoning.
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification.
DroidCall: A Dataset for LLM-powered Android Intent Invocation.
More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models.
BAMAS: Structuring Budget-Aware Multi-Agent Systems.
HiveMind: Contribution-Guided Online Prompt Optimization of LLM Multi-Agent Systems.
DIMGen: Dynamic Intent Macro Generation for Efficient LLM-Driven Mobile Automation.
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls.
MELTing Point: Mobile Evaluation of Language Transformers (MobiCom 2024)
Your Data, Your Model: A Framework for Training and Deploying Foundational Language Models for Embedded Devices. (MobiCom 2024)
Federated Black-box Prompt Tuning System for Large Language Models on the Edge. (MobiCom 2024)
A Framework for Training and Deploying Foundational Language Models for Embedded Sensing. (MobiCom 2024)
EdgeFM: Leveraging Foundation Model for Open-set Learning on the Edge (SenSys 2023)
Confidant: Customizing Transformer-based LLMs via Collaborative Training on Mobile Devices (MobiCom 2025)
Dynamic Sparse Attention on Mobile SoCs
MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
PhoneLM: an Efficient and Capable Small Language Model Family through Principled Pre-training
LLM as a System Service on Mobile Devices
LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades
ELMS: Elasticized Large Language Models On Mobile Devices
Serving MoE Models on Resource-constrained Edge Devices via Dynamic Expert Swapping
HAPE: Hardware-Aware LLM Pruning For Efficient On-Device Inference Optimization
Demystifying Small Language Models for Edge Deployment
TinyLLM: A Framework for Training and Deploying Language Models at the Edge Computerss
Confidant: Customizing Transformer-based LLMs via Collaborative Edge Training
Modality Plug-and-Play: Runtime Modality Adaptation in LLM-Driven Autonomous Mobile Systems
D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
Improving On-Device LLMs' Sensory Understanding with Embedding Interpolations (MobiCom 2024)
Penetrative AI: Making LLMs Comprehend the Physical World (ACL 2024)
Exploring the Capabilities of LLMs for IMU-based Fine-grained Human Activity Understanding (ACM FMSys 2025)
ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs. (HotMobile 2025)
SensorBench: Benchmarking LLMs in Coding-Based Sensor Processing (HotMobile 2025)
Making Sensing Interactive and Descriptive with LLMs: Context Reasoning from Multi-Sensor Data (HotMobile 2025)
Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality Alignment
Empowering Agentic Video Analytics Systems with Video Language Models
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
ChainStream: An LLM-based Framework for Unified Synthetic Sensing
MASTER: A Multi-modal Foundation Model for Human Activity Recognition
LLM-CoSen: Revisiting Collaborative Sensing With Large Language Models (LLMs)
High Resolution Millimeter Wave Imaging Based on FMCW Radar Systems at W-Band
SHADE-AD: An LLM-Based Framework for Synthesizing Activity Data of Alzheimer's Patients. (SenSys 2025)
DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs
Spider: Any-to-Many Multimodal LLM
CCC: cross-modal contrastive creator for end-to-end sign language generation
AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT Applications (MobiCom 2025)
GPIoT: Tailoring Small Language Models for IoT Program Synthesis and Development (SenSys 2025)
CheckMate: LLM-Powered Approximate Intermittent Computing (SenSys 2025)
LLM for Complex Signal Processing in FPGA-based Software Defined Radios: A Case Study on FFT
WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
DeepFeature: Iterative Context-aware Feature Generation for Wearable Biosignals
TransCompressor: LLM-Powered Multimodal Data Compression for Smart Transportation
Llambda: An LLM-Empowered Low-Resolution Vision System for On-Device Human Behavior Understanding
IoT-LLM: Enhancing Real-World IoT Task Reasoning with Large Language Models (Arxiv) Paper
LightLLM: A Versatile Large Language Model for Predictive Light Sensing (SenSys 2025)
FlexiFly: Interfacing the Physical World with Foundation Models Empowered by Reconfigurable Drone Systems (SenSys 2025)
Socialmind: LLM-based proactive ar social assistive system with human-like perception for in-situ live interactions (IMWUT 2025)
TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable Platforms (IMWUT 2024)
Exploring Foundation Models in Detecting Concerning Daily Functioning in Psychotherapeutic Context Based on Images from Smart Home Devices (ACM FMSys 2024)
Sensor2Scene: Foundation Model-Driven Interactive Realities (ACM FMSys 2024)
See Where You Read with Eye Gaze Tracking and Large Language Model
Empower Vision Applications with LoRA LMM
Can we make FCC Experts out of LLMs? (HotMobile 2025)
RouteLLM: A Large Language Model with Native Route Context Understanding to Enable Context-Aware Reasoning (IMWUT 2025)
Congestion Control System Optimization with Large Language Models
Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
Toward Foundation Models for Online Complex Event Detection in CPS-IoT: A Case Study
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
Large Language Model-Guided Disentangled Belief Representation Learning on Polarized Social Graphs
Decoding the Silent Majority: Inducing Belief Augmented Social Graph with Large Language Model for Response Forecasting
SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
Leveraging AI Agents for Autonomous Networks: A Reference Architecture and Empirical Studies
Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
Large Language Models for Wireless Communications: From Adaptation to Autonomy
ThermiKit: Edge-Optimized LWIR Analytics with Agent-Driven Interactions
All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles
SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
LLM-Assisted IoT Testing: Finding Conformance Bugs in Matter SDKs
CBM-RAG: Demonstrating Enhanced Interpretability in Radiology Report Generation with Multi-Agent RAG and Concept Bottleneck Models
MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models
Reasoning Visual Language Model for Chest X-Ray Analysis
DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge (IMWUT 2024)
LLM-based Conversational AI Therapist for Daily Functioning Screening and Psychotherapeutic Intervention via Everyday Smart Devices
AutoLife: Automatic Life Journaling with Smartphones and LLMs
MDTeamGPT: A Self-Evolving LLM-based Multi-Agent Framework for Multi-Disciplinary Team Medical Consultation
CataractBot: An LLM-powered Expert-in-the-Loop Chatbot for Cataract Patients
GLOSS: Group of LLMs for Open-ended Sensemaking of Passive Sensing Data for Health and Wellbeing.
Demo: Myotrainer: Muscle-Aware Motion Analysis and Feedback System for In-Home Resistance Training
A Disease-Centric Vision-Language Foundation Model for Precision Oncology in Kidney Cancer
DermINO: Hybrid Pretraining for a Versatile Dermatology Foundation Model
DynamiCare: A Dynamic Multi-Agent Framework for Interactive and Open-Ended Medical Decision-Making
DocCHA: Towards LLM-Augmented Interactive Online diagnosis System
MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
MedAgentSim: Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions
MegaAgent: A Large-Scale Autonomous LLM-based Multi-Agent System Without Predefined SOPs
Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
SynLLM: A Comparative Analysis of Large Language Models for Medical Tabular Synthetic Data Generation via Prompt Engineering
Myo-Trainer: A Vision-based Muscle-Aware Motion Feedback System for In-Home Resistance Training
MedAide: Towards an Omni Medical Aide via Specialized LLM-based Multi-Agent Collaboration
MedPlan: A Two-Stage RAG-Based System for Personalized Medical Plan Generation
RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
Affordances-Oriented Planning using Foundation Models for Continuous Vision-Language Navigation
FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph
Multi-robot Rigid Formation Navigation via Synchronous Motion and Discrete-time Communication-Control Optimization
Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation
Guide-LLM: An Embodied LLM Agent and Text-Based Topological Map for Robotic Guidance of People with Visual Impairments
Memoro: Using Large Language Models to Realize a Concise Interface for Real-Time Memory Augmentation
Exploring Large Language Model as an Interactive Sports Coach: Lessons from a Single-Subject Half Marathon Preparation
TransforMerger: Transformer-based Voice-Gesture Fusion for Robust Human-Robot Communication
A collection of papers on LLM applications in the IoT field.
25
117 commits
updated Jan 21, 2026
A Survey of Foundation Models for IoT: Taxonomy and Criteria-Based Analysis.
The role of Large Language Models in addressing IoT challenges: A systematic literature review.
Foundation Models for CPS-IoT: Opportunities and Challenges.
LLMs and IoT: A Comprehensive Survey on Large Language Models and the Internet of Things.
A Review on Edge Large Language Models: Design, Execution, and Applications.
Large Language Models for Network Intrusion Detection Systems: Foundations, Implementations, and Future Directions.
Toward Edge General Intelligence via Large Language Models: Opportunities and Challenges.
Large Language Models in Smart Grid: Applications and Risks.
Small Language Models: Survey, Measurements, and Insights.
From Specific-MLLMs to Omni-MLLMs: A Survey on MLLMs Aligned with Multi-modalities.
AutoDroid: LLM-powered Task Automation in Android (MobiCom 2024)
MobileGPT: Augmenting LLM with Human-like App Memory for Mobile Task Automation (MobiCom 2024)
Poster: Enabling Agent-centric Interaction on Smartphones with LLM-based UI Reassembling (MobiSys 2024)
LLMind: Orchestrating AI and IoT with LLM for Complex Task Execution.
LLMind 2.0: Distributed IoT Automation with Natural Language M2M Communication and Lightweight LLM Agents.
TaskSense: A Translation-like Approach for Tasking Heterogeneous Sensor Systems with LLMs. (SenSys 2025)
Toward Sensor-In-the-Loop LLM Agent: Benchmarks and Implications. (SenSys 2025)
ContextAgent: Context-Aware Proactive LLM Agents with Open-World Sensory Perceptions. (NeurIPS 2025)
SensorMCP: A Model Context Protocol Server for Custom Sensor Tool Creation. (NetAISys 2025)
Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment.
Empowering Agentic Video Analytics Systems with Video Language Models.
AutoBridge: Automating Smart Device Integration with Centralized Platform.
LLM-based Question-Answer Framework for Sensor-driven HVAC System Interaction.
Towards Privacy-Preserving and Personalized Smart Homes via Tailored Small Language Models.
UIShift: Enhancing VLM-based GUI Agents through Self-supervised Reinforcement Learning.
MCPWorld: A Unified Benchmarking Testbed for API, GUI, and Hybrid Computer Use Agents.
MobileViews: A Large-Scale Mobile GUI Dataset.
AutoDroid-V2: Boosting SLM-based GUI Agents via Code Generation.
LLM-Explorer: Towards Efficient and Affordable LLM-based Exploration for Mobile Apps.
Process-Supervised Reinforcement Learning for Interactive Multimodal Tool-Use Agents.
InfiGUIAgent: A Multimodal Generalist GUI Agent with Native Reasoning and Reflection.
AgentCPM-GUI: Building Mobile-Use Agents with Reinforcement Fine-Tuning.
Advancing Mobile GUI Agents: A Verifier-Driven Approach to Practical Deployment.
FingerTip 20K: A Benchmark for Proactive and Personalized Mobile LLM Agents.
AgentSense: LLMs Empower Generalizable and Explainable Web-Based Participatory Urban Sensing.
IndusGCC: A Data Benchmark and Evaluation Framework for GUI-Based General Computer Control in Industrial Automation.
IoT-MCP: Bridging LLMs and IoT Systems Through Model Context Protocol.
AsyncVoice Agent: Real-Time Explanation for LLM Planning and Reasoning.
VeriSafe Agent: Safeguarding Mobile GUI Agent via Logic-based Action Verification.
DroidCall: A Dataset for LLM-powered Android Intent Invocation.
More Than Meets the Eye? Uncovering the Reasoning-Planning Disconnect in Training Vision-Language Driving Models.
BAMAS: Structuring Budget-Aware Multi-Agent Systems.
HiveMind: Contribution-Guided Online Prompt Optimization of LLM Multi-Agent Systems.
DIMGen: Dynamic Intent Macro Generation for Efficient LLM-Driven Mobile Automation.
NESTFUL: A Benchmark for Evaluating LLMs on Nested Sequences of API Calls.
MELTing Point: Mobile Evaluation of Language Transformers (MobiCom 2024)
Your Data, Your Model: A Framework for Training and Deploying Foundational Language Models for Embedded Devices. (MobiCom 2024)
Federated Black-box Prompt Tuning System for Large Language Models on the Edge. (MobiCom 2024)
A Framework for Training and Deploying Foundational Language Models for Embedded Sensing. (MobiCom 2024)
EdgeFM: Leveraging Foundation Model for Open-set Learning on the Edge (SenSys 2023)
Confidant: Customizing Transformer-based LLMs via Collaborative Training on Mobile Devices (MobiCom 2025)
Dynamic Sparse Attention on Mobile SoCs
MobiEdit: Resource-efficient Knowledge Editing for Personalized On-device LLMs
PhoneLM: an Efficient and Capable Small Language Model Family through Principled Pre-training
LLM as a System Service on Mobile Devices
LoRASuite: Efficient LoRA Adaptation Across Large Language Model Upgrades
ELMS: Elasticized Large Language Models On Mobile Devices
Serving MoE Models on Resource-constrained Edge Devices via Dynamic Expert Swapping
HAPE: Hardware-Aware LLM Pruning For Efficient On-Device Inference Optimization
Demystifying Small Language Models for Edge Deployment
TinyLLM: A Framework for Training and Deploying Language Models at the Edge Computerss
Confidant: Customizing Transformer-based LLMs via Collaborative Edge Training
Modality Plug-and-Play: Runtime Modality Adaptation in LLM-Driven Autonomous Mobile Systems
D2MoE: Dual Routing and Dynamic Scheduling for Efficient On-Device MoE-based LLM Serving
PerCache: Predictive Hierarchical Cache for RAG Applications on Mobile Devices
Improving On-Device LLMs' Sensory Understanding with Embedding Interpolations (MobiCom 2024)
Penetrative AI: Making LLMs Comprehend the Physical World (ACL 2024)
Exploring the Capabilities of LLMs for IMU-based Fine-grained Human Activity Understanding (ACM FMSys 2025)
ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs. (HotMobile 2025)
SensorBench: Benchmarking LLMs in Coding-Based Sensor Processing (HotMobile 2025)
Making Sensing Interactive and Descriptive with LLMs: Context Reasoning from Multi-Sensor Data (HotMobile 2025)
Babel: A Scalable Pre-trained Model for Multi-Modal Sensing via Expandable Modality Alignment
Empowering Agentic Video Analytics Systems with Video Language Models
SensorQA: A Question Answering Benchmark for Daily-Life Monitoring
ChainStream: An LLM-based Framework for Unified Synthetic Sensing
MASTER: A Multi-modal Foundation Model for Human Activity Recognition
LLM-CoSen: Revisiting Collaborative Sensing With Large Language Models (LLMs)
High Resolution Millimeter Wave Imaging Based on FMCW Radar Systems at W-Band
SHADE-AD: An LLM-Based Framework for Synthesizing Activity Data of Alzheimer's Patients. (SenSys 2025)
DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs
Spider: Any-to-Many Multimodal LLM
CCC: cross-modal contrastive creator for end-to-end sign language generation
AutoIOT: LLM-Driven Automated Natural Language Programming for AIoT Applications (MobiCom 2025)
GPIoT: Tailoring Small Language Models for IoT Program Synthesis and Development (SenSys 2025)
CheckMate: LLM-Powered Approximate Intermittent Computing (SenSys 2025)
LLM for Complex Signal Processing in FPGA-based Software Defined Radios: A Case Study on FFT
WebGen-Agent: Enhancing Interactive Website Generation with Multi-Level Feedback and Step-Level Reinforcement Learning
DeepFeature: Iterative Context-aware Feature Generation for Wearable Biosignals
TransCompressor: LLM-Powered Multimodal Data Compression for Smart Transportation
Llambda: An LLM-Empowered Low-Resolution Vision System for On-Device Human Behavior Understanding
IoT-LLM: Enhancing Real-World IoT Task Reasoning with Large Language Models (Arxiv) Paper
LightLLM: A Versatile Large Language Model for Predictive Light Sensing (SenSys 2025)
FlexiFly: Interfacing the Physical World with Foundation Models Empowered by Reconfigurable Drone Systems (SenSys 2025)
Socialmind: LLM-based proactive ar social assistive system with human-like perception for in-situ live interactions (IMWUT 2025)
TRAMBA: A Hybrid Transformer and Mamba Architecture for Practical Audio and Bone Conduction Speech Super Resolution and Enhancement on Mobile and Wearable Platforms (IMWUT 2024)
Exploring Foundation Models in Detecting Concerning Daily Functioning in Psychotherapeutic Context Based on Images from Smart Home Devices (ACM FMSys 2024)
Sensor2Scene: Foundation Model-Driven Interactive Realities (ACM FMSys 2024)
See Where You Read with Eye Gaze Tracking and Large Language Model
Empower Vision Applications with LoRA LMM
Can we make FCC Experts out of LLMs? (HotMobile 2025)
RouteLLM: A Large Language Model with Native Route Context Understanding to Enable Context-Aware Reasoning (IMWUT 2025)
Congestion Control System Optimization with Large Language Models
Enhancing User Engagement in Socially-Driven Dialogue through Interactive LLM Alignments
SecurityLingua: Efficient Defense of LLM Jailbreak Attacks via Security-Aware Prompt Compression
Toward Foundation Models for Online Complex Event Detection in CPS-IoT: A Case Study
RealFactBench: A Benchmark for Evaluating Large Language Models in Real-World Fact-Checking
Large Language Model-Guided Disentangled Belief Representation Learning on Polarized Social Graphs
Decoding the Silent Majority: Inducing Belief Augmented Social Graph with Large Language Model for Response Forecasting
SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
Leveraging AI Agents for Autonomous Networks: A Reference Architecture and Empirical Studies
Inducing Causal World Models in LLMs for Zero-Shot Physical Reasoning
Multimodal Chain of Continuous Thought for Latent-Space Reasoning in Vision-Language Models
Large Language Models for Wireless Communications: From Adaptation to Autonomy
ThermiKit: Edge-Optimized LWIR Analytics with Agent-Driven Interactions
All You Need for Object Detection: From Pixels, Points, and Prompts to Next-Gen Fusion and Multimodal LLMs/VLMs in Autonomous Vehicles
SA-OOSC: A Multimodal LLM-Distilled Semantic Communication Framework for Enhanced Coding Efficiency with Scenario Understanding
LLM-Assisted IoT Testing: Finding Conformance Bugs in Matter SDKs
CBM-RAG: Demonstrating Enhanced Interpretability in Radiology Report Generation with Multi-Agent RAG and Concept Bottleneck Models
MedSeg-R: Reasoning Segmentation in Medical Images with Multimodal Large Language Models
Reasoning Visual Language Model for Chest X-Ray Analysis
DrHouse: An LLM-empowered Diagnostic Reasoning System through Harnessing Outcomes from Sensor Data and Expert Knowledge (IMWUT 2024)
LLM-based Conversational AI Therapist for Daily Functioning Screening and Psychotherapeutic Intervention via Everyday Smart Devices
AutoLife: Automatic Life Journaling with Smartphones and LLMs
MDTeamGPT: A Self-Evolving LLM-based Multi-Agent Framework for Multi-Disciplinary Team Medical Consultation
CataractBot: An LLM-powered Expert-in-the-Loop Chatbot for Cataract Patients
GLOSS: Group of LLMs for Open-ended Sensemaking of Passive Sensing Data for Health and Wellbeing.
Demo: Myotrainer: Muscle-Aware Motion Analysis and Feedback System for In-Home Resistance Training
A Disease-Centric Vision-Language Foundation Model for Precision Oncology in Kidney Cancer
DermINO: Hybrid Pretraining for a Versatile Dermatology Foundation Model
DynamiCare: A Dynamic Multi-Agent Framework for Interactive and Open-Ended Medical Decision-Making
DocCHA: Towards LLM-Augmented Interactive Online diagnosis System
MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
MedAgentSim: Self-Evolving Multi-Agent Simulations for Realistic Clinical Interactions
MegaAgent: A Large-Scale Autonomous LLM-based Multi-Agent System Without Predefined SOPs
Fleming-VL: Towards Universal Medical Visual Reasoning with Multimodal LLMs
Lingshu: A Generalist Foundation Model for Unified Multimodal Medical Understanding and Reasoning
SynLLM: A Comparative Analysis of Large Language Models for Medical Tabular Synthetic Data Generation via Prompt Engineering
Myo-Trainer: A Vision-based Muscle-Aware Motion Feedback System for In-Home Resistance Training
MedAide: Towards an Omni Medical Aide via Specialized LLM-based Multi-Agent Collaboration
MedPlan: A Two-Stage RAG-Based System for Personalized Medical Plan Generation
RoboInspector: Unveiling the Unreliability of Policy Code for LLM-enabled Robotic Manipulation
NaVILA: Legged Robot Vision-Language-Action Model for Navigation
Affordances-Oriented Planning using Foundation Models for Continuous Vision-Language Navigation
FSR-VLN: Fast and Slow Reasoning for Vision-Language Navigation with Hierarchical Multi-modal Scene Graph
Multi-robot Rigid Formation Navigation via Synchronous Motion and Discrete-time Communication-Control Optimization
Expertise need not monopolize: Action-Specialized Mixture of Experts for Vision-Language-Action Learning
See, Point, Fly: A Learning-Free VLM Framework for Universal Unmanned Aerial Navigation
Guide-LLM: An Embodied LLM Agent and Text-Based Topological Map for Robotic Guidance of People with Visual Impairments
Memoro: Using Large Language Models to Realize a Concise Interface for Real-Time Memory Augmentation
Exploring Large Language Model as an Interactive Sports Coach: Lessons from a Single-Subject Half Marathon Preparation
TransforMerger: Transformer-based Voice-Gesture Fusion for Robust Human-Robot Communication