33
169 commits
updated Sep 23, 2026
This repository provides a curated, continuously updated index of foundation models (FMs) in the human activity recognition (HAR) domain. The organization follows our survey’s lifecycle-based taxonomy and major development directions. It serves as a living companion to the ACM IMWUT paper “Foundation Models Defining a New Era in Human Activity Recognition: A Survey and Outlook”, offering direct access to representative works, datasets, and model resources. Our goal is to foster transparency, reproducibility, and collaboration across the HAR community as the field transitions toward large-scale, multimodal, and language-grounded sensing models.
Contributions are welcome! Whether you add new papers, improve taxonomy coverage, link open-source implementations, or update existing groups. Please help the community build a shared, evolving reference for next-generation HAR foundation models.
(To include your related work in this repository, please create a pull request with the relevant details or drop us a message through email: sizhen.bian@nwpu.edu.cn)
The paper is openly accessible at: https://dl.acm.org/doi/10.1145/3810230 (with contributors from DFKI (Germany), RPTU (Germany), GIT (America), and NWPU(China)).
Historical development of sensor-based Human Activity Recognition (HAR) models.
Historical development of sensor-based Human Activity Recognition (HAR) models. From classical machine learning with hand-crafted features and shallow classifiers to the rise of deep learning with CNNs and RNNs, the field progressed toward a phase focused on transfer and domain generalization (robustness across users, devices, and datasets). More recently, self-supervised learning (SSL) approaches have enabled pretraining on unlabeled sensor data using contrastive or masked objectives. Today, the field is moving toward foundation models, exemplified by large-scale sensor–language alignment, emphasizing scalability, generalization, and interpretability
Here are the growth trends of publications since 2022 and the model names cloud:
Left: HAR-FM papers showing a sharp acceleration with the vast majority of works emerging since 2024.
Right: Representative model name cloud.
Definition of Foundation Models and it's adaptation in different fields.
Foundation Model: Any model trained on broad data (generally using self-supervision at scale) that can be adapted (e.g., fine-tuned) to a wide range of downstream tasks.[Paper]
Foundation Model in the CV Field: A pre-trained model and its adapters capable of solving all vision tasks within the space–time–modality continuum (ranging from coarse to fine-grained, static to dynamic, and single (RGB) to multimodal sensory inputs) while supporting transferability through zero-/few-shot learning and fine-tuning. [Paper]
Foundation Model in the NLP Field: The large Pretrained Language Models (PLMs), characterized by their ability to generate fluent text, handle multiple modalities, and follow natural-language instructions to perform diverse tasks.[Paper]
Foundation Model in the HAR Field: A pretrained, sensor-grounded model and its adapters that can solve diverse activity-understanding tasks across the sensing–temporal–context continuum while generalizing across sensor modalities, body placements, users, devices, and environments.
Heuristic 1–7 scores of representative works against six HAR–FM criteria. Each radar chart profiles a model on: A) Corpus Coverage & Diversity, B) Cross-Domain Generalization, C) Modality Extensibility & Grounding, D) Label-Efficient Pretraining, E) Adaptation Surfaces & Reusability, and F) Broad Applicability & Emergent Capabilities. The “Ideal HAR-FM” panel depicts a target profile. Scores (1 = limited evidence to 7 = strong evidence) are judgment-based syntheses from reported results (compared both to the other models in this survey and to an aspirational “ideal” FM-for-HAR reference point) and are intended for qualitative comparison rather than a leaderboard.
Note: Although foundation models in HAR domain are still in their formative phase, a model should exhibit core hallmarks of the foundation model paradigm, not necessarily all at once, but in substance. These dimensions outline what defines a model as foundational: the ability to scale across data and users, generalize beyond training domains, adapt efficiently to new tasks, and support reuse across modalities and contexts. Collectively, they set a directional standard rather than a checklist, marking the shift from task-specific modeling toward unified, adaptable representations of human activity.
"A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)". Xiong et al.. IEEE BIBM 2024. [Paper]
"Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]
"Scaling Wearable Foundation Models". Narayanswamy et al.. arXiv 2024. [Paper]
"RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]
"HAR-DoReMi: Optimizing Data Mixture for Self-Supervised Human Activity Recognition Across Heterogeneous IMU Datasets". Ban et al.. arXiv 2025. [Paper]
"Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions". Erturk et al.. arXiv 2025. [Paper]
"LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing Applications". Xu et al.. SenSys 2021. [Paper][Code]
"Cosmo: Contrastive Fusion Learning with Small Data for Multimodal Human Activity Recognition". Ouyang et al.. MobiCom 2022. [Paper][Code]
"Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications across Lab and Field Settings". Saha et al.. IMWUT 2025. [Paper]
"RobustHAR: Multi-scale Spatial-temporal Masked Self-supervised Pre-training for Robust Human Activity Recognition". Liu et al.. IJCAI 2025. [Paper]
"Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices". Minghui Qiu et al. IMWUT 2025. [Paper]
"HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]
| Dataset | Sensors | Datasize | Subjects | Activities |
|---|---|---|---|---|
| CAPTURE-24 | acc | 3883 h | 151 | 200 unique labels |
| TNDA-HAR | acc, gyro | 5.7 h | 23 | 8 daily activities |
| HAR70+ | acc | 12.6 h | 18 | 8 daily activities |
| WISDM | acc, gyro | 91.8 h | 51 | 18 daily activities |
| MotionSense | acc, gyro | — | 24 | 6 daily activities |
| SHL Challenge | acc, gyro, mag | 2812 h | 3 | 8 transport modes |
| Shoaib | acc, gyro, mag | 6.5 h | 10 | 13 daily activities |
| HHAR | acc, gyro | — | 9 | 6 daily activities |
| WHARF | acc | — | 16 | 8 motion primitives |
| DSADS | acc, gyro, mag | 12.7 h | 8 | 19 daily/sports activities |
| UCI-HAR | acc, gyro | — | 30 | 6 daily activities |
| USC-HAD | acc, gyro, mag | — | 14 | 12 daily activities |
| Daphnet FoG | acc | 8.3 h | 10 | 3 walking activities |
| HAPT | acc, gyro | — | 30 | 6 static, dynamic activities |
| REALDISP | acc, gyro, mag | — | 17 | 33 daily, fitness activities |
| UniMiB SHAR | acc | — | 30 | 17 daily, fall activities |
| UMAFall | acc, gyro, mag | 2.2 h | 17 | 11 daily, fall activities |
| MobiAct | acc, gyro | — | 57 | 13 daily, fall activities |
| Skoda Mini Checkpoint | acc (3D) | — | 1 | 10 assembly-line activities |
| Dataset | Sensors | Datasize | Subjects | Activities |
|---|---|---|---|---|
| RecGym | acc, gyro, human body capacitance | 50h | 10 | 12 fitness activities |
| WEAR | acc, video | 19 h | 22 | 18 sports activities |
| iSPL | acc, gyro, stretch | — | 1 | 9 daily activities |
| HARTH | acc, video | 35.9 h | 22 | 12 daily activities |
| w-HAR | acc, gyro, stretch | 3 h | 22 | 7 daily activities |
| RealLifeHAR | acc, gyro, mag, GPS | — | 19 | 4 daily activities |
| MMAct | RGB, keypoints, acc, gyro, ori, Wi‑Fi, pressure | — | 40 | 37 activities |
| HuGaDB | acc, gyro, EMG | 10 h | 18 | 12 activities |
| RealWorld HAR | acc, gyro, mag, GPS, light, sound level | 124.3 h | 15 | 8 daily activities |
| ExtraSensory | acc, gyro, mag, location, audio, additional | — | 60 | 51 activities |
| UTD-MHAD | RGB, depth, skeleton, acc, gyro | — | 8 | 27 activities |
| MHEALTH | acc, gyro, mag, ECG | — | 10 | 12 daily activities |
| Berkeley MHAD | acc, optical capture, video, depth, audio | 1.37 h | 12 | 11 daily activities |
| PAMAP2 | acc, gyro, mag, HR | 10 h | 9 | 18 daily activities |
| Opportunity | acc, gyro, mag, ambient sensors | 25 h | 4 | 9 kitchen + 9 gestures |
| MRI | mmWave, RGB-D, IMU | 5.3 h | 20 | pose estimation |
| NORMWEAR | PPG, ECG, EEG, GSR, IMU | 14,943 h | 20 | pose estimation |
| SensorLM [Paper ] | PPG, EDA, ACC, TEMP, ALT | 59,749 h | 103,731 | sensor-language study |
| Apple Study (AHMS; WBM) [Paper] | HealthKit metrics (27) | > 2.5 B h | 162 K | 57 health tasks |
| WESAD | EDA/PPG/Temp + Acc | 45 h | 15 | stress detection |
| PPG-Dalia | ECG, PPG, IMU, GSR | 36 h | 15 | daily activities |
| Dataset | Sensors | Datasize | Subjects | Activities |
|---|---|---|---|---|
| Sleep-EDF | EEG/EOG/EMG/ECG | 1,576 h | 197 | sleep stages |
| MIT-BIH Arrhythmia | 2‑lead ECG | 1128 h | 47 | ambulatory ECG |
| PTB-XL | 12‑lead ECG | 6.06 h | 18,885 | ECG status |
| TUH EEG | multi‑channel EEG | 1476 h | 675 | seizure activity |
| Cuff‑Less‑BP | ECG, PPG | 72 h | — | blood‑pressure estimation |
| Auditory‑EEG | EEG | 23 h | — | auditory attention |
| PhyAAt | EEG | 33 h | 25 | auditory attention |
| MAUS | ECG, PPG, GSR | 22 h | 22 | cognitive workload/stress |
| Mendeley‑YAAD | ECG, GSR | 5 h | — | affect/stress elicitation |
| Brain‑Cognitive | EEG | 85 h | 20 | cognitive state regulation |
| EPHNOGRAM | ECG, PCG | 61 h | 24 | cardiac auscultation |
| BIDMC | ECG, PPG | 14 h | 53 | clinical monitoring |
| MOODS [Paper] | PPG | 54 K h | 122 | stress monitoring |
| SleepFM | BAS, ECG, respiratory | 112,544 h | 14,068 | sleep quality monitoring |
| RecGym | acc, gyro, human body capacitance | 50h | 10 | 12 fitness activities |
| PPG-Dalia | ECG, PPG, IMU, GSR | 36 h | 15 | daily activities |
| Dataset | Sensors | Datasize | Subjects | Activities |
|---|---|---|---|---|
| CASAS (Aruba/Milan/...) | Ambient binary sensors (motion, doors) | — | — | smart home activities |
| Dataset | Sensors | Datasize | Subjects | Activities |
|---|---|---|---|---|
| mmWave (var.)[Paper] | Range–Doppler / RF point clouds | 5 h | 10 | daily activities (10 scenes) |
| MM-Fi | mmWave, LiDAR, Wi‑Fi, RGB‑D | 10.6 h | 40 | 27 daily activities |
Four base computation graphs for sensor foundation models. Encoder-only stacks (top left) focus on representation learning using a single sensor encoder (e.g., ViT or SSM) with lightweight heads for recognition, retrieval, or forecasting. Dual encoders (top right) independently embed sensor and text/vision streams and align them via a shared latent projection (CLIP-style) for retrieval/zero-shot transfer. Encoder–decoder stacks (bottom left) condition a language/multimodal decoder on encoded sensor tokens (cross-attention) to produce captions, rationales, or structured outputs. Language-model stacks (bottom right), either encoder–decoder or decoder-only, treat sensing as a token sequence using projection/quantization interfaces for forecasting, analysis, and reasoning.
"A Novel Human Activity Recognition Framework Based on Pre‑Trained Foundation Model (Chronos HAR Adapters)". Xiong et al.. IEEE BIBM 2024. [Paper]
"Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions". Erturk et al.. ICML 2025. [Paper]
"Cosmo: Contrastive Fusion Learning with Small Data for Multimodal Human Activity Recognition". Ouyang et al.. MobiCom 2022. [Paper][Code]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"HAR‑DoReMi: Optimizing Data Mixture for Self‑Supervised Human Activity Recognition Across Heterogeneous IMU Datasets". Ban et al.. arXiv 2025. [Paper]
"Layout‑Agnostic Human Activity Recognition in Smart Homes through Textual Descriptions Of Sensor Triggers (TDOST)". Thukral et al.. IMWUT 2025. [Paper]
"LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing Applications". Xu et al.. SenSys 2021. [Paper][Code]
"MASTER: A Multi‑modal Foundation Model for Human Activity Recognition". Zhu et al.. IMWUT 2025. [Paper]
"Pulse‑PPG: An Open‑Source Field‑Trained PPG Foundation Model for Wearable Applications across Lab and Field Settings". Saha et al.. IMWUT 2025. [Paper]
"RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]
"Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]
"SelfPAB: Large‑Scale Pre‑training on Accelerometer Data for Human Activity Recognition". Logacjov et al.. Applied Intelligence 2024. [Paper][Code]
"Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data". Narain et al.. arXiv 2025. [Paper]
"Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices (HAR‑FM)". Qiu et al.. IMWUT 2025. [Paper]
"RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]
"One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation". Wei et al. IMWUT 2025. [Paper]
"Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]
"Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications". Saha et al. IMWUT 2025. [Paper]
"One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation". Wei et al. IMWUT 2025. [Paper]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]
"Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals". Luo et al. arXiv 2024. [Paper]
"SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
Thapa et al. ICML 2024. [Paper][Code]
"Scaling Wearable Foundation Models". Narayanswamy et al. arXiv 2024. [Paper]
"DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs". Tian et al. arXiv 2025. [Paper]
"IMU2CLIP: Language-Grounded Motion Sensor Translation with Multimodal Contrastive Learning". Moon et al. EMNLP Findings 2023. [Paper][Code]
"SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]
"On the Benefit of Generative Foundation Models for Human Activity Recognition". Leng et al. arXiv 2023. [Paper]
"FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition". Weng et al. IEEE TMC 2025. [Paper]
"TxP: Reciprocal Generation of Ground Pressure Dynamics and Activity Descriptions for Improving Human Activity Recognition". Lala et al. IEEE TMC 2025. [Paper]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"IMUGPT 2.0: Language-Based Cross Modality Transfer for Sensor-Based Human Activity Recognition". Leng et al. IMWUT 2024. [Paper][Code]
"AI-Generated Fall Data: Assessing LLMs and Diffusion Model for Wearable Fall Detection". Sana et al. MDPI Sensors 2025. [Paper][Code]
"Weak-Annotation of HAR Datasets using Vision Foundation Models". Bock et al.. ISWC 2024. [Paper][Code]
Tokenization and representation for sensor-based HAR. Single-stream token formation converts raw signals (e.g., IMU, PPG, ambient/RF) into windows, statistical features, spectrograms, or quantized codes; cross-stream scaffolding then synchronizes modalities with positional/meta encodings and performs token fusion or cross-modal projection. The resulting tokens feed pretraining/training backbones (Transformer/ViT/LLM) for tasks such as classification, captioning, retrieval, and forecasting.
"RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]
"Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data". Jaya Narain et al. arXiv 2025. [Paper]
"SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
Thapa et al. ICML 2024. [Paper][Code]
"Leveraging Foundation Models for Zero-Shot IoT Sensing". Dinghao Xue et al. arXiv 2024. [Paper]
"Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)". Christoph Wieland, Victor Pankratius. IEEE Sensors Journal 2025. [Paper]
"A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)". Xiong et al. IEEE BIBM 2024. [Paper]
"SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]
"Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions". Eray Erturk et al. arXiv 2025. [Paper]
"ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs". Post et al. arXiv 2025. [Paper]
"SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
Li et al. arXiv 2024. [Paper][Code]
"Scaling Wearable Foundation Models". Narayanswamy et al. arXiv 2024. [Paper]
"A Personal Health Large Language Model for Sleep and Fitness Coaching (PH-LLM)". Khasentino et al. Nature Medicine 2025. [Paper][Code]
"BioSignal Copilot: Leveraging the Power of LLMs in Drafting Reports for Biomedical Signals". Liu et al. arXiv 2023. [Paper] [Code]
"SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]
"Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals".
Luo et al. arXiv 2024. [Paper]
"Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]
"Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors". Chen et al.. IMWUT 2024. [Paper]
"HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]
"Towards Learning Discrete Representations via Self-Supervision for Wearables-Based Human Activity Recognition".
Harish Haresamudram et al. Sensors 2024. [Paper]
"Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices".
Minghui Qiu et al. ACM IMWUT 2025. [Paper]
"Chronos: Learning the Language of Time Series".
Ansari et al. arXiv 2024. [Paper]
"Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
Pillai et al. PMLR 2025. [Paper][Code]
"SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
Li et al. arXiv 2024. [Paper][Code]
"Visible Light Human Activity Recognition Driven by Generative Language Model".
Yang et al. Elsevier, Information Fusion 2025. [Paper]
"LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
Ouyang et al. arXiv 2024. [Paper]
Pretraining paradigms for sensor-based HAR. Contrastive learns cross-view/cross-modal alignment in a shared latent space (CLIP-like), enabling zero-/few-shot transfer and retrieval. Generative uses masked reconstruction or causal prediction to model temporal continuity and support imputation and text-conditioned decoding. Hybrid / Self-supervised combines contrastive and generative objectives (often at scale) and introduces semantic grounding via language/distillation, improving robustness across users and devices.
"Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)".
Weng et al. ACM 2024. [Paper]
"FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition".
Weng et al. IEEE TMC 2025. [Paper]
"RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]
"IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text".
Moon et al. EMNLP Findings 2023. [Paper][Code]
"HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]
"Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]
"Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)".
Wieland & Pankratius. IEEE Sensors Journal 2025. [Paper]
"LLaSA: Large Multimodal Agent for Human Activity Analysis Through Wearable Sensors".
Imran et al. arXiv 2024. [Paper][Code]
"Scaling Wearable Foundation Models (LSM)".
Narayanswamy et al. arXiv 2024. [Paper]
"SelfPAB: Large-Scale Pre-training on Accelerometer Recordings with Masked Spectrogram Reconstruction".
Logacjov et al. Springer 2024. [Paper][Code]
"HAR-DoReMi: Optimizing Data Mixture for Self-Supervised Pretraining in Human Activity Recognition".
Ban et al. arXiv 2025. [Paper]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"A Personal Health Large Language Model for Sleep and Fitness Coaching (PH-LLM)".
Khasentino et al. Nature Medicine 2025. [Paper][Code]
"Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals (NORMWEAR)".
Luo et al. arXiv 2024. [Paper][Code]
"Scaling Wearable Foundation Models (LSM)".
Narayanswamy et al. arXiv 2024. [Paper]
"SensorLM: Learning the Language of Wearable Sensors".
Zhang et al. arXiv 2025. [Paper][Code]
"SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
Li et al. arXiv 2024. [Paper][Code]
"SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]
Mechanism-centric view of adaptation in HAR foundation models. We emphasize how behavior is changed: PEFT keeps the backbone frozen and learns small add-ons (adapters, LoRA/QLoRA, learnable prefixes); Full/Partial fine-tuning updates all or selected layers for tighter task coupling; and Instruction-tuning & alignment spans prompt-only in-context learning (zero-update) and supervised formatting (SFT/PEFT on curated sensor–text exemplars and task templates) to ensure format adherence and faithful generation.
"LLaSA: Large Multimodal Agent for Human Activity Analysis Through Wearable Sensors".
Imran et al. arXiv 2024. [Paper][Code]
"A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)".
Xiong et al. IEEE BIBM 2024. (No verified link available in file.)
"Leveraging Large Language Models for Digital Phenotyping and Health Forecasting".
Yuan et al. bioRxiv 2025. [Paper]
"Large Language Models for Wearable Sensor-Based Activity Understanding".
Liu et al. Sensors 2024. [Paper]
"Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
Pillai et al. PMLR 2025. [Paper][Code]
"PhysLLM: Harnessing Large Language Models for Physiological Understanding".
Xie et al. arXiv 2025. [Paper]
"MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model".
Bandyopadhyay et al. arXiv 2025. [Paper]
"GOAT: A Generalized Cross-Dataset Activity Recognition Framework".
Miao et al. IMWUT 2024. [Paper]
"SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]
"LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing Applications". Xu et al.. SenSys 2021. [Paper][Code]
"Scaling Wearable Foundation Models (LSM)".
Narayanswamy et al. arXiv 2024. [Paper]
"LLM4HAR: Generalizable On-device Human Activity Recognition with Large Language Models".
Hong et al. IMWUT 2025. [Paper]
"RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]
"SelfPAB: large-scale pre-training on accelerometer data for human activity recognition".
Logacjov et al. ACM 2024. [Paper][Code]
"Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]
"Leveraging Large Language Models for Digital Phenotyping and Health Forecasting".
Yuan et al. bioRxiv 2025. [Paper]
"HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]
"StressLLM: Large Language Models for Stress Prediction and Biomarker Reasoning".
Thapa et al. IEEE ICCE 2025. [Paper]
"Leveraging Large Language Models for Digital Phenotyping: Detecting Depressive State Changes for Patients with Depressive Episodes".
Yuan et al. arXiv 2025. [Paper]
"LAHAR: Leveraging Language Models for Human Activity Recognition".
Chen et al. IEEE Access 2024. [Paper]
"DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs".
Tian et al. arXiv 2025. [Paper]
"Enabling On-Device LLMs Personalization with Sensor Prompts".
Zhang et al. arXiv 2024. [Paper]
"Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors".
Chen et al. IMWUT 2024. [Paper]
"SensorLM: Learning the Language of Wearable Sensors".
Zhang et al. arXiv 2025. [Paper][Code]
"A Personal Health Large Language Model for Sleep and Fitness Coaching (PH-LLM)".
Khasentino et al. Nature Medicine 2025. [Paper][Code]
"Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text Data".
Xu et al. ACM IMWUT 2024. [Paper][Code]
"The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition".
Georgios et al. IEEE ACII 2024. [Paper][Code]
Downstream capabilities and the accompanying generalization protocols used in sensor-based HAR. Top row (left→right): Zero-/few-shot & open-set: recognize unseen activities with 𝑘-shot label budgets and open-set rejection; Cross-dataset / device / user: train on dataset A and test on unseen datasets/devices/users with leave-one-out and cross-position splits; Cross-modal retrieval & search: sensortext/video retrieval evaluated by Recall@K and mAP under cross-domain splits with a shared embedding space. Bottom row: Captioning, Q&A, reasoning: sensor-conditioned decoding (prompts/PEFT), measured by caption/Q&A accuracy and human/expert ratings; Reconstruction, forecasting, imputation: masked-reconstruction/denoising and short/long-horizon forecasting under distribution shift; Federated & on-device evaluation: client-level personalization with communication rounds, reporting edge latency/energy and privacy constraints.
"Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition".
Weng et al. ACM 2024. [Paper]
"FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition".
Weng et al. IEEE TMC 2025. [Paper]
"ZARA: Zero-Shot Motion Time-Series Analysis via LLM-Guided Knowledge Retrieval and Reasoning".
Li et al. arXiv 2025. [Paper][Code]
"HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?".
Ji et al. arXiv 2024. [Paper]
"EEG-GPT: Exploring Capabilities of Large Language Models for EEG-Based Abnormality Detection".
Kim et al. arXiv 2024. [Paper]
"HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]
"Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices (HAR-FM)".
Qiu et al. IMWUT 2025. [Paper]
"One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation".
Wei et al. ACM IMEUT 2025. [Paper]
"MASTER: A Multi-Modal Foundation Model for Human Activity Recognition".
Zhu et al. IMWUT 2025. [Paper]
"A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)".
Xiong et al. IEEE BIBM 2024. (Link unavailable in verified file)
"Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]
"Multimodal Foundation Model for Cross-Modal Retrieval and Recognition (AURA-MFM)".
Matsuishi et al. arXiv 2025. [Paper]
"GLOSS: Group of LLMs for Open-ended Sensemaking of Passive Sensing Data for Health and Wellbeing".
Choube et al. IMWUT 2025. [Paper][Code]
"Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLM".
Li et al. IMWUT 2025. [Paper][Code]
"PhysioLLM: Supporting Personalized Health Insights with Wearables and Large Language Models".
Fang et al. arXiv 2024. [Paper]
"SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
Thapa et al. ICML 2024. [Paper][Code]
"Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors".
Chen et al. IMWUT 2024. [Paper]
"SensorLM: Learning the Language of Wearable Sensors".
Zhang et al. arXiv 2025. [Paper][Code]
"SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
Li et al. arXiv 2024. [Paper][Code]
"Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
Pillai et al. PMLR 2025. [Paper][Code]
"Visible Light Human Activity Recognition Driven by Generative Language Model".
Yang et al. Information Fusion 2025. [Paper]
"LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
Ouyang et al. SenSys-ML 2024. [Paper]
"Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"Scaling Wearable Foundation Models (LSM)".
Narayanswamy et al. arXiv 2024. [Paper]
"Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)".
Wieland & Pankratius. IEEE Sensors Journal 2025. [Paper]
"RobustHAR: Multi-Scale Spatial-Temporal Masked Autoencoder for Robust Human Activity Recognition".
Liu et al. IJCAI 2025. [Paper]
"UniMTS – Unified Pre-training for Motion Time-Series Forecasting and Recognition".
Zhang et al. arXiv 2024. [Paper][Code]
"LLM4HAR: Generalizable On-device Human Activity Recognition with Large Language Models".
Hong et al. KDD 2025. [Paper]
"DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs".
Tian et al. arXiv 2025. [Paper]
"MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model".
Bandyopadhyay et al. arXiv 2025. [Paper]
"ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs".
Post et al. ACM HotMobile 2025. [Paper]
"Enabling On-Device LLMs Personalization with Smartphone Sensing".
Zhang et al. arXiv 2024. [Paper]
"Activity transitions for semi-supervised federated learning in sensor-based human activity recognition".
Bukit et al. Elsevier, Applied Soft Computing 2025. [Paper]
"Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)". Weng et al.. SenSys 2024. [Paper]
"FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition".
Weng et al. IEEE TMC 2025. [Paper]
"Scaling Wearable Foundation Models (LSM)".
Narayanswamy et al. arXiv 2024. [Paper]
"LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
Ouyang et al. SenSys-ML 2024. [Paper]
"Leveraging foundation models for zero-shot IoT sensing".
Xue et al. arXiv 2024. [Paper][Code]
"Large Language Model-Guided Semantic Alignment for Human Activity Recognition".
Yan et al. arXiv 2024. [Paper][Code]
"Enabling On-Device LLMs Personalization with Sensor Prompts".
Zhang et al. arXiv 2024. [Paper]
"LLM4HAR: Generalizable On-Device Human Activity Recognition with Large Language Models".
Hong et al. IMWUT 2025. [Paper]
"Few-Shot Human Activity Recognition Using Lightweight Language Models".
Cruciani et al. IEEE ICCCN 2025. [Paper]
"On-device Foundation Models for Wearable Signals".
Simon et al. 2025. [Paper]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"Enabling Efficient RF Sensing with Small Language Models via Functional Data Analysis and Parameter Efficient Tuning".
Yujie Sun et al. IEEE IoTJ 2026. [Paper]
"HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]
"EdgeHAR: An Edge-Native Compact Sensor Foundation Model for Human Activity Recognition". He Zhang et al. Arxiv 2026. [Paper]
Application domains for sensor-based HAR foundation models. The radial layout highlights four commonly targeted areas: general-purpose HAR / daily living, healthcare and wellbeing, smart-home and context-aware environments, and interactive/agentic assistants.
"One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation".
Wei et al. ACM IMWUT 2025. [Paper]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices (HAR-FM)".
Qiu et al. IMWUT 2025. [Paper]
"PhysioLLM: Supporting Personalized Health Insights with Wearables and Large Language Models".
Fang et al. arXiv 2024. [Paper]
"Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLM".
Li et al. IMWUT 2025. [Paper][Code]
"Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications across Lab and Field Settings".
Saha et al. IMWUT 2025. [Paper]
"SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
Thapa et al. ICML 2024. [Paper][Code]
"The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition".
Georgios et al. IEEE ACII 2024. [Paper][Code]
"Game of LLMs: Discovering Structural Constructs in Activities using Large Language Models". Hiremath et al.. UbiComp Companion 2024. [Paper]
"A Synergistic Large Language Model and Supervised Learning Approach to Zero-Shot and Continual Activity Recognition in Smart Homes". Naoto et al.. ICBDA 2024. [Paper]
"Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)". Weng et al.. SenSys 2024. [Paper]
"Visible light human activity recognition driven by generative language model".
Yang et al. Information Fusion 2025. [Paper]
"LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
Ouyang et al. SenSys-ML 2024. [Paper]
"Large Language Model-Guided Semantic Alignment for Human Activity Recognition".
Yan et al. arXiv 2024. [Paper]
"DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs".
Tian et al. arXiv 2025. [Paper]
"SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]
"You Don't Need To Train: Agentic Heuristic Learning Studio for Executable Human Activity Recognition". Siyu et al. Arxiv 2026. [Paper][Code]
169 commits
33
169 commits
updated Sep 23, 2026
This repository provides a curated, continuously updated index of foundation models (FMs) in the human activity recognition (HAR) domain. The organization follows our survey’s lifecycle-based taxonomy and major development directions. It serves as a living companion to the ACM IMWUT paper “Foundation Models Defining a New Era in Human Activity Recognition: A Survey and Outlook”, offering direct access to representative works, datasets, and model resources. Our goal is to foster transparency, reproducibility, and collaboration across the HAR community as the field transitions toward large-scale, multimodal, and language-grounded sensing models.
Contributions are welcome! Whether you add new papers, improve taxonomy coverage, link open-source implementations, or update existing groups. Please help the community build a shared, evolving reference for next-generation HAR foundation models.
(To include your related work in this repository, please create a pull request with the relevant details or drop us a message through email: sizhen.bian@nwpu.edu.cn)
The paper is openly accessible at: https://dl.acm.org/doi/10.1145/3810230 (with contributors from DFKI (Germany), RPTU (Germany), GIT (America), and NWPU(China)).
Historical development of sensor-based Human Activity Recognition (HAR) models.
Historical development of sensor-based Human Activity Recognition (HAR) models. From classical machine learning with hand-crafted features and shallow classifiers to the rise of deep learning with CNNs and RNNs, the field progressed toward a phase focused on transfer and domain generalization (robustness across users, devices, and datasets). More recently, self-supervised learning (SSL) approaches have enabled pretraining on unlabeled sensor data using contrastive or masked objectives. Today, the field is moving toward foundation models, exemplified by large-scale sensor–language alignment, emphasizing scalability, generalization, and interpretability
Here are the growth trends of publications since 2022 and the model names cloud:
Left: HAR-FM papers showing a sharp acceleration with the vast majority of works emerging since 2024.
Right: Representative model name cloud.
Definition of Foundation Models and it's adaptation in different fields.
Foundation Model: Any model trained on broad data (generally using self-supervision at scale) that can be adapted (e.g., fine-tuned) to a wide range of downstream tasks.[Paper]
Foundation Model in the CV Field: A pre-trained model and its adapters capable of solving all vision tasks within the space–time–modality continuum (ranging from coarse to fine-grained, static to dynamic, and single (RGB) to multimodal sensory inputs) while supporting transferability through zero-/few-shot learning and fine-tuning. [Paper]
Foundation Model in the NLP Field: The large Pretrained Language Models (PLMs), characterized by their ability to generate fluent text, handle multiple modalities, and follow natural-language instructions to perform diverse tasks.[Paper]
Foundation Model in the HAR Field: A pretrained, sensor-grounded model and its adapters that can solve diverse activity-understanding tasks across the sensing–temporal–context continuum while generalizing across sensor modalities, body placements, users, devices, and environments.
Heuristic 1–7 scores of representative works against six HAR–FM criteria. Each radar chart profiles a model on: A) Corpus Coverage & Diversity, B) Cross-Domain Generalization, C) Modality Extensibility & Grounding, D) Label-Efficient Pretraining, E) Adaptation Surfaces & Reusability, and F) Broad Applicability & Emergent Capabilities. The “Ideal HAR-FM” panel depicts a target profile. Scores (1 = limited evidence to 7 = strong evidence) are judgment-based syntheses from reported results (compared both to the other models in this survey and to an aspirational “ideal” FM-for-HAR reference point) and are intended for qualitative comparison rather than a leaderboard.
Note: Although foundation models in HAR domain are still in their formative phase, a model should exhibit core hallmarks of the foundation model paradigm, not necessarily all at once, but in substance. These dimensions outline what defines a model as foundational: the ability to scale across data and users, generalize beyond training domains, adapt efficiently to new tasks, and support reuse across modalities and contexts. Collectively, they set a directional standard rather than a checklist, marking the shift from task-specific modeling toward unified, adaptable representations of human activity.
"A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)". Xiong et al.. IEEE BIBM 2024. [Paper]
"Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]
"Scaling Wearable Foundation Models". Narayanswamy et al.. arXiv 2024. [Paper]
"RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]
"HAR-DoReMi: Optimizing Data Mixture for Self-Supervised Human Activity Recognition Across Heterogeneous IMU Datasets". Ban et al.. arXiv 2025. [Paper]
"Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions". Erturk et al.. arXiv 2025. [Paper]
"LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing Applications". Xu et al.. SenSys 2021. [Paper][Code]
"Cosmo: Contrastive Fusion Learning with Small Data for Multimodal Human Activity Recognition". Ouyang et al.. MobiCom 2022. [Paper][Code]
"Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications across Lab and Field Settings". Saha et al.. IMWUT 2025. [Paper]
"RobustHAR: Multi-scale Spatial-temporal Masked Self-supervised Pre-training for Robust Human Activity Recognition". Liu et al.. IJCAI 2025. [Paper]
"Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices". Minghui Qiu et al. IMWUT 2025. [Paper]
"HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]
| Dataset | Sensors | Datasize | Subjects | Activities |
|---|---|---|---|---|
| CAPTURE-24 | acc | 3883 h | 151 | 200 unique labels |
| TNDA-HAR | acc, gyro | 5.7 h | 23 | 8 daily activities |
| HAR70+ | acc | 12.6 h | 18 | 8 daily activities |
| WISDM | acc, gyro | 91.8 h | 51 | 18 daily activities |
| MotionSense | acc, gyro | — | 24 | 6 daily activities |
| SHL Challenge | acc, gyro, mag | 2812 h | 3 | 8 transport modes |
| Shoaib | acc, gyro, mag | 6.5 h | 10 | 13 daily activities |
| HHAR | acc, gyro | — | 9 | 6 daily activities |
| WHARF | acc | — | 16 | 8 motion primitives |
| DSADS | acc, gyro, mag | 12.7 h | 8 | 19 daily/sports activities |
| UCI-HAR | acc, gyro | — | 30 | 6 daily activities |
| USC-HAD | acc, gyro, mag | — | 14 | 12 daily activities |
| Daphnet FoG | acc | 8.3 h | 10 | 3 walking activities |
| HAPT | acc, gyro | — | 30 | 6 static, dynamic activities |
| REALDISP | acc, gyro, mag | — | 17 | 33 daily, fitness activities |
| UniMiB SHAR | acc | — | 30 | 17 daily, fall activities |
| UMAFall | acc, gyro, mag | 2.2 h | 17 | 11 daily, fall activities |
| MobiAct | acc, gyro | — | 57 | 13 daily, fall activities |
| Skoda Mini Checkpoint | acc (3D) | — | 1 | 10 assembly-line activities |
| Dataset | Sensors | Datasize | Subjects | Activities |
|---|---|---|---|---|
| RecGym | acc, gyro, human body capacitance | 50h | 10 | 12 fitness activities |
| WEAR | acc, video | 19 h | 22 | 18 sports activities |
| iSPL | acc, gyro, stretch | — | 1 | 9 daily activities |
| HARTH | acc, video | 35.9 h | 22 | 12 daily activities |
| w-HAR | acc, gyro, stretch | 3 h | 22 | 7 daily activities |
| RealLifeHAR | acc, gyro, mag, GPS | — | 19 | 4 daily activities |
| MMAct | RGB, keypoints, acc, gyro, ori, Wi‑Fi, pressure | — | 40 | 37 activities |
| HuGaDB | acc, gyro, EMG | 10 h | 18 | 12 activities |
| RealWorld HAR | acc, gyro, mag, GPS, light, sound level | 124.3 h | 15 | 8 daily activities |
| ExtraSensory | acc, gyro, mag, location, audio, additional | — | 60 | 51 activities |
| UTD-MHAD | RGB, depth, skeleton, acc, gyro | — | 8 | 27 activities |
| MHEALTH | acc, gyro, mag, ECG | — | 10 | 12 daily activities |
| Berkeley MHAD | acc, optical capture, video, depth, audio | 1.37 h | 12 | 11 daily activities |
| PAMAP2 | acc, gyro, mag, HR | 10 h | 9 | 18 daily activities |
| Opportunity | acc, gyro, mag, ambient sensors | 25 h | 4 | 9 kitchen + 9 gestures |
| MRI | mmWave, RGB-D, IMU | 5.3 h | 20 | pose estimation |
| NORMWEAR | PPG, ECG, EEG, GSR, IMU | 14,943 h | 20 | pose estimation |
| SensorLM [Paper ] | PPG, EDA, ACC, TEMP, ALT | 59,749 h | 103,731 | sensor-language study |
| Apple Study (AHMS; WBM) [Paper] | HealthKit metrics (27) | > 2.5 B h | 162 K | 57 health tasks |
| WESAD | EDA/PPG/Temp + Acc | 45 h | 15 | stress detection |
| PPG-Dalia | ECG, PPG, IMU, GSR | 36 h | 15 | daily activities |
| Dataset | Sensors | Datasize | Subjects | Activities |
|---|---|---|---|---|
| Sleep-EDF | EEG/EOG/EMG/ECG | 1,576 h | 197 | sleep stages |
| MIT-BIH Arrhythmia | 2‑lead ECG | 1128 h | 47 | ambulatory ECG |
| PTB-XL | 12‑lead ECG | 6.06 h | 18,885 | ECG status |
| TUH EEG | multi‑channel EEG | 1476 h | 675 | seizure activity |
| Cuff‑Less‑BP | ECG, PPG | 72 h | — | blood‑pressure estimation |
| Auditory‑EEG | EEG | 23 h | — | auditory attention |
| PhyAAt | EEG | 33 h | 25 | auditory attention |
| MAUS | ECG, PPG, GSR | 22 h | 22 | cognitive workload/stress |
| Mendeley‑YAAD | ECG, GSR | 5 h | — | affect/stress elicitation |
| Brain‑Cognitive | EEG | 85 h | 20 | cognitive state regulation |
| EPHNOGRAM | ECG, PCG | 61 h | 24 | cardiac auscultation |
| BIDMC | ECG, PPG | 14 h | 53 | clinical monitoring |
| MOODS [Paper] | PPG | 54 K h | 122 | stress monitoring |
| SleepFM | BAS, ECG, respiratory | 112,544 h | 14,068 | sleep quality monitoring |
| RecGym | acc, gyro, human body capacitance | 50h | 10 | 12 fitness activities |
| PPG-Dalia | ECG, PPG, IMU, GSR | 36 h | 15 | daily activities |
| Dataset | Sensors | Datasize | Subjects | Activities |
|---|---|---|---|---|
| CASAS (Aruba/Milan/...) | Ambient binary sensors (motion, doors) | — | — | smart home activities |
| Dataset | Sensors | Datasize | Subjects | Activities |
|---|---|---|---|---|
| mmWave (var.)[Paper] | Range–Doppler / RF point clouds | 5 h | 10 | daily activities (10 scenes) |
| MM-Fi | mmWave, LiDAR, Wi‑Fi, RGB‑D | 10.6 h | 40 | 27 daily activities |
Four base computation graphs for sensor foundation models. Encoder-only stacks (top left) focus on representation learning using a single sensor encoder (e.g., ViT or SSM) with lightweight heads for recognition, retrieval, or forecasting. Dual encoders (top right) independently embed sensor and text/vision streams and align them via a shared latent projection (CLIP-style) for retrieval/zero-shot transfer. Encoder–decoder stacks (bottom left) condition a language/multimodal decoder on encoded sensor tokens (cross-attention) to produce captions, rationales, or structured outputs. Language-model stacks (bottom right), either encoder–decoder or decoder-only, treat sensing as a token sequence using projection/quantization interfaces for forecasting, analysis, and reasoning.
"A Novel Human Activity Recognition Framework Based on Pre‑Trained Foundation Model (Chronos HAR Adapters)". Xiong et al.. IEEE BIBM 2024. [Paper]
"Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions". Erturk et al.. ICML 2025. [Paper]
"Cosmo: Contrastive Fusion Learning with Small Data for Multimodal Human Activity Recognition". Ouyang et al.. MobiCom 2022. [Paper][Code]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"HAR‑DoReMi: Optimizing Data Mixture for Self‑Supervised Human Activity Recognition Across Heterogeneous IMU Datasets". Ban et al.. arXiv 2025. [Paper]
"Layout‑Agnostic Human Activity Recognition in Smart Homes through Textual Descriptions Of Sensor Triggers (TDOST)". Thukral et al.. IMWUT 2025. [Paper]
"LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing Applications". Xu et al.. SenSys 2021. [Paper][Code]
"MASTER: A Multi‑modal Foundation Model for Human Activity Recognition". Zhu et al.. IMWUT 2025. [Paper]
"Pulse‑PPG: An Open‑Source Field‑Trained PPG Foundation Model for Wearable Applications across Lab and Field Settings". Saha et al.. IMWUT 2025. [Paper]
"RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]
"Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]
"SelfPAB: Large‑Scale Pre‑training on Accelerometer Data for Human Activity Recognition". Logacjov et al.. Applied Intelligence 2024. [Paper][Code]
"Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data". Narain et al.. arXiv 2025. [Paper]
"Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices (HAR‑FM)". Qiu et al.. IMWUT 2025. [Paper]
"RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]
"One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation". Wei et al. IMWUT 2025. [Paper]
"Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]
"Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications". Saha et al. IMWUT 2025. [Paper]
"One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation". Wei et al. IMWUT 2025. [Paper]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]
"Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals". Luo et al. arXiv 2024. [Paper]
"SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
Thapa et al. ICML 2024. [Paper][Code]
"Scaling Wearable Foundation Models". Narayanswamy et al. arXiv 2024. [Paper]
"DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs". Tian et al. arXiv 2025. [Paper]
"IMU2CLIP: Language-Grounded Motion Sensor Translation with Multimodal Contrastive Learning". Moon et al. EMNLP Findings 2023. [Paper][Code]
"SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]
"On the Benefit of Generative Foundation Models for Human Activity Recognition". Leng et al. arXiv 2023. [Paper]
"FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition". Weng et al. IEEE TMC 2025. [Paper]
"TxP: Reciprocal Generation of Ground Pressure Dynamics and Activity Descriptions for Improving Human Activity Recognition". Lala et al. IEEE TMC 2025. [Paper]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"IMUGPT 2.0: Language-Based Cross Modality Transfer for Sensor-Based Human Activity Recognition". Leng et al. IMWUT 2024. [Paper][Code]
"AI-Generated Fall Data: Assessing LLMs and Diffusion Model for Wearable Fall Detection". Sana et al. MDPI Sensors 2025. [Paper][Code]
"Weak-Annotation of HAR Datasets using Vision Foundation Models". Bock et al.. ISWC 2024. [Paper][Code]
Tokenization and representation for sensor-based HAR. Single-stream token formation converts raw signals (e.g., IMU, PPG, ambient/RF) into windows, statistical features, spectrograms, or quantized codes; cross-stream scaffolding then synchronizes modalities with positional/meta encodings and performs token fusion or cross-modal projection. The resulting tokens feed pretraining/training backbones (Transformer/ViT/LLM) for tasks such as classification, captioning, retrieval, and forecasting.
"RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]
"Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data". Jaya Narain et al. arXiv 2025. [Paper]
"SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
Thapa et al. ICML 2024. [Paper][Code]
"Leveraging Foundation Models for Zero-Shot IoT Sensing". Dinghao Xue et al. arXiv 2024. [Paper]
"Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)". Christoph Wieland, Victor Pankratius. IEEE Sensors Journal 2025. [Paper]
"A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)". Xiong et al. IEEE BIBM 2024. [Paper]
"SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]
"Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions". Eray Erturk et al. arXiv 2025. [Paper]
"ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs". Post et al. arXiv 2025. [Paper]
"SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
Li et al. arXiv 2024. [Paper][Code]
"Scaling Wearable Foundation Models". Narayanswamy et al. arXiv 2024. [Paper]
"A Personal Health Large Language Model for Sleep and Fitness Coaching (PH-LLM)". Khasentino et al. Nature Medicine 2025. [Paper][Code]
"BioSignal Copilot: Leveraging the Power of LLMs in Drafting Reports for Biomedical Signals". Liu et al. arXiv 2023. [Paper] [Code]
"SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]
"Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals".
Luo et al. arXiv 2024. [Paper]
"Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]
"Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors". Chen et al.. IMWUT 2024. [Paper]
"HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]
"Towards Learning Discrete Representations via Self-Supervision for Wearables-Based Human Activity Recognition".
Harish Haresamudram et al. Sensors 2024. [Paper]
"Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices".
Minghui Qiu et al. ACM IMWUT 2025. [Paper]
"Chronos: Learning the Language of Time Series".
Ansari et al. arXiv 2024. [Paper]
"Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
Pillai et al. PMLR 2025. [Paper][Code]
"SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
Li et al. arXiv 2024. [Paper][Code]
"Visible Light Human Activity Recognition Driven by Generative Language Model".
Yang et al. Elsevier, Information Fusion 2025. [Paper]
"LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
Ouyang et al. arXiv 2024. [Paper]
Pretraining paradigms for sensor-based HAR. Contrastive learns cross-view/cross-modal alignment in a shared latent space (CLIP-like), enabling zero-/few-shot transfer and retrieval. Generative uses masked reconstruction or causal prediction to model temporal continuity and support imputation and text-conditioned decoding. Hybrid / Self-supervised combines contrastive and generative objectives (often at scale) and introduces semantic grounding via language/distillation, improving robustness across users and devices.
"Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)".
Weng et al. ACM 2024. [Paper]
"FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition".
Weng et al. IEEE TMC 2025. [Paper]
"RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]
"IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text".
Moon et al. EMNLP Findings 2023. [Paper][Code]
"HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]
"Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]
"Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)".
Wieland & Pankratius. IEEE Sensors Journal 2025. [Paper]
"LLaSA: Large Multimodal Agent for Human Activity Analysis Through Wearable Sensors".
Imran et al. arXiv 2024. [Paper][Code]
"Scaling Wearable Foundation Models (LSM)".
Narayanswamy et al. arXiv 2024. [Paper]
"SelfPAB: Large-Scale Pre-training on Accelerometer Recordings with Masked Spectrogram Reconstruction".
Logacjov et al. Springer 2024. [Paper][Code]
"HAR-DoReMi: Optimizing Data Mixture for Self-Supervised Pretraining in Human Activity Recognition".
Ban et al. arXiv 2025. [Paper]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"A Personal Health Large Language Model for Sleep and Fitness Coaching (PH-LLM)".
Khasentino et al. Nature Medicine 2025. [Paper][Code]
"Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals (NORMWEAR)".
Luo et al. arXiv 2024. [Paper][Code]
"Scaling Wearable Foundation Models (LSM)".
Narayanswamy et al. arXiv 2024. [Paper]
"SensorLM: Learning the Language of Wearable Sensors".
Zhang et al. arXiv 2025. [Paper][Code]
"SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
Li et al. arXiv 2024. [Paper][Code]
"SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]
Mechanism-centric view of adaptation in HAR foundation models. We emphasize how behavior is changed: PEFT keeps the backbone frozen and learns small add-ons (adapters, LoRA/QLoRA, learnable prefixes); Full/Partial fine-tuning updates all or selected layers for tighter task coupling; and Instruction-tuning & alignment spans prompt-only in-context learning (zero-update) and supervised formatting (SFT/PEFT on curated sensor–text exemplars and task templates) to ensure format adherence and faithful generation.
"LLaSA: Large Multimodal Agent for Human Activity Analysis Through Wearable Sensors".
Imran et al. arXiv 2024. [Paper][Code]
"A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)".
Xiong et al. IEEE BIBM 2024. (No verified link available in file.)
"Leveraging Large Language Models for Digital Phenotyping and Health Forecasting".
Yuan et al. bioRxiv 2025. [Paper]
"Large Language Models for Wearable Sensor-Based Activity Understanding".
Liu et al. Sensors 2024. [Paper]
"Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
Pillai et al. PMLR 2025. [Paper][Code]
"PhysLLM: Harnessing Large Language Models for Physiological Understanding".
Xie et al. arXiv 2025. [Paper]
"MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model".
Bandyopadhyay et al. arXiv 2025. [Paper]
"GOAT: A Generalized Cross-Dataset Activity Recognition Framework".
Miao et al. IMWUT 2024. [Paper]
"SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]
"LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing Applications". Xu et al.. SenSys 2021. [Paper][Code]
"Scaling Wearable Foundation Models (LSM)".
Narayanswamy et al. arXiv 2024. [Paper]
"LLM4HAR: Generalizable On-device Human Activity Recognition with Large Language Models".
Hong et al. IMWUT 2025. [Paper]
"RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]
"SelfPAB: large-scale pre-training on accelerometer data for human activity recognition".
Logacjov et al. ACM 2024. [Paper][Code]
"Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]
"Leveraging Large Language Models for Digital Phenotyping and Health Forecasting".
Yuan et al. bioRxiv 2025. [Paper]
"HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]
"StressLLM: Large Language Models for Stress Prediction and Biomarker Reasoning".
Thapa et al. IEEE ICCE 2025. [Paper]
"Leveraging Large Language Models for Digital Phenotyping: Detecting Depressive State Changes for Patients with Depressive Episodes".
Yuan et al. arXiv 2025. [Paper]
"LAHAR: Leveraging Language Models for Human Activity Recognition".
Chen et al. IEEE Access 2024. [Paper]
"DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs".
Tian et al. arXiv 2025. [Paper]
"Enabling On-Device LLMs Personalization with Sensor Prompts".
Zhang et al. arXiv 2024. [Paper]
"Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors".
Chen et al. IMWUT 2024. [Paper]
"SensorLM: Learning the Language of Wearable Sensors".
Zhang et al. arXiv 2025. [Paper][Code]
"A Personal Health Large Language Model for Sleep and Fitness Coaching (PH-LLM)".
Khasentino et al. Nature Medicine 2025. [Paper][Code]
"Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text Data".
Xu et al. ACM IMWUT 2024. [Paper][Code]
"The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition".
Georgios et al. IEEE ACII 2024. [Paper][Code]
Downstream capabilities and the accompanying generalization protocols used in sensor-based HAR. Top row (left→right): Zero-/few-shot & open-set: recognize unseen activities with 𝑘-shot label budgets and open-set rejection; Cross-dataset / device / user: train on dataset A and test on unseen datasets/devices/users with leave-one-out and cross-position splits; Cross-modal retrieval & search: sensortext/video retrieval evaluated by Recall@K and mAP under cross-domain splits with a shared embedding space. Bottom row: Captioning, Q&A, reasoning: sensor-conditioned decoding (prompts/PEFT), measured by caption/Q&A accuracy and human/expert ratings; Reconstruction, forecasting, imputation: masked-reconstruction/denoising and short/long-horizon forecasting under distribution shift; Federated & on-device evaluation: client-level personalization with communication rounds, reporting edge latency/energy and privacy constraints.
"Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition".
Weng et al. ACM 2024. [Paper]
"FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition".
Weng et al. IEEE TMC 2025. [Paper]
"ZARA: Zero-Shot Motion Time-Series Analysis via LLM-Guided Knowledge Retrieval and Reasoning".
Li et al. arXiv 2025. [Paper][Code]
"HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?".
Ji et al. arXiv 2024. [Paper]
"EEG-GPT: Exploring Capabilities of Large Language Models for EEG-Based Abnormality Detection".
Kim et al. arXiv 2024. [Paper]
"HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]
"Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices (HAR-FM)".
Qiu et al. IMWUT 2025. [Paper]
"One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation".
Wei et al. ACM IMEUT 2025. [Paper]
"MASTER: A Multi-Modal Foundation Model for Human Activity Recognition".
Zhu et al. IMWUT 2025. [Paper]
"A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)".
Xiong et al. IEEE BIBM 2024. (Link unavailable in verified file)
"Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]
"Multimodal Foundation Model for Cross-Modal Retrieval and Recognition (AURA-MFM)".
Matsuishi et al. arXiv 2025. [Paper]
"GLOSS: Group of LLMs for Open-ended Sensemaking of Passive Sensing Data for Health and Wellbeing".
Choube et al. IMWUT 2025. [Paper][Code]
"Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLM".
Li et al. IMWUT 2025. [Paper][Code]
"PhysioLLM: Supporting Personalized Health Insights with Wearables and Large Language Models".
Fang et al. arXiv 2024. [Paper]
"SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
Thapa et al. ICML 2024. [Paper][Code]
"Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors".
Chen et al. IMWUT 2024. [Paper]
"SensorLM: Learning the Language of Wearable Sensors".
Zhang et al. arXiv 2025. [Paper][Code]
"SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
Li et al. arXiv 2024. [Paper][Code]
"Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
Pillai et al. PMLR 2025. [Paper][Code]
"Visible Light Human Activity Recognition Driven by Generative Language Model".
Yang et al. Information Fusion 2025. [Paper]
"LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
Ouyang et al. SenSys-ML 2024. [Paper]
"Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"Scaling Wearable Foundation Models (LSM)".
Narayanswamy et al. arXiv 2024. [Paper]
"Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)".
Wieland & Pankratius. IEEE Sensors Journal 2025. [Paper]
"RobustHAR: Multi-Scale Spatial-Temporal Masked Autoencoder for Robust Human Activity Recognition".
Liu et al. IJCAI 2025. [Paper]
"UniMTS – Unified Pre-training for Motion Time-Series Forecasting and Recognition".
Zhang et al. arXiv 2024. [Paper][Code]
"LLM4HAR: Generalizable On-device Human Activity Recognition with Large Language Models".
Hong et al. KDD 2025. [Paper]
"DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs".
Tian et al. arXiv 2025. [Paper]
"MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model".
Bandyopadhyay et al. arXiv 2025. [Paper]
"ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs".
Post et al. ACM HotMobile 2025. [Paper]
"Enabling On-Device LLMs Personalization with Smartphone Sensing".
Zhang et al. arXiv 2024. [Paper]
"Activity transitions for semi-supervised federated learning in sensor-based human activity recognition".
Bukit et al. Elsevier, Applied Soft Computing 2025. [Paper]
"Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)". Weng et al.. SenSys 2024. [Paper]
"FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition".
Weng et al. IEEE TMC 2025. [Paper]
"Scaling Wearable Foundation Models (LSM)".
Narayanswamy et al. arXiv 2024. [Paper]
"LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
Ouyang et al. SenSys-ML 2024. [Paper]
"Leveraging foundation models for zero-shot IoT sensing".
Xue et al. arXiv 2024. [Paper][Code]
"Large Language Model-Guided Semantic Alignment for Human Activity Recognition".
Yan et al. arXiv 2024. [Paper][Code]
"Enabling On-Device LLMs Personalization with Sensor Prompts".
Zhang et al. arXiv 2024. [Paper]
"LLM4HAR: Generalizable On-Device Human Activity Recognition with Large Language Models".
Hong et al. IMWUT 2025. [Paper]
"Few-Shot Human Activity Recognition Using Lightweight Language Models".
Cruciani et al. IEEE ICCCN 2025. [Paper]
"On-device Foundation Models for Wearable Signals".
Simon et al. 2025. [Paper]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"Enabling Efficient RF Sensing with Small Language Models via Functional Data Analysis and Parameter Efficient Tuning".
Yujie Sun et al. IEEE IoTJ 2026. [Paper]
"HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]
"EdgeHAR: An Edge-Native Compact Sensor Foundation Model for Human Activity Recognition". He Zhang et al. Arxiv 2026. [Paper]
Application domains for sensor-based HAR foundation models. The radial layout highlights four commonly targeted areas: general-purpose HAR / daily living, healthcare and wellbeing, smart-home and context-aware environments, and interactive/agentic assistants.
"One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation".
Wei et al. ACM IMWUT 2025. [Paper]
"CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
Hong et al. IMWUT 2024. [Paper][Code]
"Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices (HAR-FM)".
Qiu et al. IMWUT 2025. [Paper]
"PhysioLLM: Supporting Personalized Health Insights with Wearables and Large Language Models".
Fang et al. arXiv 2024. [Paper]
"Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLM".
Li et al. IMWUT 2025. [Paper][Code]
"Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications across Lab and Field Settings".
Saha et al. IMWUT 2025. [Paper]
"SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
Thapa et al. ICML 2024. [Paper][Code]
"The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition".
Georgios et al. IEEE ACII 2024. [Paper][Code]
"Game of LLMs: Discovering Structural Constructs in Activities using Large Language Models". Hiremath et al.. UbiComp Companion 2024. [Paper]
"A Synergistic Large Language Model and Supervised Learning Approach to Zero-Shot and Continual Activity Recognition in Smart Homes". Naoto et al.. ICBDA 2024. [Paper]
"Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)". Weng et al.. SenSys 2024. [Paper]
"Visible light human activity recognition driven by generative language model".
Yang et al. Information Fusion 2025. [Paper]
"LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
Ouyang et al. SenSys-ML 2024. [Paper]
"Large Language Model-Guided Semantic Alignment for Human Activity Recognition".
Yan et al. arXiv 2024. [Paper]
"DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs".
Tian et al. arXiv 2025. [Paper]
"SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]
"You Don't Need To Train: Agentic Heuristic Learning Studio for Executable Human Activity Recognition". Siyu et al. Arxiv 2026. [Paper][Code]
169 commits