zhaxidele/Foundation-Models-Defining-A-New-Era-In-Human-Activity-Recognition

33

169 commits

updated Sep 23, 2026

See the code

README

Description

Foundation Models Defining A New Era In Sensor-based Human Activity Recognition

This repository provides a curated, continuously updated index of foundation models (FMs) in the human activity recognition (HAR) domain. The organization follows our survey’s lifecycle-based taxonomy and major development directions. It serves as a living companion to the ACM IMWUT paper “Foundation Models Defining a New Era in Human Activity Recognition: A Survey and Outlook”, offering direct access to representative works, datasets, and model resources. Our goal is to foster transparency, reproducibility, and collaboration across the HAR community as the field transitions toward large-scale, multimodal, and language-grounded sensing models.

Contributions are welcome! Whether you add new papers, improve taxonomy coverage, link open-source implementations, or update existing groups. Please help the community build a shared, evolving reference for next-generation HAR foundation models.

(To include your related work in this repository, please create a pull request with the relevant details or drop us a message through email: sizhen.bian@nwpu.edu.cn)

The paper is openly accessible at: https://dl.acm.org/doi/10.1145/3810230 (with contributors from DFKI (Germany), RPTU (Germany), GIT (America), and NWPU(China)).

The brief history of sensor-based HAR

Description

Historical development of sensor-based Human Activity Recognition (HAR) models.

Historical development of sensor-based Human Activity Recognition (HAR) models. From classical machine learning with hand-crafted features and shallow classifiers to the rise of deep learning with CNNs and RNNs, the field progressed toward a phase focused on transfer and domain generalization (robustness across users, devices, and datasets). More recently, self-supervised learning (SSL) approaches have enabled pretraining on unlabeled sensor data using contrastive or masked objectives. Today, the field is moving toward foundation models, exemplified by large-scale sensor–language alignment, emphasizing scalability, generalization, and interpretability

Here are the growth trends of publications since 2022 and the model names cloud:

Left: HAR-FM papers showing a sharp acceleration with the vast majority of works emerging since 2024.
Right: Representative model name cloud.

Table of Contents

Definition of FM in the HAR Domain

Description

Definition of Foundation Models and it's adaptation in different fields.

Foundation Model: Any model trained on broad data (generally using self-supervision at scale) that can be adapted (e.g., fine-tuned) to a wide range of downstream tasks.[Paper]

Foundation Model in the CV Field: A pre-trained model and its adapters capable of solving all vision tasks within the space–time–modality continuum (ranging from coarse to fine-grained, static to dynamic, and single (RGB) to multimodal sensory inputs) while supporting transferability through zero-/few-shot learning and fine-tuning. [Paper]

Foundation Model in the NLP Field: The large Pretrained Language Models (PLMs), characterized by their ability to generate fluent text, handle multiple modalities, and follow natural-language instructions to perform diverse tasks.[Paper]

Foundation Model in the HAR Field: A pretrained, sensor-grounded model and its adapters that can solve diverse activity-understanding tasks across the sensing–temporal–context continuum while generalizing across sensor modalities, body placements, users, devices, and environments.

HAR-FM Criteria

Description

Heuristic 1–7 scores of representative works against six HAR–FM criteria. Each radar chart profiles a model on: A) Corpus Coverage & Diversity, B) Cross-Domain Generalization, C) Modality Extensibility & Grounding, D) Label-Efficient Pretraining, E) Adaptation Surfaces & Reusability, and F) Broad Applicability & Emergent Capabilities. The “Ideal HAR-FM” panel depicts a target profile. Scores (1 = limited evidence to 7 = strong evidence) are judgment-based syntheses from reported results (compared both to the other models in this survey and to an aspirational “ideal” FM-for-HAR reference point) and are intended for qualitative comparison rather than a leaderboard.

Note: Although foundation models in HAR domain are still in their formative phase, a model should exhibit core hallmarks of the foundation model paradigm, not necessarily all at once, but in substance. These dimensions outline what defines a model as foundational: the ability to scale across data and users, generalize beyond training domains, adapt efficiently to new tasks, and support reuse across modalities and contexts. Collectively, they set a directional standard rather than a checklist, marking the shift from task-specific modeling toward unified, adaptable representations of human activity.

Major Directions

Developing HAR-Specific Foundation Models from Scratch

  1. "A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)". Xiong et al.. IEEE BIBM 2024. [Paper]

  2. "Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]

  3. "Scaling Wearable Foundation Models". Narayanswamy et al.. arXiv 2024. [Paper]

  4. "RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
    Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]

  5. "HAR-DoReMi: Optimizing Data Mixture for Self-Supervised Human Activity Recognition Across Heterogeneous IMU Datasets". Ban et al.. arXiv 2025. [Paper]

  6. "Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions". Erturk et al.. arXiv 2025. [Paper]

  7. "LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing Applications". Xu et al.. SenSys 2021. [Paper][Code]

  8. "Cosmo: Contrastive Fusion Learning with Small Data for Multimodal Human Activity Recognition". Ouyang et al.. MobiCom 2022. [Paper][Code]

  9. "Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]

  10. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  11. "Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications across Lab and Field Settings". Saha et al.. IMWUT 2025. [Paper]

  12. "RobustHAR: Multi-scale Spatial-temporal Masked Self-supervised Pre-training for Robust Human Activity Recognition". Liu et al.. IJCAI 2025. [Paper]

  13. "Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices". Minghui Qiu et al. IMWUT 2025. [Paper]

  14. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Adapting General Time-Series and Multimodal Foundation Models to HAR

  1. "A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model". Fuhai Xiong et al. BIBM 2024. [Paper]
  2. "Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)". Wieland & Pankratius. IEEE Sensors Journal 2025 (accepted). [Paper]
  3. "IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text Narrations". Moon et al.. EMNLP Findings 2023. [Paper][Code]
  4. "GOAT: A Generalized Cross-Dataset Activity Recognition Framework with Natural Language Supervision". Miao & Chen. IMWUT 2024. [Paper]
  5. "GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images". Lan et al.. arXiv 2025. [Paper][Code]
  6. "UniMTS: Unified Pre-training for Motion Time Series". Zhang et al.. NeurIPS 2024. [Paper][Code]
  7. "Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
    Pillai et al. PMLR 2025. [Paper][Code]

Leveraging Large Language Models for Human Activity Recognition

  1. "LanHAR: Language-centered Human Activity Recognition". Yan et al.. arXiv 2025. [Paper][Code]
  2. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]
  3. "Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
    Pillai et al. PMLR 2025. [Paper][Code]
  4. "DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs". Tian et al.. arXiv 2025. [Paper]
  5. "ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs". Post et al.. HotMobile 2025. [Paper]
  6. "HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?". Ji et al.. arXiv 2024. [Paper]
  7. "LLaSA: Large Multimodal Agent for Human Activity Analysis Through Wearable Sensors". Imran et al.. arXiv 2024. [Paper][Code]
  8. "StressLLM: Large Language Models for Stress Prediction via Wearable Sensor Data". Thapa et al.. IEEE ICCE 2025. [Paper]
  9. "SensorLM: Learning the Language of Wearable Sensors". Zhang et al.. arXiv 2025. [Paper][Code]
  10. "Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors". Chen et al.. IMWUT 2024. [Paper]
  11. "ZARA: Training-Free Motion Time-Series Reasoning via Evidence-Grounded LLM Agents". Zechen Li, et al.. ACL 2026. [Paper][Code]
  12. "RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge Deployment". Hansi Karunarathna, et al.. Arxiv 2026. [Paper]

Datasets Applied in FM Pretraining and Downstream Tasks

IMU-only Datasets

DatasetSensorsDatasizeSubjectsActivities
CAPTURE-24acc3883 h151200 unique labels
TNDA-HARacc, gyro5.7 h238 daily activities
HAR70+acc12.6 h188 daily activities
WISDMacc, gyro91.8 h5118 daily activities
MotionSenseacc, gyro—246 daily activities
SHL Challengeacc, gyro, mag2812 h38 transport modes
Shoaibacc, gyro, mag6.5 h1013 daily activities
HHARacc, gyro—96 daily activities
WHARFacc—168 motion primitives
DSADSacc, gyro, mag12.7 h819 daily/sports activities
UCI-HARacc, gyro—306 daily activities
USC-HADacc, gyro, mag—1412 daily activities
Daphnet FoGacc8.3 h103 walking activities
HAPTacc, gyro—306 static, dynamic activities
REALDISPacc, gyro, mag—1733 daily, fitness activities
UniMiB SHARacc—3017 daily, fall activities
UMAFallacc, gyro, mag2.2 h1711 daily, fall activities
MobiActacc, gyro—5713 daily, fall activities
Skoda Mini Checkpointacc (3D)—110 assembly-line activities

IMU+ Multimodal Datasets

DatasetSensorsDatasizeSubjectsActivities
RecGymacc, gyro, human body capacitance50h1012 fitness activities
WEARacc, video19 h2218 sports activities
iSPLacc, gyro, stretch—19 daily activities
HARTHacc, video35.9 h2212 daily activities
w-HARacc, gyro, stretch3 h227 daily activities
RealLifeHARacc, gyro, mag, GPS—194 daily activities
MMActRGB, keypoints, acc, gyro, ori, Wi‑Fi, pressure—4037 activities
HuGaDBacc, gyro, EMG10 h1812 activities
RealWorld HARacc, gyro, mag, GPS, light, sound level124.3 h158 daily activities
ExtraSensoryacc, gyro, mag, location, audio, additional—6051 activities
UTD-MHADRGB, depth, skeleton, acc, gyro—827 activities
MHEALTHacc, gyro, mag, ECG—1012 daily activities
Berkeley MHADacc, optical capture, video, depth, audio1.37 h1211 daily activities
PAMAP2acc, gyro, mag, HR10 h918 daily activities
Opportunityacc, gyro, mag, ambient sensors25 h49 kitchen + 9 gestures
MRImmWave, RGB-D, IMU5.3 h20pose estimation
NORMWEARPPG, ECG, EEG, GSR, IMU14,943 h20pose estimation
SensorLM [Paper ]PPG, EDA, ACC, TEMP, ALT59,749 h103,731sensor-language study
Apple Study (AHMS; WBM) [Paper]HealthKit metrics (27)> 2.5 B h162 K57 health tasks
WESADEDA/PPG/Temp + Acc45 h15stress detection
PPG-DaliaECG, PPG, IMU, GSR36 h15daily activities

Physiology Datasets

DatasetSensorsDatasizeSubjectsActivities
Sleep-EDFEEG/EOG/EMG/ECG1,576 h197sleep stages
MIT-BIH Arrhythmia2‑lead ECG1128 h47ambulatory ECG
PTB-XL12‑lead ECG6.06 h18,885ECG status
TUH EEGmulti‑channel EEG1476 h675seizure activity
Cuff‑Less‑BPECG, PPG72 h—blood‑pressure estimation
Auditory‑EEGEEG23 h—auditory attention
PhyAAtEEG33 h25auditory attention
MAUSECG, PPG, GSR22 h22cognitive workload/stress
Mendeley‑YAADECG, GSR5 h—affect/stress elicitation
Brain‑CognitiveEEG85 h20cognitive state regulation
EPHNOGRAMECG, PCG61 h24cardiac auscultation
BIDMCECG, PPG14 h53clinical monitoring
MOODS [Paper]PPG54 K h122stress monitoring
SleepFMBAS, ECG, respiratory112,544 h14,068sleep quality monitoring
RecGymacc, gyro, human body capacitance50h1012 fitness activities
PPG-DaliaECG, PPG, IMU, GSR36 h15daily activities

Smart Home Datasets

DatasetSensorsDatasizeSubjectsActivities
CASAS (Aruba/Milan/...)Ambient binary sensors (motion, doors)——smart home activities

RF Datasets

DatasetSensorsDatasizeSubjectsActivities
mmWave (var.)[Paper]Range–Doppler / RF point clouds5 h10daily activities (10 scenes)
MM-FimmWave, LiDAR, Wi‑Fi, RGB‑D10.6 h4027 daily activities

Base Architecture

Description

Four base computation graphs for sensor foundation models. Encoder-only stacks (top left) focus on representation learning using a single sensor encoder (e.g., ViT or SSM) with lightweight heads for recognition, retrieval, or forecasting. Dual encoders (top right) independently embed sensor and text/vision streams and align them via a shared latent projection (CLIP-style) for retrieval/zero-shot transfer. Encoder–decoder stacks (bottom left) condition a language/multimodal decoder on encoded sensor tokens (cross-attention) to produce captions, rationales, or structured outputs. Language-model stacks (bottom right), either encoder–decoder or decoder-only, treat sensing as a token sequence using projection/quantization interfaces for forecasting, analysis, and reasoning.

Encoder-only stacks

  1. "A Novel Human Activity Recognition Framework Based on Pre‑Trained Foundation Model (Chronos HAR Adapters)". Xiong et al.. IEEE BIBM 2024. [Paper]

  2. "Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions". Erturk et al.. ICML 2025. [Paper]

  3. "Cosmo: Contrastive Fusion Learning with Small Data for Multimodal Human Activity Recognition". Ouyang et al.. MobiCom 2022. [Paper][Code]

  4. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  5. "HAR‑DoReMi: Optimizing Data Mixture for Self‑Supervised Human Activity Recognition Across Heterogeneous IMU Datasets". Ban et al.. arXiv 2025. [Paper]

  6. "Layout‑Agnostic Human Activity Recognition in Smart Homes through Textual Descriptions Of Sensor Triggers (TDOST)". Thukral et al.. IMWUT 2025. [Paper]

  7. "LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing Applications". Xu et al.. SenSys 2021. [Paper][Code]

  8. "MASTER: A Multi‑modal Foundation Model for Human Activity Recognition". Zhu et al.. IMWUT 2025. [Paper]

  9. "Pulse‑PPG: An Open‑Source Field‑Trained PPG Foundation Model for Wearable Applications across Lab and Field Settings". Saha et al.. IMWUT 2025. [Paper]

  10. "RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
    Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]

  11. "Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]

  12. "SelfPAB: Large‑Scale Pre‑training on Accelerometer Data for Human Activity Recognition". Logacjov et al.. Applied Intelligence 2024. [Paper][Code]

  13. "Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data". Narain et al.. arXiv 2025. [Paper]

  14. "Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices (HAR‑FM)". Qiu et al.. IMWUT 2025. [Paper]

Dual-encoder

  1. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition". Weng et al.. IEEE TMC 2025. [Paper]
  2. "IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text Narrations". Moon et al.. EMNLP 2023. [Paper][Code]
  3. "Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)". Wieland et al.. Sensors 2025. [Paper]
  4. "Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)". Weng et al.. SenSys 2024. [Paper]
  5. "Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition – And Ways to Overcome Them". Haresamudram et al.. AAAI 2025.
  6. "Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition (AURA-MFM)". Matsuishi et al.. arXiv 2025. [Paper]
  7. "PRimuS: Pretraining IMU Encoders with Multimodal Self-Supervision". Das et al.. ICASSP 2025. [Paper][[Code](https://arxiv.org/abs/2506.03174]
  8. "SensorLM: Learning the Language of Wearable Sensors". Zhang et al.. arXiv 2025. [Paper][Code]
  9. "SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
    Thapa et al. ICML 2024. [Paper][Code]
  10. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals (NORMWEAR)".
    Luo et al. arXiv 2024. [Paper][Code]
  11. "UniMTS – Unified Pre-training for Motion Time-Series". Zhang et al.. NeurIPS 2024. [Paper][Code]
  12. "GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images". Lan et al.. arXiv 2025. [Paper][Code]
  13. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Encoder-decoder stacks

  1. "LSM: Large-Scale Masked Modeling of Daily Summaries for Population-Level Behavior Modeling". —. arXiv 2025. [Paper]
  2. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals (NORMWEAR)".
    Luo et al. arXiv 2024. [Paper][Code]
  3. "RobustHAR: Multi-Scale Spatial-Temporal Masked Autoencoder for Robust Human Activity Recognition".
    Liu et al. IJCAI 2025. [Paper]
  4. "Scaling Wearable Foundation Models". Narayanswamy et al. arXiv 2024. [Paper]
  5. "SelfPAB: Large-Scale Pre-training on Accelerometer Data for Human Activity Recognition". Logacjov et al.. Applied Intelligence 2024. [Paper][Code]
  6. "Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors". Chen et al.. IMWUT 2024. [Paper]
  7. "SensorLM: Learning the Language of Wearable Sensors". Zhang et al.. arXiv 2025. [Paper][Code]
  8. "Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]
  9. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals (NORMWEAR)".
    Luo et al. arXiv 2024. [Paper][Code]
  10. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Language-model stacks

  1. "SensorLM: Learning the Language of Wearable Sensors". Zhang et al.. arXiv 2025. [Paper][Code]
  2. "Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
    Pillai et al. PMLR 2025. [Paper][Code]
  3. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]
  4. "LLaSA: A Multimodal LLM for Human Activity Analysis Through Wearable and Smartphone Sensors". Imran et al.. arXiv 2024. [Paper][Code]
  5. "HARGPT: A Language‑Conditioned Foundation Model for Human Activity Recognition". Ji et al.. arXiv 2024. [Paper]
  6. "LanHAR: Language‑Driven Human Activity Recognition". Yan et al.. arXiv 2024. [Paper][Code]
  7. "StressLLM: Large Language Models for Wearable Stress Detection". Thapa et al.. arXiv 2025. [Paper]
  8. "DailyLLM: Large Language Models for Daily Behavior Understanding". Kang et al.. arXiv 2025. [Paper]
  9. "ContextLLM: Multimodal Context Understanding from Wearable Devices". Wang et al.. arXiv 2025. [Paper]
  10. "Health‑LLM: Aligning Large Language Models with Wearable Sensor Health Data". Liu et al.. Information Fusion 2025. [Paper][Code]
  11. "SensorGPT: Generative Pretraining for Wearable Sensing". Sharma et al.. arXiv 2025. [Paper]
  12. "Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors". Chen et al.. IMWUT 2024. [Paper]

Modality Scope

Unimodal foundation models

  1. "RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
    Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]

  2. "One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation". Wei et al. IMWUT 2025. [Paper]

  3. "Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]

  4. "Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications". Saha et al. IMWUT 2025. [Paper]

Multimodal foundation models

  1. "Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices". Qiu et al. arXiv 2025. [Paper]
  2. "MASTER: A Multi-modal Foundation Model for Human Activity Recognition". Zhu et al. IEEE TMC 2025. [Paper]
  3. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals (NORMWEAR)".
    Luo et al. arXiv 2024. [Paper][Code]
  4. "MuJo: Multimodal Joint Feature Space Learning for Human Activity Recognition". Fritsch et al. arXiv 2024. [Paper]
  5. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Cross-modal foundation models

  1. "SensorLM: Learning the Language of Wearable Sensors". Zhang et al. arXiv 2025. [Paper][Code]
  2. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]
  3. "IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text". Moon et al. EMNLP Findings 2023. [Paper][Code]
  4. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition". Weng et al. IEEE TMC 2025. [Paper]
  5. "Leveraging foundation models for zero-shot IoT sensing". Xue et al. arXiv 2024. [Paper]

Data Landscape: Collected and Generated Corpora

Corpora of sensor data collected in the wild

  1. "One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation". Wei et al. IMWUT 2025. [Paper]

  2. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  3. "Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]

  4. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals". Luo et al. arXiv 2024. [Paper]

  5. "SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
    Thapa et al. ICML 2024. [Paper][Code]

  6. "Scaling Wearable Foundation Models". Narayanswamy et al. arXiv 2024. [Paper]

  7. "DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs". Tian et al. arXiv 2025. [Paper]

  8. "IMU2CLIP: Language-Grounded Motion Sensor Translation with Multimodal Contrastive Learning". Moon et al. EMNLP Findings 2023. [Paper][Code]

  9. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Generated datasets and augmentation

  1. "On the Benefit of Generative Foundation Models for Human Activity Recognition". Leng et al. arXiv 2023. [Paper]

  2. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition". Weng et al. IEEE TMC 2025. [Paper]

  3. "TxP: Reciprocal Generation of Ground Pressure Dynamics and Activity Descriptions for Improving Human Activity Recognition". Lala et al. IEEE TMC 2025. [Paper]

  4. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  5. "IMUGPT 2.0: Language-Based Cross Modality Transfer for Sensor-Based Human Activity Recognition". Leng et al. IMWUT 2024. [Paper][Code]

  6. "AI-Generated Fall Data: Assessing LLMs and Diffusion Model for Wearable Fall Detection". Sana et al. MDPI Sensors 2025. [Paper][Code]

  7. "Weak-Annotation of HAR Datasets using Vision Foundation Models". Bock et al.. ISWC 2024. [Paper][Code]

Tokenization and Representation Strategies

Description

Tokenization and representation for sensor-based HAR. Single-stream token formation converts raw signals (e.g., IMU, PPG, ambient/RF) into windows, statistical features, spectrograms, or quantized codes; cross-stream scaffolding then synchronizes modalities with positional/meta encodings and performs token fusion or cross-modal projection. The resulting tokens feed pretraining/training backbones (Transformer/ViT/LLM) for tasks such as classification, captioning, retrieval, and forecasting.

Window-based and patch-level segmentation

  1. "RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
    Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]

  2. "Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data". Jaya Narain et al. arXiv 2025. [Paper]

  3. "SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
    Thapa et al. ICML 2024. [Paper][Code]

  4. "Leveraging Foundation Models for Zero-Shot IoT Sensing". Dinghao Xue et al. arXiv 2024. [Paper]

  5. "Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)". Christoph Wieland, Victor Pankratius. IEEE Sensors Journal 2025. [Paper]

  6. "A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)". Xiong et al. IEEE BIBM 2024. [Paper]

  7. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Feature-based aggregation and statistical embeddings

  1. "Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions". Eray Erturk et al. arXiv 2025. [Paper]

  2. "ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs". Post et al. arXiv 2025. [Paper]

  3. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]

  4. "Scaling Wearable Foundation Models". Narayanswamy et al. arXiv 2024. [Paper]

  5. "A Personal Health Large Language Model for Sleep and Fitness Coaching (PH-LLM)". Khasentino et al. Nature Medicine 2025. [Paper][Code]

  6. "BioSignal Copilot: Leveraging the Power of LLMs in Drafting Reports for Biomedical Signals". Liu et al. arXiv 2023. [Paper] [Code]

  7. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Spectrogram and frequency-domain embeddings

  1. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals".
    Luo et al. arXiv 2024. [Paper]

  2. "Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]

  3. "Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors". Chen et al.. IMWUT 2024. [Paper]

  4. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Discrete and quantized sensor tokens

  1. "Towards Learning Discrete Representations via Self-Supervision for Wearables-Based Human Activity Recognition".
    Harish Haresamudram et al. Sensors 2024. [Paper]

  2. "Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices".
    Minghui Qiu et al. ACM IMWUT 2025. [Paper]

  3. "Chronos: Learning the Language of Time Series".
    Ansari et al. arXiv 2024. [Paper]

  4. "Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
    Pillai et al. PMLR 2025. [Paper][Code]

Multimodal alignment and positional encoding

  1. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]

  2. "Visible Light Human Activity Recognition Driven by Generative Language Model".
    Yang et al. Elsevier, Information Fusion 2025. [Paper]

  3. "LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
    Ouyang et al. arXiv 2024. [Paper]

Token fusion and cross-modal projection

  1. "IMU2CLIP: Language-Grounded Motion Sensor Translation with Multimodal Contrastive Learning". Moon et al. EMNLP Findings 2023. [Paper][Code]
  2. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition". Weng et al. IEEE TMC 2025. [Paper]
  3. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Pretraining Paradigms

Description

Pretraining paradigms for sensor-based HAR. Contrastive learns cross-view/cross-modal alignment in a shared latent space (CLIP-like), enabling zero-/few-shot transfer and retrieval. Generative uses masked reconstruction or causal prediction to model temporal continuity and support imputation and text-conditioned decoding. Hybrid / Self-supervised combines contrastive and generative objectives (often at scale) and introduces semantic grounding via language/distillation, improving robustness across users and devices.

Contrastive pretraining

  1. "Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)".
    Weng et al. ACM 2024. [Paper]

  2. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition".
    Weng et al. IEEE TMC 2025. [Paper]

  3. "RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
    Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]

  4. "IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text".
    Moon et al. EMNLP Findings 2023. [Paper][Code]

  5. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Generative pretraining

  1. "Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]

  2. "Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)".
    Wieland & Pankratius. IEEE Sensors Journal 2025. [Paper]

  3. "LLaSA: Large Multimodal Agent for Human Activity Analysis Through Wearable Sensors".
    Imran et al. arXiv 2024. [Paper][Code]

  4. "Scaling Wearable Foundation Models (LSM)".
    Narayanswamy et al. arXiv 2024. [Paper]

  5. "SelfPAB: Large-Scale Pre-training on Accelerometer Recordings with Masked Spectrogram Reconstruction".
    Logacjov et al. Springer 2024. [Paper][Code]

Hybrid and self-supervised pretraining

  1. "HAR-DoReMi: Optimizing Data Mixture for Self-Supervised Pretraining in Human Activity Recognition".
    Ban et al. arXiv 2025. [Paper]

  2. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  3. "A Personal Health Large Language Model for Sleep and Fitness Coaching (PH-LLM)".
    Khasentino et al. Nature Medicine 2025. [Paper][Code]

  4. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals (NORMWEAR)".
    Luo et al. arXiv 2024. [Paper][Code]

  5. "Scaling Wearable Foundation Models (LSM)".
    Narayanswamy et al. arXiv 2024. [Paper]

  6. "SensorLM: Learning the Language of Wearable Sensors".
    Zhang et al. arXiv 2025. [Paper][Code]

  7. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]

  8. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Adaptation Strategies

Description

Mechanism-centric view of adaptation in HAR foundation models. We emphasize how behavior is changed: PEFT keeps the backbone frozen and learns small add-ons (adapters, LoRA/QLoRA, learnable prefixes); Full/Partial fine-tuning updates all or selected layers for tighter task coupling; and Instruction-tuning & alignment spans prompt-only in-context learning (zero-update) and supervised formatting (SFT/PEFT on curated sensor–text exemplars and task templates) to ensure format adherence and faithful generation.

Parameter-efficient fine-tuning (PEFT)

  1. "LLaSA: Large Multimodal Agent for Human Activity Analysis Through Wearable Sensors".
    Imran et al. arXiv 2024. [Paper][Code]

  2. "A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)".
    Xiong et al. IEEE BIBM 2024. (No verified link available in file.)

  3. "Leveraging Large Language Models for Digital Phenotyping and Health Forecasting".
    Yuan et al. bioRxiv 2025. [Paper]

  4. "Large Language Models for Wearable Sensor-Based Activity Understanding".
    Liu et al. Sensors 2024. [Paper]

  5. "Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
    Pillai et al. PMLR 2025. [Paper][Code]

  6. "PhysLLM: Harnessing Large Language Models for Physiological Understanding".
    Xie et al. arXiv 2025. [Paper]

  7. "MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model".
    Bandyopadhyay et al. arXiv 2025. [Paper]

  8. "GOAT: A Generalized Cross-Dataset Activity Recognition Framework".
    Miao et al. IMWUT 2024. [Paper]

  9. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Full or partial fine-tuning

  1. "LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing Applications". Xu et al.. SenSys 2021. [Paper][Code]

  2. "Scaling Wearable Foundation Models (LSM)".
    Narayanswamy et al. arXiv 2024. [Paper]

  3. "LLM4HAR: Generalizable On-device Human Activity Recognition with Large Language Models".
    Hong et al. IMWUT 2025. [Paper]

  4. "RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
    Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]

  5. "SelfPAB: large-scale pre-training on accelerometer data for human activity recognition".
    Logacjov et al. ACM 2024. [Paper][Code]

  6. "Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]

  7. "Leveraging Large Language Models for Digital Phenotyping and Health Forecasting".
    Yuan et al. bioRxiv 2025. [Paper]

  8. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Instruction-tuning & alignment

  1. "StressLLM: Large Language Models for Stress Prediction and Biomarker Reasoning".
    Thapa et al. IEEE ICCE 2025. [Paper]

  2. "Leveraging Large Language Models for Digital Phenotyping: Detecting Depressive State Changes for Patients with Depressive Episodes".
    Yuan et al. arXiv 2025. [Paper]

  3. "LAHAR: Leveraging Language Models for Human Activity Recognition".
    Chen et al. IEEE Access 2024. [Paper]

  4. "DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs".
    Tian et al. arXiv 2025. [Paper]

  5. "Enabling On-Device LLMs Personalization with Sensor Prompts".
    Zhang et al. arXiv 2024. [Paper]

  6. "Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors".
    Chen et al. IMWUT 2024. [Paper]

  7. "SensorLM: Learning the Language of Wearable Sensors".
    Zhang et al. arXiv 2025. [Paper][Code]

  8. "A Personal Health Large Language Model for Sleep and Fitness Coaching (PH-LLM)".
    Khasentino et al. Nature Medicine 2025. [Paper][Code]

  9. "Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text Data".
    Xu et al. ACM IMWUT 2024. [Paper][Code]

  10. "The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition".
    Georgios et al. IEEE ACII 2024. [Paper][Code]

Downstream Capabilities

Description

Downstream capabilities and the accompanying generalization protocols used in sensor-based HAR. Top row (left→right): Zero-/few-shot & open-set: recognize unseen activities with 𝑘-shot label budgets and open-set rejection; Cross-dataset / device / user: train on dataset A and test on unseen datasets/devices/users with leave-one-out and cross-position splits; Cross-modal retrieval & search: sensortext/video retrieval evaluated by Recall@K and mAP under cross-domain splits with a shared embedding space. Bottom row: Captioning, Q&A, reasoning: sensor-conditioned decoding (prompts/PEFT), measured by caption/Q&A accuracy and human/expert ratings; Reconstruction, forecasting, imputation: masked-reconstruction/denoising and short/long-horizon forecasting under distribution shift; Federated & on-device evaluation: client-level personalization with communication rounds, reporting edge latency/energy and privacy constraints.

Zero-/few-shot & open-set recognition

  1. "Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition".
    Weng et al. ACM 2024. [Paper]

  2. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition".
    Weng et al. IEEE TMC 2025. [Paper]

  3. "ZARA: Zero-Shot Motion Time-Series Analysis via LLM-Guided Knowledge Retrieval and Reasoning".
    Li et al. arXiv 2025. [Paper][Code]

  4. "HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?".
    Ji et al. arXiv 2024. [Paper]

  5. "EEG-GPT: Exploring Capabilities of Large Language Models for EEG-Based Abnormality Detection".
    Kim et al. arXiv 2024. [Paper]

  6. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Cross-dataset / cross-device / cross-user generalization

  1. "Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices (HAR-FM)".
    Qiu et al. IMWUT 2025. [Paper]

  2. "One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation".
    Wei et al. ACM IMEUT 2025. [Paper]

  3. "MASTER: A Multi-Modal Foundation Model for Human Activity Recognition".
    Zhu et al. IMWUT 2025. [Paper]

  4. "A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)".
    Xiong et al. IEEE BIBM 2024. (Link unavailable in verified file)

  5. "Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]

  1. "Multimodal Foundation Model for Cross-Modal Retrieval and Recognition (AURA-MFM)".
    Matsuishi et al. arXiv 2025. [Paper]

  2. "GLOSS: Group of LLMs for Open-ended Sensemaking of Passive Sensing Data for Health and Wellbeing".
    Choube et al. IMWUT 2025. [Paper][Code]

  3. "Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLM".
    Li et al. IMWUT 2025. [Paper][Code]

  4. "PhysioLLM: Supporting Personalized Health Insights with Wearables and Large Language Models".
    Fang et al. arXiv 2024. [Paper]

  5. "SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
    Thapa et al. ICML 2024. [Paper][Code]

Language-grounded captioning, Q&A, and reasoning

  1. "Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors".
    Chen et al. IMWUT 2024. [Paper]

  2. "SensorLM: Learning the Language of Wearable Sensors".
    Zhang et al. arXiv 2025. [Paper][Code]

  3. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]

  4. "Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
    Pillai et al. PMLR 2025. [Paper][Code]

  5. "Visible Light Human Activity Recognition Driven by Generative Language Model".
    Yang et al. Information Fusion 2025. [Paper]

  6. "LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
    Ouyang et al. SenSys-ML 2024. [Paper]

Generative reconstruction, forecasting, and imputation

  1. "Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]

  2. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  3. "Scaling Wearable Foundation Models (LSM)".
    Narayanswamy et al. arXiv 2024. [Paper]

  4. "Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)".
    Wieland & Pankratius. IEEE Sensors Journal 2025. [Paper]

  5. "RobustHAR: Multi-Scale Spatial-Temporal Masked Autoencoder for Robust Human Activity Recognition".
    Liu et al. IJCAI 2025. [Paper]

  6. "UniMTS – Unified Pre-training for Motion Time-Series Forecasting and Recognition".
    Zhang et al. arXiv 2024. [Paper][Code]

On-device, federated, and online adaptation

  1. "LLM4HAR: Generalizable On-device Human Activity Recognition with Large Language Models".
    Hong et al. KDD 2025. [Paper]

  2. "DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs".
    Tian et al. arXiv 2025. [Paper]

  3. "MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model".
    Bandyopadhyay et al. arXiv 2025. [Paper]

  4. "ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs".
    Post et al. ACM HotMobile 2025. [Paper]

  5. "Enabling On-Device LLMs Personalization with Smartphone Sensing".
    Zhang et al. arXiv 2024. [Paper]

  6. "Activity transitions for semi-supervised federated learning in sensor-based human activity recognition".
    Bukit et al. Elsevier, Applied Soft Computing 2025. [Paper]

Deployment Settings

Cloud-scale training and centralized evaluation

  1. "Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)". Weng et al.. SenSys 2024. [Paper]

  2. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition".
    Weng et al. IEEE TMC 2025. [Paper]

  3. "Scaling Wearable Foundation Models (LSM)".
    Narayanswamy et al. arXiv 2024. [Paper]

On-device and mobile execution

  1. "LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
    Ouyang et al. SenSys-ML 2024. [Paper]

  2. "Leveraging foundation models for zero-shot IoT sensing".
    Xue et al. arXiv 2024. [Paper][Code]

  3. "Large Language Model-Guided Semantic Alignment for Human Activity Recognition".
    Yan et al. arXiv 2024. [Paper][Code]

  4. "Enabling On-Device LLMs Personalization with Sensor Prompts".
    Zhang et al. arXiv 2024. [Paper]

  5. "LLM4HAR: Generalizable On-Device Human Activity Recognition with Large Language Models".
    Hong et al. IMWUT 2025. [Paper]

  6. "Few-Shot Human Activity Recognition Using Lightweight Language Models".
    Cruciani et al. IEEE ICCCN 2025. [Paper]

  7. "On-device Foundation Models for Wearable Signals".
    Simon et al. 2025. [Paper]

  8. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  9. "Enabling Efficient RF Sensing with Small Language Models via Functional Data Analysis and Parameter Efficient Tuning".
    Yujie Sun et al. IEEE IoTJ 2026. [Paper]

  10. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

  11. "EdgeHAR: An Edge-Native Compact Sensor Foundation Model for Human Activity Recognition". He Zhang et al. Arxiv 2026. [Paper]

Edge-Cloud cooperation

  1. "PFHAR: Practically Adopting Multi-Modal Foundation Model for Human Activity Recognition through Edge-cloud Collaborative Learning".
    Zhengyuan Zhang et al. TMC 2026. [Paper]

Application Domains

Description

Application domains for sensor-based HAR foundation models. The radial layout highlights four commonly targeted areas: general-purpose HAR / daily living, healthcare and wellbeing, smart-home and context-aware environments, and interactive/agentic assistants.

General-purpose HAR / ADL

  1. "One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation".
    Wei et al. ACM IMWUT 2025. [Paper]

  2. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  3. "Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices (HAR-FM)".
    Qiu et al. IMWUT 2025. [Paper]

Healthcare & wellbeing

  1. "PhysioLLM: Supporting Personalized Health Insights with Wearables and Large Language Models".
    Fang et al. arXiv 2024. [Paper]

  2. "Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLM".
    Li et al. IMWUT 2025. [Paper][Code]

  3. "Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications across Lab and Field Settings".
    Saha et al. IMWUT 2025. [Paper]

  4. "SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
    Thapa et al. ICML 2024. [Paper][Code]

  5. "The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition".
    Georgios et al. IEEE ACII 2024. [Paper][Code]

Smart-home & context-aware environments

  1. "Game of LLMs: Discovering Structural Constructs in Activities using Large Language Models". Hiremath et al.. UbiComp Companion 2024. [Paper]

  2. "A Synergistic Large Language Model and Supervised Learning Approach to Zero-Shot and Continual Activity Recognition in Smart Homes". Naoto et al.. ICBDA 2024. [Paper]

  3. "Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)". Weng et al.. SenSys 2024. [Paper]

  4. "Visible light human activity recognition driven by generative language model".
    Yang et al. Information Fusion 2025. [Paper]

Interactive & agentic assistants

  1. "LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
    Ouyang et al. SenSys-ML 2024. [Paper]

  2. "Large Language Model-Guided Semantic Alignment for Human Activity Recognition".
    Yan et al. arXiv 2024. [Paper]

  3. "DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs".
    Tian et al. arXiv 2025. [Paper]

  4. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

  5. "You Don't Need To Train: Agentic Heuristic Learning Studio for Executable Human Activity Recognition". Siyu et al. Arxiv 2026. [Paper][Code]

Contributors

zhaxidele

169 commits

zhaxidele/Foundation-Models-Defining-A-New-Era-In-Human-Activity-Recognition

33

169 commits

updated Sep 23, 2026

See the code

README

Description

Foundation Models Defining A New Era In Sensor-based Human Activity Recognition

This repository provides a curated, continuously updated index of foundation models (FMs) in the human activity recognition (HAR) domain. The organization follows our survey’s lifecycle-based taxonomy and major development directions. It serves as a living companion to the ACM IMWUT paper “Foundation Models Defining a New Era in Human Activity Recognition: A Survey and Outlook”, offering direct access to representative works, datasets, and model resources. Our goal is to foster transparency, reproducibility, and collaboration across the HAR community as the field transitions toward large-scale, multimodal, and language-grounded sensing models.

Contributions are welcome! Whether you add new papers, improve taxonomy coverage, link open-source implementations, or update existing groups. Please help the community build a shared, evolving reference for next-generation HAR foundation models.

(To include your related work in this repository, please create a pull request with the relevant details or drop us a message through email: sizhen.bian@nwpu.edu.cn)

The paper is openly accessible at: https://dl.acm.org/doi/10.1145/3810230 (with contributors from DFKI (Germany), RPTU (Germany), GIT (America), and NWPU(China)).

The brief history of sensor-based HAR

Description

Historical development of sensor-based Human Activity Recognition (HAR) models.

Historical development of sensor-based Human Activity Recognition (HAR) models. From classical machine learning with hand-crafted features and shallow classifiers to the rise of deep learning with CNNs and RNNs, the field progressed toward a phase focused on transfer and domain generalization (robustness across users, devices, and datasets). More recently, self-supervised learning (SSL) approaches have enabled pretraining on unlabeled sensor data using contrastive or masked objectives. Today, the field is moving toward foundation models, exemplified by large-scale sensor–language alignment, emphasizing scalability, generalization, and interpretability

Here are the growth trends of publications since 2022 and the model names cloud:

Left: HAR-FM papers showing a sharp acceleration with the vast majority of works emerging since 2024.
Right: Representative model name cloud.

Table of Contents

Definition of FM in the HAR Domain

Description

Definition of Foundation Models and it's adaptation in different fields.

Foundation Model: Any model trained on broad data (generally using self-supervision at scale) that can be adapted (e.g., fine-tuned) to a wide range of downstream tasks.[Paper]

Foundation Model in the CV Field: A pre-trained model and its adapters capable of solving all vision tasks within the space–time–modality continuum (ranging from coarse to fine-grained, static to dynamic, and single (RGB) to multimodal sensory inputs) while supporting transferability through zero-/few-shot learning and fine-tuning. [Paper]

Foundation Model in the NLP Field: The large Pretrained Language Models (PLMs), characterized by their ability to generate fluent text, handle multiple modalities, and follow natural-language instructions to perform diverse tasks.[Paper]

Foundation Model in the HAR Field: A pretrained, sensor-grounded model and its adapters that can solve diverse activity-understanding tasks across the sensing–temporal–context continuum while generalizing across sensor modalities, body placements, users, devices, and environments.

HAR-FM Criteria

Description

Heuristic 1–7 scores of representative works against six HAR–FM criteria. Each radar chart profiles a model on: A) Corpus Coverage & Diversity, B) Cross-Domain Generalization, C) Modality Extensibility & Grounding, D) Label-Efficient Pretraining, E) Adaptation Surfaces & Reusability, and F) Broad Applicability & Emergent Capabilities. The “Ideal HAR-FM” panel depicts a target profile. Scores (1 = limited evidence to 7 = strong evidence) are judgment-based syntheses from reported results (compared both to the other models in this survey and to an aspirational “ideal” FM-for-HAR reference point) and are intended for qualitative comparison rather than a leaderboard.

Note: Although foundation models in HAR domain are still in their formative phase, a model should exhibit core hallmarks of the foundation model paradigm, not necessarily all at once, but in substance. These dimensions outline what defines a model as foundational: the ability to scale across data and users, generalize beyond training domains, adapt efficiently to new tasks, and support reuse across modalities and contexts. Collectively, they set a directional standard rather than a checklist, marking the shift from task-specific modeling toward unified, adaptable representations of human activity.

Major Directions

Developing HAR-Specific Foundation Models from Scratch

  1. "A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)". Xiong et al.. IEEE BIBM 2024. [Paper]

  2. "Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]

  3. "Scaling Wearable Foundation Models". Narayanswamy et al.. arXiv 2024. [Paper]

  4. "RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
    Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]

  5. "HAR-DoReMi: Optimizing Data Mixture for Self-Supervised Human Activity Recognition Across Heterogeneous IMU Datasets". Ban et al.. arXiv 2025. [Paper]

  6. "Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions". Erturk et al.. arXiv 2025. [Paper]

  7. "LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing Applications". Xu et al.. SenSys 2021. [Paper][Code]

  8. "Cosmo: Contrastive Fusion Learning with Small Data for Multimodal Human Activity Recognition". Ouyang et al.. MobiCom 2022. [Paper][Code]

  9. "Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]

  10. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  11. "Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications across Lab and Field Settings". Saha et al.. IMWUT 2025. [Paper]

  12. "RobustHAR: Multi-scale Spatial-temporal Masked Self-supervised Pre-training for Robust Human Activity Recognition". Liu et al.. IJCAI 2025. [Paper]

  13. "Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices". Minghui Qiu et al. IMWUT 2025. [Paper]

  14. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Adapting General Time-Series and Multimodal Foundation Models to HAR

  1. "A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model". Fuhai Xiong et al. BIBM 2024. [Paper]
  2. "Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)". Wieland & Pankratius. IEEE Sensors Journal 2025 (accepted). [Paper]
  3. "IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text Narrations". Moon et al.. EMNLP Findings 2023. [Paper][Code]
  4. "GOAT: A Generalized Cross-Dataset Activity Recognition Framework with Natural Language Supervision". Miao & Chen. IMWUT 2024. [Paper]
  5. "GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images". Lan et al.. arXiv 2025. [Paper][Code]
  6. "UniMTS: Unified Pre-training for Motion Time Series". Zhang et al.. NeurIPS 2024. [Paper][Code]
  7. "Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
    Pillai et al. PMLR 2025. [Paper][Code]

Leveraging Large Language Models for Human Activity Recognition

  1. "LanHAR: Language-centered Human Activity Recognition". Yan et al.. arXiv 2025. [Paper][Code]
  2. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]
  3. "Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
    Pillai et al. PMLR 2025. [Paper][Code]
  4. "DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs". Tian et al.. arXiv 2025. [Paper]
  5. "ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs". Post et al.. HotMobile 2025. [Paper]
  6. "HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?". Ji et al.. arXiv 2024. [Paper]
  7. "LLaSA: Large Multimodal Agent for Human Activity Analysis Through Wearable Sensors". Imran et al.. arXiv 2024. [Paper][Code]
  8. "StressLLM: Large Language Models for Stress Prediction via Wearable Sensor Data". Thapa et al.. IEEE ICCE 2025. [Paper]
  9. "SensorLM: Learning the Language of Wearable Sensors". Zhang et al.. arXiv 2025. [Paper][Code]
  10. "Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors". Chen et al.. IMWUT 2024. [Paper]
  11. "ZARA: Training-Free Motion Time-Series Reasoning via Evidence-Grounded LLM Agents". Zechen Li, et al.. ACL 2026. [Paper][Code]
  12. "RAG-HAR+: Towards Cost-Efficient LLM-Based Human Activity Recognition for Edge Deployment". Hansi Karunarathna, et al.. Arxiv 2026. [Paper]

Datasets Applied in FM Pretraining and Downstream Tasks

IMU-only Datasets

DatasetSensorsDatasizeSubjectsActivities
CAPTURE-24acc3883 h151200 unique labels
TNDA-HARacc, gyro5.7 h238 daily activities
HAR70+acc12.6 h188 daily activities
WISDMacc, gyro91.8 h5118 daily activities
MotionSenseacc, gyro—246 daily activities
SHL Challengeacc, gyro, mag2812 h38 transport modes
Shoaibacc, gyro, mag6.5 h1013 daily activities
HHARacc, gyro—96 daily activities
WHARFacc—168 motion primitives
DSADSacc, gyro, mag12.7 h819 daily/sports activities
UCI-HARacc, gyro—306 daily activities
USC-HADacc, gyro, mag—1412 daily activities
Daphnet FoGacc8.3 h103 walking activities
HAPTacc, gyro—306 static, dynamic activities
REALDISPacc, gyro, mag—1733 daily, fitness activities
UniMiB SHARacc—3017 daily, fall activities
UMAFallacc, gyro, mag2.2 h1711 daily, fall activities
MobiActacc, gyro—5713 daily, fall activities
Skoda Mini Checkpointacc (3D)—110 assembly-line activities

IMU+ Multimodal Datasets

DatasetSensorsDatasizeSubjectsActivities
RecGymacc, gyro, human body capacitance50h1012 fitness activities
WEARacc, video19 h2218 sports activities
iSPLacc, gyro, stretch—19 daily activities
HARTHacc, video35.9 h2212 daily activities
w-HARacc, gyro, stretch3 h227 daily activities
RealLifeHARacc, gyro, mag, GPS—194 daily activities
MMActRGB, keypoints, acc, gyro, ori, Wi‑Fi, pressure—4037 activities
HuGaDBacc, gyro, EMG10 h1812 activities
RealWorld HARacc, gyro, mag, GPS, light, sound level124.3 h158 daily activities
ExtraSensoryacc, gyro, mag, location, audio, additional—6051 activities
UTD-MHADRGB, depth, skeleton, acc, gyro—827 activities
MHEALTHacc, gyro, mag, ECG—1012 daily activities
Berkeley MHADacc, optical capture, video, depth, audio1.37 h1211 daily activities
PAMAP2acc, gyro, mag, HR10 h918 daily activities
Opportunityacc, gyro, mag, ambient sensors25 h49 kitchen + 9 gestures
MRImmWave, RGB-D, IMU5.3 h20pose estimation
NORMWEARPPG, ECG, EEG, GSR, IMU14,943 h20pose estimation
SensorLM [Paper ]PPG, EDA, ACC, TEMP, ALT59,749 h103,731sensor-language study
Apple Study (AHMS; WBM) [Paper]HealthKit metrics (27)> 2.5 B h162 K57 health tasks
WESADEDA/PPG/Temp + Acc45 h15stress detection
PPG-DaliaECG, PPG, IMU, GSR36 h15daily activities

Physiology Datasets

DatasetSensorsDatasizeSubjectsActivities
Sleep-EDFEEG/EOG/EMG/ECG1,576 h197sleep stages
MIT-BIH Arrhythmia2‑lead ECG1128 h47ambulatory ECG
PTB-XL12‑lead ECG6.06 h18,885ECG status
TUH EEGmulti‑channel EEG1476 h675seizure activity
Cuff‑Less‑BPECG, PPG72 h—blood‑pressure estimation
Auditory‑EEGEEG23 h—auditory attention
PhyAAtEEG33 h25auditory attention
MAUSECG, PPG, GSR22 h22cognitive workload/stress
Mendeley‑YAADECG, GSR5 h—affect/stress elicitation
Brain‑CognitiveEEG85 h20cognitive state regulation
EPHNOGRAMECG, PCG61 h24cardiac auscultation
BIDMCECG, PPG14 h53clinical monitoring
MOODS [Paper]PPG54 K h122stress monitoring
SleepFMBAS, ECG, respiratory112,544 h14,068sleep quality monitoring
RecGymacc, gyro, human body capacitance50h1012 fitness activities
PPG-DaliaECG, PPG, IMU, GSR36 h15daily activities

Smart Home Datasets

DatasetSensorsDatasizeSubjectsActivities
CASAS (Aruba/Milan/...)Ambient binary sensors (motion, doors)——smart home activities

RF Datasets

DatasetSensorsDatasizeSubjectsActivities
mmWave (var.)[Paper]Range–Doppler / RF point clouds5 h10daily activities (10 scenes)
MM-FimmWave, LiDAR, Wi‑Fi, RGB‑D10.6 h4027 daily activities

Base Architecture

Description

Four base computation graphs for sensor foundation models. Encoder-only stacks (top left) focus on representation learning using a single sensor encoder (e.g., ViT or SSM) with lightweight heads for recognition, retrieval, or forecasting. Dual encoders (top right) independently embed sensor and text/vision streams and align them via a shared latent projection (CLIP-style) for retrieval/zero-shot transfer. Encoder–decoder stacks (bottom left) condition a language/multimodal decoder on encoded sensor tokens (cross-attention) to produce captions, rationales, or structured outputs. Language-model stacks (bottom right), either encoder–decoder or decoder-only, treat sensing as a token sequence using projection/quantization interfaces for forecasting, analysis, and reasoning.

Encoder-only stacks

  1. "A Novel Human Activity Recognition Framework Based on Pre‑Trained Foundation Model (Chronos HAR Adapters)". Xiong et al.. IEEE BIBM 2024. [Paper]

  2. "Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions". Erturk et al.. ICML 2025. [Paper]

  3. "Cosmo: Contrastive Fusion Learning with Small Data for Multimodal Human Activity Recognition". Ouyang et al.. MobiCom 2022. [Paper][Code]

  4. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  5. "HAR‑DoReMi: Optimizing Data Mixture for Self‑Supervised Human Activity Recognition Across Heterogeneous IMU Datasets". Ban et al.. arXiv 2025. [Paper]

  6. "Layout‑Agnostic Human Activity Recognition in Smart Homes through Textual Descriptions Of Sensor Triggers (TDOST)". Thukral et al.. IMWUT 2025. [Paper]

  7. "LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing Applications". Xu et al.. SenSys 2021. [Paper][Code]

  8. "MASTER: A Multi‑modal Foundation Model for Human Activity Recognition". Zhu et al.. IMWUT 2025. [Paper]

  9. "Pulse‑PPG: An Open‑Source Field‑Trained PPG Foundation Model for Wearable Applications across Lab and Field Settings". Saha et al.. IMWUT 2025. [Paper]

  10. "RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
    Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]

  11. "Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]

  12. "SelfPAB: Large‑Scale Pre‑training on Accelerometer Data for Human Activity Recognition". Logacjov et al.. Applied Intelligence 2024. [Paper][Code]

  13. "Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data". Narain et al.. arXiv 2025. [Paper]

  14. "Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices (HAR‑FM)". Qiu et al.. IMWUT 2025. [Paper]

Dual-encoder

  1. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition". Weng et al.. IEEE TMC 2025. [Paper]
  2. "IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text Narrations". Moon et al.. EMNLP 2023. [Paper][Code]
  3. "Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)". Wieland et al.. Sensors 2025. [Paper]
  4. "Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)". Weng et al.. SenSys 2024. [Paper]
  5. "Limitations in Employing Natural Language Supervision for Sensor-Based Human Activity Recognition – And Ways to Overcome Them". Haresamudram et al.. AAAI 2025.
  6. "Multimodal Foundation Model for Cross-Modal Retrieval and Activity Recognition (AURA-MFM)". Matsuishi et al.. arXiv 2025. [Paper]
  7. "PRimuS: Pretraining IMU Encoders with Multimodal Self-Supervision". Das et al.. ICASSP 2025. [Paper][[Code](https://arxiv.org/abs/2506.03174]
  8. "SensorLM: Learning the Language of Wearable Sensors". Zhang et al.. arXiv 2025. [Paper][Code]
  9. "SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
    Thapa et al. ICML 2024. [Paper][Code]
  10. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals (NORMWEAR)".
    Luo et al. arXiv 2024. [Paper][Code]
  11. "UniMTS – Unified Pre-training for Motion Time-Series". Zhang et al.. NeurIPS 2024. [Paper][Code]
  12. "GEM: Empowering MLLM for Grounded ECG Understanding with Time Series and Images". Lan et al.. arXiv 2025. [Paper][Code]
  13. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Encoder-decoder stacks

  1. "LSM: Large-Scale Masked Modeling of Daily Summaries for Population-Level Behavior Modeling". —. arXiv 2025. [Paper]
  2. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals (NORMWEAR)".
    Luo et al. arXiv 2024. [Paper][Code]
  3. "RobustHAR: Multi-Scale Spatial-Temporal Masked Autoencoder for Robust Human Activity Recognition".
    Liu et al. IJCAI 2025. [Paper]
  4. "Scaling Wearable Foundation Models". Narayanswamy et al. arXiv 2024. [Paper]
  5. "SelfPAB: Large-Scale Pre-training on Accelerometer Data for Human Activity Recognition". Logacjov et al.. Applied Intelligence 2024. [Paper][Code]
  6. "Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors". Chen et al.. IMWUT 2024. [Paper]
  7. "SensorLM: Learning the Language of Wearable Sensors". Zhang et al.. arXiv 2025. [Paper][Code]
  8. "Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]
  9. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals (NORMWEAR)".
    Luo et al. arXiv 2024. [Paper][Code]
  10. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Language-model stacks

  1. "SensorLM: Learning the Language of Wearable Sensors". Zhang et al.. arXiv 2025. [Paper][Code]
  2. "Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
    Pillai et al. PMLR 2025. [Paper][Code]
  3. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]
  4. "LLaSA: A Multimodal LLM for Human Activity Analysis Through Wearable and Smartphone Sensors". Imran et al.. arXiv 2024. [Paper][Code]
  5. "HARGPT: A Language‑Conditioned Foundation Model for Human Activity Recognition". Ji et al.. arXiv 2024. [Paper]
  6. "LanHAR: Language‑Driven Human Activity Recognition". Yan et al.. arXiv 2024. [Paper][Code]
  7. "StressLLM: Large Language Models for Wearable Stress Detection". Thapa et al.. arXiv 2025. [Paper]
  8. "DailyLLM: Large Language Models for Daily Behavior Understanding". Kang et al.. arXiv 2025. [Paper]
  9. "ContextLLM: Multimodal Context Understanding from Wearable Devices". Wang et al.. arXiv 2025. [Paper]
  10. "Health‑LLM: Aligning Large Language Models with Wearable Sensor Health Data". Liu et al.. Information Fusion 2025. [Paper][Code]
  11. "SensorGPT: Generative Pretraining for Wearable Sensing". Sharma et al.. arXiv 2025. [Paper]
  12. "Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors". Chen et al.. IMWUT 2024. [Paper]

Modality Scope

Unimodal foundation models

  1. "RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
    Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]

  2. "One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation". Wei et al. IMWUT 2025. [Paper]

  3. "Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]

  4. "Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications". Saha et al. IMWUT 2025. [Paper]

Multimodal foundation models

  1. "Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices". Qiu et al. arXiv 2025. [Paper]
  2. "MASTER: A Multi-modal Foundation Model for Human Activity Recognition". Zhu et al. IEEE TMC 2025. [Paper]
  3. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals (NORMWEAR)".
    Luo et al. arXiv 2024. [Paper][Code]
  4. "MuJo: Multimodal Joint Feature Space Learning for Human Activity Recognition". Fritsch et al. arXiv 2024. [Paper]
  5. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Cross-modal foundation models

  1. "SensorLM: Learning the Language of Wearable Sensors". Zhang et al. arXiv 2025. [Paper][Code]
  2. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]
  3. "IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text". Moon et al. EMNLP Findings 2023. [Paper][Code]
  4. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition". Weng et al. IEEE TMC 2025. [Paper]
  5. "Leveraging foundation models for zero-shot IoT sensing". Xue et al. arXiv 2024. [Paper]

Data Landscape: Collected and Generated Corpora

Corpora of sensor data collected in the wild

  1. "One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation". Wei et al. IMWUT 2025. [Paper]

  2. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  3. "Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]

  4. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals". Luo et al. arXiv 2024. [Paper]

  5. "SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
    Thapa et al. ICML 2024. [Paper][Code]

  6. "Scaling Wearable Foundation Models". Narayanswamy et al. arXiv 2024. [Paper]

  7. "DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs". Tian et al. arXiv 2025. [Paper]

  8. "IMU2CLIP: Language-Grounded Motion Sensor Translation with Multimodal Contrastive Learning". Moon et al. EMNLP Findings 2023. [Paper][Code]

  9. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Generated datasets and augmentation

  1. "On the Benefit of Generative Foundation Models for Human Activity Recognition". Leng et al. arXiv 2023. [Paper]

  2. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition". Weng et al. IEEE TMC 2025. [Paper]

  3. "TxP: Reciprocal Generation of Ground Pressure Dynamics and Activity Descriptions for Improving Human Activity Recognition". Lala et al. IEEE TMC 2025. [Paper]

  4. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  5. "IMUGPT 2.0: Language-Based Cross Modality Transfer for Sensor-Based Human Activity Recognition". Leng et al. IMWUT 2024. [Paper][Code]

  6. "AI-Generated Fall Data: Assessing LLMs and Diffusion Model for Wearable Fall Detection". Sana et al. MDPI Sensors 2025. [Paper][Code]

  7. "Weak-Annotation of HAR Datasets using Vision Foundation Models". Bock et al.. ISWC 2024. [Paper][Code]

Tokenization and Representation Strategies

Description

Tokenization and representation for sensor-based HAR. Single-stream token formation converts raw signals (e.g., IMU, PPG, ambient/RF) into windows, statistical features, spectrograms, or quantized codes; cross-stream scaffolding then synchronizes modalities with positional/meta encodings and performs token fusion or cross-modal projection. The resulting tokens feed pretraining/training backbones (Transformer/ViT/LLM) for tasks such as classification, captioning, retrieval, and forecasting.

Window-based and patch-level segmentation

  1. "RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
    Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]

  2. "Speech Foundation Models Generalize to Time Series Tasks from Wearable Sensor Data". Jaya Narain et al. arXiv 2025. [Paper]

  3. "SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
    Thapa et al. ICML 2024. [Paper][Code]

  4. "Leveraging Foundation Models for Zero-Shot IoT Sensing". Dinghao Xue et al. arXiv 2024. [Paper]

  5. "Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)". Christoph Wieland, Victor Pankratius. IEEE Sensors Journal 2025. [Paper]

  6. "A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)". Xiong et al. IEEE BIBM 2024. [Paper]

  7. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Feature-based aggregation and statistical embeddings

  1. "Beyond Sensor Data: Foundation Models of Behavioral Data from Wearables Improve Health Predictions". Eray Erturk et al. arXiv 2025. [Paper]

  2. "ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs". Post et al. arXiv 2025. [Paper]

  3. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]

  4. "Scaling Wearable Foundation Models". Narayanswamy et al. arXiv 2024. [Paper]

  5. "A Personal Health Large Language Model for Sleep and Fitness Coaching (PH-LLM)". Khasentino et al. Nature Medicine 2025. [Paper][Code]

  6. "BioSignal Copilot: Leveraging the Power of LLMs in Drafting Reports for Biomedical Signals". Liu et al. arXiv 2023. [Paper] [Code]

  7. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Spectrogram and frequency-domain embeddings

  1. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals".
    Luo et al. arXiv 2024. [Paper]

  2. "Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]

  3. "Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors". Chen et al.. IMWUT 2024. [Paper]

  4. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Discrete and quantized sensor tokens

  1. "Towards Learning Discrete Representations via Self-Supervision for Wearables-Based Human Activity Recognition".
    Harish Haresamudram et al. Sensors 2024. [Paper]

  2. "Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices".
    Minghui Qiu et al. ACM IMWUT 2025. [Paper]

  3. "Chronos: Learning the Language of Time Series".
    Ansari et al. arXiv 2024. [Paper]

  4. "Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
    Pillai et al. PMLR 2025. [Paper][Code]

Multimodal alignment and positional encoding

  1. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]

  2. "Visible Light Human Activity Recognition Driven by Generative Language Model".
    Yang et al. Elsevier, Information Fusion 2025. [Paper]

  3. "LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
    Ouyang et al. arXiv 2024. [Paper]

Token fusion and cross-modal projection

  1. "IMU2CLIP: Language-Grounded Motion Sensor Translation with Multimodal Contrastive Learning". Moon et al. EMNLP Findings 2023. [Paper][Code]
  2. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition". Weng et al. IEEE TMC 2025. [Paper]
  3. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Pretraining Paradigms

Description

Pretraining paradigms for sensor-based HAR. Contrastive learns cross-view/cross-modal alignment in a shared latent space (CLIP-like), enabling zero-/few-shot transfer and retrieval. Generative uses masked reconstruction or causal prediction to model temporal continuity and support imputation and text-conditioned decoding. Hybrid / Self-supervised combines contrastive and generative objectives (often at scale) and introduces semantic grounding via language/distillation, improving robustness across users and devices.

Contrastive pretraining

  1. "Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)".
    Weng et al. ACM 2024. [Paper]

  2. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition".
    Weng et al. IEEE TMC 2025. [Paper]

  3. "RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
    Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]

  4. "IMU2CLIP: Multimodal Contrastive Learning for IMU Motion Sensors from Egocentric Videos and Text".
    Moon et al. EMNLP Findings 2023. [Paper][Code]

  5. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Generative pretraining

  1. "Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]

  2. "Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)".
    Wieland & Pankratius. IEEE Sensors Journal 2025. [Paper]

  3. "LLaSA: Large Multimodal Agent for Human Activity Analysis Through Wearable Sensors".
    Imran et al. arXiv 2024. [Paper][Code]

  4. "Scaling Wearable Foundation Models (LSM)".
    Narayanswamy et al. arXiv 2024. [Paper]

  5. "SelfPAB: Large-Scale Pre-training on Accelerometer Recordings with Masked Spectrogram Reconstruction".
    Logacjov et al. Springer 2024. [Paper][Code]

Hybrid and self-supervised pretraining

  1. "HAR-DoReMi: Optimizing Data Mixture for Self-Supervised Pretraining in Human Activity Recognition".
    Ban et al. arXiv 2025. [Paper]

  2. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  3. "A Personal Health Large Language Model for Sleep and Fitness Coaching (PH-LLM)".
    Khasentino et al. Nature Medicine 2025. [Paper][Code]

  4. "Toward Foundation Model for Multivariate Wearable Sensing of Physiological Signals (NORMWEAR)".
    Luo et al. arXiv 2024. [Paper][Code]

  5. "Scaling Wearable Foundation Models (LSM)".
    Narayanswamy et al. arXiv 2024. [Paper]

  6. "SensorLM: Learning the Language of Wearable Sensors".
    Zhang et al. arXiv 2025. [Paper][Code]

  7. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]

  8. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Adaptation Strategies

Description

Mechanism-centric view of adaptation in HAR foundation models. We emphasize how behavior is changed: PEFT keeps the backbone frozen and learns small add-ons (adapters, LoRA/QLoRA, learnable prefixes); Full/Partial fine-tuning updates all or selected layers for tighter task coupling; and Instruction-tuning & alignment spans prompt-only in-context learning (zero-update) and supervised formatting (SFT/PEFT on curated sensor–text exemplars and task templates) to ensure format adherence and faithful generation.

Parameter-efficient fine-tuning (PEFT)

  1. "LLaSA: Large Multimodal Agent for Human Activity Analysis Through Wearable Sensors".
    Imran et al. arXiv 2024. [Paper][Code]

  2. "A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)".
    Xiong et al. IEEE BIBM 2024. (No verified link available in file.)

  3. "Leveraging Large Language Models for Digital Phenotyping and Health Forecasting".
    Yuan et al. bioRxiv 2025. [Paper]

  4. "Large Language Models for Wearable Sensor-Based Activity Understanding".
    Liu et al. Sensors 2024. [Paper]

  5. "Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
    Pillai et al. PMLR 2025. [Paper][Code]

  6. "PhysLLM: Harnessing Large Language Models for Physiological Understanding".
    Xie et al. arXiv 2025. [Paper]

  7. "MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model".
    Bandyopadhyay et al. arXiv 2025. [Paper]

  8. "GOAT: A Generalized Cross-Dataset Activity Recognition Framework".
    Miao et al. IMWUT 2024. [Paper]

  9. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

Full or partial fine-tuning

  1. "LIMU-BERT: Unleashing the Potential of Unlabeled Data for IMU Sensing Applications". Xu et al.. SenSys 2021. [Paper][Code]

  2. "Scaling Wearable Foundation Models (LSM)".
    Narayanswamy et al. arXiv 2024. [Paper]

  3. "LLM4HAR: Generalizable On-device Human Activity Recognition with Large Language Models".
    Hong et al. IMWUT 2025. [Paper]

  4. "RelCon: Relative Contrastive Learning for a Motion Foundation Model for Wearable Data".
    Xu et al. arXiv 2024 (ICLR 2025). [Paper][Code]

  5. "SelfPAB: large-scale pre-training on accelerometer data for human activity recognition".
    Logacjov et al. ACM 2024. [Paper][Code]

  6. "Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]

  7. "Leveraging Large Language Models for Digital Phenotyping and Health Forecasting".
    Yuan et al. bioRxiv 2025. [Paper]

  8. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Instruction-tuning & alignment

  1. "StressLLM: Large Language Models for Stress Prediction and Biomarker Reasoning".
    Thapa et al. IEEE ICCE 2025. [Paper]

  2. "Leveraging Large Language Models for Digital Phenotyping: Detecting Depressive State Changes for Patients with Depressive Episodes".
    Yuan et al. arXiv 2025. [Paper]

  3. "LAHAR: Leveraging Language Models for Human Activity Recognition".
    Chen et al. IEEE Access 2024. [Paper]

  4. "DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs".
    Tian et al. arXiv 2025. [Paper]

  5. "Enabling On-Device LLMs Personalization with Sensor Prompts".
    Zhang et al. arXiv 2024. [Paper]

  6. "Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors".
    Chen et al. IMWUT 2024. [Paper]

  7. "SensorLM: Learning the Language of Wearable Sensors".
    Zhang et al. arXiv 2025. [Paper][Code]

  8. "A Personal Health Large Language Model for Sleep and Fitness Coaching (PH-LLM)".
    Khasentino et al. Nature Medicine 2025. [Paper][Code]

  9. "Mental-LLM: Leveraging Large Language Models for Mental Health Prediction via Online Text Data".
    Xu et al. ACM IMWUT 2024. [Paper][Code]

  10. "The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition".
    Georgios et al. IEEE ACII 2024. [Paper][Code]

Downstream Capabilities

Description

Downstream capabilities and the accompanying generalization protocols used in sensor-based HAR. Top row (left→right): Zero-/few-shot & open-set: recognize unseen activities with 𝑘-shot label budgets and open-set rejection; Cross-dataset / device / user: train on dataset A and test on unseen datasets/devices/users with leave-one-out and cross-position splits; Cross-modal retrieval & search: sensortext/video retrieval evaluated by Recall@K and mAP under cross-domain splits with a shared embedding space. Bottom row: Captioning, Q&A, reasoning: sensor-conditioned decoding (prompts/PEFT), measured by caption/Q&A accuracy and human/expert ratings; Reconstruction, forecasting, imputation: masked-reconstruction/denoising and short/long-horizon forecasting under distribution shift; Federated & on-device evaluation: client-level personalization with communication rounds, reporting edge latency/energy and privacy constraints.

Zero-/few-shot & open-set recognition

  1. "Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition".
    Weng et al. ACM 2024. [Paper]

  2. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition".
    Weng et al. IEEE TMC 2025. [Paper]

  3. "ZARA: Zero-Shot Motion Time-Series Analysis via LLM-Guided Knowledge Retrieval and Reasoning".
    Li et al. arXiv 2025. [Paper][Code]

  4. "HARGPT: Are LLMs Zero-Shot Human Activity Recognizers?".
    Ji et al. arXiv 2024. [Paper]

  5. "EEG-GPT: Exploring Capabilities of Large Language Models for EEG-Based Abnormality Detection".
    Kim et al. arXiv 2024. [Paper]

  6. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

Cross-dataset / cross-device / cross-user generalization

  1. "Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices (HAR-FM)".
    Qiu et al. IMWUT 2025. [Paper]

  2. "One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation".
    Wei et al. ACM IMEUT 2025. [Paper]

  3. "MASTER: A Multi-Modal Foundation Model for Human Activity Recognition".
    Zhu et al. IMWUT 2025. [Paper]

  4. "A Novel Human Activity Recognition Framework Based on Pre-Trained Foundation Model (Chronos HAR Adapters)".
    Xiong et al. IEEE BIBM 2024. (Link unavailable in verified file)

  5. "Self-supervised Learning for Human Activity Recognition Using 700,000 Person-days of Wearable Data". Yuan et al.. npj Digital Medicine 2024. [Paper][Code]

  1. "Multimodal Foundation Model for Cross-Modal Retrieval and Recognition (AURA-MFM)".
    Matsuishi et al. arXiv 2025. [Paper]

  2. "GLOSS: Group of LLMs for Open-ended Sensemaking of Passive Sensing Data for Health and Wellbeing".
    Choube et al. IMWUT 2025. [Paper][Code]

  3. "Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLM".
    Li et al. IMWUT 2025. [Paper][Code]

  4. "PhysioLLM: Supporting Personalized Health Insights with Wearables and Large Language Models".
    Fang et al. arXiv 2024. [Paper]

  5. "SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
    Thapa et al. ICML 2024. [Paper][Code]

Language-grounded captioning, Q&A, and reasoning

  1. "Sensor2Text: Enabling Natural Language Interactions for Daily Activity Tracking Using Wearable Sensors".
    Chen et al. IMWUT 2024. [Paper]

  2. "SensorLM: Learning the Language of Wearable Sensors".
    Zhang et al. arXiv 2025. [Paper][Code]

  3. "SensorLLM: Aligning Large Language Models with Motion Sensors for Human Activity Recognition".
    Li et al. arXiv 2024. [Paper][Code]

  4. "Time2Lang: Bridging Time-Series Foundation Models and Large Language Models for Health Sensing Beyond Prompting".
    Pillai et al. PMLR 2025. [Paper][Code]

  5. "Visible Light Human Activity Recognition Driven by Generative Language Model".
    Yang et al. Information Fusion 2025. [Paper]

  6. "LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
    Ouyang et al. SenSys-ML 2024. [Paper]

Generative reconstruction, forecasting, and imputation

  1. "Spatial-Temporal Masked Autoencoder for Multi-Device Wearable Human Activity Recognition (STMAE)". Miao et al.. IMWUT 2024. [Paper][Code]

  2. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  3. "Scaling Wearable Foundation Models (LSM)".
    Narayanswamy et al. arXiv 2024. [Paper]

  4. "Inertial Signal Forecasting with Foundation Model Techniques (Dual-View FM)".
    Wieland & Pankratius. IEEE Sensors Journal 2025. [Paper]

  5. "RobustHAR: Multi-Scale Spatial-Temporal Masked Autoencoder for Robust Human Activity Recognition".
    Liu et al. IJCAI 2025. [Paper]

  6. "UniMTS – Unified Pre-training for Motion Time-Series Forecasting and Recognition".
    Zhang et al. arXiv 2024. [Paper][Code]

On-device, federated, and online adaptation

  1. "LLM4HAR: Generalizable On-device Human Activity Recognition with Large Language Models".
    Hong et al. KDD 2025. [Paper]

  2. "DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs".
    Tian et al. arXiv 2025. [Paper]

  3. "MHARFedLLM: Multimodal Human Activity Recognition Using Federated Large Language Model".
    Bandyopadhyay et al. arXiv 2025. [Paper]

  4. "ContextLLM: Meaningful Context Reasoning from Multi-Sensor and Multi-Device Data Using LLMs".
    Post et al. ACM HotMobile 2025. [Paper]

  5. "Enabling On-Device LLMs Personalization with Smartphone Sensing".
    Zhang et al. arXiv 2024. [Paper]

  6. "Activity transitions for semi-supervised federated learning in sensor-based human activity recognition".
    Bukit et al. Elsevier, Applied Soft Computing 2025. [Paper]

Deployment Settings

Cloud-scale training and centralized evaluation

  1. "Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)". Weng et al.. SenSys 2024. [Paper]

  2. "FM-Fi 2.0: Foundation Model for Cross-Modal Multi-Person Human Activity Recognition".
    Weng et al. IEEE TMC 2025. [Paper]

  3. "Scaling Wearable Foundation Models (LSM)".
    Narayanswamy et al. arXiv 2024. [Paper]

On-device and mobile execution

  1. "LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
    Ouyang et al. SenSys-ML 2024. [Paper]

  2. "Leveraging foundation models for zero-shot IoT sensing".
    Xue et al. arXiv 2024. [Paper][Code]

  3. "Large Language Model-Guided Semantic Alignment for Human Activity Recognition".
    Yan et al. arXiv 2024. [Paper][Code]

  4. "Enabling On-Device LLMs Personalization with Sensor Prompts".
    Zhang et al. arXiv 2024. [Paper]

  5. "LLM4HAR: Generalizable On-Device Human Activity Recognition with Large Language Models".
    Hong et al. IMWUT 2025. [Paper]

  6. "Few-Shot Human Activity Recognition Using Lightweight Language Models".
    Cruciani et al. IEEE ICCCN 2025. [Paper]

  7. "On-device Foundation Models for Wearable Signals".
    Simon et al. 2025. [Paper]

  8. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  9. "Enabling Efficient RF Sensing with Small Language Models via Functional Data Analysis and Parameter Efficient Tuning".
    Yujie Sun et al. IEEE IoTJ 2026. [Paper]

  10. "HALO: A Heterogeneity-Aware Language-Aligned IMU Foundation Model for Open-Set Human Activity Recognition". Zihan Ding et al. Arxiv 2026. [Paper]

  11. "EdgeHAR: An Edge-Native Compact Sensor Foundation Model for Human Activity Recognition". He Zhang et al. Arxiv 2026. [Paper]

Edge-Cloud cooperation

  1. "PFHAR: Practically Adopting Multi-Modal Foundation Model for Human Activity Recognition through Edge-cloud Collaborative Learning".
    Zhengyuan Zhang et al. TMC 2026. [Paper]

Application Domains

Description

Application domains for sensor-based HAR foundation models. The radial layout highlights four commonly targeted areas: general-purpose HAR / daily living, healthcare and wellbeing, smart-home and context-aware environments, and interactive/agentic assistants.

General-purpose HAR / ADL

  1. "One Model to Fit Them All: Universal IMU-based Human Activity Recognition with LLM-assisted Cross-dataset Representation".
    Wei et al. ACM IMWUT 2025. [Paper]

  2. "CrossHAR: Generalizing Cross-Dataset Human Activity Recognition via Hierarchical Self-Supervised Pretraining".
    Hong et al. IMWUT 2024. [Paper][Code]

  3. "Towards Customizable Foundation Models for Human Activity Recognition with Wearable Devices (HAR-FM)".
    Qiu et al. IMWUT 2025. [Paper]

Healthcare & wellbeing

  1. "PhysioLLM: Supporting Personalized Health Insights with Wearables and Large Language Models".
    Fang et al. arXiv 2024. [Paper]

  2. "Vital Insight: Assisting Experts' Context-Driven Sensemaking of Multi-modal Personal Tracking Data Using Visualization and Human-in-the-Loop LLM".
    Li et al. IMWUT 2025. [Paper][Code]

  3. "Pulse-PPG: An Open-Source Field-Trained PPG Foundation Model for Wearable Applications across Lab and Field Settings".
    Saha et al. IMWUT 2025. [Paper]

  4. "SleepFM: Multi-modal Representation Learning for Sleep Across Brain Activity, ECG and Respiratory Signals".
    Thapa et al. ICML 2024. [Paper][Code]

  5. "The Strong Pull of Prior Knowledge in Large Language Models and Its Impact on Emotion Recognition".
    Georgios et al. IEEE ACII 2024. [Paper][Code]

Smart-home & context-aware environments

  1. "Game of LLMs: Discovering Structural Constructs in Activities using Large Language Models". Hiremath et al.. UbiComp Companion 2024. [Paper]

  2. "A Synergistic Large Language Model and Supervised Learning Approach to Zero-Shot and Continual Activity Recognition in Smart Homes". Naoto et al.. ICBDA 2024. [Paper]

  3. "Large Model for Small Data: Foundation Model for Cross-Modal RF Human Activity Recognition (FM-Fi)". Weng et al.. SenSys 2024. [Paper]

  4. "Visible light human activity recognition driven by generative language model".
    Yang et al. Information Fusion 2025. [Paper]

Interactive & agentic assistants

  1. "LLMSense: Harnessing LLMs for High-Level Reasoning over Spatiotemporal Sensor Traces".
    Ouyang et al. SenSys-ML 2024. [Paper]

  2. "Large Language Model-Guided Semantic Alignment for Human Activity Recognition".
    Yan et al. arXiv 2024. [Paper]

  3. "DailyLLM: Context-Aware Activity Log Generation Using Multi-Modal Sensors and LLMs".
    Tian et al. arXiv 2025. [Paper]

  4. "SensorFM: Towards a General Intelligence and Interface for Wearable Health Data". Narayanswamy et al. Google Deepmind 2026. [Paper]

  5. "You Don't Need To Train: Agentic Heuristic Learning Studio for Executable Human Activity Recognition". Siyu et al. Arxiv 2026. [Paper][Code]

Contributors

zhaxidele

169 commits