zjunlp/LightMem-Ego

LightMem-Ego: Your AI Memory for Everyday Life

Python

108

60 commits

updated Sep 22, 2026

See the code

README

LightMem-Ego: Your AI Memory for Everyday Life

An open-source, self-hostable multimodal memory system for smart glasses and phones.

arXiv Hugging Face Paper EM²Mem paper EMNLP 2026 Findings License: MIT

🌐 Try in Browser  ·  🚀 Quick Start  ·  🎬 Watch the Demo  ·  📱 Glasses APK

Python 3.10+ FastAPI React 19 Vite Android Docker

LightMem-Ego is the end-to-end system; its long-term tier (M_lt) is powered by EM²Mem (EMNLP 2026 Findings), part of the ZJUNLP LightMem project family.

EM²Mem versus the strongest baseline: 76.8 Video-MME (L), 67.7 Ego-R1 Bench, 66.0 EgoLifeQA, 4.67 times faster per query
📑 Table of contents

📢 News


🎬 Demo

Ask the glasses a question in the middle of your day, and get an answer grounded in what you actually saw and heard. Prefer typing? Join the same live session from the web page.

[!TIP] No hardware? Try it right now. The live web demo runs the full LightMem-Ego workflow in your browser — on a phone too, where it captures from the phone's own camera and microphone. No glasses, no local installation.

The glasses record; the user asks where a plastic bottle was placed; the answer comes back with the timestamped evidence

Ask on the glasses, get a memory-grounded answer on the HUD — watch the full demo.

Watch the full demo on YouTube   Watch the full demo on Bilibili

Asking a question on Rokid AI Glasses
Hands-free on Rokid AI Glasses
Ask by voice or preset question
Memory-grounded answer over a real-world scene
Answers grounded in memory
Timestamps + visual evidence
Asking typed questions about the same live session from the web
Same session on the web
Type questions when speaking isn't convenient
Runs on glasses and phones; self-hostable; timestamped evidence; three-tier memory
⭐ If LightMem-Ego is useful to you, a star helps more people find it.

🎯 Why LightMem-Ego

  • 🎥 Always-on egocentric capture — streams first-person camera frames and microphone audio from Rokid AI Glasses or a phone.
  • 🧠 Three-tier memory — a rolling current memory, short-term micro-events, and consolidated long-term episodes, routines, and preferences.
  • ⏱️ One aligned timeline — frames, audio chunks, ASR transcripts and metadata all share a single session timeline.
  • 🔍 Memory-grounded answers — every answer ships with the timestamped visual and transcript evidence behind it.
  • 👓 Glasses and web, one session — start capture on the glasses, keep asking from the web page in the same live session.
  • 🐳 Self-hostable — docker compose up --build brings up the web UI and the full backend worker pipeline.

Text memory systems only know what you typed, live assistants only know the current scene, and video systems only let you search afterwards. LightMem-Ego covers all three — the system-by-system table is in How It Compares.


🚀 Quick Start

[!IMPORTANT] Every path except the hosted demo needs an OpenAI-compatible LLM endpoint (base URL, API key, model names) and Xfyun ASR credentials for speech. The default stack also expects Qwen3-Embedding-4B weights under docker-data/models/.

PathSetupWhat it addsGPU
🌐 Live demononethe full hosted workflow—
🐳 DockerDocker + your LLM endpointthe standard self-hosted setupno
👁 + visual retrieval--profile models + VLM2Vec weightsframe-level visual matchingyes
⚡ + local QwenvLLM server on your GPUlower first-token latencyyes
👓 Rokid glassesAPK + any backend abovehands-free capture, HUD answersno

[!NOTE] The last two are independent add-ons, not requirements. The plain Docker setup is what we run day to day — add visual retrieval when caption and transcript evidence is not enough, and local Qwen when a remote API feels slow.

🐳 Docker

git clone https://github.com/zjunlp/LightMem-Ego.git
cd LightMem-Ego
cp deploy/.env.example .env     # LLM endpoint, keys, model names, Xfyun credentials
docker compose up --build

Open http://localhost:8080. The web container proxies /api to the backend, so no CORS setup is needed. The first build takes a few minutes.

By default EM2MEM_VISUAL_BACKEND=mock, so retrieval runs on captions and transcripts. Current memory, short-term micro-events, long-term consolidation and evidence-grounded answers all behave as in the demo — only frame-level visual matching is off.

👁 Add visual retrieval (optional)

Put VLM2Vec-V2.0 and Qwen3-Embedding-4B under docker-data/models/, then start the model services:

docker compose --profile models up --build

Point the backend at them in .env:

EM2MEM_VISUAL_BACKEND=remote
EM2MEM_TEXT_EMBED_BACKEND=remote

On a GPU host, add the override so the workers get the GPU as well:

docker compose -f compose.yaml -f compose.gpu.yaml --profile models up --build

This needs the NVIDIA Container Toolkit and model directories matching the paths in .env — see deploy/DOCKER.md for details.

⚡ Local Qwen for lower latency (optional)

This is where a GPU helps most. Every memory write and every answer otherwise round-trips to a remote API; serving the LLM locally cuts first-token latency noticeably. The scripts build an isolated vLLM environment and switch the backend onto it:

cd src/backend
scripts/setup_local_qwen35_env.sh
scripts/download_local_qwen35_model.sh
scripts/select_llm_profile.sh local-qwen35
scripts/stop_server_and_workers.sh --keep-api --force

Details and the smoke test: src/backend/README.md.

👓 Rokid AI Glasses

Install the released APK:

adb install -r app-release.apk

Or build it (JDK + Android SDK):

cd src/ai_glass_app
./gradlew assembleDebug        # Windows: .\gradlew.bat assembleDebug

Set API_BASE_URL in LightMemEgoConfig.kt to your own backend — it points at our demo server by default. Details: src/ai_glass_app/README.md.

Building from source

Web frontend (Node.js + npm)
cd src/frontend/online_web
npm install
npm run dev

Point it at your backend by creating online_web/.env.local:

VITE_API_BASE_URL=http://127.0.0.1:8000

Details: src/frontend/README.md

Backend (Python 3.10+, ffmpeg/ffprobe)
cd src/backend
python -m venv .venv && source .venv/bin/activate
python -m pip install --upgrade pip && python -m pip install -e .
cp .env.example .env            # configure model paths and API credentials
scripts/start_api.sh
scripts/start_online_all_workers.sh

Details: src/backend/README.md and DEPLOYMENT.md.


💬 What You Can Ask

ScenarioExample questionMemory used
Object finding"Where did I leave my badge?"Current + short-term
Conversation recall"What did the doctor tell me after checking the report?"Short-term + transcript
Day summarization"What did I do this afternoon?"Short-term + long-term
Routine discovery"What do I usually do after arriving at the office?"Long-term semantic
Live assistance"What am I looking at right now?"Current

🏗️ How It Works

LightMem-Ego system design
Rokid AI Glasses ─┐
                  ├─► Stream API ─► M_cur ─► M_st ─► M_lt ─► Retrieval ─► Answer + Evidence
Browser (web) ────┘                 current  short   long
Memory tierScopeExample
M_cur current memoryThe ongoing scene, updated as frames arrive"What am I looking at?"
M_st short-term memoryRecent micro-events, actions, and conversations"What did she just tell me?"
M_lt long-term memoryConsolidated episodes, routines, preferences, semantic facts"What do I usually do on Fridays?"

The backend divides each session into short event anchors and stores multimodal evidence per anchor. The long-term tier (M_lt) is built by EM²Mem, our event-centric multimodal memory framework (EMNLP 2026 Findings, arXiv:2609.00551): events are the retrieval unit, and episodic and semantic graphs link them across a session. At query time the system retrieves aligned event-level evidence — captions, transcripts, frames, timestamps — instead of reconstructing context at inference.

EM²Mem architecture: event-centric memory schema, event-linked graph construction, and lightweight retrieval

EM²Mem in one picture. A video is segmented into 30-second event anchors, and each anchor becomes a memory cell holding dense captions, transcripts, keyframes, and metadata. Episodic and semantic graphs link those cells, and retrieval reads grounded evidence from them instead of re-aligning raw fragments at query time.


📊 Results

End-to-end system — LightMem-Ego

[!NOTE] These numbers come from a small, intentionally balanced set — 27 queries (nine per scenario) over five everyday-life videos, about 45.7 minutes of footage — collected with the phone and glasses client profiles used in the paper. In the open-source release the phone client is the web frontend running in a mobile browser, capturing from the phone's own camera and microphone — a native app is on the roadmap. It measures the current prototype rather than a public leaderboard. Reported in the LightMem-Ego paper.

Retrieval accuracy — Recall@k over the retrieved memory entries, with MRR for the first relevant hit:

ScenarioR@1R@3R@5MRR
Object finding22.266.777.80.454
Conversation recall44.455.655.60.481
Life summarization88.9100.0100.00.944
Overall51.974.177.80.627

Answer accuracy — experience QA over daily scenarios:

ScenarioLLM-JudgeHuman
Object finding44.455.6
Conversation recall33.333.3
Life summarization77.877.8
Overall51.955.6

Latency — P50 / P90 across two client profiles:

StagePhone P50Phone P90Glasses P50Glasses P90
Short-term memory QA
Retrieval76 ms131 ms44 ms87 ms
Time to first token532 ms643 ms423 ms494 ms
Answer generation6.13 s10.11 s6.81 s9.14 s
End-to-end6.42 s10.34 s6.95 s9.31 s
Long-term memory QA
Retrieval2.99 s3.84 s3.06 s3.44 s
Time to first token4.64 s5.44 s4.74 s5.07 s
Answer generation5.78 s9.56 s4.37 s9.16 s
End-to-end10.57 s13.93 s8.61 s13.60 s

Glasses columns are the glasses-style client profile. Short-term queries stay near-interactive; long-term ones trade latency for temporal coverage.

Long-term memory engine — EM²Mem

The long-term tier (M_lt) is built by EM²Mem. Average accuracy (%) across three long-video and egocentric benchmarks, as reported in the EM²Mem paper:

MethodEgoLifeQAEgo-R1 BenchVideo-MME (L)
GPT-548.646.374.3
HippoRAG59.656.052.1
M3-Agent53.552.055.3
Ego-R153.052.042.7
WorldMM65.665.376.6
EM²Mem66.067.776.8

Against the strongest baseline (WorldMM, reproduced under the same evaluation setting):

MetricEM²MemWorldMMGain
Avg. latency per query98.21 s459.00 s4.67× faster
Wall-clock evaluation time6,138 s229,502 s37.4× faster
Total tokens15.27M42.03M63.7% fewer

EM²Mem moves multimodal alignment and graph organization into offline memory construction, so inference reads from pre-built event-indexed memory cells instead of re-aligning isolated fragments. Full per-category tables are in the backend README; reproduction scripts in experiments/egolife.


🆚 How It Compares

Representative commercial assistants, text-based memory systems, and egocentric multimodal assistants. This compares publicly described capabilities rather than measured performance.

SystemPlatform & inputReal-time A/V streamCurrent / short-term MM memoryLong-term episodicLong-term semanticTimestamped evidence
ChatGPT MemoryText chat———Partial—
Mem0-style memoryText and agent memory—Partial—✓Partial
Memories.aiVideo archives and visual memoryPartialPartial✓PartialPartial
Gemini LivePhone✓Partial———
Ray-Ban Meta AI GlassesGlassesPartialPartial———
VinciPhone or wearable camera✓✓PartialPartialPartial
VisualClawStreaming video with agent workspacePartialPartial—PartialPartial
VisionClawSmart glasses✓Partial——Partial
Egocentric Co-PilotSmart glasses with web agents✓✓PartialPartialPartial
EgoButlerAI-glasses egocentric video and audioPartialPartialPartialPartial✓
LightMem-EgoPhone and glasses-style client✓✓✓✓✓

✓ implemented as an explicit first-class component · Partial limited, implicit, offline, session-level, or modality-restricted · — not explicitly supported or not publicly described. Adapted from the LightMem-Ego paper. Phone was the paper's evaluation client — the open-source clients are the browser frontend and the Rokid AI Glasses app.


📦 Repository Layout

PathWhat's insideDocs
src/ai_glass_app/Android app for Rokid AI Glasses (Kotlin, Jetpack Compose, CameraX)README
src/frontend/Vite + React web UI for capture, sessions, QA, and evidence reviewREADME
src/backend/FastAPI service plus the online worker pipeline (ASR, memory, retrieval, QA)README
compose.yaml, deploy/Docker Compose stack and deployment notesDOCKER.md

🗺️ Roadmap

  • Ship a native phone app — today the phone client is the web frontend in a mobile browser.
  • Release the end-to-end evaluation dataset and reproducibility scripts.
  • Pluggable ASR, VLM, and embedding backends beyond the current defaults.
  • Support wearable devices beyond Rokid AI Glass.
  • On-device filtering and user-controlled memory editing for privacy-sensitive capture.
  • One-click deployment template for a full cloud deployment.

📄 Citation

If you find LightMem-Ego useful, please cite our paper:

@article{chen2026lightmemego,
  title={LightMem-Ego: Your AI Memory for Everyday Life},
  author={Chen, Yijun and Xiao, Boyi and Zhao, Yixian and Xia, Haoting and Xu, Buqiang and Fang, Jizhan and Li, Yanya and Zheng, Yaqi and Wang, Xuehai and Xue, Zirui and others},
  journal={arXiv preprint arXiv:2607.11487},
  year={2026}
}

The long-term memory tier (M_lt) of the backend is built by EM²Mem, which has been accepted to EMNLP 2026 Findings. Please cite it as well when you use that module:

@article{chen2026em2mem,
  title={EM$^{2}$Mem: Event-Centric Multimodal Memory for Large Language Models},
  author={Chen, Yijun and Zheng, Yaqi and Li, Yanya and Xiao, Boyi and Xu, Buqiang and Qiao, Shuofei and Fang, Jizhan and Deng, Xinle and Yao, Yunzhi and Wang, Xuehai and others},
  journal={arXiv preprint arXiv:2609.00551},
  year={2026}
}

This repository belongs to the ZJUNLP LightMem series, which targets context bloat, excessive token consumption, and low cache utilization in long-running LLM agents:

  • LightMem — a lightweight and efficient memory management framework for LLMs and AI agents
  • LightRSI — a modular framework for recursive improvement in long-horizon LLM agents
  • EM²Mem (EMNLP 2026 Findings) — event-centric multimodal memory for long-video QA, and the long-term memory engine behind this system (code overview)

🙏 Acknowledgements

LightMem-Ego builds on the broader line of work on memory-augmented agents, egocentric multimodal understanding, and wearable AI assistants. We thank all contributors and collaborators who helped develop the system.


⚖️ License

Released under the MIT License.


🔐 Privacy

LightMem-Ego processes camera frames, microphone audio, transcripts, and generated memories. It is released for research and demonstration; a production deployment needs HTTPS, access control, encryption at rest, a data retention/deletion policy, and explicit user consent. Runtime media and memory are written to local, Git-ignored directories.


⭐ Star History

GitHub Stars

Star history chart
agent-memory
artificial-intelligence
ego-centric
lifelogging
lightmem
lightmem-ego
long-video-qa
multimodal-agent
multimodal-ai
multimodal-memory
rokid
smart-glasses
wearable

Significant stargazers

lichuang

1,833 followers · starred Sep 2026

zjunlp/LightMem-Ego

LightMem-Ego: Your AI Memory for Everyday Life

Python

108

60 commits

updated Sep 22, 2026

See the code

README

LightMem-Ego: Your AI Memory for Everyday Life

An open-source, self-hostable multimodal memory system for smart glasses and phones.

arXiv Hugging Face Paper EM²Mem paper EMNLP 2026 Findings License: MIT

🌐 Try in Browser  ·  🚀 Quick Start  ·  🎬 Watch the Demo  ·  📱 Glasses APK

Python 3.10+ FastAPI React 19 Vite Android Docker

LightMem-Ego is the end-to-end system; its long-term tier (M_lt) is powered by EM²Mem (EMNLP 2026 Findings), part of the ZJUNLP LightMem project family.

EM²Mem versus the strongest baseline: 76.8 Video-MME (L), 67.7 Ego-R1 Bench, 66.0 EgoLifeQA, 4.67 times faster per query
📑 Table of contents

📢 News


🎬 Demo

Ask the glasses a question in the middle of your day, and get an answer grounded in what you actually saw and heard. Prefer typing? Join the same live session from the web page.

[!TIP] No hardware? Try it right now. The live web demo runs the full LightMem-Ego workflow in your browser — on a phone too, where it captures from the phone's own camera and microphone. No glasses, no local installation.

The glasses record; the user asks where a plastic bottle was placed; the answer comes back with the timestamped evidence

Ask on the glasses, get a memory-grounded answer on the HUD — watch the full demo.

Watch the full demo on YouTube   Watch the full demo on Bilibili

Asking a question on Rokid AI Glasses
Hands-free on Rokid AI Glasses
Ask by voice or preset question
Memory-grounded answer over a real-world scene
Answers grounded in memory
Timestamps + visual evidence
Asking typed questions about the same live session from the web
Same session on the web
Type questions when speaking isn't convenient
Runs on glasses and phones; self-hostable; timestamped evidence; three-tier memory
⭐ If LightMem-Ego is useful to you, a star helps more people find it.

🎯 Why LightMem-Ego

  • 🎥 Always-on egocentric capture — streams first-person camera frames and microphone audio from Rokid AI Glasses or a phone.
  • 🧠 Three-tier memory — a rolling current memory, short-term micro-events, and consolidated long-term episodes, routines, and preferences.
  • ⏱️ One aligned timeline — frames, audio chunks, ASR transcripts and metadata all share a single session timeline.
  • 🔍 Memory-grounded answers — every answer ships with the timestamped visual and transcript evidence behind it.
  • 👓 Glasses and web, one session — start capture on the glasses, keep asking from the web page in the same live session.
  • 🐳 Self-hostable — docker compose up --build brings up the web UI and the full backend worker pipeline.

Text memory systems only know what you typed, live assistants only know the current scene, and video systems only let you search afterwards. LightMem-Ego covers all three — the system-by-system table is in How It Compares.


🚀 Quick Start

[!IMPORTANT] Every path except the hosted demo needs an OpenAI-compatible LLM endpoint (base URL, API key, model names) and Xfyun ASR credentials for speech. The default stack also expects Qwen3-Embedding-4B weights under docker-data/models/.

PathSetupWhat it addsGPU
🌐 Live demononethe full hosted workflow—
🐳 DockerDocker + your LLM endpointthe standard self-hosted setupno
👁 + visual retrieval--profile models + VLM2Vec weightsframe-level visual matchingyes
⚡ + local QwenvLLM server on your GPUlower first-token latencyyes
👓 Rokid glassesAPK + any backend abovehands-free capture, HUD answersno

[!NOTE] The last two are independent add-ons, not requirements. The plain Docker setup is what we run day to day — add visual retrieval when caption and transcript evidence is not enough, and local Qwen when a remote API feels slow.

🐳 Docker

git clone https://github.com/zjunlp/LightMem-Ego.git
cd LightMem-Ego
cp deploy/.env.example .env     # LLM endpoint, keys, model names, Xfyun credentials
docker compose up --build

Open http://localhost:8080. The web container proxies /api to the backend, so no CORS setup is needed. The first build takes a few minutes.

By default EM2MEM_VISUAL_BACKEND=mock, so retrieval runs on captions and transcripts. Current memory, short-term micro-events, long-term consolidation and evidence-grounded answers all behave as in the demo — only frame-level visual matching is off.

👁 Add visual retrieval (optional)

Put VLM2Vec-V2.0 and Qwen3-Embedding-4B under docker-data/models/, then start the model services:

docker compose --profile models up --build

Point the backend at them in .env:

EM2MEM_VISUAL_BACKEND=remote
EM2MEM_TEXT_EMBED_BACKEND=remote

On a GPU host, add the override so the workers get the GPU as well:

docker compose -f compose.yaml -f compose.gpu.yaml --profile models up --build

This needs the NVIDIA Container Toolkit and model directories matching the paths in .env — see deploy/DOCKER.md for details.

⚡ Local Qwen for lower latency (optional)

This is where a GPU helps most. Every memory write and every answer otherwise round-trips to a remote API; serving the LLM locally cuts first-token latency noticeably. The scripts build an isolated vLLM environment and switch the backend onto it:

cd src/backend
scripts/setup_local_qwen35_env.sh
scripts/download_local_qwen35_model.sh
scripts/select_llm_profile.sh local-qwen35
scripts/stop_server_and_workers.sh --keep-api --force

Details and the smoke test: src/backend/README.md.

👓 Rokid AI Glasses

Install the released APK:

adb install -r app-release.apk

Or build it (JDK + Android SDK):

cd src/ai_glass_app
./gradlew assembleDebug        # Windows: .\gradlew.bat assembleDebug

Set API_BASE_URL in LightMemEgoConfig.kt to your own backend — it points at our demo server by default. Details: src/ai_glass_app/README.md.

Building from source

Web frontend (Node.js + npm)
cd src/frontend/online_web
npm install
npm run dev

Point it at your backend by creating online_web/.env.local:

VITE_API_BASE_URL=http://127.0.0.1:8000

Details: src/frontend/README.md

Backend (Python 3.10+, ffmpeg/ffprobe)
cd src/backend
python -m venv .venv && source .venv/bin/activate
python -m pip install --upgrade pip && python -m pip install -e .
cp .env.example .env            # configure model paths and API credentials
scripts/start_api.sh
scripts/start_online_all_workers.sh

Details: src/backend/README.md and DEPLOYMENT.md.


💬 What You Can Ask

ScenarioExample questionMemory used
Object finding"Where did I leave my badge?"Current + short-term
Conversation recall"What did the doctor tell me after checking the report?"Short-term + transcript
Day summarization"What did I do this afternoon?"Short-term + long-term
Routine discovery"What do I usually do after arriving at the office?"Long-term semantic
Live assistance"What am I looking at right now?"Current

🏗️ How It Works

LightMem-Ego system design
Rokid AI Glasses ─┐
                  ├─► Stream API ─► M_cur ─► M_st ─► M_lt ─► Retrieval ─► Answer + Evidence
Browser (web) ────┘                 current  short   long
Memory tierScopeExample
M_cur current memoryThe ongoing scene, updated as frames arrive"What am I looking at?"
M_st short-term memoryRecent micro-events, actions, and conversations"What did she just tell me?"
M_lt long-term memoryConsolidated episodes, routines, preferences, semantic facts"What do I usually do on Fridays?"

The backend divides each session into short event anchors and stores multimodal evidence per anchor. The long-term tier (M_lt) is built by EM²Mem, our event-centric multimodal memory framework (EMNLP 2026 Findings, arXiv:2609.00551): events are the retrieval unit, and episodic and semantic graphs link them across a session. At query time the system retrieves aligned event-level evidence — captions, transcripts, frames, timestamps — instead of reconstructing context at inference.

EM²Mem architecture: event-centric memory schema, event-linked graph construction, and lightweight retrieval

EM²Mem in one picture. A video is segmented into 30-second event anchors, and each anchor becomes a memory cell holding dense captions, transcripts, keyframes, and metadata. Episodic and semantic graphs link those cells, and retrieval reads grounded evidence from them instead of re-aligning raw fragments at query time.


📊 Results

End-to-end system — LightMem-Ego

[!NOTE] These numbers come from a small, intentionally balanced set — 27 queries (nine per scenario) over five everyday-life videos, about 45.7 minutes of footage — collected with the phone and glasses client profiles used in the paper. In the open-source release the phone client is the web frontend running in a mobile browser, capturing from the phone's own camera and microphone — a native app is on the roadmap. It measures the current prototype rather than a public leaderboard. Reported in the LightMem-Ego paper.

Retrieval accuracy — Recall@k over the retrieved memory entries, with MRR for the first relevant hit:

ScenarioR@1R@3R@5MRR
Object finding22.266.777.80.454
Conversation recall44.455.655.60.481
Life summarization88.9100.0100.00.944
Overall51.974.177.80.627

Answer accuracy — experience QA over daily scenarios:

ScenarioLLM-JudgeHuman
Object finding44.455.6
Conversation recall33.333.3
Life summarization77.877.8
Overall51.955.6

Latency — P50 / P90 across two client profiles:

StagePhone P50Phone P90Glasses P50Glasses P90
Short-term memory QA
Retrieval76 ms131 ms44 ms87 ms
Time to first token532 ms643 ms423 ms494 ms
Answer generation6.13 s10.11 s6.81 s9.14 s
End-to-end6.42 s10.34 s6.95 s9.31 s
Long-term memory QA
Retrieval2.99 s3.84 s3.06 s3.44 s
Time to first token4.64 s5.44 s4.74 s5.07 s
Answer generation5.78 s9.56 s4.37 s9.16 s
End-to-end10.57 s13.93 s8.61 s13.60 s

Glasses columns are the glasses-style client profile. Short-term queries stay near-interactive; long-term ones trade latency for temporal coverage.

Long-term memory engine — EM²Mem

The long-term tier (M_lt) is built by EM²Mem. Average accuracy (%) across three long-video and egocentric benchmarks, as reported in the EM²Mem paper:

MethodEgoLifeQAEgo-R1 BenchVideo-MME (L)
GPT-548.646.374.3
HippoRAG59.656.052.1
M3-Agent53.552.055.3
Ego-R153.052.042.7
WorldMM65.665.376.6
EM²Mem66.067.776.8

Against the strongest baseline (WorldMM, reproduced under the same evaluation setting):

MetricEM²MemWorldMMGain
Avg. latency per query98.21 s459.00 s4.67× faster
Wall-clock evaluation time6,138 s229,502 s37.4× faster
Total tokens15.27M42.03M63.7% fewer

EM²Mem moves multimodal alignment and graph organization into offline memory construction, so inference reads from pre-built event-indexed memory cells instead of re-aligning isolated fragments. Full per-category tables are in the backend README; reproduction scripts in experiments/egolife.


🆚 How It Compares

Representative commercial assistants, text-based memory systems, and egocentric multimodal assistants. This compares publicly described capabilities rather than measured performance.

SystemPlatform & inputReal-time A/V streamCurrent / short-term MM memoryLong-term episodicLong-term semanticTimestamped evidence
ChatGPT MemoryText chat———Partial—
Mem0-style memoryText and agent memory—Partial—✓Partial
Memories.aiVideo archives and visual memoryPartialPartial✓PartialPartial
Gemini LivePhone✓Partial———
Ray-Ban Meta AI GlassesGlassesPartialPartial———
VinciPhone or wearable camera✓✓PartialPartialPartial
VisualClawStreaming video with agent workspacePartialPartial—PartialPartial
VisionClawSmart glasses✓Partial——Partial
Egocentric Co-PilotSmart glasses with web agents✓✓PartialPartialPartial
EgoButlerAI-glasses egocentric video and audioPartialPartialPartialPartial✓
LightMem-EgoPhone and glasses-style client✓✓✓✓✓

✓ implemented as an explicit first-class component · Partial limited, implicit, offline, session-level, or modality-restricted · — not explicitly supported or not publicly described. Adapted from the LightMem-Ego paper. Phone was the paper's evaluation client — the open-source clients are the browser frontend and the Rokid AI Glasses app.


📦 Repository Layout

PathWhat's insideDocs
src/ai_glass_app/Android app for Rokid AI Glasses (Kotlin, Jetpack Compose, CameraX)README
src/frontend/Vite + React web UI for capture, sessions, QA, and evidence reviewREADME
src/backend/FastAPI service plus the online worker pipeline (ASR, memory, retrieval, QA)README
compose.yaml, deploy/Docker Compose stack and deployment notesDOCKER.md

🗺️ Roadmap

  • Ship a native phone app — today the phone client is the web frontend in a mobile browser.
  • Release the end-to-end evaluation dataset and reproducibility scripts.
  • Pluggable ASR, VLM, and embedding backends beyond the current defaults.
  • Support wearable devices beyond Rokid AI Glass.
  • On-device filtering and user-controlled memory editing for privacy-sensitive capture.
  • One-click deployment template for a full cloud deployment.

📄 Citation

If you find LightMem-Ego useful, please cite our paper:

@article{chen2026lightmemego,
  title={LightMem-Ego: Your AI Memory for Everyday Life},
  author={Chen, Yijun and Xiao, Boyi and Zhao, Yixian and Xia, Haoting and Xu, Buqiang and Fang, Jizhan and Li, Yanya and Zheng, Yaqi and Wang, Xuehai and Xue, Zirui and others},
  journal={arXiv preprint arXiv:2607.11487},
  year={2026}
}

The long-term memory tier (M_lt) of the backend is built by EM²Mem, which has been accepted to EMNLP 2026 Findings. Please cite it as well when you use that module:

@article{chen2026em2mem,
  title={EM$^{2}$Mem: Event-Centric Multimodal Memory for Large Language Models},
  author={Chen, Yijun and Zheng, Yaqi and Li, Yanya and Xiao, Boyi and Xu, Buqiang and Qiao, Shuofei and Fang, Jizhan and Deng, Xinle and Yao, Yunzhi and Wang, Xuehai and others},
  journal={arXiv preprint arXiv:2609.00551},
  year={2026}
}

This repository belongs to the ZJUNLP LightMem series, which targets context bloat, excessive token consumption, and low cache utilization in long-running LLM agents:

  • LightMem — a lightweight and efficient memory management framework for LLMs and AI agents
  • LightRSI — a modular framework for recursive improvement in long-horizon LLM agents
  • EM²Mem (EMNLP 2026 Findings) — event-centric multimodal memory for long-video QA, and the long-term memory engine behind this system (code overview)

🙏 Acknowledgements

LightMem-Ego builds on the broader line of work on memory-augmented agents, egocentric multimodal understanding, and wearable AI assistants. We thank all contributors and collaborators who helped develop the system.


⚖️ License

Released under the MIT License.


🔐 Privacy

LightMem-Ego processes camera frames, microphone audio, transcripts, and generated memories. It is released for research and demonstration; a production deployment needs HTTPS, access control, encryption at rest, a data retention/deletion policy, and explicit user consent. Runtime media and memory are written to local, Git-ignored directories.


⭐ Star History

GitHub Stars

Star history chart
agent-memory
artificial-intelligence
ego-centric
lifelogging
lightmem
lightmem-ego
long-video-qa
multimodal-agent
multimodal-ai
multimodal-memory
rokid
smart-glasses
wearable

Significant stargazers

lichuang

1,833 followers · starred Sep 2026

Languages

Python

84.4%

JavaScript

8.9%

Kotlin

3.6%

Shell

1.6%

CSS

1.4%