Teach a local AI what your camera sees and turn it into Home Assistant sensors — garage door open/closed, gate shut, car parked. 100% local, any CPU.
Python
20
70 commits
updated Oct 3, 2026
Teach a local AI what your camera sees — and turn it into a Home Assistant sensor.
Garage door open, closed or halfway? Gate shut? Click a few examples and you have a sensor.
Person at the door, car in the driveway, cat on the lawn? Pick the objects — no training needed.
kWh on the meter, minutes left on the washer? Read the number.
Your cameras already see whether the garage door is half open, the gate is shut or the lights in the shed are still on. VisionState turns what they see into states you can use in Home Assistant — and because every home is different, you teach it yourself, in minutes, without writing code or leaving Home Assistant:
1–9).
The model retrains in about a second after every click.camera.* entity in Home Assistant (ESP32-CAM, IP cameras, NVRs…),
or a direct RTSP / HTTP snapshot URL.![]() | ![]() |
| Objects — people, cars, animals and more, found without any training. Each one becomes an on/off sensor and a count in Home Assistant. | Reading — the number on a display or on the rolling wheels of a water or gas meter, checked before it is published: a counter never goes down. |
![]() | |
| New sensor — pick a camera, draw the region and choose what to detect: your own states, objects or a number. The wizard tests it on a fresh frame right away. | |
Every sensor is a normal Home Assistant device, set up automatically through MQTT discovery — ready for dashboards, automations and the history graph.
![]() | ![]() |
State sensor — the state (closed, open, …) with its confidence, the last frame, a button to check now and a switch to pause it. | Object sensor — an on/off sensor and a count for every object you picked, plus the last frame with the boxes drawn in. |
Requirements: Home Assistant OS or Supervised, the Mosquitto broker app and the MQTT integration, and at least one camera.
https://github.com/oleost/VisionState.New features are released as beta first and move to the stable app once they are tested. To
help test them, add https://github.com/oleost/VisionState#beta as a repository and install
VisionState (beta).
That's it — sensor.visionstate_garage_door is now in Home Assistant:
| Entity | What it is |
|---|---|
sensor.visionstate_<name> | The state (open, closed, …) — or unknown when the AI isn't sure |
sensor.visionstate_<name>_confidence | How sure the AI is, in % |
image.visionstate_<name>_frame | The region that was classified |
button.visionstate_<name>_classify | Check right now (handy in automations) |
switch.visionstate_<name>_enabled | Pause / resume |
sensor.visionstate_review_queue | Frames waiting for review (all sensors) |
# Example: notify when the garage door has been open for 10 minutes
alias: Garage door left open
triggers:
- trigger: state
entity_id: sensor.visionstate_garage_door
to: open
for: "00:10:00"
actions:
- action: notify.mobile_app_phone
data:
message: The garage door has been open for 10 minutes.
The full user guide is in visionstate/DOCS.md (also on the app's
Documentation tab).
camera ─► crop to your region ─► DINOv2 (ONNX, 8-bit) ─► feature vector ─► tiny per-sensor classifier
│
Home Assistant ◄── MQTT discovery ◄── debounce + "unknown" below threshold ◄─────────┘
A pre-trained vision model (DINOv2) turns the region into a feature vector. On top of that, each sensor gets its own small classifier trained on your labelled images — which is why a handful of examples is enough and training takes about a second. Results are debounced so someone walking past doesn't flip the state.
Object sensors use a pretrained detector instead (D-FINE, Apache-2.0, trained on the COCO objects). It finds every object in the region; each object you picked is reported with a count and cleared a while after it was last seen.
Reading sensors read the digits in the region with a small text recognizer (PaddleOCR, Apache-2.0) that may only output digits, then check the value (a counter never goes down) before publishing it.
Everything stays on your machine: images live in /media/visionstate, models and settings in
the app's data folder (included in Home Assistant backups).
visionstate/ Home Assistant app (config.yaml, Dockerfile, docs)
backend/ Python 3.14 · FastAPI · ONNX Runtime · scikit-learn
frontend/ Svelte 5 · Vite · TypeScript
scripts/channel.py Switches the app config between the stable and beta channel
scripts/fake_camera.py A fake camera (garage door, real photos, drawn displays and counters) for local testing
docs/SCOPE.md Design as built, decisions and roadmap
CLAUDE.md Contributor guide: conventions, how to test and verify, release steps
New features land on the beta branch first and reach main (stable) only after testing;
please open pull requests against beta. CLAUDE.md explains the conventions, how to
test changes locally (fake camera, MQTT, UI tests on desktop and phone) and the pitfalls we ran
into — it is written for people and AI coding assistants alike.
Backend:
cd visionstate/backend
python3.14 -m venv .venv && . .venv/bin/activate # Windows: py -3.14 -m venv .venv; .venv\Scripts\activate
pip install -r requirements-dev.txt
python -m visionstate.backbones models # download the bundled model once
pytest
VISIONSTATE_DATA=./dev/data VISIONSTATE_MEDIA=./dev/media VISIONSTATE_BUNDLED_MODELS=./models \
HA_URL=http://homeassistant.local:8123 HA_TOKEN=<long-lived token> \
VISIONSTATE_MQTT_HOST=<broker> python -m visionstate
⚠️ Outside Home Assistant the app has no login: anyone who can reach port 8099 can use it. Only run it like this on your own machine or behind a reverse proxy with authentication.
Frontend (proxies /api to the backend on port 8099):
cd visionstate/frontend
npm install
npm run dev
UI tests (Playwright) start the backend and a fake camera by themselves and run every page on a desktop browser and on an emulated phone with touch:
cd visionstate/frontend
npx playwright install chromium # once
npm run build
VS_PYTHON=../backend/.venv/bin/python npm run e2e
Issues and ideas are welcome in GitHub Issues.
Apache-2.0. The bundled models are Apache-2.0 as well: DINOv2 (Meta), D-FINE and PaddleOCR PP-OCRv6.
If VisionState is useful to you, you can buy me a coffee ☕
Python
53.3%
Svelte
33.2%
TypeScript
11.8%
CSS
1.4%
Teach a local AI what your camera sees and turn it into Home Assistant sensors — garage door open/closed, gate shut, car parked. 100% local, any CPU.
Python
20
70 commits
updated Oct 3, 2026
Teach a local AI what your camera sees — and turn it into a Home Assistant sensor.
Garage door open, closed or halfway? Gate shut? Click a few examples and you have a sensor.
Person at the door, car in the driveway, cat on the lawn? Pick the objects — no training needed.
kWh on the meter, minutes left on the washer? Read the number.
Your cameras already see whether the garage door is half open, the gate is shut or the lights in the shed are still on. VisionState turns what they see into states you can use in Home Assistant — and because every home is different, you teach it yourself, in minutes, without writing code or leaving Home Assistant:
1–9).
The model retrains in about a second after every click.camera.* entity in Home Assistant (ESP32-CAM, IP cameras, NVRs…),
or a direct RTSP / HTTP snapshot URL.![]() | ![]() |
| Objects — people, cars, animals and more, found without any training. Each one becomes an on/off sensor and a count in Home Assistant. | Reading — the number on a display or on the rolling wheels of a water or gas meter, checked before it is published: a counter never goes down. |
![]() | |
| New sensor — pick a camera, draw the region and choose what to detect: your own states, objects or a number. The wizard tests it on a fresh frame right away. | |
Every sensor is a normal Home Assistant device, set up automatically through MQTT discovery — ready for dashboards, automations and the history graph.
![]() | ![]() |
State sensor — the state (closed, open, …) with its confidence, the last frame, a button to check now and a switch to pause it. | Object sensor — an on/off sensor and a count for every object you picked, plus the last frame with the boxes drawn in. |
Requirements: Home Assistant OS or Supervised, the Mosquitto broker app and the MQTT integration, and at least one camera.
https://github.com/oleost/VisionState.New features are released as beta first and move to the stable app once they are tested. To
help test them, add https://github.com/oleost/VisionState#beta as a repository and install
VisionState (beta).
That's it — sensor.visionstate_garage_door is now in Home Assistant:
| Entity | What it is |
|---|---|
sensor.visionstate_<name> | The state (open, closed, …) — or unknown when the AI isn't sure |
sensor.visionstate_<name>_confidence | How sure the AI is, in % |
image.visionstate_<name>_frame | The region that was classified |
button.visionstate_<name>_classify | Check right now (handy in automations) |
switch.visionstate_<name>_enabled | Pause / resume |
sensor.visionstate_review_queue | Frames waiting for review (all sensors) |
# Example: notify when the garage door has been open for 10 minutes
alias: Garage door left open
triggers:
- trigger: state
entity_id: sensor.visionstate_garage_door
to: open
for: "00:10:00"
actions:
- action: notify.mobile_app_phone
data:
message: The garage door has been open for 10 minutes.
The full user guide is in visionstate/DOCS.md (also on the app's
Documentation tab).
camera ─► crop to your region ─► DINOv2 (ONNX, 8-bit) ─► feature vector ─► tiny per-sensor classifier
│
Home Assistant ◄── MQTT discovery ◄── debounce + "unknown" below threshold ◄─────────┘
A pre-trained vision model (DINOv2) turns the region into a feature vector. On top of that, each sensor gets its own small classifier trained on your labelled images — which is why a handful of examples is enough and training takes about a second. Results are debounced so someone walking past doesn't flip the state.
Object sensors use a pretrained detector instead (D-FINE, Apache-2.0, trained on the COCO objects). It finds every object in the region; each object you picked is reported with a count and cleared a while after it was last seen.
Reading sensors read the digits in the region with a small text recognizer (PaddleOCR, Apache-2.0) that may only output digits, then check the value (a counter never goes down) before publishing it.
Everything stays on your machine: images live in /media/visionstate, models and settings in
the app's data folder (included in Home Assistant backups).
visionstate/ Home Assistant app (config.yaml, Dockerfile, docs)
backend/ Python 3.14 · FastAPI · ONNX Runtime · scikit-learn
frontend/ Svelte 5 · Vite · TypeScript
scripts/channel.py Switches the app config between the stable and beta channel
scripts/fake_camera.py A fake camera (garage door, real photos, drawn displays and counters) for local testing
docs/SCOPE.md Design as built, decisions and roadmap
CLAUDE.md Contributor guide: conventions, how to test and verify, release steps
New features land on the beta branch first and reach main (stable) only after testing;
please open pull requests against beta. CLAUDE.md explains the conventions, how to
test changes locally (fake camera, MQTT, UI tests on desktop and phone) and the pitfalls we ran
into — it is written for people and AI coding assistants alike.
Backend:
cd visionstate/backend
python3.14 -m venv .venv && . .venv/bin/activate # Windows: py -3.14 -m venv .venv; .venv\Scripts\activate
pip install -r requirements-dev.txt
python -m visionstate.backbones models # download the bundled model once
pytest
VISIONSTATE_DATA=./dev/data VISIONSTATE_MEDIA=./dev/media VISIONSTATE_BUNDLED_MODELS=./models \
HA_URL=http://homeassistant.local:8123 HA_TOKEN=<long-lived token> \
VISIONSTATE_MQTT_HOST=<broker> python -m visionstate
⚠️ Outside Home Assistant the app has no login: anyone who can reach port 8099 can use it. Only run it like this on your own machine or behind a reverse proxy with authentication.
Frontend (proxies /api to the backend on port 8099):
cd visionstate/frontend
npm install
npm run dev
UI tests (Playwright) start the backend and a fake camera by themselves and run every page on a desktop browser and on an emulated phone with touch:
cd visionstate/frontend
npx playwright install chromium # once
npm run build
VS_PYTHON=../backend/.venv/bin/python npm run e2e
Issues and ideas are welcome in GitHub Issues.
Apache-2.0. The bundled models are Apache-2.0 as well: DINOv2 (Meta), D-FINE and PaddleOCR PP-OCRv6.
If VisionState is useful to you, you can buy me a coffee ☕
Python
53.3%
Svelte
33.2%
TypeScript
11.8%
CSS
1.4%