PrimeEmre/Jarvis-Hardware-Program--local

Python

0

5 commits

updated Jul 22, 2026

See the code

See what people are saying

SourceMessageScoreDate

Local voice assistant that picks its Ollama model based on your VRAM (and shows its status on an Arduino) (r/LocalLLaMA)

Sharing a side project in case the setup is useful to someone. It's a Flask app that answers questions with a local model through Ollama, grounds the answer with a quick DuckDuckGo search, reads the reply out loud with a local TTS server, and drives an Arduino so you can see when it's working. The…

0

Oct 4, 2026

README

J.A.R.V.I.S. Hardware Program

A local, self-hosted J.A.R.V.I.S.-style AI assistant: a Flask web app that answers questions through a local LLM grounded with live web search, speaks its replies with a local text-to-speech engine, and drives a physical Arduino rig (LEDs + servo) in real time to visualize what it's doing — researching, debating, or done.

Status Python Platform

Features

  • Conversational chat UI — clean, single-page interface with a live status "orb" and message log.
  • Local LLM inference via Ollama, with automatic model/context selection based on detected GPU VRAM (nvidia-smi), so the same install scales from a small laptop GPU to a high-end card.
  • Live web search grounding using DDGS, so answers reflect current information instead of only the model's training data.
  • Optional CrewAI integration — point it at a deployed CrewAI crew and it will kick off runs and poll for results, falling back to direct search if unset or if a run fails/times out.
  • Local text-to-speech via OmniVoice Studio, an OpenAI-compatible TTS server that's auto-launched if it isn't already running. No cloud API key required.
  • Physical hardware feedback — an Arduino sketch drives two status LEDs and a sweeping servo that mirror the assistant's state (idle / researching / debating / complete) in the real world, synced to the UI over WebSockets.
  • Graceful degradation — runs in software-only mode with no Arduino attached, auto-reconnects if the board drops mid-session, and falls back cleanly if CrewAI/OmniVoice are unavailable.
  • Built-in rate limiting and response/audio caching to keep repeated queries cheap and fast.

How it works

Browser (chat UI)
    │  POST /chat
    ▼
Flask app (app.py)
    │
    ├─► DDGS web search  ─┐
    ├─► CrewAI crew (opt.)│──► context ──► Ollama (local LLM) ──► reply
    │                     │
    ├─► Flask-SocketIO ───┴──► pushes live state updates to the browser
    │
    └─► Arduino (serial)  ──► LEDs + servo react to the current state
    │
    ▼
POST /tts  ──► OmniVoice Studio ──► spoken audio streamed back to the browser

Every chat request walks through four visible states — idle → researching → debating → complete — each broadcast to the browser via WebSocket and mirrored on the Arduino via a one-character serial signal, so the physical hardware and the on-screen orb always agree.

SignalStateHardware behavior
0IdleLEDs off, servo parked at 90°
BResearchingBlue LED (pin 12) pulses, servo sweeps
YDebatingYellow LED (pin 13) pulses faster, servo sweeps
WCompleteLEDs off, servo parks

Requirements

  • Python 3.10+
  • An Ollama instance running locally (or reachable via OLLAMA_URL)
  • OmniVoice Studio for text-to-speech (optional but recommended — the app will try to auto-launch it)
  • An Arduino (Uno/Nano or similar) with 2 LEDs + a servo wired up, flashed with jarvis_haedware/jarvis_hardware.ino (optional — the app runs fine without one, just without the physical light show)

Getting started

  1. Install dependencies

    pip install flask flask-socketio python-dotenv requests pyserial ddgs
    
  2. Configure environment variables

    Create a .env file in the project root:

    # Flask
    SECRET_KEY=<random_key>
    FLASK_DEBUG=1                  # 1 to enable debug/auto-reload, 0 for production
    
    # Arduino serial (optional — auto-detected if omitted)
    ARDUINO_PORT=COM3
    
    # CrewAI (optional — falls back to direct web search if unset)
    CREWAI_CREW_URL=
    CREWAI_CREW_TOKEN=
    CREWAI_INPUT_KEY=query
    
    # Ollama (local LLM)
    OLLAMA_URL=http://localhost:11434/api/chat
    
    # OmniVoice Studio (local TTS)
    OMNIVOICE_URL=http://localhost:3900/v1
    OMNIVOICE_VOICE=alloy
    OMNIVOICE_APP_PATH=C:\Program Files\OmniVoice Studio\omnivoice-studio.exe
    
  3. (Optional) Flash the Arduino

    Upload jarvis_haedware/jarvis_hardware.ino to your board using the Arduino IDE or CLI. Wiring: blue LED on pin 12, yellow LED on pin 13, servo signal on pin 8.

  4. Run the app

    python app.py
    

    Then open http://localhost:5000 in your browser.

Verifying hardware separately

To test LED/servo wiring without booting the full Flask app:

python test_connection.py

This connects directly over serial and cycles through every state (4 seconds each) so you can confirm the board reacts correctly.

Configuration reference

VariablePurposeDefault
SECRET_KEYFlask session secretrandom, generated at startup
FLASK_DEBUGEnable Flask debug/reload mode1
ARDUINO_PORTForce a specific serial portauto-detected
CREWAI_CREW_URLBase URL of a deployed CrewAI crewunset (uses direct search)
CREWAI_CREW_TOKENBearer token for the CrewAI APIunset
CREWAI_INPUT_KEYInput key your crew's tasks expectquery
OLLAMA_URLOllama chat endpointhttp://localhost:11434/api/chat
OMNIVOICE_URLOmniVoice Studio API base URLhttp://localhost:3900/v1
OMNIVOICE_VOICEVoice name to use for synthesisalloy
OMNIVOICE_APP_PATHPath to the OmniVoice executable, for auto-launchWindows default install path
CACHE_TTLResponse/audio cache lifetime, in seconds3600
OVERRIDE_MODELForce a specific Ollama model regardless of detected VRAMunset

Troubleshooting

IssueLikely causeFix
"No Arduino found" on startupCable disconnected or board in bootloaderCheck the physical connection; restart the board
Hardware stops responding mid-sessionServo brownout on the board's own 5V railGive the servo a separate 5V supply
OmniVoice errors in the logsTTS service not running yetThe app auto-launches it; check OMNIVOICE_APP_PATH
Requests always rate-limitedRapid repeated requests from the same clientRate limit is 10 requests/60s per IP by default
Stale-looking repliesCache TTL too longLower CACHE_TTL or restart the app

Project layout

app.py                              Flask app, LLM/search/TTS orchestration, hardware control
test_connection.py                  Standalone Arduino connectivity check
jarvis_haedware/jarvis_hardware.ino Arduino firmware (LED + servo state machine)
templates/index.html                Chat UI
static/script.js                    Client-side chat, TTS playback, WebSocket handling
static/style.css                    Styling
.env                                Local configuration (not committed)

See AGENTS.md for a deeper architectural walkthrough intended for AI coding agents working in this repo.

PrimeEmre/Jarvis-Hardware-Program--local

Python

0

5 commits

updated Jul 22, 2026

See the code

See what people are saying

SourceMessageScoreDate

Local voice assistant that picks its Ollama model based on your VRAM (and shows its status on an Arduino) (r/LocalLLaMA)

Sharing a side project in case the setup is useful to someone. It's a Flask app that answers questions with a local model through Ollama, grounds the answer with a quick DuckDuckGo search, reads the reply out loud with a local TTS server, and drives an Arduino so you can see when it's working. The…

0

Oct 4, 2026

README

J.A.R.V.I.S. Hardware Program

A local, self-hosted J.A.R.V.I.S.-style AI assistant: a Flask web app that answers questions through a local LLM grounded with live web search, speaks its replies with a local text-to-speech engine, and drives a physical Arduino rig (LEDs + servo) in real time to visualize what it's doing — researching, debating, or done.

Status Python Platform

Features

  • Conversational chat UI — clean, single-page interface with a live status "orb" and message log.
  • Local LLM inference via Ollama, with automatic model/context selection based on detected GPU VRAM (nvidia-smi), so the same install scales from a small laptop GPU to a high-end card.
  • Live web search grounding using DDGS, so answers reflect current information instead of only the model's training data.
  • Optional CrewAI integration — point it at a deployed CrewAI crew and it will kick off runs and poll for results, falling back to direct search if unset or if a run fails/times out.
  • Local text-to-speech via OmniVoice Studio, an OpenAI-compatible TTS server that's auto-launched if it isn't already running. No cloud API key required.
  • Physical hardware feedback — an Arduino sketch drives two status LEDs and a sweeping servo that mirror the assistant's state (idle / researching / debating / complete) in the real world, synced to the UI over WebSockets.
  • Graceful degradation — runs in software-only mode with no Arduino attached, auto-reconnects if the board drops mid-session, and falls back cleanly if CrewAI/OmniVoice are unavailable.
  • Built-in rate limiting and response/audio caching to keep repeated queries cheap and fast.

How it works

Browser (chat UI)
    │  POST /chat
    ▼
Flask app (app.py)
    │
    ├─► DDGS web search  ─┐
    ├─► CrewAI crew (opt.)│──► context ──► Ollama (local LLM) ──► reply
    │                     │
    ├─► Flask-SocketIO ───┴──► pushes live state updates to the browser
    │
    └─► Arduino (serial)  ──► LEDs + servo react to the current state
    │
    ▼
POST /tts  ──► OmniVoice Studio ──► spoken audio streamed back to the browser

Every chat request walks through four visible states — idle → researching → debating → complete — each broadcast to the browser via WebSocket and mirrored on the Arduino via a one-character serial signal, so the physical hardware and the on-screen orb always agree.

SignalStateHardware behavior
0IdleLEDs off, servo parked at 90°
BResearchingBlue LED (pin 12) pulses, servo sweeps
YDebatingYellow LED (pin 13) pulses faster, servo sweeps
WCompleteLEDs off, servo parks

Requirements

  • Python 3.10+
  • An Ollama instance running locally (or reachable via OLLAMA_URL)
  • OmniVoice Studio for text-to-speech (optional but recommended — the app will try to auto-launch it)
  • An Arduino (Uno/Nano or similar) with 2 LEDs + a servo wired up, flashed with jarvis_haedware/jarvis_hardware.ino (optional — the app runs fine without one, just without the physical light show)

Getting started

  1. Install dependencies

    pip install flask flask-socketio python-dotenv requests pyserial ddgs
    
  2. Configure environment variables

    Create a .env file in the project root:

    # Flask
    SECRET_KEY=<random_key>
    FLASK_DEBUG=1                  # 1 to enable debug/auto-reload, 0 for production
    
    # Arduino serial (optional — auto-detected if omitted)
    ARDUINO_PORT=COM3
    
    # CrewAI (optional — falls back to direct web search if unset)
    CREWAI_CREW_URL=
    CREWAI_CREW_TOKEN=
    CREWAI_INPUT_KEY=query
    
    # Ollama (local LLM)
    OLLAMA_URL=http://localhost:11434/api/chat
    
    # OmniVoice Studio (local TTS)
    OMNIVOICE_URL=http://localhost:3900/v1
    OMNIVOICE_VOICE=alloy
    OMNIVOICE_APP_PATH=C:\Program Files\OmniVoice Studio\omnivoice-studio.exe
    
  3. (Optional) Flash the Arduino

    Upload jarvis_haedware/jarvis_hardware.ino to your board using the Arduino IDE or CLI. Wiring: blue LED on pin 12, yellow LED on pin 13, servo signal on pin 8.

  4. Run the app

    python app.py
    

    Then open http://localhost:5000 in your browser.

Verifying hardware separately

To test LED/servo wiring without booting the full Flask app:

python test_connection.py

This connects directly over serial and cycles through every state (4 seconds each) so you can confirm the board reacts correctly.

Configuration reference

VariablePurposeDefault
SECRET_KEYFlask session secretrandom, generated at startup
FLASK_DEBUGEnable Flask debug/reload mode1
ARDUINO_PORTForce a specific serial portauto-detected
CREWAI_CREW_URLBase URL of a deployed CrewAI crewunset (uses direct search)
CREWAI_CREW_TOKENBearer token for the CrewAI APIunset
CREWAI_INPUT_KEYInput key your crew's tasks expectquery
OLLAMA_URLOllama chat endpointhttp://localhost:11434/api/chat
OMNIVOICE_URLOmniVoice Studio API base URLhttp://localhost:3900/v1
OMNIVOICE_VOICEVoice name to use for synthesisalloy
OMNIVOICE_APP_PATHPath to the OmniVoice executable, for auto-launchWindows default install path
CACHE_TTLResponse/audio cache lifetime, in seconds3600
OVERRIDE_MODELForce a specific Ollama model regardless of detected VRAMunset

Troubleshooting

IssueLikely causeFix
"No Arduino found" on startupCable disconnected or board in bootloaderCheck the physical connection; restart the board
Hardware stops responding mid-sessionServo brownout on the board's own 5V railGive the servo a separate 5V supply
OmniVoice errors in the logsTTS service not running yetThe app auto-launches it; check OMNIVOICE_APP_PATH
Requests always rate-limitedRapid repeated requests from the same clientRate limit is 10 requests/60s per IP by default
Stale-looking repliesCache TTL too longLower CACHE_TTL or restart the app

Project layout

app.py                              Flask app, LLM/search/TTS orchestration, hardware control
test_connection.py                  Standalone Arduino connectivity check
jarvis_haedware/jarvis_hardware.ino Arduino firmware (LED + servo state machine)
templates/index.html                Chat UI
static/script.js                    Client-side chat, TTS playback, WebSocket handling
static/style.css                    Styling
.env                                Local configuration (not committed)

See AGENTS.md for a deeper architectural walkthrough intended for AI coding agents working in this repo.

Languages

Python

53.0%

CSS

16.6%

JavaScript

16.4%

HTML

7.6%

C++

6.5%