Say "kenning." Talk. Get answers in a custom voice. Everything runs on your GPU.
Commit 816df7c is the designated stable, restorable baseline of the gaming
teammate-relay assistant (lean boot ยท gaming engaged ยท anticheat default) โ git tag
ultron-0.1, pinned branch release/ultron-0.1, a GitHub
release, and a
standalone launchable backup at E:\Ultron-0.1\. Restore the dev tree with
git checkout ultron-0.1; launch the backup (a known-good build that stays
streamable while the dev version is under maintenance) with
E:\Ultron-0.1\launch_ultron_0_1.ps1. Runbook: docs/ultron_0_1_baseline.md.
The active dev line is Ultron 1.0: every response is now authored by the loaded local LLM by default,
with a voice command to switch back to the deterministic curated pools at any time (verbatim "repeat" lines and
known-fact answers always stay exact). The deterministic snap matchers are retired-not-removed and repurposed as
routers that pick a curated prompt and inject the matching snap lines + per-agent libraries as in-context
exemplars. Verbosity is now set at the prompt level on two independent, voice-switchable axes โ a callout
axis (none / low / medium / high / max flavor-tail length on a tactical callout; "none" โ the old deterministic
snap) and a conversation axis (low / medium / high / max reply length for private and social lines) โ each
with strict per-level guidelines so wording and personality stay varied without losing facts or overrunning the
length. Plus an optional-wakeword always-listening four-class intent gate wired into the run loop
(addressing.always_listening, default off) and private "me-only" replies that route to the desktop channel. A
stop-window FLAG button logs the last turn โ a disliked response, one that shouldn't have fired, or one that
should have but didn't โ to a review log for later refinement. Each increment is proven regression-clean against
the frozen control baseline; the stable Ultron 0.1 baseline above is untouched. Architecture, live status, and
the regression baseline live under docs/ultron_1_0/ (start at 00_process_log/STATUS.md).
Live calibration of the gate thresholds and the latency pass remain before the flags flip on by default.
A fully-local Twitch presence layer (flag-gated, default OFF โ main runtime is byte-identical until
enabled) lets Ultron co-host a stream without touching the anticheat contract: every Twitch network call lives in
out-of-process sidecars (chat read ยท Helix moderation write ยท a Llama-Guard safety model ยท an OBS overlay), so
the anticheat-pinned main process never imports a network or automation library for any of it.
@viewer, no TTS), and every viewer-directed reply is @-tagged to them."Ultron, ban / timeout / unban / delete <viewer>" (two-phase voice confirm with a
spoken read-back) plus chat-settings (slow mode ยท followers only ยท subscribers only ยท emote only ยท
clear chat); all writes go through the broadcaster-token sidecar with idempotency + a mass-action breaker.!slots / !wheel / !heist / !duel / !leaderboard and the
channel-point game redeems are provably-fair, house-funded, and EventSub-replay-safe (a local idempotency layer
over the SE points API).script-src CSP
grant lets it run inside OBS's embedded browser) with a small replay buffer so a freshly-connected source catches
up; ?demo=1 previews every card./shoutouts the raider and vocally welcomes the
raiders (thanks them, invites them to stick around, points them to chat) โ idempotent + fail-open, riding the
broadcaster-token EventSub session."open moderation panel") with a button per mod action (+ a username box) and a fuzzy-match confirm popup, and a dev
test panel ("show me the test panel") that fires any redeem / chat game / raid / speak locally so you can
rehearse without real viewers โ in-process tkinter (a click is the app's own window message, not input monitoring).!ultron posts the command list on demand; an auto commands-panel and a
"talk to Ultron" hint both link a one-page viewer guide, plus periodic auto-trivia rounds.Built, unit-tested offline (full tests/twitch green), and live on-stream. Go-live runbook:
docs/twitch_integration/FIRST_STREAM_CHECKLIST.md.
What would a voice assistant feel like if it lived entirely on your GPU instead of in someone else's data center?
| You say | Kenning does |
|---|---|
๐ฆ๏ธ ย "kenning, what's the weather in Paris?" | Detects fresh-data intent โ SearxNG โ reads result โ speaks the forecast |
๐ป ย "kenning, write me a script that converts PDFs to Docx" | Spawns isolated AI coding agent โ scaffolds project โ runs tests โ narrates progress |
๐ฎ ย "kenning, engage gaming mode" | Swaps LLM โ kills GPU services โ frees ~2.3 GB VRAM for your game |
๐ ย "kenning, take me to HBO Max" | Recognizes navigate intent โ opens Chrome to the best-matching domain |
๐ ย "kenning, what time is it in Tokyo?" | Hits local zoneinfo cache โ speaks the answer in ~5 ms (no LLM, no search) |
๐งญ ย "kenning, switch to the 8B" | Hot-swaps the local LLM preset mid-conversation |
โก ย "kenning, switch the model to the GPU" / "โฆback to the CPU" | Hot-moves the 3B between CPU and GPU mid-game with a device-optimized config (GPU: full layer offload + CUDA flash-attention + quantized KV + large batches; CPU: zero GPU layers + F16 KV + a smaller micro-batch so prefill never steals game frames) โ borrow the card for faster replies between rounds, hand it back for the next fight. Only the model compute location changes (VRAM vs system RAM), so it's anticheat-irrelevant |
๐ฃ๏ธ ย "kenning, tell my team they are pushing B" | Valorant teammate-relay: tactical callouts resolve deterministically (subject-exact, every count/agent/location/ability preserved, never the LLM) โ a fact-preserving fallback relays the literal rather than let the model drop or invert a callout. Nearly every line then carries a short, in-character Ultron flavor tail (a faithful Avengers: Age of Ultron clone) selected for the callout โ agent-specific for a named enemy ("Their Neon has ult. Overdrive. A finite surge."), plural for a group, owner-aware (contempt at enemies, cold command for your orders, stoic for your own status โ never mocking you) โ from a ~1,628-entry character-tailored library covering all 29 agents, hand-curated line-by-line (every agent's ult cell uses its real ultimate, every utility cell is ability-tagged, filler/off-topic/wrong-kit lines cut; 4,147 drafts โ 1,628 tight entries). Tail selection is a hybrid keyed-coarse + tagged-pool pipeline: a coarse route (agent + one of 16 enemy situations) picks the right pool, a 4-tier tag filter (location / damage / ability tags) narrows it to the tails that fit this callout, and for large pools an opt-in semantic re-ranker on the loopback embedder sidecar can fine-select with anti-repeat โ but for small curated cells (under 5 candidates) the sidecar embed is skipped entirely and a deterministic LRU pick is used, so every curated callout gets zero-latency contextual routing by default (KENNING_ENABLE_TAIL_SELECTOR off). An ult keyword (ulted/ultimate) always lifts the situation to the correct ult pool regardless of parse path; a callout verb (mollied/walled/darted) routes to the right per-ability cell via _VERB_TO_ABILITY. A semantic relay-intent gate (_relay_intent.py, reusing the router's sidecar) vetoes the bare-callout "tell my team" prepend for narration, banter, questions, and Marvel/identity talk before they reach the relay path โ with a fast-path narration regex (_NARRATION_MUSING_RE) that fires even when the sidecar is down โ hardened by a 25,000-case audit so the deterministic layer alone (no embedder) reframes "let my team know X" callouts, accepts terse "they are A" position calls, keeps a directive after a reported-context clause, and refuses first-person musings/recounts ("I told my team โฆ and they โฆ", "part of me wants to tell them โฆ") rather than relaying them. Ask-form team questions invert to natural spoken order โ a trailing copula or negated auxiliary fronts after the wh-word ("ask my team why they aren't smoking" โ "Why aren't they smoking?"), and a yes/no question gets the auxiliary inserted so the inflection survives TTS where a bare "?" would not ("ask Sage if she has a heal" โ "Sage, do you have a heal?"). A baked common-English-word set (4,771 words, _common_words.py, generated offline) protects real words from the STT gazetteer snapper so common tokens are never corrupted into agent names. The pre-routing normalizer's disfluency resolver was rewritten to preserve the relay lead across restarts/corrections while correctly stripping upstream fillers. All the ML stays in the sidecar or in offline build/audit scripts (never imported at runtime); the anticheat-pinned main process imports only numpy + urllib for this path. Off-snap lines (banter, economy, opinions, identity, Marvel, answers, greets) get the full persona โ plays on a VoiceMeeter strip so your voice chat hears it. When you don't trust the model to improvise, 73 explicit fallback commands (refuse a dumb question, criticize/praise a teammate by name, call out a throw, status reports, strategy with map callouts) each resolve to one of up to 40 curated full-Ultron lines. The set-piece and register pools were also de-biblicalized (machine/evolution/immortal/superior register throughout). Covers all agents & maps and holds real conversation (the wake word is required for every turn โ the wake-free follow-up window is off by default after it was observed firing on un-addressed room/stream speech); validated against a 20,000-case adversarial corpus (matcher 99.4% clean after corpus-loop hardening, deterministic path ~0.15 ms) |
๐ ย "kenning, repeat to my team watermelon" | The soundboard check โ when teammates ask you to say a specific word to prove a human's on comms, Ultron speaks the exact phrase verbatim (no LLM, any literal word) in his trained voice |
๐๏ธ ย "kenning, pull up your settings" | Spawns a detached dark-theme control panel โ edit knobs at a glance โ every toggle hot-applies live (no restart) โ CLOSE leaves zero residue |
๐ต ย "kenning, play some Daft Punk" | Full hands-free Spotify control by voice โ play / queue / "play X next" / pause / resume / skip / previous / restart / "what's playing" / volume upยทdownยท"set it to 40"ยท"lower it by 10%" / muteยทunmute / shuffle / repeat / likeยทunlike โ understanding dozens of natural phrasings, with confirmations in Ultron's cold machine register. Web API over HTTPS only (no GPU, no LLM) so it stays live even in gaming/anticheat mode |
๐ก๏ธ ย "kenning, engage gaming mode" | Frees VRAM/Docker and keeps every desktop-interaction surface (input injection, screen capture, UIA, windows) entirely out of RAM โ not just call-blocked but never imported (pinned by clean-subprocess tests + a boot posture self-audit); zero foreign-process memory reads / injection / hooks / raw-input anywhere in source. Default-ON and safe-by-default. Only shared-mode audio (the team relay + Spotify, the well-trodden voice-changer class) stays live โ kernel-anticheat-safe (Vanguard/EAC/BattlEye) |
| ๐งช ย Tests | 12188 passing ยท 39 skipped ยท 27 pre-existing baseline/env fails (~240 s sweep) |
| โก ย Latency (TTFA) | ~266 ms composite cache-hit turn (LLM TTFT 172 ms, TTS synth 78 ms, STT 16 ms) |
| ๐ง ย VRAM | ~6.3 GB standby on RTX 4070 Ti (peak ~6.7 GB) โ ~2.1 GB in gaming mode |
| ๐ ๏ธ ย Active stack | Parakeet TDT STT (CUDA) ยท Qwen 3.5 4B Q4_K_M (CUDA) ยท Kokoro StyleTTS2 (CUDA, fine-tuned voice) ยท OpenClaw bridge live |
| ๐ ย License | MIT |
๐ค Voice pipeline
|
๐ง Reasoning
|
๐ Web + tools
|
๐ก๏ธ Safety + ops
|
mic โ wake "kenning" OR addressing classifier (WARM)
โ Silero VAD + Smart Turn V3 (CPU, ~12 ms)
โ STT: DualSTTRegistry (moonshine | parakeet | whisper)
โ Intent recognizer (Gemma-300M CPU): short-circuits gaming / fresh-data intents
โ Local clock reply for bare time/date queries (~5 ms, no LLM)
โ classify_routing() โ 23 RoutingIntentKind dispatches
โโ coding kinds โ AI coding agent subprocess (optional supervisor stack)
โโ OPEN_LAST_SOURCE โ opens cited URL from prior search
โโ NAVIGATE_TO_SITE โ SearxNG top-10 โ domain-score โ opens best
โโ APP_LAUNCH โ native Chrome/Discord/Spotify launcher
โโ GAMING_MODE โ VRAM reclaim chain (~2.3 GB freed)
โโ conversational โ LLM (Qwen 3.5 4B) with optional:
โ ยท web-search gate (rules + preflight LLM)
โ ยท multi-pass RAG retrieval
โ ยท news-category SearxNG routing
โโ stream tokens โ Kokoro TTS (CUDA, fine-tuned voice)
โ typed-bus events publish at every stage
โ async-write conversation turn to Qdrant
โ (follow-up window OFF by default โ every turn requires the wake word)
๐ Full per-module reference:
docs/codebase_structure.mdโ the binding single-source map of the system.
# 1. Clone
git clone https://github.com/1v9Khan/ultronPrototype.git
cd ultronPrototype
# 2. Python 3.11 + deps (~7 GB; PyTorch CUDA, llama-cpp, faster-whisper, Kokoro, ...)
python -m venv .venv
.venv\Scripts\activate # Windows
# source .venv/bin/activate # Linux/macOS (untested)
pip install -e .
# 3. Models (~5 GB; wake word, Smart Turn, Moonshine, Kokoro, Qwen GGUFs)
python scripts/download_models.py
# 4. Configure
copy .env.example .env # optional: add Brave API key for web search
# tune config.yaml for your mic / monitors / preferences
# 5. Launch
python -m kenning
Then say: "kenning" โ and start talking.
โ ๏ธ This is a research prototype, not a turn-key product. It targets one developer's specific hardware (RTX 4070 Ti, AMD CPU, Windows 11) and use case. Treat the setup as a recipe to adapt, not a one-click install. Some optional integrations (OpenClaw, Telegram, ComfyUI media gen, mobile node) require additional credential-dependent setup โ see the docs below.
| Recommended | Minimum | |
|---|---|---|
| GPU | RTX 4070 Ti (12 GB) | RTX 3060 (12 GB) โ untested, expect higher latency |
| CPU | AMD Ryzen 7 5800X+ (8c/16t) | 4 cores / 8 threads |
| RAM | 32 GB | 16 GB (constrained) |
| Disk | 30 GB free | 20 GB free |
| OS | Windows 11 | Windows 10 / Linux (untested) |
| Python | 3.11 | 3.10+ |
| CUDA | 12.4+ | 11.8 |
All tunables live in config.yaml at the project root โ schema-validated by Pydantic in src/kenning/config.py. The top of that file lists the ~12 actively-tuned knobs.
Key sections:
| Section | What it controls |
|---|---|
audio | Mic input device + output device + ring buffer |
vad ยท stt | VAD silence thresholds + STT engine selector + gaming fallback |
llm | Preset + n_ctx + speculative decoding + KV cache |
tts | Engine + voicepack + boundary smoothing + cadence |
memory | Qdrant store + RAG top-k + min-relevance + contextual retrieval |
web_search | Provider chain + reader chain + ranker dispatch |
safety | 141 rule toggles + sandbox roots + audit log path |
coding.supervisor | opencode-inspired project digest stack (default OFF) |
gaming_mode | VRAM reclaim chain triggers + targets |
Override via KENNING_* env vars; see .env.example. Restart after any change.
๐ Start here:
docs/codebase_structure.mdโ the binding single-source map of every module, script, test, and runtime artifact. Maintenance contract enforced per commit.
| Doc | Topic |
|---|---|
docs/architecture.md | Pipeline overview + hardware target |
docs/configuration.md | Per-key config reference |
docs/operations.md | Day-to-day running + recovery |
docs/development.md | Test layout + debugging recipes |
docs/routing.md | Capability routing |
docs/error_handling.md | Error catalog |
docs/4b_optimization_plan.md | 4B LLM migration (complete) |
| Doc | Topic |
|---|---|
docs/openclaw_integration_final_summary.md | Cross-phase summary + setup checklist |
docs/openclaw_telegram_setup.md | Telegram channel (bot token) |
docs/openclaw_heartbeat_setup.md | Heartbeat agents block |
docs/openclaw_browser_setup.md | Browser tool (Playwright + Chromium) |
docs/openclaw_cron_setup.md | Cron jobs (Task Scheduler fallback) |
docs/openclaw_hooks_setup.md | Bundled hooks |
docs/openclaw_memory_wiki_setup.md | Memory Wiki plugin |
docs/openclaw_media_generation_setup.md | Local ComfyUI media generation |
docs/mobile_node_setup.md | iOS / Android pairing |
| Doc | Topic |
|---|---|
docs/comprehensive_test_plan.md / comprehensive_test_report.md | Functional pass (16 phases, 38 dimensions) |
docs/comprehensive_quality_plan.md / comprehensive_quality_report.md | Quality pass (Q0โQ13, 38 dimensions, prompt-injection defense audit) |
docs/smoke_test.md | 16-step interactive smoke procedure |
This is a research prototype, not a production product. It evolves through many tight iteration cycles. Behavior-changing features land behind feature flags (default OFF) until live-validated. The voice-quality baseline is treated as a strict latency / VRAM contract โ any hot-path change re-runs scripts/measure_baseline.py and documents the delta.
If you're reading the source, the highest-leverage entry point is src/kenning/pipeline/orchestrator.py โ the main event loop everything else hangs off.
If you find Kenning interesting, a star helps it surface to other folks who want a local voice assistant.
MIT โ see LICENSE.
Built on the shoulders of these open-source projects:
bge-small ยท DuckDuckGo ยท faster-whisper ยท flan-t5-small ยท Kokoro ยท llama.cpp ยท moondream2 ยท Moonshine ยท opencode ยท openWakeWord ยท Parakeet TDT ยท Piper ยท pywinauto ยท Qdrant ยท Qwen 3.5 ยท RVC ยท SearxNG ยท Silero VAD ยท Smart Turn V3 ยท Trafilatura ยท XTTS v2
Built for one developer's RTX 4070 Ti, then shared.
699 commits
Python
99.5%
Say "kenning." Talk. Get answers in a custom voice. Everything runs on your GPU.
Commit 816df7c is the designated stable, restorable baseline of the gaming
teammate-relay assistant (lean boot ยท gaming engaged ยท anticheat default) โ git tag
ultron-0.1, pinned branch release/ultron-0.1, a GitHub
release, and a
standalone launchable backup at E:\Ultron-0.1\. Restore the dev tree with
git checkout ultron-0.1; launch the backup (a known-good build that stays
streamable while the dev version is under maintenance) with
E:\Ultron-0.1\launch_ultron_0_1.ps1. Runbook: docs/ultron_0_1_baseline.md.
The active dev line is Ultron 1.0: every response is now authored by the loaded local LLM by default,
with a voice command to switch back to the deterministic curated pools at any time (verbatim "repeat" lines and
known-fact answers always stay exact). The deterministic snap matchers are retired-not-removed and repurposed as
routers that pick a curated prompt and inject the matching snap lines + per-agent libraries as in-context
exemplars. Verbosity is now set at the prompt level on two independent, voice-switchable axes โ a callout
axis (none / low / medium / high / max flavor-tail length on a tactical callout; "none" โ the old deterministic
snap) and a conversation axis (low / medium / high / max reply length for private and social lines) โ each
with strict per-level guidelines so wording and personality stay varied without losing facts or overrunning the
length. Plus an optional-wakeword always-listening four-class intent gate wired into the run loop
(addressing.always_listening, default off) and private "me-only" replies that route to the desktop channel. A
stop-window FLAG button logs the last turn โ a disliked response, one that shouldn't have fired, or one that
should have but didn't โ to a review log for later refinement. Each increment is proven regression-clean against
the frozen control baseline; the stable Ultron 0.1 baseline above is untouched. Architecture, live status, and
the regression baseline live under docs/ultron_1_0/ (start at 00_process_log/STATUS.md).
Live calibration of the gate thresholds and the latency pass remain before the flags flip on by default.
A fully-local Twitch presence layer (flag-gated, default OFF โ main runtime is byte-identical until
enabled) lets Ultron co-host a stream without touching the anticheat contract: every Twitch network call lives in
out-of-process sidecars (chat read ยท Helix moderation write ยท a Llama-Guard safety model ยท an OBS overlay), so
the anticheat-pinned main process never imports a network or automation library for any of it.
@viewer, no TTS), and every viewer-directed reply is @-tagged to them."Ultron, ban / timeout / unban / delete <viewer>" (two-phase voice confirm with a
spoken read-back) plus chat-settings (slow mode ยท followers only ยท subscribers only ยท emote only ยท
clear chat); all writes go through the broadcaster-token sidecar with idempotency + a mass-action breaker.!slots / !wheel / !heist / !duel / !leaderboard and the
channel-point game redeems are provably-fair, house-funded, and EventSub-replay-safe (a local idempotency layer
over the SE points API).script-src CSP
grant lets it run inside OBS's embedded browser) with a small replay buffer so a freshly-connected source catches
up; ?demo=1 previews every card./shoutouts the raider and vocally welcomes the
raiders (thanks them, invites them to stick around, points them to chat) โ idempotent + fail-open, riding the
broadcaster-token EventSub session."open moderation panel") with a button per mod action (+ a username box) and a fuzzy-match confirm popup, and a dev
test panel ("show me the test panel") that fires any redeem / chat game / raid / speak locally so you can
rehearse without real viewers โ in-process tkinter (a click is the app's own window message, not input monitoring).!ultron posts the command list on demand; an auto commands-panel and a
"talk to Ultron" hint both link a one-page viewer guide, plus periodic auto-trivia rounds.Built, unit-tested offline (full tests/twitch green), and live on-stream. Go-live runbook:
docs/twitch_integration/FIRST_STREAM_CHECKLIST.md.
What would a voice assistant feel like if it lived entirely on your GPU instead of in someone else's data center?
| You say | Kenning does |
|---|---|
๐ฆ๏ธ ย "kenning, what's the weather in Paris?" | Detects fresh-data intent โ SearxNG โ reads result โ speaks the forecast |
๐ป ย "kenning, write me a script that converts PDFs to Docx" | Spawns isolated AI coding agent โ scaffolds project โ runs tests โ narrates progress |
๐ฎ ย "kenning, engage gaming mode" | Swaps LLM โ kills GPU services โ frees ~2.3 GB VRAM for your game |
๐ ย "kenning, take me to HBO Max" | Recognizes navigate intent โ opens Chrome to the best-matching domain |
๐ ย "kenning, what time is it in Tokyo?" | Hits local zoneinfo cache โ speaks the answer in ~5 ms (no LLM, no search) |
๐งญ ย "kenning, switch to the 8B" | Hot-swaps the local LLM preset mid-conversation |
โก ย "kenning, switch the model to the GPU" / "โฆback to the CPU" | Hot-moves the 3B between CPU and GPU mid-game with a device-optimized config (GPU: full layer offload + CUDA flash-attention + quantized KV + large batches; CPU: zero GPU layers + F16 KV + a smaller micro-batch so prefill never steals game frames) โ borrow the card for faster replies between rounds, hand it back for the next fight. Only the model compute location changes (VRAM vs system RAM), so it's anticheat-irrelevant |
๐ฃ๏ธ ย "kenning, tell my team they are pushing B" | Valorant teammate-relay: tactical callouts resolve deterministically (subject-exact, every count/agent/location/ability preserved, never the LLM) โ a fact-preserving fallback relays the literal rather than let the model drop or invert a callout. Nearly every line then carries a short, in-character Ultron flavor tail (a faithful Avengers: Age of Ultron clone) selected for the callout โ agent-specific for a named enemy ("Their Neon has ult. Overdrive. A finite surge."), plural for a group, owner-aware (contempt at enemies, cold command for your orders, stoic for your own status โ never mocking you) โ from a ~1,628-entry character-tailored library covering all 29 agents, hand-curated line-by-line (every agent's ult cell uses its real ultimate, every utility cell is ability-tagged, filler/off-topic/wrong-kit lines cut; 4,147 drafts โ 1,628 tight entries). Tail selection is a hybrid keyed-coarse + tagged-pool pipeline: a coarse route (agent + one of 16 enemy situations) picks the right pool, a 4-tier tag filter (location / damage / ability tags) narrows it to the tails that fit this callout, and for large pools an opt-in semantic re-ranker on the loopback embedder sidecar can fine-select with anti-repeat โ but for small curated cells (under 5 candidates) the sidecar embed is skipped entirely and a deterministic LRU pick is used, so every curated callout gets zero-latency contextual routing by default (KENNING_ENABLE_TAIL_SELECTOR off). An ult keyword (ulted/ultimate) always lifts the situation to the correct ult pool regardless of parse path; a callout verb (mollied/walled/darted) routes to the right per-ability cell via _VERB_TO_ABILITY. A semantic relay-intent gate (_relay_intent.py, reusing the router's sidecar) vetoes the bare-callout "tell my team" prepend for narration, banter, questions, and Marvel/identity talk before they reach the relay path โ with a fast-path narration regex (_NARRATION_MUSING_RE) that fires even when the sidecar is down โ hardened by a 25,000-case audit so the deterministic layer alone (no embedder) reframes "let my team know X" callouts, accepts terse "they are A" position calls, keeps a directive after a reported-context clause, and refuses first-person musings/recounts ("I told my team โฆ and they โฆ", "part of me wants to tell them โฆ") rather than relaying them. Ask-form team questions invert to natural spoken order โ a trailing copula or negated auxiliary fronts after the wh-word ("ask my team why they aren't smoking" โ "Why aren't they smoking?"), and a yes/no question gets the auxiliary inserted so the inflection survives TTS where a bare "?" would not ("ask Sage if she has a heal" โ "Sage, do you have a heal?"). A baked common-English-word set (4,771 words, _common_words.py, generated offline) protects real words from the STT gazetteer snapper so common tokens are never corrupted into agent names. The pre-routing normalizer's disfluency resolver was rewritten to preserve the relay lead across restarts/corrections while correctly stripping upstream fillers. All the ML stays in the sidecar or in offline build/audit scripts (never imported at runtime); the anticheat-pinned main process imports only numpy + urllib for this path. Off-snap lines (banter, economy, opinions, identity, Marvel, answers, greets) get the full persona โ plays on a VoiceMeeter strip so your voice chat hears it. When you don't trust the model to improvise, 73 explicit fallback commands (refuse a dumb question, criticize/praise a teammate by name, call out a throw, status reports, strategy with map callouts) each resolve to one of up to 40 curated full-Ultron lines. The set-piece and register pools were also de-biblicalized (machine/evolution/immortal/superior register throughout). Covers all agents & maps and holds real conversation (the wake word is required for every turn โ the wake-free follow-up window is off by default after it was observed firing on un-addressed room/stream speech); validated against a 20,000-case adversarial corpus (matcher 99.4% clean after corpus-loop hardening, deterministic path ~0.15 ms) |
๐ ย "kenning, repeat to my team watermelon" | The soundboard check โ when teammates ask you to say a specific word to prove a human's on comms, Ultron speaks the exact phrase verbatim (no LLM, any literal word) in his trained voice |
๐๏ธ ย "kenning, pull up your settings" | Spawns a detached dark-theme control panel โ edit knobs at a glance โ every toggle hot-applies live (no restart) โ CLOSE leaves zero residue |
๐ต ย "kenning, play some Daft Punk" | Full hands-free Spotify control by voice โ play / queue / "play X next" / pause / resume / skip / previous / restart / "what's playing" / volume upยทdownยท"set it to 40"ยท"lower it by 10%" / muteยทunmute / shuffle / repeat / likeยทunlike โ understanding dozens of natural phrasings, with confirmations in Ultron's cold machine register. Web API over HTTPS only (no GPU, no LLM) so it stays live even in gaming/anticheat mode |
๐ก๏ธ ย "kenning, engage gaming mode" | Frees VRAM/Docker and keeps every desktop-interaction surface (input injection, screen capture, UIA, windows) entirely out of RAM โ not just call-blocked but never imported (pinned by clean-subprocess tests + a boot posture self-audit); zero foreign-process memory reads / injection / hooks / raw-input anywhere in source. Default-ON and safe-by-default. Only shared-mode audio (the team relay + Spotify, the well-trodden voice-changer class) stays live โ kernel-anticheat-safe (Vanguard/EAC/BattlEye) |
| ๐งช ย Tests | 12188 passing ยท 39 skipped ยท 27 pre-existing baseline/env fails (~240 s sweep) |
| โก ย Latency (TTFA) | ~266 ms composite cache-hit turn (LLM TTFT 172 ms, TTS synth 78 ms, STT 16 ms) |
| ๐ง ย VRAM | ~6.3 GB standby on RTX 4070 Ti (peak ~6.7 GB) โ ~2.1 GB in gaming mode |
| ๐ ๏ธ ย Active stack | Parakeet TDT STT (CUDA) ยท Qwen 3.5 4B Q4_K_M (CUDA) ยท Kokoro StyleTTS2 (CUDA, fine-tuned voice) ยท OpenClaw bridge live |
| ๐ ย License | MIT |
๐ค Voice pipeline
|
๐ง Reasoning
|
๐ Web + tools
|
๐ก๏ธ Safety + ops
|
mic โ wake "kenning" OR addressing classifier (WARM)
โ Silero VAD + Smart Turn V3 (CPU, ~12 ms)
โ STT: DualSTTRegistry (moonshine | parakeet | whisper)
โ Intent recognizer (Gemma-300M CPU): short-circuits gaming / fresh-data intents
โ Local clock reply for bare time/date queries (~5 ms, no LLM)
โ classify_routing() โ 23 RoutingIntentKind dispatches
โโ coding kinds โ AI coding agent subprocess (optional supervisor stack)
โโ OPEN_LAST_SOURCE โ opens cited URL from prior search
โโ NAVIGATE_TO_SITE โ SearxNG top-10 โ domain-score โ opens best
โโ APP_LAUNCH โ native Chrome/Discord/Spotify launcher
โโ GAMING_MODE โ VRAM reclaim chain (~2.3 GB freed)
โโ conversational โ LLM (Qwen 3.5 4B) with optional:
โ ยท web-search gate (rules + preflight LLM)
โ ยท multi-pass RAG retrieval
โ ยท news-category SearxNG routing
โโ stream tokens โ Kokoro TTS (CUDA, fine-tuned voice)
โ typed-bus events publish at every stage
โ async-write conversation turn to Qdrant
โ (follow-up window OFF by default โ every turn requires the wake word)
๐ Full per-module reference:
docs/codebase_structure.mdโ the binding single-source map of the system.
# 1. Clone
git clone https://github.com/1v9Khan/ultronPrototype.git
cd ultronPrototype
# 2. Python 3.11 + deps (~7 GB; PyTorch CUDA, llama-cpp, faster-whisper, Kokoro, ...)
python -m venv .venv
.venv\Scripts\activate # Windows
# source .venv/bin/activate # Linux/macOS (untested)
pip install -e .
# 3. Models (~5 GB; wake word, Smart Turn, Moonshine, Kokoro, Qwen GGUFs)
python scripts/download_models.py
# 4. Configure
copy .env.example .env # optional: add Brave API key for web search
# tune config.yaml for your mic / monitors / preferences
# 5. Launch
python -m kenning
Then say: "kenning" โ and start talking.
โ ๏ธ This is a research prototype, not a turn-key product. It targets one developer's specific hardware (RTX 4070 Ti, AMD CPU, Windows 11) and use case. Treat the setup as a recipe to adapt, not a one-click install. Some optional integrations (OpenClaw, Telegram, ComfyUI media gen, mobile node) require additional credential-dependent setup โ see the docs below.
| Recommended | Minimum | |
|---|---|---|
| GPU | RTX 4070 Ti (12 GB) | RTX 3060 (12 GB) โ untested, expect higher latency |
| CPU | AMD Ryzen 7 5800X+ (8c/16t) | 4 cores / 8 threads |
| RAM | 32 GB | 16 GB (constrained) |
| Disk | 30 GB free | 20 GB free |
| OS | Windows 11 | Windows 10 / Linux (untested) |
| Python | 3.11 | 3.10+ |
| CUDA | 12.4+ | 11.8 |
All tunables live in config.yaml at the project root โ schema-validated by Pydantic in src/kenning/config.py. The top of that file lists the ~12 actively-tuned knobs.
Key sections:
| Section | What it controls |
|---|---|
audio | Mic input device + output device + ring buffer |
vad ยท stt | VAD silence thresholds + STT engine selector + gaming fallback |
llm | Preset + n_ctx + speculative decoding + KV cache |
tts | Engine + voicepack + boundary smoothing + cadence |
memory | Qdrant store + RAG top-k + min-relevance + contextual retrieval |
web_search | Provider chain + reader chain + ranker dispatch |
safety | 141 rule toggles + sandbox roots + audit log path |
coding.supervisor | opencode-inspired project digest stack (default OFF) |
gaming_mode | VRAM reclaim chain triggers + targets |
Override via KENNING_* env vars; see .env.example. Restart after any change.
๐ Start here:
docs/codebase_structure.mdโ the binding single-source map of every module, script, test, and runtime artifact. Maintenance contract enforced per commit.
| Doc | Topic |
|---|---|
docs/architecture.md | Pipeline overview + hardware target |
docs/configuration.md | Per-key config reference |
docs/operations.md | Day-to-day running + recovery |
docs/development.md | Test layout + debugging recipes |
docs/routing.md | Capability routing |
docs/error_handling.md | Error catalog |
docs/4b_optimization_plan.md | 4B LLM migration (complete) |
| Doc | Topic |
|---|---|
docs/openclaw_integration_final_summary.md | Cross-phase summary + setup checklist |
docs/openclaw_telegram_setup.md | Telegram channel (bot token) |
docs/openclaw_heartbeat_setup.md | Heartbeat agents block |
docs/openclaw_browser_setup.md | Browser tool (Playwright + Chromium) |
docs/openclaw_cron_setup.md | Cron jobs (Task Scheduler fallback) |
docs/openclaw_hooks_setup.md | Bundled hooks |
docs/openclaw_memory_wiki_setup.md | Memory Wiki plugin |
docs/openclaw_media_generation_setup.md | Local ComfyUI media generation |
docs/mobile_node_setup.md | iOS / Android pairing |
| Doc | Topic |
|---|---|
docs/comprehensive_test_plan.md / comprehensive_test_report.md | Functional pass (16 phases, 38 dimensions) |
docs/comprehensive_quality_plan.md / comprehensive_quality_report.md | Quality pass (Q0โQ13, 38 dimensions, prompt-injection defense audit) |
docs/smoke_test.md | 16-step interactive smoke procedure |
This is a research prototype, not a production product. It evolves through many tight iteration cycles. Behavior-changing features land behind feature flags (default OFF) until live-validated. The voice-quality baseline is treated as a strict latency / VRAM contract โ any hot-path change re-runs scripts/measure_baseline.py and documents the delta.
If you're reading the source, the highest-leverage entry point is src/kenning/pipeline/orchestrator.py โ the main event loop everything else hangs off.
If you find Kenning interesting, a star helps it surface to other folks who want a local voice assistant.
MIT โ see LICENSE.
Built on the shoulders of these open-source projects:
bge-small ยท DuckDuckGo ยท faster-whisper ยท flan-t5-small ยท Kokoro ยท llama.cpp ยท moondream2 ยท Moonshine ยท opencode ยท openWakeWord ยท Parakeet TDT ยท Piper ยท pywinauto ยท Qdrant ยท Qwen 3.5 ยท RVC ยท SearxNG ยท Silero VAD ยท Smart Turn V3 ยท Trafilatura ยท XTTS v2
Built for one developer's RTX 4070 Ti, then shared.
699 commits
Python
99.5%