A Unity package for building fully local, offline AI companions and NPCs — a complete LLM + retrieval-augmented generation (RAG) + neural text-to-speech + character-animation pipeline that runs entirely on-device, with no cloud service, no API key, and no recurring cost.
Built as part of an M.Sc. thesis in Data Science at Università degli Studi di Milano-Bicocca, comparing this local pipeline against Convai, a leading commercial cloud platform, across latency, dialogue quality, voice quality, and knowledge-base capacity.
⚠️ Demo assets note: this repository includes the package and a reference demo scene. The demo scene's character model (VRoid Studio, imported via UniVRM) and animations (Mixamo).
llama.cpp server that the package launches and manages itself — no Ollama, no Python, no manual setup required by the end player.llama-server infrastructure as dialogue) and decodes the result in-engine via Unity's Inference Engine (Sentis) — no external TTS process, no cloud TTS API.NpcPersona, NpcVoice, RagSourceAsset), editable from the Inspector.| Reference demo scene | Embedded Server Setup |
|---|---|
![]() | ![]() |
| Test Console | Character configuration (no code) |
|---|---|
![]() | ![]() |
(Drop the corresponding image files into docs/images/ — see Adding the screenshots below.)
NPCDialogueAgent (central orchestrator — exposes events only, no direct
dependency on UI/animation)
│
├── LLM layer LlmClient (OpenAI-compatible) + EmbeddedLlmServer /
│ LlamaServerHost (launches & supervises llama-server
│ as a managed child process)
│
├── RAG layer RagRetriever — hybrid semantic + BM25 search,
│ fused with Reciprocal Rank Fusion; indices are
│ baked offline in the Editor, never at runtime
│
├── TTS layer SpeechSynthesizer → NeuTtsClient (speech-as-
│ language-modeling via llama-server) → in-engine
│ NeuCodec decoder (ONNX via Sentis)
│
└── UnityEvents → NPCAnimController (Idle / Thinking / Talking state
machine), subtitle UI, or any other presentation
layer — fully decoupled from the backend in use
Configuration is entirely data-driven: DynamicNpcSettings (global connection/runtime settings), NpcPersona (personality, world context, attached RAG sources), NpcVoice (cloneable voice + baked reference codes), and RagSourceAsset (one knowledge source + its chunking parameters).
llama-server binary, the dialogue GGUF model, the NeuTTS backbone and NeuCodec decoder, and espeak-ng.Platform note: the bundled binaries currently used in development are Windows builds (
llama-server.exe,espeak-ng.exe). If you need macOS/Linux support, you'll need to source or build compatible binaries for those platforms.
NpcPersona asset (Create → Dynamic NPCs → Persona) and fill in its name, personality, and world context.NpcVoice asset from a short (10–20s) reference audio sample and transcript, then bake it via the Voice Baker (requires a local Python environment, provisioned automatically).RagSourceAssets pointing at your own text files, set a chunk size/overlap, and bake the index via the RAG Baker.NpcPersona to an NPCDialogueAgent in your scene and you're done — no code required.Assets/
├── Scripts/
│ ├── NPCDialogueAgent.cs # central orchestrator
│ ├── NPCAnimController.cs # reference animation/presentation layer
│ ├── Runtime/
│ │ ├── Llm/ # LlmClient, EmbeddedLlmServer, LlamaServerHost
│ │ ├── Rag/ # RagRetriever, RagBm25Index, RagSourceAsset
│ │ ├── Tts/ # SpeechSynthesizer, NeuTtsClient, NeuCodecDecoder
│ │ ├── Config/ # DynamicNpcSettings, NpcPersona, NpcVoice
│ │ └── Internal/ # SentenceChunker, WavUtility, SseDownloadHandler, ...
│ ├── Editor/ # Setup Wizard, Test Console, bakers, diagnostics
│ └── Tools/ # Python voice-baking script (dev-machine only)
└── ...
As part of the accompanying thesis, this pipeline was compared against Convai, a commercial cloud-hosted conversational AI platform. Summary of findings:
| This project (local) | Convai (cloud) | |
|---|---|---|
| Latency | Comparable; faster on short exchanges, more variable on complex ones | Comparable; more consistent across question complexity |
| Voice quality | Judged more natural and expressive (subjective) | More monotone (subjective) |
| Knowledge base | No size limit (local, offline indexing) | 1 MB cap on Free tier; full feature gated behind Enterprise plan |
| Cost | One-time, offline | Subscription, usage-metered |
| Animation/lip-sync tooling | Basic (state-machine driven) | More polished out of the box |
| Standalone VR/AR | Not currently supported (see below) | Supported (cloud-mediated, no local inference needed) |
(Full methodology and results in the thesis.)
llama-server as a child process via System.Diagnostics.Process, which does not work on Android (the platform Meta Quest's OS is built on). A RemoteServer configuration — pointing a standalone build at a llama-server instance running elsewhere on the local network — should work in principle (it only requires UnityWebRequest, which is supported on Android) but has not yet been implemented or tested.C#
94.2%
ShaderLab
4.2%
HLSL
1.3%
A Unity package for building fully local, offline AI companions and NPCs — a complete LLM + retrieval-augmented generation (RAG) + neural text-to-speech + character-animation pipeline that runs entirely on-device, with no cloud service, no API key, and no recurring cost.
Built as part of an M.Sc. thesis in Data Science at Università degli Studi di Milano-Bicocca, comparing this local pipeline against Convai, a leading commercial cloud platform, across latency, dialogue quality, voice quality, and knowledge-base capacity.
⚠️ Demo assets note: this repository includes the package and a reference demo scene. The demo scene's character model (VRoid Studio, imported via UniVRM) and animations (Mixamo).
llama.cpp server that the package launches and manages itself — no Ollama, no Python, no manual setup required by the end player.llama-server infrastructure as dialogue) and decodes the result in-engine via Unity's Inference Engine (Sentis) — no external TTS process, no cloud TTS API.NpcPersona, NpcVoice, RagSourceAsset), editable from the Inspector.| Reference demo scene | Embedded Server Setup |
|---|---|
![]() | ![]() |
| Test Console | Character configuration (no code) |
|---|---|
![]() | ![]() |
(Drop the corresponding image files into docs/images/ — see Adding the screenshots below.)
NPCDialogueAgent (central orchestrator — exposes events only, no direct
dependency on UI/animation)
│
├── LLM layer LlmClient (OpenAI-compatible) + EmbeddedLlmServer /
│ LlamaServerHost (launches & supervises llama-server
│ as a managed child process)
│
├── RAG layer RagRetriever — hybrid semantic + BM25 search,
│ fused with Reciprocal Rank Fusion; indices are
│ baked offline in the Editor, never at runtime
│
├── TTS layer SpeechSynthesizer → NeuTtsClient (speech-as-
│ language-modeling via llama-server) → in-engine
│ NeuCodec decoder (ONNX via Sentis)
│
└── UnityEvents → NPCAnimController (Idle / Thinking / Talking state
machine), subtitle UI, or any other presentation
layer — fully decoupled from the backend in use
Configuration is entirely data-driven: DynamicNpcSettings (global connection/runtime settings), NpcPersona (personality, world context, attached RAG sources), NpcVoice (cloneable voice + baked reference codes), and RagSourceAsset (one knowledge source + its chunking parameters).
llama-server binary, the dialogue GGUF model, the NeuTTS backbone and NeuCodec decoder, and espeak-ng.Platform note: the bundled binaries currently used in development are Windows builds (
llama-server.exe,espeak-ng.exe). If you need macOS/Linux support, you'll need to source or build compatible binaries for those platforms.
NpcPersona asset (Create → Dynamic NPCs → Persona) and fill in its name, personality, and world context.NpcVoice asset from a short (10–20s) reference audio sample and transcript, then bake it via the Voice Baker (requires a local Python environment, provisioned automatically).RagSourceAssets pointing at your own text files, set a chunk size/overlap, and bake the index via the RAG Baker.NpcPersona to an NPCDialogueAgent in your scene and you're done — no code required.Assets/
├── Scripts/
│ ├── NPCDialogueAgent.cs # central orchestrator
│ ├── NPCAnimController.cs # reference animation/presentation layer
│ ├── Runtime/
│ │ ├── Llm/ # LlmClient, EmbeddedLlmServer, LlamaServerHost
│ │ ├── Rag/ # RagRetriever, RagBm25Index, RagSourceAsset
│ │ ├── Tts/ # SpeechSynthesizer, NeuTtsClient, NeuCodecDecoder
│ │ ├── Config/ # DynamicNpcSettings, NpcPersona, NpcVoice
│ │ └── Internal/ # SentenceChunker, WavUtility, SseDownloadHandler, ...
│ ├── Editor/ # Setup Wizard, Test Console, bakers, diagnostics
│ └── Tools/ # Python voice-baking script (dev-machine only)
└── ...
As part of the accompanying thesis, this pipeline was compared against Convai, a commercial cloud-hosted conversational AI platform. Summary of findings:
| This project (local) | Convai (cloud) | |
|---|---|---|
| Latency | Comparable; faster on short exchanges, more variable on complex ones | Comparable; more consistent across question complexity |
| Voice quality | Judged more natural and expressive (subjective) | More monotone (subjective) |
| Knowledge base | No size limit (local, offline indexing) | 1 MB cap on Free tier; full feature gated behind Enterprise plan |
| Cost | One-time, offline | Subscription, usage-metered |
| Animation/lip-sync tooling | Basic (state-machine driven) | More polished out of the box |
| Standalone VR/AR | Not currently supported (see below) | Supported (cloud-mediated, no local inference needed) |
(Full methodology and results in the thesis.)
llama-server as a child process via System.Diagnostics.Process, which does not work on Android (the platform Meta Quest's OS is built on). A RemoteServer configuration — pointing a standalone build at a llama-server instance running elsewhere on the local network — should work in principle (it only requires UnityWebRequest, which is supported on Android) but has not yet been implemented or tested.C#
94.2%
ShaderLab
4.2%
HLSL
1.3%