[[TOC]]
There are two scenes in this project which showcase current LLM capabilities and let you play around with them by changing prompts and other settings.
Showcases a small detective game with LLM characters that have a knowledge base.
The scene is a changed variant of an example scene from the Unity plugin LLMUnity. LLMUnity allows you to use local LLM models in Unity. Local means that an internet connection is not required to run the model and the model is stored on the device.
LLMUnity also implements RAG (Retrieval-Augmented Generation) which allows you to query a knowledge base and generate a response based on the query.
Each character in the scene has their own knowledge base and considers it when responding. Through the character's prompt you decide how the character deals with this knowledge.
I have extended the plugin to also support a connection to OpenAI's GPT API, so you can use that if you prefer. However, to use this in an app that you intend to ship, you must not include the API key directly, which requires you to set up a separate server to deal with user authentication and change the code in this demo.
You could extend the demo to work with audio input by using a Speech-to-Text (STT) and Text-to-Speech (TTS) plugin or service to work more naturally.
For local STT, you could start with whisper.unity which is based on the open-source Whisper model.
Be aware that you might require a Voice Activity Detector (VAD) to detect when the user is speaking.
For TTS you can start with the Piper implementation in .
LLMUnity is also planning a TTS extension (OuteTTS).
There are implementations of Kokoro TTS for C# but these are possibly not easy to get to work in Unity. (Sherpa-onnx (also includes STT) and Kokorosharp).
Showcases OpenAI's speech-to-speech LLM which responds to user input in real-time.
The scene features a rudimentary setup of a "character" that you can approach. When you enter the colored area, a connection with GPT Realtime is established and the area turns green. Press a button (left mouse button) to unmute the microphone. If you leave the area, the connection is severed and the area turns red.
As here too an API key is required, shipping it in an app is not recommended.
To use the scene, you have to create a file "Assets/StreamingAssets/api-keys.json" which should contain
{
"OPENAI_API_KEY": "your-api-key"
}
You can also just copy the sample file "Assets/StreamingAssets/api-keys.json.sample" and rename it.
Watch out for the costs of using the API! Currently, it is rather expensive, especially if you do not use the "mini" model. Set a hard limit in your OpenAI account to prevent getting out of this poor.
If you are not using a mute button or headphones, you have to deal with the audio coming from the speakers being picked up by the microphone. This can be done with Acoustic Echo Cancellation (AEC). As of now, no AEC plugin for Unity exists. Smartphones implement hardware AEC but I am unsure on how to use it.
If you want to animate a character from audio input, you can start with the following plugins.
If you want to create avatars in Unity that can talk but use a more comprehensive solution, you can take a look at the following plugins. There may be high running costs associated with these plugins.
C#
86.8%
ShaderLab
11.0%
HLSL
2.2%
[[TOC]]
There are two scenes in this project which showcase current LLM capabilities and let you play around with them by changing prompts and other settings.
Showcases a small detective game with LLM characters that have a knowledge base.
The scene is a changed variant of an example scene from the Unity plugin LLMUnity. LLMUnity allows you to use local LLM models in Unity. Local means that an internet connection is not required to run the model and the model is stored on the device.
LLMUnity also implements RAG (Retrieval-Augmented Generation) which allows you to query a knowledge base and generate a response based on the query.
Each character in the scene has their own knowledge base and considers it when responding. Through the character's prompt you decide how the character deals with this knowledge.
I have extended the plugin to also support a connection to OpenAI's GPT API, so you can use that if you prefer. However, to use this in an app that you intend to ship, you must not include the API key directly, which requires you to set up a separate server to deal with user authentication and change the code in this demo.
You could extend the demo to work with audio input by using a Speech-to-Text (STT) and Text-to-Speech (TTS) plugin or service to work more naturally.
For local STT, you could start with whisper.unity which is based on the open-source Whisper model.
Be aware that you might require a Voice Activity Detector (VAD) to detect when the user is speaking.
For TTS you can start with the Piper implementation in .
LLMUnity is also planning a TTS extension (OuteTTS).
There are implementations of Kokoro TTS for C# but these are possibly not easy to get to work in Unity. (Sherpa-onnx (also includes STT) and Kokorosharp).
Showcases OpenAI's speech-to-speech LLM which responds to user input in real-time.
The scene features a rudimentary setup of a "character" that you can approach. When you enter the colored area, a connection with GPT Realtime is established and the area turns green. Press a button (left mouse button) to unmute the microphone. If you leave the area, the connection is severed and the area turns red.
As here too an API key is required, shipping it in an app is not recommended.
To use the scene, you have to create a file "Assets/StreamingAssets/api-keys.json" which should contain
{
"OPENAI_API_KEY": "your-api-key"
}
You can also just copy the sample file "Assets/StreamingAssets/api-keys.json.sample" and rename it.
Watch out for the costs of using the API! Currently, it is rather expensive, especially if you do not use the "mini" model. Set a hard limit in your OpenAI account to prevent getting out of this poor.
If you are not using a mute button or headphones, you have to deal with the audio coming from the speakers being picked up by the microphone. This can be done with Acoustic Echo Cancellation (AEC). As of now, no AEC plugin for Unity exists. Smartphones implement hardware AEC but I am unsure on how to use it.
If you want to animate a character from audio input, you can start with the following plugins.
If you want to create avatars in Unity that can talk but use a more comprehensive solution, you can take a look at the following plugins. There may be high running costs associated with these plugins.
C#
86.8%
ShaderLab
11.0%
HLSL
2.2%