Create characters in Unity with LLMs!
See the code
LLM for Unity enables seamless integration of Large Language Models (LLMs) within the Unity engine.
It allows to create intelligent AI characters that players can interact with for an immersive experience.
The package includes a Retrieval-Augmented Generation (RAG) system for semantic search across your data, which can be used to enhance the character's knowledge.
The LLM backend, LlamaLib, is built on top of the awesome llama.cpp library and provided as a standalone C++/C# library.
At a glance • How to help • Games / Projects using LLM for Unity • Setup • Quick start • Advanced usage • RAG • LLM model management • Examples • License🧪 Tested on Unity: 2021 LTS, 2022 LTS, 2023, Unity 6
🚦 Upcoming Releases
For business inquiries you can reach out at hello@undream.ai.
Contact hello@undream.ai to add your project!
Method 1: Install using the asset store
Add to My AssetsWindow > Package ManagerPackages: My Assets option from the drop-downLLM for Unity package, click Download and then ImportMethod 2: Install using the GitHub repo:
Window > Package Manager+ button and select Add package from git URLhttps://github.com/undreamai/LLMUnity.git and click Add
First you will setup the LLM for your game:
Add Component and select the LLM script.Download Model button (~GBs).Load model button (see LLM model management).Then you can setup each of your characters as follows:
Add Component and select the LLMAgent script.System Prompt.LLM field if you have more than one LLM GameObjects.In your script you can then use it as follows:
using LLMUnity;
public class MyScript {
public LLMAgent llmAgent;
void HandleReply(string replySoFar){
// do something with the reply from the model as it is being produced
Debug.Log(replySoFar);
}
void Game(){
// handle the response as it is being produced
...
_ = llmAgent.Chat("Hello bot!", HandleReply);
...
}
async void GameAsync(){
// or handle the entire response in one go
...
string reply = await llmAgent.Chat("Hello bot!");
Debug.Log(reply);
...
}
}
You can also specify a function to call when the model reply has been completed:
void ReplyCompleted(){
// do something when the reply from the model is complete
Debug.Log("The AI has finished replying");
}
void Game(){
// your game function
...
_ = llmAgent.Chat("Hello bot!", HandleReply, ReplyCompleted);
...
}
To stop the chat without waiting for its completion you can use:
llmAgent.CancelRequests();
That's it! Your AI character is ready to chat! ✨
For mobile apps you can use models with up to 1-2 billion parameters ("Tiny models" in the LLM model manager).
Larger models will typically not work due to the limited mobile hardware.
iOS iOS can be built with the default player settings.
Android
On Android you need to specify the IL2CPP scripting backend and the ARM64 as the target architecture in the player settings.
These settings can be accessed from the Edit > Project Settings menu within the Player > Other Settings section.

Since mobile app sizes are typically small, you can download the LLM model the first time the app launches.
This functionality is enabled with the Download on Build option.
In your project you can wait until the model download is complete with:
await LLM.WaitUntilModelSetup();
You can also receive calls the download progress during the model download:
await LLM.WaitUntilModelSetup(SetProgress);
void SetProgress(float progress){
string progressPercent = ((int)(progress * 100)).ToString() + "%";
Debug.Log($"Download progress: {progressPercent}");
}
This is useful to present e.g. a progress bar. The MobileDemo demonstrates an example application for Android / iOS.
To restrict the output of the LLM you can use a grammar, read more here.
The grammar can edited directly in the Grammar field of the LLMAgent or saved in a gbnf / json schema file and loaded with the Load Grammar button (Advanced options).
For instance to receive replies in json format you can use the json.gbnf grammar.
Alternatively you can set the grammar directly in your script:
llmAgent.grammar = "your grammar here";
For function calling you can define similarly a grammar that allows only the function names as output, and then call the respective function.
You can look into the FunctionCalling sample for an example implementation.
To add new messages you can do:
_ = llmAgent.AddUserMessage("your user message");
_ = llmAgent.AddAssistantMessage("your assistant reply");
To automatically save / load your chat history, you can specify the Save parameter of the LLMAgent to the filename (or relative path) of your choice.
The chat history is saved in the persistentDataPath folder of Unity as a json object.
To manually save your chat history, you can use:
_ = llmAgent.SaveHistory();
and to load the history:
_ = llmAgent.Loadistory();
void WarmupCompleted(){
// do something when the warmup is complete
Debug.Log("The AI is nice and ready");
}
void Game(){
// your game function
...
_ = llmAgent.Warmup(WarmupCompleted);
...
}
The last argument of the Chat function is a boolean that specifies whether to add the message to the history (default: true):
void Game(){
// your game function
...
string message = "Hello bot!";
_ = llmAgent.Chat(message, HandleReply, ReplyCompleted, false);
...
}
void Game(){
// your game function
...
string message = "The cat is away";
_ = llmAgent.Completion(message, HandleReply, ReplyCompleted);
...
}
using UnityEngine;
using LLMUnity;
public class MyScript : MonoBehaviour
{
LLM llm;
LLMAgent llmAgent;
async void Start()
{
// disable gameObject so that theAwake is not called immediately
gameObject.SetActive(false);
// Add an LLM object
llm = gameObject.AddComponent<LLM>();
// set the model using the filename of the model.
// The model needs to be added to the LLM model manager (see LLM model management) by loading or downloading it.
// Otherwise the model file can be copied directly inside the StreamingAssets folder.
llm.model = "Qwen3-4B-Q4_K_M.gguf";
// optional: you can also set loras in a similar fashion and set their weights (if needed)
llm.AddLora("my-lora.gguf");
llm.AddLora("my-lora-2.gguf", 0.5f);
// optional: set number of threads
llm.numThreads = -1;
// optional: enable GPU by setting the number of model layers to offload to it
llm.numGPULayers = 10;
// Add an LLMAgent object
llmAgent = gameObject.AddComponent<LLMAgent>();
// set the LLM object that handles the model
llmAgent.llm = llm;
// set the character prompt
llmAgent.systemPrompt = "A chat between a curious human and an artificial intelligence assistant.";
// set the AI and player name
llmAgent.assistantRole = "AI";
llmAgent.userRole = "Human";
// optional: set a save path
llmAgent.save = "AICharacter1.json";
// optional: set a grammar
llmAgent.grammar = "your grammar here";
// re-enable gameObject
gameObject.SetActive(true);
}
}
You can use a remote server to carry out the processing and implement characters that interact with it.
Create the server
To create the server:
LLM script as described aboveRemote option of the LLM and optionally configure the server port and API keyAlternatively you can use a server binary for easier deployment:
servers folder selected and start the server by running the command copied from above.Create the characters
Create a second project with the game characters using the LLMAgent script as described above.
Enable the Remote option and configure the host with the IP address (starting with "http://") and port / API key of the server.
The Embeddings function can be used to obtain the emdeddings of a phrase:
List<float> embeddings = await llmAgent.Embeddings("hi, how are you?");
A detailed documentation on function level can be found here:
LLM for Unity implements a super-fast similarity search functionality with a Retrieval-Augmented Generation (RAG) system.
It is based on the LLM embeddings, and the Approximate Nearest Neighbors (ANN) search from the usearch library.
Semantic search works as follows.
Building the data You provide text inputs (a phrase, paragraph, document) to add to the data.
Each input is split into chunks (optional) and encoded into embeddings with a LLM.
Searching You can then search for a query text input.
The input is again encoded and the most similar text inputs or chunks in the data are retrieved.
To use semantic serch:
Add Component and select the RAG script.SimpleSearch is a simple brute-force search, whileDBSearch is a fast ANN method that should be preferred in most cases.Alternatively, you can create the RAG from code (where llm is your LLM):
RAG rag = gameObject.AddComponent<RAG>();
rag.Init(SearchMethods.DBSearch, ChunkingMethods.SentenceSplitter, llm);
In your script you can then use it as follows :unicorn::
using LLMUnity;
public class MyScript : MonoBehaviour
{
RAG rag;
async void Game(){
...
string[] inputs = new string[]{
"Hi! I'm a search system.",
"the weather is nice. I like it.",
"I'm a RAG system"
};
// add the inputs to the RAG
foreach (string input in inputs) await rag.Add(input);
// get the 2 most similar inputs and their distance (dissimilarity) to the search query
(string[] results, float[] distances) = await rag.Search("hello!", 2);
// to get the most similar text parts (chunks), instead of full input, you can enable the returnChunks option
rag.ReturnChunks(true);
(results, distances) = await rag.Search("hello!", 2);
...
}
}
You can also add / search text inputs for groups of data e.g. for a specific character or scene:
// add the inputs to the RAG for a group of data e.g. an orc character
foreach (string input in inputs) await rag.Add(input, "orc");
// get the 2 most similar inputs for the group of data e.g. the orc character
(string[] results, float[] distances) = await rag.Search("how do you feel?", 2, "orc");
...
You can save the RAG state (stored in the `Assets/StreamingAssets` folder):
``` c#
rag.Save("rag.zip");
and load it from disk:
await rag.Load("rag.zip");
You can use the RAG to feed relevant data to the LLM based on a user message:
string message = "How is the weather?";
(string[] similarPhrases, float[] distances) = await rag.Search(message, 3);
string prompt = "Answer the user query based on the provided data.\n\n";
prompt += $"User query: {message}\n\n";
prompt += $"Data:\n";
foreach (string similarPhrase in similarPhrases) prompt += $"\n- {similarPhrase}";
_ = llmAgent.Chat(prompt, HandleReply, ReplyCompleted);
The RAG sample includes an example RAG implementation as well as an example RAG-LLM integration.
That's all :sparkles:!
LLM for Unity includes a built-in model manager for easy model handling.
The model manager allows to load or download LLMs and can be found as part of the LLM GameObject:

You can download models with the Download model button.
LLM for Unity includes different state of the art models built-in for different model sizes, quantised with the Q4_K_M method.
Alternative models can be downloaded from HuggingFace in .gguf format.
You can download a model locally and load it with the Load model button, or copy the URL in the Download model > Custom URL field to directly download it.
If a HuggingFace model does not provide a gguf file, it can be converted to gguf with this online converter.
You can create lighter builds by selecting the Download on Build option.
The models will be downloaded the first time the game starts instead of bundled in the build.
If you have loaded a model locally you need to set its URL through the expanded view, otherwise it will be copied in the build.
❕ Before using any model make sure you check their license ❕
The Samples~ folder contains several examples of interaction 🤖:
To install a sample:
Window > Package ManagerLLM for Unity Package. From the Samples Tab, click Import next to the sample you want to install.The samples can be run with the Scene.unity scene they contain inside their folder.
In the scene, select the LLM GameObject and specify the LLM of your choice (see LLM model management).
Save the scene, run and enjoy!
The license of LLM for Unity is Apache 2.0 (LICENSE.md) and uses third-party software with MIT and Apache licenses. Some models included in the asset define their own license terms, please review them before using each model. Third-party licenses can be found in the (Third Party Notices.md).
12,069 followers · starred Mar 2024
216 followers · starred Aug 2026
3 followers · starred Apr 2025
250 followers · starred May 2026
C#
100.0%
Create characters in Unity with LLMs!
See the code
LLM for Unity enables seamless integration of Large Language Models (LLMs) within the Unity engine.
It allows to create intelligent AI characters that players can interact with for an immersive experience.
The package includes a Retrieval-Augmented Generation (RAG) system for semantic search across your data, which can be used to enhance the character's knowledge.
The LLM backend, LlamaLib, is built on top of the awesome llama.cpp library and provided as a standalone C++/C# library.
At a glance • How to help • Games / Projects using LLM for Unity • Setup • Quick start • Advanced usage • RAG • LLM model management • Examples • License🧪 Tested on Unity: 2021 LTS, 2022 LTS, 2023, Unity 6
🚦 Upcoming Releases
For business inquiries you can reach out at hello@undream.ai.
Contact hello@undream.ai to add your project!
Method 1: Install using the asset store
Add to My AssetsWindow > Package ManagerPackages: My Assets option from the drop-downLLM for Unity package, click Download and then ImportMethod 2: Install using the GitHub repo:
Window > Package Manager+ button and select Add package from git URLhttps://github.com/undreamai/LLMUnity.git and click Add
First you will setup the LLM for your game:
Add Component and select the LLM script.Download Model button (~GBs).Load model button (see LLM model management).Then you can setup each of your characters as follows:
Add Component and select the LLMAgent script.System Prompt.LLM field if you have more than one LLM GameObjects.In your script you can then use it as follows:
using LLMUnity;
public class MyScript {
public LLMAgent llmAgent;
void HandleReply(string replySoFar){
// do something with the reply from the model as it is being produced
Debug.Log(replySoFar);
}
void Game(){
// handle the response as it is being produced
...
_ = llmAgent.Chat("Hello bot!", HandleReply);
...
}
async void GameAsync(){
// or handle the entire response in one go
...
string reply = await llmAgent.Chat("Hello bot!");
Debug.Log(reply);
...
}
}
You can also specify a function to call when the model reply has been completed:
void ReplyCompleted(){
// do something when the reply from the model is complete
Debug.Log("The AI has finished replying");
}
void Game(){
// your game function
...
_ = llmAgent.Chat("Hello bot!", HandleReply, ReplyCompleted);
...
}
To stop the chat without waiting for its completion you can use:
llmAgent.CancelRequests();
That's it! Your AI character is ready to chat! ✨
For mobile apps you can use models with up to 1-2 billion parameters ("Tiny models" in the LLM model manager).
Larger models will typically not work due to the limited mobile hardware.
iOS iOS can be built with the default player settings.
Android
On Android you need to specify the IL2CPP scripting backend and the ARM64 as the target architecture in the player settings.
These settings can be accessed from the Edit > Project Settings menu within the Player > Other Settings section.

Since mobile app sizes are typically small, you can download the LLM model the first time the app launches.
This functionality is enabled with the Download on Build option.
In your project you can wait until the model download is complete with:
await LLM.WaitUntilModelSetup();
You can also receive calls the download progress during the model download:
await LLM.WaitUntilModelSetup(SetProgress);
void SetProgress(float progress){
string progressPercent = ((int)(progress * 100)).ToString() + "%";
Debug.Log($"Download progress: {progressPercent}");
}
This is useful to present e.g. a progress bar. The MobileDemo demonstrates an example application for Android / iOS.
To restrict the output of the LLM you can use a grammar, read more here.
The grammar can edited directly in the Grammar field of the LLMAgent or saved in a gbnf / json schema file and loaded with the Load Grammar button (Advanced options).
For instance to receive replies in json format you can use the json.gbnf grammar.
Alternatively you can set the grammar directly in your script:
llmAgent.grammar = "your grammar here";
For function calling you can define similarly a grammar that allows only the function names as output, and then call the respective function.
You can look into the FunctionCalling sample for an example implementation.
To add new messages you can do:
_ = llmAgent.AddUserMessage("your user message");
_ = llmAgent.AddAssistantMessage("your assistant reply");
To automatically save / load your chat history, you can specify the Save parameter of the LLMAgent to the filename (or relative path) of your choice.
The chat history is saved in the persistentDataPath folder of Unity as a json object.
To manually save your chat history, you can use:
_ = llmAgent.SaveHistory();
and to load the history:
_ = llmAgent.Loadistory();
void WarmupCompleted(){
// do something when the warmup is complete
Debug.Log("The AI is nice and ready");
}
void Game(){
// your game function
...
_ = llmAgent.Warmup(WarmupCompleted);
...
}
The last argument of the Chat function is a boolean that specifies whether to add the message to the history (default: true):
void Game(){
// your game function
...
string message = "Hello bot!";
_ = llmAgent.Chat(message, HandleReply, ReplyCompleted, false);
...
}
void Game(){
// your game function
...
string message = "The cat is away";
_ = llmAgent.Completion(message, HandleReply, ReplyCompleted);
...
}
using UnityEngine;
using LLMUnity;
public class MyScript : MonoBehaviour
{
LLM llm;
LLMAgent llmAgent;
async void Start()
{
// disable gameObject so that theAwake is not called immediately
gameObject.SetActive(false);
// Add an LLM object
llm = gameObject.AddComponent<LLM>();
// set the model using the filename of the model.
// The model needs to be added to the LLM model manager (see LLM model management) by loading or downloading it.
// Otherwise the model file can be copied directly inside the StreamingAssets folder.
llm.model = "Qwen3-4B-Q4_K_M.gguf";
// optional: you can also set loras in a similar fashion and set their weights (if needed)
llm.AddLora("my-lora.gguf");
llm.AddLora("my-lora-2.gguf", 0.5f);
// optional: set number of threads
llm.numThreads = -1;
// optional: enable GPU by setting the number of model layers to offload to it
llm.numGPULayers = 10;
// Add an LLMAgent object
llmAgent = gameObject.AddComponent<LLMAgent>();
// set the LLM object that handles the model
llmAgent.llm = llm;
// set the character prompt
llmAgent.systemPrompt = "A chat between a curious human and an artificial intelligence assistant.";
// set the AI and player name
llmAgent.assistantRole = "AI";
llmAgent.userRole = "Human";
// optional: set a save path
llmAgent.save = "AICharacter1.json";
// optional: set a grammar
llmAgent.grammar = "your grammar here";
// re-enable gameObject
gameObject.SetActive(true);
}
}
You can use a remote server to carry out the processing and implement characters that interact with it.
Create the server
To create the server:
LLM script as described aboveRemote option of the LLM and optionally configure the server port and API keyAlternatively you can use a server binary for easier deployment:
servers folder selected and start the server by running the command copied from above.Create the characters
Create a second project with the game characters using the LLMAgent script as described above.
Enable the Remote option and configure the host with the IP address (starting with "http://") and port / API key of the server.
The Embeddings function can be used to obtain the emdeddings of a phrase:
List<float> embeddings = await llmAgent.Embeddings("hi, how are you?");
A detailed documentation on function level can be found here:
LLM for Unity implements a super-fast similarity search functionality with a Retrieval-Augmented Generation (RAG) system.
It is based on the LLM embeddings, and the Approximate Nearest Neighbors (ANN) search from the usearch library.
Semantic search works as follows.
Building the data You provide text inputs (a phrase, paragraph, document) to add to the data.
Each input is split into chunks (optional) and encoded into embeddings with a LLM.
Searching You can then search for a query text input.
The input is again encoded and the most similar text inputs or chunks in the data are retrieved.
To use semantic serch:
Add Component and select the RAG script.SimpleSearch is a simple brute-force search, whileDBSearch is a fast ANN method that should be preferred in most cases.Alternatively, you can create the RAG from code (where llm is your LLM):
RAG rag = gameObject.AddComponent<RAG>();
rag.Init(SearchMethods.DBSearch, ChunkingMethods.SentenceSplitter, llm);
In your script you can then use it as follows :unicorn::
using LLMUnity;
public class MyScript : MonoBehaviour
{
RAG rag;
async void Game(){
...
string[] inputs = new string[]{
"Hi! I'm a search system.",
"the weather is nice. I like it.",
"I'm a RAG system"
};
// add the inputs to the RAG
foreach (string input in inputs) await rag.Add(input);
// get the 2 most similar inputs and their distance (dissimilarity) to the search query
(string[] results, float[] distances) = await rag.Search("hello!", 2);
// to get the most similar text parts (chunks), instead of full input, you can enable the returnChunks option
rag.ReturnChunks(true);
(results, distances) = await rag.Search("hello!", 2);
...
}
}
You can also add / search text inputs for groups of data e.g. for a specific character or scene:
// add the inputs to the RAG for a group of data e.g. an orc character
foreach (string input in inputs) await rag.Add(input, "orc");
// get the 2 most similar inputs for the group of data e.g. the orc character
(string[] results, float[] distances) = await rag.Search("how do you feel?", 2, "orc");
...
You can save the RAG state (stored in the `Assets/StreamingAssets` folder):
``` c#
rag.Save("rag.zip");
and load it from disk:
await rag.Load("rag.zip");
You can use the RAG to feed relevant data to the LLM based on a user message:
string message = "How is the weather?";
(string[] similarPhrases, float[] distances) = await rag.Search(message, 3);
string prompt = "Answer the user query based on the provided data.\n\n";
prompt += $"User query: {message}\n\n";
prompt += $"Data:\n";
foreach (string similarPhrase in similarPhrases) prompt += $"\n- {similarPhrase}";
_ = llmAgent.Chat(prompt, HandleReply, ReplyCompleted);
The RAG sample includes an example RAG implementation as well as an example RAG-LLM integration.
That's all :sparkles:!
LLM for Unity includes a built-in model manager for easy model handling.
The model manager allows to load or download LLMs and can be found as part of the LLM GameObject:

You can download models with the Download model button.
LLM for Unity includes different state of the art models built-in for different model sizes, quantised with the Q4_K_M method.
Alternative models can be downloaded from HuggingFace in .gguf format.
You can download a model locally and load it with the Load model button, or copy the URL in the Download model > Custom URL field to directly download it.
If a HuggingFace model does not provide a gguf file, it can be converted to gguf with this online converter.
You can create lighter builds by selecting the Download on Build option.
The models will be downloaded the first time the game starts instead of bundled in the build.
If you have loaded a model locally you need to set its URL through the expanded view, otherwise it will be copied in the build.
❕ Before using any model make sure you check their license ❕
The Samples~ folder contains several examples of interaction 🤖:
To install a sample:
Window > Package ManagerLLM for Unity Package. From the Samples Tab, click Import next to the sample you want to install.The samples can be run with the Scene.unity scene they contain inside their folder.
In the scene, select the LLM GameObject and specify the LLM of your choice (see LLM model management).
Save the scene, run and enjoy!
The license of LLM for Unity is Apache 2.0 (LICENSE.md) and uses third-party software with MIT and Apache licenses. Some models included in the asset define their own license terms, please review them before using each model. Third-party licenses can be found in the (Third Party Notices.md).
12,069 followers · starred Mar 2024
216 followers · starred Aug 2026
3 followers · starred Apr 2025
250 followers · starred May 2026
C#
100.0%