DeepUnity is an add-on framework that provides tensor computation [with GPU acceleration support] and deep neural networks, along with reinforcement learning tools that enable training for intelligent agents within Unity environments using state-of-the-art algorithms.
using UnityEngine;
using DeepUnity;
using DeepUnity.Optimizers;
using DeepUnity.Activations;
using DeepUnity.Modules;
using DeepUnity.Models;
public class Tutorial : MonoBehaviour
{
[SerializeField] private Sequential network;
private Optimizer optim;
private Tensor x;
private Tensor y;
public void Start()
{
network = new Sequential(
new Dense(512, 256, device: Device.GPU),
new Swish(),
new Dropout(0.1f),
new Dense(256, 64, device: Device.GPU),
new Swish(),
new RMSNorm(64),
new Dense(64, 32)).CreateAsset("TutorialModel");
optim = new AdamW(network.Parameters());
x = Tensor.RandomNormal(64, 512);
y = Tensor.RandomNormal(64, 32);
}
public void Update()
{
Tensor yHat = network.Forward(x);
Loss loss = Loss.MSE(yHat, y);
optim.ZeroGrad();
network.Backward(loss.Grad);
optim.Step();
print($"Epoch: {Time.frameCount} - Train Loss: {loss.Item}");
network.Save();
}
}
DeepUnity provides full-GPU inference (HLSL compute shaders, streaming weight upload, KV caching) for the following LLMs, along with SFT scripts w/ QLoRA:
| Model | Sizes | Quantization |
|---|---|---|
| Qwen3.5 | 0.8B, 2B | FP16, INT8, INT4 |
| Gemma3 | 270M | FP16, INT8, INT4 |
| MiniCPM5 | 1B | FP16, INT8, INT4 |
All weights are exported into the engine's own format with import_params.py (--quant fp16|int8|int4) into Assets/Resources/Weights/. Every ported model registers itself in LLMRegistry, so it automatically shows up in the NPC inspector dropdowns.
using UnityEngine;
using DeepUnity;
using UnityEngine.UI;
public class ChatWithQwen : MonoBehaviour
{
[SerializeField] private Text display;
[SerializeField] private Text uiText;
[SerializeField] private float temperature = 0.7f;
[SerializeField] private float top_p = 0.95f;
private Qwen3_5ForCausalLM model;
public void Start()
{
// INT8 weights: ~half the VRAM/disk of FP16, ~lossless quality. Export with
// `python import_params.py Qwen/Qwen3.5-0.8B --quant int8`.
model = new Qwen3_5ForCausalLM(Qwen3_5Size.B0_8, LLMQuant.INT8);
StartCoroutine(model.Generate(
prompt: "Explain quantum entanglement in just 5 words.",
onTokenGenerated: (x) => { display.text += x; },
temperature: temperature,
top_p: top_p));
}
}
--quant fp16|int8|int4).DeepUnity runs TTS fully on the GPU as well, with real-time streaming synthesis (speech starts while the reply is still generating) and per-NPC voices:
| Model | Size | Quantization | Notes |
|---|---|---|---|
| pocket-tts (default) | 100M | FP16, INT8 | autoregressive flow-matching + Mimi codec, RTF ~0.15 on an RTX 4060, ~100 ms to first audio; runtime voice cloning from a 10 s reference clip (one-click precompute + disk cache — cloned voices load instantly, in builds too) |
| Kokoro-82M | 82M | FP16, INT8 | non-autoregressive, RTF ~0.3 on an RTX 4060 — real-time with headroom; 15 voicepacks + blends |
| Fun-CosyVoice3 | 0.5B | FP16, INT8 | autoregressive LM + causal DiT flow, token-level streaming, voices baked offline from any reference clip |
| Chatterbox-Turbo | 0.5B | FP16, INT8 | offline-style synthesis with clause-level streaming |
Speech-to-text runs fully on the GPU too: Qwen3-ASR 0.6B/1.7B (RTF 0.22-0.66) and Parakeet-TDT 0.6B v2/v3 (RTF ~0.08), both validated transcript-exact against their reference implementations.
using UnityEngine;
using DeepUnity;
// Attach next to an AudioSource. The engine is shared across all voices in the
// scene; weights stream to the GPU budgeted, without frame drops.
public class TalkingNpc : MonoBehaviour
{
[SerializeField] private PocketTTSVoice voice; // voiceName, pitch set in the inspector
[SerializeField] private AudioClip referenceClip; // ~10 s clip to clone a voice at runtime
public void Start()
{
// one-shot utterance
voice.Say("Welcome, traveler! What brings you to the village?");
// or clone a voice from a reference clip instead of a baked voiceName
// voice.SetClonedVoice(referenceClip);
}
// ...or stream an LLM reply as it generates: feed token deltas and flush at the end.
// Each completed sentence is synthesized and played while the next one is still being written.
public void OnLlmToken(string delta) => voice.FeedText(delta);
public void OnLlmDone() => voice.FlushText();
}
DeepUnity/TTS/Build VoiceLab Scene): pick an engine + voicepack, tune pitch/speed live and save presets.NPCChatBase is a drop-in MonoBehaviour that turns the LLM + TTS stack into a talking game character — everything is set up in the inspector (persona, model, voice), the demo subclasses (NPCInteractor3D/NPCInteractor2D) only add presentation.
Features:
<think> reasoning with an animated Thinking… indicator.AudioClip, audio-synced talk animation (or text-only mode).HISTORY: block in the system prompt) that lets a chat continue indefinitely.StartInteraction()/AskNPC()/CloseInteraction() API.In order to work with Reinforcement Learning tools, you must create a 2D or 3D agent using Unity provided GameObjects and Components. The setup flow works similary to ML Agents, so you must create a new behaviour script (e.g. ReachGoal) that must inherit the Agent class. Attach the new behaviour script to the agent GameObject (automatically DecisionRequester script is attached too) [Optionally, a TrainingStatistics script can be attached]. Choose the space size and number of continuous/discrete actions, then override the following methods in the behavior script:
Also in order to decide the reward function and episode's terminal state, use the following calls:
When the setup is ready, press the Bake button; a behaviour along with all neural networks and hyperparameters assets are created inside a folder with the behaviour's name, located in Assets/ folder. From this point everything is ready to go.
To get into advanced training, check out the following assets created:
using UnityEngine;
using DeepUnity.ReinforcementLearning;
public class MoveToGoal : Agent
{
public Transform apple;
public override void OnEpisodeBegin()
{
float xrand = Random.Range(-8, 8);
float zrand = Random.Range(-8, 8);
apple.localPosition = new Vector3(xrand, 2.25f, zrand);
xrand = Random.Range(-8, 8);
zrand = Random.Range(-8, 8);
transform.localPosition = new Vector3(xrand, 2.25f, zrand);
}
public override void CollectObservations(StateVector sensorBuffer)
{
sensorBuffer.AddObservation(transform.localPosition.x);
sensorBuffer.AddObservation(transform.localPosition.z);
sensorBuffer.AddObservation(apple.transform.localPosition.x);
sensorBuffer.AddObservation(apple.transform.localPosition.z);
}
public override void OnActionReceived(ActionBuffer actionBuffer)
{
float xmov = actionBuffer.ContinuousActions[0];
float zmov = actionBuffer.ContinuousActions[1];
transform.position += new Vector3(xmov, 0, zmov) * Time.fixedDeltaTime * 10f;
AddReward(-0.0025f); // Step penalty
}
private void OnTriggerEnter(Collider other)
{
if (other.CompareTag("Apple"))
{
SetReward(1f);
EndEpisode();
}
if (other.CompareTag("Wall"))
{
SetReward(-1f);
EndEpisode();
}
}
}

Parallel training is one option to use your device at maximum efficiency. After inserting your agent inside an Environment GameObject, you can duplicate that environment several times along the scene before starting the training session; this method is necessary for multi-agent co-op or adversarial training. Note that DeepUnity dynamically adapts the timescale of the simulation to get the maximum efficiency out of your machine.
In order to properly get use of AddReward() and EndEpisode() consult the diagram below. These methods work well being called inside OnTriggerXXX() or OnCollisionXXX(), as well as inside OnActionReceived() rightafter actions are performed.
Decision Period high values increases overall performance of the training session, but lacks when it comes to agent inference accuracy. Typically, use a higher value for broader parallel environments, then decrease this value to 1 to fine-tune the agent.
Input Normalization plays a huge role in policy convergence. To outcome this problem, observations can be auto-normalized by checking the corresponding box inside behaviour asset, but instead, is highly recommended to manually normalize all input values before adding them to the SensorBuffer. Scalar values can be normalized within [0, 1] or [-1, 1] ranges by using the formula normalized_value = (value - min) / (max - min). Note that inputs are clipped for network scability (default [-5, 5]).
The following MonoBehaviour methods: Awake(), Start(), FixedUpdate(), Update() and LateUpdate() are virtual. If neccesary, in order to override them, call the their base each time, respecting the logic of the diagram below.
Training inside the Editor is a bit more cumbersome comparing to the built version. Building the application and open it up to start up the training enables faster inference, and the framework was adapted for this.
Whenever you want to stop the training, close the .exe file. The trained behavior is automatically saved and serialized in .json format on your desktop. Go back in Unity and check your behavior asset, and press on the newly button to overwrite the editor behavior with the trained weights from .json.
The previous built app, along with the trained weights in .json format are now disposable (remove them and replace the build with a new one).



All tutorial's scripts are included inside Assets/DeepUnity/Tutorials folder, containing all the features provided by the framework and RL environments inspired from ML-Agents examples (note that not all of them have trained models attached).
DeepUnity is an add-on framework that provides tensor computation [with GPU acceleration support] and deep neural networks, along with reinforcement learning tools that enable training for intelligent agents within Unity environments using state-of-the-art algorithms.
using UnityEngine;
using DeepUnity;
using DeepUnity.Optimizers;
using DeepUnity.Activations;
using DeepUnity.Modules;
using DeepUnity.Models;
public class Tutorial : MonoBehaviour
{
[SerializeField] private Sequential network;
private Optimizer optim;
private Tensor x;
private Tensor y;
public void Start()
{
network = new Sequential(
new Dense(512, 256, device: Device.GPU),
new Swish(),
new Dropout(0.1f),
new Dense(256, 64, device: Device.GPU),
new Swish(),
new RMSNorm(64),
new Dense(64, 32)).CreateAsset("TutorialModel");
optim = new AdamW(network.Parameters());
x = Tensor.RandomNormal(64, 512);
y = Tensor.RandomNormal(64, 32);
}
public void Update()
{
Tensor yHat = network.Forward(x);
Loss loss = Loss.MSE(yHat, y);
optim.ZeroGrad();
network.Backward(loss.Grad);
optim.Step();
print($"Epoch: {Time.frameCount} - Train Loss: {loss.Item}");
network.Save();
}
}
DeepUnity provides full-GPU inference (HLSL compute shaders, streaming weight upload, KV caching) for the following LLMs, along with SFT scripts w/ QLoRA:
| Model | Sizes | Quantization |
|---|---|---|
| Qwen3.5 | 0.8B, 2B | FP16, INT8, INT4 |
| Gemma3 | 270M | FP16, INT8, INT4 |
| MiniCPM5 | 1B | FP16, INT8, INT4 |
All weights are exported into the engine's own format with import_params.py (--quant fp16|int8|int4) into Assets/Resources/Weights/. Every ported model registers itself in LLMRegistry, so it automatically shows up in the NPC inspector dropdowns.
using UnityEngine;
using DeepUnity;
using UnityEngine.UI;
public class ChatWithQwen : MonoBehaviour
{
[SerializeField] private Text display;
[SerializeField] private Text uiText;
[SerializeField] private float temperature = 0.7f;
[SerializeField] private float top_p = 0.95f;
private Qwen3_5ForCausalLM model;
public void Start()
{
// INT8 weights: ~half the VRAM/disk of FP16, ~lossless quality. Export with
// `python import_params.py Qwen/Qwen3.5-0.8B --quant int8`.
model = new Qwen3_5ForCausalLM(Qwen3_5Size.B0_8, LLMQuant.INT8);
StartCoroutine(model.Generate(
prompt: "Explain quantum entanglement in just 5 words.",
onTokenGenerated: (x) => { display.text += x; },
temperature: temperature,
top_p: top_p));
}
}
--quant fp16|int8|int4).DeepUnity runs TTS fully on the GPU as well, with real-time streaming synthesis (speech starts while the reply is still generating) and per-NPC voices:
| Model | Size | Quantization | Notes |
|---|---|---|---|
| pocket-tts (default) | 100M | FP16, INT8 | autoregressive flow-matching + Mimi codec, RTF ~0.15 on an RTX 4060, ~100 ms to first audio; runtime voice cloning from a 10 s reference clip (one-click precompute + disk cache — cloned voices load instantly, in builds too) |
| Kokoro-82M | 82M | FP16, INT8 | non-autoregressive, RTF ~0.3 on an RTX 4060 — real-time with headroom; 15 voicepacks + blends |
| Fun-CosyVoice3 | 0.5B | FP16, INT8 | autoregressive LM + causal DiT flow, token-level streaming, voices baked offline from any reference clip |
| Chatterbox-Turbo | 0.5B | FP16, INT8 | offline-style synthesis with clause-level streaming |
Speech-to-text runs fully on the GPU too: Qwen3-ASR 0.6B/1.7B (RTF 0.22-0.66) and Parakeet-TDT 0.6B v2/v3 (RTF ~0.08), both validated transcript-exact against their reference implementations.
using UnityEngine;
using DeepUnity;
// Attach next to an AudioSource. The engine is shared across all voices in the
// scene; weights stream to the GPU budgeted, without frame drops.
public class TalkingNpc : MonoBehaviour
{
[SerializeField] private PocketTTSVoice voice; // voiceName, pitch set in the inspector
[SerializeField] private AudioClip referenceClip; // ~10 s clip to clone a voice at runtime
public void Start()
{
// one-shot utterance
voice.Say("Welcome, traveler! What brings you to the village?");
// or clone a voice from a reference clip instead of a baked voiceName
// voice.SetClonedVoice(referenceClip);
}
// ...or stream an LLM reply as it generates: feed token deltas and flush at the end.
// Each completed sentence is synthesized and played while the next one is still being written.
public void OnLlmToken(string delta) => voice.FeedText(delta);
public void OnLlmDone() => voice.FlushText();
}
DeepUnity/TTS/Build VoiceLab Scene): pick an engine + voicepack, tune pitch/speed live and save presets.NPCChatBase is a drop-in MonoBehaviour that turns the LLM + TTS stack into a talking game character — everything is set up in the inspector (persona, model, voice), the demo subclasses (NPCInteractor3D/NPCInteractor2D) only add presentation.
Features:
<think> reasoning with an animated Thinking… indicator.AudioClip, audio-synced talk animation (or text-only mode).HISTORY: block in the system prompt) that lets a chat continue indefinitely.StartInteraction()/AskNPC()/CloseInteraction() API.In order to work with Reinforcement Learning tools, you must create a 2D or 3D agent using Unity provided GameObjects and Components. The setup flow works similary to ML Agents, so you must create a new behaviour script (e.g. ReachGoal) that must inherit the Agent class. Attach the new behaviour script to the agent GameObject (automatically DecisionRequester script is attached too) [Optionally, a TrainingStatistics script can be attached]. Choose the space size and number of continuous/discrete actions, then override the following methods in the behavior script:
Also in order to decide the reward function and episode's terminal state, use the following calls:
When the setup is ready, press the Bake button; a behaviour along with all neural networks and hyperparameters assets are created inside a folder with the behaviour's name, located in Assets/ folder. From this point everything is ready to go.
To get into advanced training, check out the following assets created:
using UnityEngine;
using DeepUnity.ReinforcementLearning;
public class MoveToGoal : Agent
{
public Transform apple;
public override void OnEpisodeBegin()
{
float xrand = Random.Range(-8, 8);
float zrand = Random.Range(-8, 8);
apple.localPosition = new Vector3(xrand, 2.25f, zrand);
xrand = Random.Range(-8, 8);
zrand = Random.Range(-8, 8);
transform.localPosition = new Vector3(xrand, 2.25f, zrand);
}
public override void CollectObservations(StateVector sensorBuffer)
{
sensorBuffer.AddObservation(transform.localPosition.x);
sensorBuffer.AddObservation(transform.localPosition.z);
sensorBuffer.AddObservation(apple.transform.localPosition.x);
sensorBuffer.AddObservation(apple.transform.localPosition.z);
}
public override void OnActionReceived(ActionBuffer actionBuffer)
{
float xmov = actionBuffer.ContinuousActions[0];
float zmov = actionBuffer.ContinuousActions[1];
transform.position += new Vector3(xmov, 0, zmov) * Time.fixedDeltaTime * 10f;
AddReward(-0.0025f); // Step penalty
}
private void OnTriggerEnter(Collider other)
{
if (other.CompareTag("Apple"))
{
SetReward(1f);
EndEpisode();
}
if (other.CompareTag("Wall"))
{
SetReward(-1f);
EndEpisode();
}
}
}

Parallel training is one option to use your device at maximum efficiency. After inserting your agent inside an Environment GameObject, you can duplicate that environment several times along the scene before starting the training session; this method is necessary for multi-agent co-op or adversarial training. Note that DeepUnity dynamically adapts the timescale of the simulation to get the maximum efficiency out of your machine.
In order to properly get use of AddReward() and EndEpisode() consult the diagram below. These methods work well being called inside OnTriggerXXX() or OnCollisionXXX(), as well as inside OnActionReceived() rightafter actions are performed.
Decision Period high values increases overall performance of the training session, but lacks when it comes to agent inference accuracy. Typically, use a higher value for broader parallel environments, then decrease this value to 1 to fine-tune the agent.
Input Normalization plays a huge role in policy convergence. To outcome this problem, observations can be auto-normalized by checking the corresponding box inside behaviour asset, but instead, is highly recommended to manually normalize all input values before adding them to the SensorBuffer. Scalar values can be normalized within [0, 1] or [-1, 1] ranges by using the formula normalized_value = (value - min) / (max - min). Note that inputs are clipped for network scability (default [-5, 5]).
The following MonoBehaviour methods: Awake(), Start(), FixedUpdate(), Update() and LateUpdate() are virtual. If neccesary, in order to override them, call the their base each time, respecting the logic of the diagram below.
Training inside the Editor is a bit more cumbersome comparing to the built version. Building the application and open it up to start up the training enables faster inference, and the framework was adapted for this.
Whenever you want to stop the training, close the .exe file. The trained behavior is automatically saved and serialized in .json format on your desktop. Go back in Unity and check your behavior asset, and press on the newly button to overwrite the editor behavior with the trained weights from .json.
The previous built app, along with the trained weights in .json format are now disposable (remove them and replace the build with a new one).



All tutorial's scripts are included inside Assets/DeepUnity/Tutorials folder, containing all the features provided by the framework and RL environments inspired from ML-Agents examples (note that not all of them have trained models attached).