Open-source Unity package for local LLM-driven NPC and game-system decisions, with configurable actions and playable samples.
See the code
Watch the demo (69 seconds) · Quick start · Samples · How it works · User guide · Changelog · Discussions
BehaviorLLM is an open-source Unity package that lets a game object decide for itself, using a local language model. You describe what it can do (a discrete action set) and what it can perceive (observation modules that turn game state into text); the package builds the prompt, constrains the model to a valid JSON decision, and dispatches the chosen action to your UnityEvents.
It is not only for characters. The same component drives an enemy, a companion, or a game system such as a director adjusting difficulty, which is why the component is called a Decision Maker rather than an agent.
observe IObservationModule.GetObservation() -> "--- Vision ---\n- [3.2m] ID: Intruder | Type: Hostile"
think system prompt (cached) + state block -> llama-server /v1/chat/completions + JSON Schema
act {"action":"Chase","arg":"Intruder"} -> actionBindings["Chase"].Invoke(args) // args.First == "Intruder"
llama-server over HTTP. No cloud, no API keys, no data leaving the machine.Deliberative profile; enemies and allies answer as fast as the model can write about 20 tokens.The historical September 4–10 baseline numbers below come from the harness and recorded runs in Experiments~/, on a mid-range PC (Ryzen 5 7600X, RX 6650 XT 8 GB) with 4-bit models:
| Structurally valid decisions | 192 of 192 across three models, two profiles, schema on and off |
| Expected action chosen | 62–94% with the schema enabled, on 16 labelled situations per configuration; the full matrix includes a 100% schema-disabled row |
| One decision, warm | 112 ms on Qwen3.5-2B, against 1,956 ms for the same decision cold |
| Eight characters on one model | 444 decisions in five minutes on Granite 4.1-3B, median 1,686 ms, no failed requests |
| Cost of the schema | none measurable: 189 ms with it against 184 ms without |
The folder holds the scripts, the exact prompts and schemas they send, the per-decision CSVs and three telemetry runs, so the matrix and live-run figures can be inspected. The cold/warm pair is a historical illustration without its original raw capture; run_coldwarm.py now supplies an explicit repeatable protocol. New validation runs live in dated subfolders and do not replace these baseline claims.
llama-server (the package can download it) and a GGUF model (the package can download one).com.unity.ai.navigation and com.unity.inputsystem.In Unity, open Window > Package Manager > Add package from git URL… and paste one of these URLs.
Release (stable, recommended):
https://github.com/faelion/BehaviorLLM.git#release
Dev (latest development on main; may include breaking changes):
https://github.com/faelion/BehaviorLLM.git#main
These URLs stay the same across versions. Unity locks the installed commit; to update, add the same Git URL again through Package Manager. The release branch advances when a stable GitHub release is published.
Optional: pin a specific release. Replace the branch suffix with a tag from Releases, for example:
https://github.com/faelion/BehaviorLLM.git#v0.6.0
Manual. Download the latest release and copy the folder into your project's Packages/ or Assets/ folder.
BehaviorLLM > Model Manager, pick one of the three models listed (Granite 4.1 3B is the best starting point for reactive characters; Qwen3.5 2B is faster, Qwen3.5 4B more accurate) and click Download. Then, in the Installed tab, click Set active in current scene: it assigns the model's config to the server and clients in the open scene and writes StreamingAssets/behaviorllm_backend_config.json. The Catalog also searches Hugging Face live, filtered to text-generation models whose architecture the bundled llama-server loads, with a size cap that keeps results runnable on one machine.BehaviorLLM > Model Manager (Server tab) and click Download Server, or install llama.cpp yourself (winget install llama.cpp, brew install llama.cpp); a llama-server on your PATH is picked up automatically.BehaviorLLMServer component to the scene and tick Auto Start On Awake on its server config, or run llama-server -m model.gguf --jinja -np 4 yourself.
DecisionMaker and BehaviorLLMClient. Both arrive already pointing at the presets the package ships, so they work with no further setup.ModularVisionModule (sees LLMContextObjects in a sphere or cone), BasicMemory (the last few executed actions), and for self-status an LLMContextObject plus SelfObservationModule on the object itself.ActionConfig asset (Create > BehaviorLLM > Action Config) and list the actions: name, description, parameter type, optional example and allowed arguments.On Execute (UnityEvent<ActionArguments>; args.First carries a single value, args["name"] a named one).LLMContextObject (name, type, description, live data bindings) and a collider on the same GameObject.
Open BehaviorLLM > Prompt Inspector. Every decision appears as a row with its latency and token counts; select one to see the reply, the state block that was rebuilt for it, the cached system prompt beneath, and the schema the reply was constrained to.

Nothing that tunes behaviour lives on the components. Four asset types hold it instead, and each can be shared by as many objects as you like:
| Asset | Holds | Read by |
|---|---|---|
| Decision Maker Config | Decision interval, profile, reason and thinking budgets, structured output, prompt budget, console noise | DecisionMaker |
| Server Config | Address, transport, port, slots, launch flags, timeouts, retry policy, server log verbosity | BehaviorLLMClient, BehaviorLLMServer |
| Model Config | Model file, context size, GPU layers, sampling, and what that model was measured to do well | BehaviorLLMClient, BehaviorLLMServer |
| Perception Config | Vision shape, range, field of view, layers, scan rate, sight interrupts, memory capacity | ModularVisionModule, BasicMemory |
Presets ship under Runtime/Defaults and are wired automatically when you add a component. Create your own with Create > BehaviorLLM > …, or duplicate a preset and edit it. To vary one object at runtime without affecting the others sharing an asset, clone it with Instantiate and call ApplyConfig.
| Profile | When | What it does |
|---|---|---|
Reactive (default) | Enemies, allies, anything that must respond within a beat | No thinking, no reason field, about a 50-token budget |
Deliberative | Managers, directors, planners that decide every few seconds | A short capped reason before the action, and optionally a native thinking budget |
Both profiles share one server; the choice travels with each request. Deliberation cost 50 to 90 percent more latency in the measurements above and improved schema-enabled expected-action matches from 13/16 to 15/16 on Qwen3.5-4B and 10/16 to 11/16 on Qwen3.5-2B, while Granite fell from 13/16 to 12/16. These small samples are not a general ranking; measure before enabling it. Each shipped model config records the profile it was measured to prefer, and the package warns at startup if you pick the other one.
Two runnable samples ship inside the package under Samples/. Each is self-contained: scenes, scripts, config assets and art in its own folder, referencing nothing outside itself and the package.
They need com.unity.ai.navigation and com.unity.inputsystem; the package core needs neither, and the sample assemblies are gated on those packages so a project without them still installs and works normally.
Open BehaviorLLM > Readme to delete either sample, the experiment data, or all of them plus the readme itself, once you have finished with them. Nothing in the package references them. Deleting in place needs the package to be writable, so the buttons work for a package copied into your Assets folder; installed from the git URL it is read-only, and the page says so and points at the manifest instead.
DecisionMaker runs the loop on the configured interval, or when a module reports an interrupt. An interrupt cancels any in-flight request, discards its answer, and starts a fresh decision. It builds the prompt with PromptBuilder, the schema with ActionSchemaBuilder, parses with DecisionParser, dispatches, and publishes a DecisionTelemetry record per decision.ActionConfig is the schema (what exists), actionBindings is the wiring (who handles it). The inspector keeps the two in step, so you never type an action name twice.ILLMBackend receives an LLMRequest (prompt sections, schema, slot, token budget, thinking policy) and returns an LLMResponse (text, reasoning, token counts, latency). BehaviorLLMClient is the llama-server implementation; any OpenAI-compatible local server should work.IActionAvailabilityProvider decides which actions may be chosen each turn, and the ones it withholds are dropped from the schema so the model cannot produce them. Rules written as prose were followed unreliably by every model measured; withholding the action worked.| BehaviorLLM | LLMUnity | Cloud-API prototypes | Behaviour trees | |
|---|---|---|---|---|
| Runs | local only | local, remote optional | remote | local |
| Built for | decisions: which action, with which argument | conversation and RAG | one game each | decisions |
| Output | schema-constrained during generation | free text (grammar optional) | free text, validated after | deterministic |
| Package dependencies | none | bundles its own inference library | game-specific | none |
| Authoring | an action set asset and observation components | prompts and chat components | prompts and parsers | hand-built trees |
BehaviorLLM is a complement to behaviour trees and state machines, not a replacement: use it where the decision is genuinely open, and keep the deterministic parts deterministic.
allowedArguments on an action for values that never change, or add a component implementing IArgumentOptionsProvider for scene-derived ones (routes, nearby characters). Provider values are printed in the state block, never in the cached half. An optional IActionArgumentPolicy component can still normalise or reject arguments after parsing.IActionAvailabilityProvider returns the actions legal this turn; the schema is rebuilt from it every decision, and the choice is rechecked at dispatch so a reply that went stale in flight is diverted to the fallback.RebuildBindingLookup(). Changing action definitions: call RebuildSchema(). Swapping a config asset on a live component: call ApplyConfig(...).DecisionTelemetryRecorder to a scene and every decision appends a JSON record (source, latency, prompt and cached token counts, action, outcome) under the persistent data folder, with a per-run summary CSV. The multi-character numbers above come from it; the runs themselves are in Experiments~/telemetry/.BehaviorLLMSettings asset in Resources/ sets the log level in the Editor and in a build, whether telemetry records at all, and whether applied schemas are dumped for reproduction. It is a ceiling, never an override: it can silence a decision maker, never make one louder.reason field are set per request, so one server serves reactive guards and a deliberative director at once.Documentation~/USER_GUIDE.md, step-by-step setup and troubleshootingDocumentation~/BehaviorLLM_Architecture_Analysis.md, design and data flowSamples/README.md, what each sample demonstratesExperiments~/README.md, how the numbers above were measured and how to re-run themCHANGELOG.mdCONTRIBUTING.md for the checklist (tooltips on every field, no Debug.Log in Runtime/, changelog entry for anything a consumer sees).MIT. The sample art is CC0 and credited beside it: KayKit in PrisonYard and Kenney in StealthGuard.
C#
96.7%
Python
3.3%
Open-source Unity package for local LLM-driven NPC and game-system decisions, with configurable actions and playable samples.
See the code
Watch the demo (69 seconds) · Quick start · Samples · How it works · User guide · Changelog · Discussions
BehaviorLLM is an open-source Unity package that lets a game object decide for itself, using a local language model. You describe what it can do (a discrete action set) and what it can perceive (observation modules that turn game state into text); the package builds the prompt, constrains the model to a valid JSON decision, and dispatches the chosen action to your UnityEvents.
It is not only for characters. The same component drives an enemy, a companion, or a game system such as a director adjusting difficulty, which is why the component is called a Decision Maker rather than an agent.
observe IObservationModule.GetObservation() -> "--- Vision ---\n- [3.2m] ID: Intruder | Type: Hostile"
think system prompt (cached) + state block -> llama-server /v1/chat/completions + JSON Schema
act {"action":"Chase","arg":"Intruder"} -> actionBindings["Chase"].Invoke(args) // args.First == "Intruder"
llama-server over HTTP. No cloud, no API keys, no data leaving the machine.Deliberative profile; enemies and allies answer as fast as the model can write about 20 tokens.The historical September 4–10 baseline numbers below come from the harness and recorded runs in Experiments~/, on a mid-range PC (Ryzen 5 7600X, RX 6650 XT 8 GB) with 4-bit models:
| Structurally valid decisions | 192 of 192 across three models, two profiles, schema on and off |
| Expected action chosen | 62–94% with the schema enabled, on 16 labelled situations per configuration; the full matrix includes a 100% schema-disabled row |
| One decision, warm | 112 ms on Qwen3.5-2B, against 1,956 ms for the same decision cold |
| Eight characters on one model | 444 decisions in five minutes on Granite 4.1-3B, median 1,686 ms, no failed requests |
| Cost of the schema | none measurable: 189 ms with it against 184 ms without |
The folder holds the scripts, the exact prompts and schemas they send, the per-decision CSVs and three telemetry runs, so the matrix and live-run figures can be inspected. The cold/warm pair is a historical illustration without its original raw capture; run_coldwarm.py now supplies an explicit repeatable protocol. New validation runs live in dated subfolders and do not replace these baseline claims.
llama-server (the package can download it) and a GGUF model (the package can download one).com.unity.ai.navigation and com.unity.inputsystem.In Unity, open Window > Package Manager > Add package from git URL… and paste one of these URLs.
Release (stable, recommended):
https://github.com/faelion/BehaviorLLM.git#release
Dev (latest development on main; may include breaking changes):
https://github.com/faelion/BehaviorLLM.git#main
These URLs stay the same across versions. Unity locks the installed commit; to update, add the same Git URL again through Package Manager. The release branch advances when a stable GitHub release is published.
Optional: pin a specific release. Replace the branch suffix with a tag from Releases, for example:
https://github.com/faelion/BehaviorLLM.git#v0.6.0
Manual. Download the latest release and copy the folder into your project's Packages/ or Assets/ folder.
BehaviorLLM > Model Manager, pick one of the three models listed (Granite 4.1 3B is the best starting point for reactive characters; Qwen3.5 2B is faster, Qwen3.5 4B more accurate) and click Download. Then, in the Installed tab, click Set active in current scene: it assigns the model's config to the server and clients in the open scene and writes StreamingAssets/behaviorllm_backend_config.json. The Catalog also searches Hugging Face live, filtered to text-generation models whose architecture the bundled llama-server loads, with a size cap that keeps results runnable on one machine.BehaviorLLM > Model Manager (Server tab) and click Download Server, or install llama.cpp yourself (winget install llama.cpp, brew install llama.cpp); a llama-server on your PATH is picked up automatically.BehaviorLLMServer component to the scene and tick Auto Start On Awake on its server config, or run llama-server -m model.gguf --jinja -np 4 yourself.
DecisionMaker and BehaviorLLMClient. Both arrive already pointing at the presets the package ships, so they work with no further setup.ModularVisionModule (sees LLMContextObjects in a sphere or cone), BasicMemory (the last few executed actions), and for self-status an LLMContextObject plus SelfObservationModule on the object itself.ActionConfig asset (Create > BehaviorLLM > Action Config) and list the actions: name, description, parameter type, optional example and allowed arguments.On Execute (UnityEvent<ActionArguments>; args.First carries a single value, args["name"] a named one).LLMContextObject (name, type, description, live data bindings) and a collider on the same GameObject.
Open BehaviorLLM > Prompt Inspector. Every decision appears as a row with its latency and token counts; select one to see the reply, the state block that was rebuilt for it, the cached system prompt beneath, and the schema the reply was constrained to.

Nothing that tunes behaviour lives on the components. Four asset types hold it instead, and each can be shared by as many objects as you like:
| Asset | Holds | Read by |
|---|---|---|
| Decision Maker Config | Decision interval, profile, reason and thinking budgets, structured output, prompt budget, console noise | DecisionMaker |
| Server Config | Address, transport, port, slots, launch flags, timeouts, retry policy, server log verbosity | BehaviorLLMClient, BehaviorLLMServer |
| Model Config | Model file, context size, GPU layers, sampling, and what that model was measured to do well | BehaviorLLMClient, BehaviorLLMServer |
| Perception Config | Vision shape, range, field of view, layers, scan rate, sight interrupts, memory capacity | ModularVisionModule, BasicMemory |
Presets ship under Runtime/Defaults and are wired automatically when you add a component. Create your own with Create > BehaviorLLM > …, or duplicate a preset and edit it. To vary one object at runtime without affecting the others sharing an asset, clone it with Instantiate and call ApplyConfig.
| Profile | When | What it does |
|---|---|---|
Reactive (default) | Enemies, allies, anything that must respond within a beat | No thinking, no reason field, about a 50-token budget |
Deliberative | Managers, directors, planners that decide every few seconds | A short capped reason before the action, and optionally a native thinking budget |
Both profiles share one server; the choice travels with each request. Deliberation cost 50 to 90 percent more latency in the measurements above and improved schema-enabled expected-action matches from 13/16 to 15/16 on Qwen3.5-4B and 10/16 to 11/16 on Qwen3.5-2B, while Granite fell from 13/16 to 12/16. These small samples are not a general ranking; measure before enabling it. Each shipped model config records the profile it was measured to prefer, and the package warns at startup if you pick the other one.
Two runnable samples ship inside the package under Samples/. Each is self-contained: scenes, scripts, config assets and art in its own folder, referencing nothing outside itself and the package.
They need com.unity.ai.navigation and com.unity.inputsystem; the package core needs neither, and the sample assemblies are gated on those packages so a project without them still installs and works normally.
Open BehaviorLLM > Readme to delete either sample, the experiment data, or all of them plus the readme itself, once you have finished with them. Nothing in the package references them. Deleting in place needs the package to be writable, so the buttons work for a package copied into your Assets folder; installed from the git URL it is read-only, and the page says so and points at the manifest instead.
DecisionMaker runs the loop on the configured interval, or when a module reports an interrupt. An interrupt cancels any in-flight request, discards its answer, and starts a fresh decision. It builds the prompt with PromptBuilder, the schema with ActionSchemaBuilder, parses with DecisionParser, dispatches, and publishes a DecisionTelemetry record per decision.ActionConfig is the schema (what exists), actionBindings is the wiring (who handles it). The inspector keeps the two in step, so you never type an action name twice.ILLMBackend receives an LLMRequest (prompt sections, schema, slot, token budget, thinking policy) and returns an LLMResponse (text, reasoning, token counts, latency). BehaviorLLMClient is the llama-server implementation; any OpenAI-compatible local server should work.IActionAvailabilityProvider decides which actions may be chosen each turn, and the ones it withholds are dropped from the schema so the model cannot produce them. Rules written as prose were followed unreliably by every model measured; withholding the action worked.| BehaviorLLM | LLMUnity | Cloud-API prototypes | Behaviour trees | |
|---|---|---|---|---|
| Runs | local only | local, remote optional | remote | local |
| Built for | decisions: which action, with which argument | conversation and RAG | one game each | decisions |
| Output | schema-constrained during generation | free text (grammar optional) | free text, validated after | deterministic |
| Package dependencies | none | bundles its own inference library | game-specific | none |
| Authoring | an action set asset and observation components | prompts and chat components | prompts and parsers | hand-built trees |
BehaviorLLM is a complement to behaviour trees and state machines, not a replacement: use it where the decision is genuinely open, and keep the deterministic parts deterministic.
allowedArguments on an action for values that never change, or add a component implementing IArgumentOptionsProvider for scene-derived ones (routes, nearby characters). Provider values are printed in the state block, never in the cached half. An optional IActionArgumentPolicy component can still normalise or reject arguments after parsing.IActionAvailabilityProvider returns the actions legal this turn; the schema is rebuilt from it every decision, and the choice is rechecked at dispatch so a reply that went stale in flight is diverted to the fallback.RebuildBindingLookup(). Changing action definitions: call RebuildSchema(). Swapping a config asset on a live component: call ApplyConfig(...).DecisionTelemetryRecorder to a scene and every decision appends a JSON record (source, latency, prompt and cached token counts, action, outcome) under the persistent data folder, with a per-run summary CSV. The multi-character numbers above come from it; the runs themselves are in Experiments~/telemetry/.BehaviorLLMSettings asset in Resources/ sets the log level in the Editor and in a build, whether telemetry records at all, and whether applied schemas are dumped for reproduction. It is a ceiling, never an override: it can silence a decision maker, never make one louder.reason field are set per request, so one server serves reactive guards and a deliberative director at once.Documentation~/USER_GUIDE.md, step-by-step setup and troubleshootingDocumentation~/BehaviorLLM_Architecture_Analysis.md, design and data flowSamples/README.md, what each sample demonstratesExperiments~/README.md, how the numbers above were measured and how to re-run themCHANGELOG.mdCONTRIBUTING.md for the checklist (tooltips on every field, no Debug.Log in Runtime/, changelog entry for anything a consumer sees).MIT. The sample art is CC0 and credited beside it: KayKit in PrisonYard and Kenney in StealthGuard.
C#
96.7%
Python
3.3%