README with repository metadata
Browse files
README.md
ADDED
|
@@ -0,0 +1,183 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
tags:
|
| 4 |
+
- unity
|
| 5 |
+
- upm
|
| 6 |
+
- unity-package
|
| 7 |
+
- xr
|
| 8 |
+
- virtual-reality
|
| 9 |
+
- llama.cpp
|
| 10 |
+
- offline
|
| 11 |
+
- voice-control
|
| 12 |
+
- agent
|
| 13 |
+
---
|
| 14 |
+
|
| 15 |
+
# Maverick XR Agent for Unity
|
| 16 |
+
|
| 17 |
+
The Unity package for [Maverick 4B](https://huggingface.co/ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF).
|
| 18 |
+
|
| 19 |
+
Let people control your Unity scene in plain English, typed or spoken, with a small language model that runs **on
|
| 20 |
+
the same machine, offline**. "Put the red mug on the table", "turn on the lamp", "where is the stapler?": the model
|
| 21 |
+
reads the objects you marked as interactive, calls the right tool, and Unity executes it. When the words fit two
|
| 22 |
+
objects it asks which one; when an object is in another room it searches there; when a request is impossible it says
|
| 23 |
+
so instead of acting.
|
| 24 |
+
|
| 25 |
+
> **Status: release candidate (1.0.0-rc.2).** Tested on Windows 11 x64, Unity 2022.3 LTS, NVIDIA RTX
|
| 26 |
+
> 3050 Ti Laptop (4 GB). See [Limitations](#limitations) before building on it.
|
| 27 |
+
|
| 28 |
+
- **11 tools:** grab, release, place, move_by, rotate, set_state, press, highlight, teleport, find_objects,
|
| 29 |
+
ask_clarification.
|
| 30 |
+
- **Understands** synonyms and general words ("the fruit"), colours, spatial references ("the one next to the chair",
|
| 31 |
+
"the second from the left"), "this / that" with the user's gaze, the object in hand ("put it on the shelf"), filler
|
| 32 |
+
words and speech-recognition noise.
|
| 33 |
+
- **Safe by default:**
|
| 34 |
+
- D5b: a call to a tool or a value that does not exist is turned into a refusal before it reaches your scene; an
|
| 35 |
+
unknown object id is generated again under a grammar.
|
| 36 |
+
- Substitute guard: when the user names an object that is only in another room and the model tries to act on a
|
| 37 |
+
different object here, the call is not executed and the model is told where the named object is.
|
| 38 |
+
- **Private and free to run:** llama.cpp on `127.0.0.1`; no account, no cloud call, no per-request cost.
|
| 39 |
+
|
| 40 |
+
## Requirements
|
| 41 |
+
| | |
|
| 42 |
+
|---|---|
|
| 43 |
+
| Unity | 2022.3 LTS (tested 2022.3.62f3) |
|
| 44 |
+
| OS | Windows x64 (tested). llama.cpp also runs on Linux and macOS; `LocalModelServer` is written for them but untested |
|
| 45 |
+
| GPU | ~3.1 GB free VRAM; CPU fallback works but is slow |
|
| 46 |
+
| Disk | 2.5 GB for the model, plus ~100 MB for llama.cpp |
|
| 47 |
+
|
| 48 |
+
## Install
|
| 49 |
+
1. **Package:** Window → Package Manager → **+** → *Add package from git URL…* →
|
| 50 |
+
`https://huggingface.co/ErenAta00/Maverick-Unity.git` (or *Add package from disk…* → this folder's `package.json`).
|
| 51 |
+
2. **Model files:** Tools → Scene Agent → **Open Model Folder** (creates `Assets/StreamingAssets/SceneAgent/`) and put
|
| 52 |
+
there:
|
| 53 |
+
- `llama-server.exe` and the DLLs next to it, from a llama.cpp release for Windows (CUDA 12 build for NVIDIA GPUs,
|
| 54 |
+
Vulkan build for other GPUs, CPU build otherwise). Copy the whole folder's contents, not only the .exe.
|
| 55 |
+
- The model: `Maverick-4B-Unity-XR-Agent-Q4_K_M.gguf` (2.5 GB) from
|
| 56 |
+
https://huggingface.co/ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF.
|
| 57 |
+
The server takes the first `.gguf` it finds there, or set `LocalModelServer.modelPath`.
|
| 58 |
+
|
| 59 |
+
StreamingAssets is copied into a build as it is, so a built game runs without installing anything.
|
| 60 |
+
|
| 61 |
+
## Quick start (5 minutes)
|
| 62 |
+
1. **Tools → Scene Agent → Set Up Scene.** Adds a *Scene Agent* object with everything wired: `SceneWorld` (your
|
| 63 |
+
objects, as the model sees them), `SceneAgent` (the conversation with the model), `LocalModelServer` (starts
|
| 64 |
+
llama.cpp), `SceneAgentSession` and `SceneAgentChatBox` (an on-screen box to try it).
|
| 65 |
+
2. **Build or open your scene.** Name objects with plain words: *Red Mug*, *Table*, *Floor Lamp*, *Book*.
|
| 66 |
+
3. **Select the objects** the agent may use → **Tools → Scene Agent → Make Selected Interactive.** Each gets a
|
| 67 |
+
`SceneItem`: an id (`mug_1`), a label (`mug`) and a colour (`red`) from its name, a room, and what it stands on (from
|
| 68 |
+
the renderers' bounds). They move under the *Scene Agent* object.
|
| 69 |
+
4. **Tools → Scene Agent → Validate Scene.** Fix what it reports (duplicate ids, labels the model does not know,
|
| 70 |
+
missing model files).
|
| 71 |
+
5. **Play.** Type `put the red mug on the table` and press Enter. When the agent asks "Which mug do you mean?", click
|
| 72 |
+
one. F1 hides the box.
|
| 73 |
+
|
| 74 |
+
## Use it from your code
|
| 75 |
+
```csharp
|
| 76 |
+
using SceneAgent;
|
| 77 |
+
|
| 78 |
+
public class MyVoiceUI : MonoBehaviour
|
| 79 |
+
{
|
| 80 |
+
public SceneAgentSession session; // on the Scene Agent object
|
| 81 |
+
|
| 82 |
+
void Start()
|
| 83 |
+
{
|
| 84 |
+
session.Line += text => Debug.Log(text); // "You: …", "→ call result", "Agent: …"
|
| 85 |
+
session.Finished += r => { if (r.clarificationQuestion != null) ShowChoices(session.Candidates); };
|
| 86 |
+
session.world.ToolExecuted += (tool, args, result) =>
|
| 87 |
+
{
|
| 88 |
+
if (tool == "teleport") MoveCameraTo(session.world.playerRoom); // react in your own way
|
| 89 |
+
};
|
| 90 |
+
}
|
| 91 |
+
|
| 92 |
+
public void OnSpeech(string text) => StartCoroutine(session.Command(text)); // typed or recognised speech
|
| 93 |
+
public void OnPick(string itemId) => StartCoroutine(session.Answer(itemId)); // the user's answer to "which one?"
|
| 94 |
+
|
| 95 |
+
void ShowChoices(System.Collections.Generic.List<string> ids) { /* your UI */ }
|
| 96 |
+
void MoveCameraTo(string room) { /* your rig */ }
|
| 97 |
+
}
|
| 98 |
+
```
|
| 99 |
+
- **Lower level:** `StartCoroutine(agent.Run(text, result => …))` returns an `AgentResult`: the calls, their results,
|
| 100 |
+
the final text or the clarification question, an error, and the time.
|
| 101 |
+
- **Rooms:** give each `SceneItem` a `room`. The model sees the player's room only; it finds objects elsewhere with
|
| 102 |
+
`find_objects` and moves the player with `teleport` (which sets `SceneWorld.playerRoom`; you move the camera in
|
| 103 |
+
`ToolExecuted`).
|
| 104 |
+
- **Gaze:** `SceneAgentGaze` (added by Set Up Scene) sets `SceneWorld.gaze` to the object under the mouse pointer, or at
|
| 105 |
+
the centre of the camera (XR head gaze), so "this / that" resolve. Objects need colliders.
|
| 106 |
+
- **Clarification:** when the agent asks "Which toolbox do you mean?", the user can click a candidate (`session.Answer(id)`)
|
| 107 |
+
or just type "the red one": a short reply while a question is pending answers it; a new command starts over.
|
| 108 |
+
|
| 109 |
+
## What objects can do
|
| 110 |
+
Each label maps to what it affords (can be picked up, is a surface, a container, has on/off or open/closed states, is
|
| 111 |
+
fixed), from the model's catalogue of about 200 everyday object types. For a label the catalogue does not know, tick
|
| 112 |
+
`SceneItem.overrideAffordances` and set the flags, or rename it to a common word; Validate Scene lists every such
|
| 113 |
+
label.
|
| 114 |
+
|
| 115 |
+
## Performance
|
| 116 |
+
| setting (4 GB laptop GPU) | Maverick 4B |
|
| 117 |
+
|---|---|
|
| 118 |
+
| model alone, first answer | 1.2 s |
|
| 119 |
+
| next to a rendering scene, per command, GPU cool | 0.5–3.6 s |
|
| 120 |
+
| next to a rendering scene, per command, GPU at its thermal limit | 5.9 s |
|
| 121 |
+
|
| 122 |
+
- **Cap the frame rate.** An uncapped renderer on the same GPU slowed the model about 5×. `LocalModelServer`
|
| 123 |
+
sets `Application.targetFrameRate` to `capFrameRate` (30) while the model runs on the GPU; set 0 in XR, where the
|
| 124 |
+
headset sets the frame rate.
|
| 125 |
+
- **Heat:** long sessions on a laptop hold the GPU at its thermal limit and slow everything; cooling helps more than
|
| 126 |
+
any setting.
|
| 127 |
+
|
| 128 |
+
## Test and QA
|
| 129 |
+
- **Session logs:** `SceneAgentSession.logDirectory` (set by Set Up Scene to `SceneAgentSessions`) writes every
|
| 130 |
+
command, the scene the model saw, the calls and the answer to `Application.persistentDataPath/SceneAgentSessions`.
|
| 131 |
+
- **Scripted runs:** add `SceneAgentScriptRunner`, give it a text file:
|
| 132 |
+
```
|
| 133 |
+
pick up the red mug
|
| 134 |
+
#answer mug_1
|
| 135 |
+
put it on the table
|
| 136 |
+
#look lamp_2
|
| 137 |
+
turn that on
|
| 138 |
+
```
|
| 139 |
+
It writes `report.json` (calls, answers, errors, time per step, optional screenshots). From a build, in any scene with
|
| 140 |
+
a Scene Agent (the runner adds itself): `Game.exe -sceneAgentScript commands.txt -sceneAgentOut results -sceneAgentShots`.
|
| 141 |
+
|
| 142 |
+
## Speech input
|
| 143 |
+
Any speech-to-text works: pass the text to `session.Command`. For a fully offline app use a local recogniser (for
|
| 144 |
+
example whisper.cpp); cloud dictation sends audio off the device. The model was trained on speech-recognition noise
|
| 145 |
+
("um put er the mug…"), not on typing errors.
|
| 146 |
+
|
| 147 |
+
## Limitations
|
| 148 |
+
- **English only; the 11 tools only.** No creating or deleting objects, materials, lighting, physics or code.
|
| 149 |
+
- **Acting on an object in another room:** it searches other rooms for fetch/find requests ("bring me the book",
|
| 150 |
+
"where is…"), but "turn on the lamp" when the lamp is elsewhere is answered with "there is no lamp here", and "put
|
| 151 |
+
the drill on the workbench" with the drill elsewhere made the model reach for a similar object here (the hammer, the
|
| 152 |
+
screwdriver). The substitute guard blocks that call; the user should say the room or go there first.
|
| 153 |
+
- **Placing puts objects at the target's centre:** two things put on the same shelf overlap. Arrange them in
|
| 154 |
+
`ToolExecuted` if it matters in your scene.
|
| 155 |
+
- **Questions about the scene** ("what is on the workbench?") are answered from the scene data but were not evaluated
|
| 156 |
+
and can be wrong (in one test it listed an object that had been moved).
|
| 157 |
+
- **Room size:** a room of 90 objects worked; 170 exceeded the 4096-token context (a clear error, no crash). Keep rooms
|
| 158 |
+
under about 60 objects for multi-step commands, or raise `contextSize` (more VRAM).
|
| 159 |
+
- **Other languages:** trained on English only. A Turkish command worked in one test; do not rely on it.
|
| 160 |
+
- **Typos** can be read as unknown objects ("lapm").
|
| 161 |
+
- **Infeasible requests (4B):** it sometimes tries the action, Unity refuses it, then it explains.
|
| 162 |
+
- **It sees the scene as data, not the camera.** Keep labels and colours accurate.
|
| 163 |
+
- Not validated for real robots or safety-critical control.
|
| 164 |
+
|
| 165 |
+
Full evaluation: 21 of 22 pre-registered capability targets, 83.6–86.6% on real human instructions (ALFRED-TC, our
|
| 166 |
+
adaptation), 94.5–97.3% on 440 commands through this C# code.
|
| 167 |
+
|
| 168 |
+
## Troubleshooting
|
| 169 |
+
| symptom | fix |
|
| 170 |
+
|---|---|
|
| 171 |
+
| "llama-server not found" | Tools → Scene Agent → Open Model Folder; copy the llama.cpp release there (all files) |
|
| 172 |
+
| "no model (.gguf) found" | put one `.gguf` in the same folder, or set `LocalModelServer.modelPath` |
|
| 173 |
+
| model loads on the CPU | the GPU build of llama.cpp does not match your GPU or VRAM is full; see the log (`LocalModelServer.LogPath`) |
|
| 174 |
+
| every answer is slow | cap the frame rate; check GPU temperature |
|
| 175 |
+
| "the scene and request need N tokens…" | too many objects in the room: split it, or raise `LocalModelServer.contextSize` |
|
| 176 |
+
| "not done: … is in the storage, not in this room" | the substitute guard: the user named an object in another room |
|
| 177 |
+
| "object … is not visible" | the object's `room` differs from `SceneWorld.playerRoom` |
|
| 178 |
+
| grab/place refused on your object | its label is not in the catalogue: `overrideAffordances`, or a common word |
|
| 179 |
+
| objects jump when first moved | `SceneWorld.initFromTransforms` must be on (Set Up Scene turns it on) |
|
| 180 |
+
|
| 181 |
+
## License
|
| 182 |
+
Package code: MIT (`LICENSE.md`). The model weights are a fine-tune of Qwen3 (Apache-2.0). llama.cpp is MIT. See
|
| 183 |
+
`Third Party Notices.md`.
|