ErenAta00 commited on
Commit
a40c877
·
verified ·
1 Parent(s): 34bec90

README with repository metadata

Browse files
Files changed (1) hide show
  1. README.md +183 -0
README.md ADDED
@@ -0,0 +1,183 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ tags:
4
+ - unity
5
+ - upm
6
+ - unity-package
7
+ - xr
8
+ - virtual-reality
9
+ - llama.cpp
10
+ - offline
11
+ - voice-control
12
+ - agent
13
+ ---
14
+
15
+ # Maverick XR Agent for Unity
16
+
17
+ The Unity package for [Maverick 4B](https://huggingface.co/ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF).
18
+
19
+ Let people control your Unity scene in plain English, typed or spoken, with a small language model that runs **on
20
+ the same machine, offline**. "Put the red mug on the table", "turn on the lamp", "where is the stapler?": the model
21
+ reads the objects you marked as interactive, calls the right tool, and Unity executes it. When the words fit two
22
+ objects it asks which one; when an object is in another room it searches there; when a request is impossible it says
23
+ so instead of acting.
24
+
25
+ > **Status: release candidate (1.0.0-rc.2).** Tested on Windows 11 x64, Unity 2022.3 LTS, NVIDIA RTX
26
+ > 3050 Ti Laptop (4 GB). See [Limitations](#limitations) before building on it.
27
+
28
+ - **11 tools:** grab, release, place, move_by, rotate, set_state, press, highlight, teleport, find_objects,
29
+ ask_clarification.
30
+ - **Understands** synonyms and general words ("the fruit"), colours, spatial references ("the one next to the chair",
31
+ "the second from the left"), "this / that" with the user's gaze, the object in hand ("put it on the shelf"), filler
32
+ words and speech-recognition noise.
33
+ - **Safe by default:**
34
+ - D5b: a call to a tool or a value that does not exist is turned into a refusal before it reaches your scene; an
35
+ unknown object id is generated again under a grammar.
36
+ - Substitute guard: when the user names an object that is only in another room and the model tries to act on a
37
+ different object here, the call is not executed and the model is told where the named object is.
38
+ - **Private and free to run:** llama.cpp on `127.0.0.1`; no account, no cloud call, no per-request cost.
39
+
40
+ ## Requirements
41
+ | | |
42
+ |---|---|
43
+ | Unity | 2022.3 LTS (tested 2022.3.62f3) |
44
+ | OS | Windows x64 (tested). llama.cpp also runs on Linux and macOS; `LocalModelServer` is written for them but untested |
45
+ | GPU | ~3.1 GB free VRAM; CPU fallback works but is slow |
46
+ | Disk | 2.5 GB for the model, plus ~100 MB for llama.cpp |
47
+
48
+ ## Install
49
+ 1. **Package:** Window → Package Manager → **+** → *Add package from git URL…* →
50
+ `https://huggingface.co/ErenAta00/Maverick-Unity.git` (or *Add package from disk…* → this folder's `package.json`).
51
+ 2. **Model files:** Tools → Scene Agent → **Open Model Folder** (creates `Assets/StreamingAssets/SceneAgent/`) and put
52
+ there:
53
+ - `llama-server.exe` and the DLLs next to it, from a llama.cpp release for Windows (CUDA 12 build for NVIDIA GPUs,
54
+ Vulkan build for other GPUs, CPU build otherwise). Copy the whole folder's contents, not only the .exe.
55
+ - The model: `Maverick-4B-Unity-XR-Agent-Q4_K_M.gguf` (2.5 GB) from
56
+ https://huggingface.co/ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF.
57
+ The server takes the first `.gguf` it finds there, or set `LocalModelServer.modelPath`.
58
+
59
+ StreamingAssets is copied into a build as it is, so a built game runs without installing anything.
60
+
61
+ ## Quick start (5 minutes)
62
+ 1. **Tools → Scene Agent → Set Up Scene.** Adds a *Scene Agent* object with everything wired: `SceneWorld` (your
63
+ objects, as the model sees them), `SceneAgent` (the conversation with the model), `LocalModelServer` (starts
64
+ llama.cpp), `SceneAgentSession` and `SceneAgentChatBox` (an on-screen box to try it).
65
+ 2. **Build or open your scene.** Name objects with plain words: *Red Mug*, *Table*, *Floor Lamp*, *Book*.
66
+ 3. **Select the objects** the agent may use → **Tools → Scene Agent → Make Selected Interactive.** Each gets a
67
+ `SceneItem`: an id (`mug_1`), a label (`mug`) and a colour (`red`) from its name, a room, and what it stands on (from
68
+ the renderers' bounds). They move under the *Scene Agent* object.
69
+ 4. **Tools → Scene Agent → Validate Scene.** Fix what it reports (duplicate ids, labels the model does not know,
70
+ missing model files).
71
+ 5. **Play.** Type `put the red mug on the table` and press Enter. When the agent asks "Which mug do you mean?", click
72
+ one. F1 hides the box.
73
+
74
+ ## Use it from your code
75
+ ```csharp
76
+ using SceneAgent;
77
+
78
+ public class MyVoiceUI : MonoBehaviour
79
+ {
80
+ public SceneAgentSession session; // on the Scene Agent object
81
+
82
+ void Start()
83
+ {
84
+ session.Line += text => Debug.Log(text); // "You: …", "→ call result", "Agent: …"
85
+ session.Finished += r => { if (r.clarificationQuestion != null) ShowChoices(session.Candidates); };
86
+ session.world.ToolExecuted += (tool, args, result) =>
87
+ {
88
+ if (tool == "teleport") MoveCameraTo(session.world.playerRoom); // react in your own way
89
+ };
90
+ }
91
+
92
+ public void OnSpeech(string text) => StartCoroutine(session.Command(text)); // typed or recognised speech
93
+ public void OnPick(string itemId) => StartCoroutine(session.Answer(itemId)); // the user's answer to "which one?"
94
+
95
+ void ShowChoices(System.Collections.Generic.List<string> ids) { /* your UI */ }
96
+ void MoveCameraTo(string room) { /* your rig */ }
97
+ }
98
+ ```
99
+ - **Lower level:** `StartCoroutine(agent.Run(text, result => …))` returns an `AgentResult`: the calls, their results,
100
+ the final text or the clarification question, an error, and the time.
101
+ - **Rooms:** give each `SceneItem` a `room`. The model sees the player's room only; it finds objects elsewhere with
102
+ `find_objects` and moves the player with `teleport` (which sets `SceneWorld.playerRoom`; you move the camera in
103
+ `ToolExecuted`).
104
+ - **Gaze:** `SceneAgentGaze` (added by Set Up Scene) sets `SceneWorld.gaze` to the object under the mouse pointer, or at
105
+ the centre of the camera (XR head gaze), so "this / that" resolve. Objects need colliders.
106
+ - **Clarification:** when the agent asks "Which toolbox do you mean?", the user can click a candidate (`session.Answer(id)`)
107
+ or just type "the red one": a short reply while a question is pending answers it; a new command starts over.
108
+
109
+ ## What objects can do
110
+ Each label maps to what it affords (can be picked up, is a surface, a container, has on/off or open/closed states, is
111
+ fixed), from the model's catalogue of about 200 everyday object types. For a label the catalogue does not know, tick
112
+ `SceneItem.overrideAffordances` and set the flags, or rename it to a common word; Validate Scene lists every such
113
+ label.
114
+
115
+ ## Performance
116
+ | setting (4 GB laptop GPU) | Maverick 4B |
117
+ |---|---|
118
+ | model alone, first answer | 1.2 s |
119
+ | next to a rendering scene, per command, GPU cool | 0.5–3.6 s |
120
+ | next to a rendering scene, per command, GPU at its thermal limit | 5.9 s |
121
+
122
+ - **Cap the frame rate.** An uncapped renderer on the same GPU slowed the model about 5×. `LocalModelServer`
123
+ sets `Application.targetFrameRate` to `capFrameRate` (30) while the model runs on the GPU; set 0 in XR, where the
124
+ headset sets the frame rate.
125
+ - **Heat:** long sessions on a laptop hold the GPU at its thermal limit and slow everything; cooling helps more than
126
+ any setting.
127
+
128
+ ## Test and QA
129
+ - **Session logs:** `SceneAgentSession.logDirectory` (set by Set Up Scene to `SceneAgentSessions`) writes every
130
+ command, the scene the model saw, the calls and the answer to `Application.persistentDataPath/SceneAgentSessions`.
131
+ - **Scripted runs:** add `SceneAgentScriptRunner`, give it a text file:
132
+ ```
133
+ pick up the red mug
134
+ #answer mug_1
135
+ put it on the table
136
+ #look lamp_2
137
+ turn that on
138
+ ```
139
+ It writes `report.json` (calls, answers, errors, time per step, optional screenshots). From a build, in any scene with
140
+ a Scene Agent (the runner adds itself): `Game.exe -sceneAgentScript commands.txt -sceneAgentOut results -sceneAgentShots`.
141
+
142
+ ## Speech input
143
+ Any speech-to-text works: pass the text to `session.Command`. For a fully offline app use a local recogniser (for
144
+ example whisper.cpp); cloud dictation sends audio off the device. The model was trained on speech-recognition noise
145
+ ("um put er the mug…"), not on typing errors.
146
+
147
+ ## Limitations
148
+ - **English only; the 11 tools only.** No creating or deleting objects, materials, lighting, physics or code.
149
+ - **Acting on an object in another room:** it searches other rooms for fetch/find requests ("bring me the book",
150
+ "where is…"), but "turn on the lamp" when the lamp is elsewhere is answered with "there is no lamp here", and "put
151
+ the drill on the workbench" with the drill elsewhere made the model reach for a similar object here (the hammer, the
152
+ screwdriver). The substitute guard blocks that call; the user should say the room or go there first.
153
+ - **Placing puts objects at the target's centre:** two things put on the same shelf overlap. Arrange them in
154
+ `ToolExecuted` if it matters in your scene.
155
+ - **Questions about the scene** ("what is on the workbench?") are answered from the scene data but were not evaluated
156
+ and can be wrong (in one test it listed an object that had been moved).
157
+ - **Room size:** a room of 90 objects worked; 170 exceeded the 4096-token context (a clear error, no crash). Keep rooms
158
+ under about 60 objects for multi-step commands, or raise `contextSize` (more VRAM).
159
+ - **Other languages:** trained on English only. A Turkish command worked in one test; do not rely on it.
160
+ - **Typos** can be read as unknown objects ("lapm").
161
+ - **Infeasible requests (4B):** it sometimes tries the action, Unity refuses it, then it explains.
162
+ - **It sees the scene as data, not the camera.** Keep labels and colours accurate.
163
+ - Not validated for real robots or safety-critical control.
164
+
165
+ Full evaluation: 21 of 22 pre-registered capability targets, 83.6–86.6% on real human instructions (ALFRED-TC, our
166
+ adaptation), 94.5–97.3% on 440 commands through this C# code.
167
+
168
+ ## Troubleshooting
169
+ | symptom | fix |
170
+ |---|---|
171
+ | "llama-server not found" | Tools → Scene Agent → Open Model Folder; copy the llama.cpp release there (all files) |
172
+ | "no model (.gguf) found" | put one `.gguf` in the same folder, or set `LocalModelServer.modelPath` |
173
+ | model loads on the CPU | the GPU build of llama.cpp does not match your GPU or VRAM is full; see the log (`LocalModelServer.LogPath`) |
174
+ | every answer is slow | cap the frame rate; check GPU temperature |
175
+ | "the scene and request need N tokens…" | too many objects in the room: split it, or raise `LocalModelServer.contextSize` |
176
+ | "not done: … is in the storage, not in this room" | the substitute guard: the user named an object in another room |
177
+ | "object … is not visible" | the object's `room` differs from `SceneWorld.playerRoom` |
178
+ | grab/place refused on your object | its label is not in the catalogue: `overrideAffordances`, or a common word |
179
+ | objects jump when first moved | `SceneWorld.initFromTransforms` must be on (Set Up Scene turns it on) |
180
+
181
+ ## License
182
+ Package code: MIT (`LICENSE.md`). The model weights are a fine-tune of Qwen3 (Apache-2.0). llama.cpp is MIT. See
183
+ `Third Party Notices.md`.