File size: 8,864 Bytes
a40c877
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
33cde86
 
 
 
 
 
 
a40c877
1bb8c83
 
a40c877
33cde86
 
 
 
 
 
 
 
 
 
 
92315be
33cde86
 
92315be
 
 
33cde86
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a40c877
33cde86
a40c877
33cde86
a40c877
33cde86
a40c877
33cde86
a40c877
 
 
33cde86
 
 
 
 
a40c877
 
33cde86
a40c877
 
 
33cde86
 
 
 
 
a40c877
33cde86
 
a40c877
 
 
33cde86
 
 
 
 
 
 
 
 
 
 
 
 
a40c877
33cde86
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
a40c877
 
33cde86
 
 
 
 
 
 
 
 
a40c877
 
33cde86
 
a40c877
33cde86
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
---
license: mit
tags:
- unity
- upm
- unity-package
- xr
- virtual-reality
- llama.cpp
- offline
- voice-control
- agent
---

# Maverick XR Agent for Unity

Control a Unity scene with typed or spoken English. This package runs
[Maverick-4B-Unity-XR-Agent](https://huggingface.co/ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF) on the player's own
machine through llama.cpp: no account, no internet connection and no API key.

You choose which objects the agent may use. For each command the model reads those objects and calls one of 11 tools,
which Unity then carries out: "put the red mug on the table", "turn on the lamp", "where is the stapler?". When a
command could mean more than one object, it asks which one. When it cannot do something, it says so.

The article [Offline Voice Control for Unity XR Scenes](https://huggingface.co/blog/ErenAta00/maverick-4b-unity-xr-agent) explains how the model was trained and tested.

## Requirements

- Unity 2022.3 LTS or later
- Windows x64 (Linux and macOS are untested)
- About 3 GB of free GPU memory. Without a GPU the model runs on the CPU, but slowly.
- About 3 GB of disk space for the model and llama.cpp

## Installation

1. In Unity, open **Window > Package Manager**, click **+**, choose **Add package from git URL** and enter:

   ```text
   https://github.com/mcbu-xrlab/Maverick-Unity.git#v1.0.0
   ```

   This package is also mirrored here on Hugging Face: `https://huggingface.co/ErenAta00/Maverick-Unity.git`.
   The [GitHub repository](https://github.com/mcbu-xrlab/Maverick-Unity) has the full integration guides.

2. Choose **Tools > Scene Agent > Open Model Folder**. This creates `Assets/StreamingAssets/SceneAgent/`.
3. Copy two things into that folder:
   - **llama.cpp for Windows** from the [llama.cpp releases](https://github.com/ggml-org/llama.cpp/releases): the CUDA 12
     build for NVIDIA GPUs, the Vulkan build for other GPUs, or the CPU build. Copy every file in the archive, not only
     `llama-server.exe`.
   - **The model** `Maverick-4B-Unity-XR-Agent-Q4_K_M.gguf` from the
     [model page](https://huggingface.co/ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF/tree/main).

Unity copies `StreamingAssets` into your build, so players do not need to install anything.

## Quick start

1. Choose **Tools > Scene Agent > Set Up Scene**. This adds a *Scene Agent* object with all components connected,
   including a chat box for testing.
2. Give your objects plain names, such as *Red Mug*, *Table* or *Floor Lamp*.
3. Select the objects the agent may use and choose **Tools > Scene Agent > Make Selected Interactive**. Each one gets a
   `SceneItem` with an id (`mug_1`), a label (`mug`), a colour (`red`), a room and the object it stands on.
4. Choose **Tools > Scene Agent > Validate Scene** and fix anything it reports.
5. Enter Play mode, type `put the red mug on the table` and press Enter. If the agent asks which mug you mean, click
   one. F1 hides the chat box.

## Using it from code

```csharp
using System.Collections.Generic;
using SceneAgent;
using UnityEngine;

public class VoiceCommands : MonoBehaviour
{
    public SceneAgentSession session;   // on the Scene Agent object

    void Start()
    {
        session.Line += text => Debug.Log(text);
        session.Finished += result =>
        {
            if (result.clarificationQuestion != null) ShowChoices(session.Candidates);
        };
        session.world.ToolExecuted += (tool, args, result) =>
        {
            if (tool == "teleport") MoveRigTo(session.world.playerRoom);
        };
    }

    // Typed text or the output of a speech recogniser.
    public void OnCommand(string text) => StartCoroutine(session.Command(text));

    // The user's answer to "Which one do you mean?".
    public void OnPick(string itemId) => StartCoroutine(session.Answer(itemId));

    void ShowChoices(List<string> itemIds) { /* your UI */ }
    void MoveRigTo(string room) { /* your camera rig */ }
}
```

- **Rooms.** Give each `SceneItem` a `room`. The model sees only the player's room. It searches other rooms with
  `find_objects` and moves the player with `teleport`, which sets `SceneWorld.playerRoom`; move your camera rig in
  `ToolExecuted`.
- **"This" and "that".** `SceneAgentGaze` points them at the object under the mouse, or at the centre of the view in
  XR. The objects need colliders.
- **Follow-up answers.** When the agent asks which object you mean, the user can pick one (`session.Answer`) or reply
  in words, such as "the red one".
- **Speech.** Pass the recognised text to `session.Command`. For a fully offline app, use a local recogniser such as
  whisper.cpp.

## Tools

| Tool | What it does |
|---|---|
| `grab`, `release` | Picks up or drops an object with the left, right or both hands |
| `place` | Puts an object on a surface or in a container |
| `move_by`, `rotate` | Moves an object by a given distance or turns it |
| `set_state` | Switches on or off, opens, closes, locks or unlocks |
| `press` | Presses a button |
| `highlight` | Highlights objects |
| `teleport` | Moves the player to another room |
| `find_objects` | Searches every room for objects |
| `ask_clarification` | Asks which object the user means |

Unity refuses a call that does not fit the object, such as picking up a table, and tells the model why. What each
object can do comes from its label and a catalogue of about 200 everyday objects. For a label outside the catalogue,
turn on `SceneItem.overrideAffordances` and set the flags yourself; **Validate Scene** lists these labels.

## Components

| Component | Purpose | Main settings |
|---|---|---|
| `SceneWorld` | Describes the scene to the model and carries out the tools | `rooms`, `playerRoom`, `ToolExecuted` |
| `SceneItem` | Marks an object the agent may use | `itemId`, `label`, `color`, `room`, `overrideAffordances` |
| `SceneAgent` | Sends commands to the model and runs its tool calls | `maxTurns`, `fallbackDecoding`, `guardSubstitutes` |
| `LocalModelServer` | Starts and stops llama.cpp with the game | `modelPath`, `contextSize`, `preferGpu`, `capFrameRate` |
| `SceneAgentSession` | Holds one conversation, including follow-up questions | `logDirectory` |
| `SceneAgentGaze` | Resolves "this" and "that" | |
| `SceneAgentChatBox` | On-screen chat for testing | `toggleKey`, `visibleLines` |
| `SceneAgentScriptRunner` | Plays a text file of commands and writes a report | `script`, `outputDirectory` |

`fallbackDecoding` checks every answer before it reaches your scene: a call to a tool or value that does not exist is
turned into a refusal. `guardSubstitutes` stops the model from acting on a similar object when the one the user named
is in another room. Leave both on.

## Performance

- **Cap the frame rate.** An uncapped renderer on the same GPU can make the model several times slower.
  `LocalModelServer` caps rendering at 30 frames per second while the model runs on the GPU. In XR, set `capFrameRate`
  to 0 and let the headset set the frame rate.
- **Keep rooms small.** Under about 60 interactive objects per room, multi-step commands fit the default 4,096-token
  context. For larger rooms, raise `LocalModelServer.contextSize`, which needs more GPU memory.

## Limitations

- English only, and only the 11 tools. The agent cannot create or delete objects or change materials, lighting or
  physics.
- It finds and fetches objects in other rooms ("bring me the book"), but when the lamp is in another room, "turn on
  the lamp" gets "there is no lamp here". The user should name the room or go there first.
- Placed objects go to the centre of the target, so two objects placed on the same shelf overlap. Arrange them in
  `ToolExecuted` if this matters in your scene.
- The model reads the scene as data, not through the camera. Keep object names and colours accurate.
- Not intended for robots or safety-critical control.

## Troubleshooting

| Problem | Fix |
|---|---|
| "llama-server not found" | Copy every file of the llama.cpp release into the model folder. |
| "no model (.gguf) found" | Put the `.gguf` file in the model folder, or set `LocalModelServer.modelPath`. |
| The model runs on the CPU | The llama.cpp build does not match your GPU, or GPU memory is full. See the log at `LocalModelServer.LogPath`. |
| Every answer is slow | Cap the frame rate and check the GPU temperature. |
| "the scene and request need N tokens" | The room has too many objects. Split it or raise `LocalModelServer.contextSize`. |
| "object ... is not visible" | The object's `room` is not `SceneWorld.playerRoom`. |
| A call on your object is refused | Its label is not in the catalogue. Use `overrideAffordances` or a more common name. |

## Licence

The package is released under the [MIT License](LICENSE.md). The model is released under Apache-2.0, like its base
model, Qwen3-4B. llama.cpp is released under the MIT License.