|
Download README.md from ErenAta00/Maverick-Unity: direct link, hf CLI and curl.
- Browser
- Download file 8.86 kB
-
https://huggingface.co/ErenAta00/Maverick-Unity/resolve/main/README.md
- Command line
-
hf download hf://ErenAta00/Maverick-Unity/README.md
-
curl -L -o README.md https://huggingface.co/ErenAta00/Maverick-Unity/resolve/main/README.md
8.86 kB
| license: mit | |
| tags: | |
| - unity | |
| - upm | |
| - unity-package | |
| - xr | |
| - virtual-reality | |
| - llama.cpp | |
| - offline | |
| - voice-control | |
| - agent | |
| # Maverick XR Agent for Unity | |
| Control a Unity scene with typed or spoken English. This package runs | |
| [Maverick-4B-Unity-XR-Agent](https://huggingface.co/ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF) on the player's own | |
| machine through llama.cpp: no account, no internet connection and no API key. | |
| You choose which objects the agent may use. For each command the model reads those objects and calls one of 11 tools, | |
| which Unity then carries out: "put the red mug on the table", "turn on the lamp", "where is the stapler?". When a | |
| command could mean more than one object, it asks which one. When it cannot do something, it says so. | |
| The article [Offline Voice Control for Unity XR Scenes](https://huggingface.co/blog/ErenAta00/maverick-4b-unity-xr-agent) explains how the model was trained and tested. | |
| ## Requirements | |
| - Unity 2022.3 LTS or later | |
| - Windows x64 (Linux and macOS are untested) | |
| - About 3 GB of free GPU memory. Without a GPU the model runs on the CPU, but slowly. | |
| - About 3 GB of disk space for the model and llama.cpp | |
| ## Installation | |
| 1. In Unity, open **Window > Package Manager**, click **+**, choose **Add package from git URL** and enter: | |
| ```text | |
| https://github.com/mcbu-xrlab/Maverick-Unity.git#v1.0.0 | |
| ``` | |
| This package is also mirrored here on Hugging Face: `https://huggingface.co/ErenAta00/Maverick-Unity.git`. | |
| The [GitHub repository](https://github.com/mcbu-xrlab/Maverick-Unity) has the full integration guides. | |
| 2. Choose **Tools > Scene Agent > Open Model Folder**. This creates `Assets/StreamingAssets/SceneAgent/`. | |
| 3. Copy two things into that folder: | |
| - **llama.cpp for Windows** from the [llama.cpp releases](https://github.com/ggml-org/llama.cpp/releases): the CUDA 12 | |
| build for NVIDIA GPUs, the Vulkan build for other GPUs, or the CPU build. Copy every file in the archive, not only | |
| `llama-server.exe`. | |
| - **The model** `Maverick-4B-Unity-XR-Agent-Q4_K_M.gguf` from the | |
| [model page](https://huggingface.co/ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF/tree/main). | |
| Unity copies `StreamingAssets` into your build, so players do not need to install anything. | |
| ## Quick start | |
| 1. Choose **Tools > Scene Agent > Set Up Scene**. This adds a *Scene Agent* object with all components connected, | |
| including a chat box for testing. | |
| 2. Give your objects plain names, such as *Red Mug*, *Table* or *Floor Lamp*. | |
| 3. Select the objects the agent may use and choose **Tools > Scene Agent > Make Selected Interactive**. Each one gets a | |
| `SceneItem` with an id (`mug_1`), a label (`mug`), a colour (`red`), a room and the object it stands on. | |
| 4. Choose **Tools > Scene Agent > Validate Scene** and fix anything it reports. | |
| 5. Enter Play mode, type `put the red mug on the table` and press Enter. If the agent asks which mug you mean, click | |
| one. F1 hides the chat box. | |
| ## Using it from code | |
| ```csharp | |
| using System.Collections.Generic; | |
| using SceneAgent; | |
| using UnityEngine; | |
| public class VoiceCommands : MonoBehaviour | |
| { | |
| public SceneAgentSession session; // on the Scene Agent object | |
| void Start() | |
| { | |
| session.Line += text => Debug.Log(text); | |
| session.Finished += result => | |
| { | |
| if (result.clarificationQuestion != null) ShowChoices(session.Candidates); | |
| }; | |
| session.world.ToolExecuted += (tool, args, result) => | |
| { | |
| if (tool == "teleport") MoveRigTo(session.world.playerRoom); | |
| }; | |
| } | |
| // Typed text or the output of a speech recogniser. | |
| public void OnCommand(string text) => StartCoroutine(session.Command(text)); | |
| // The user's answer to "Which one do you mean?". | |
| public void OnPick(string itemId) => StartCoroutine(session.Answer(itemId)); | |
| void ShowChoices(List<string> itemIds) { /* your UI */ } | |
| void MoveRigTo(string room) { /* your camera rig */ } | |
| } | |
| ``` | |
| - **Rooms.** Give each `SceneItem` a `room`. The model sees only the player's room. It searches other rooms with | |
| `find_objects` and moves the player with `teleport`, which sets `SceneWorld.playerRoom`; move your camera rig in | |
| `ToolExecuted`. | |
| - **"This" and "that".** `SceneAgentGaze` points them at the object under the mouse, or at the centre of the view in | |
| XR. The objects need colliders. | |
| - **Follow-up answers.** When the agent asks which object you mean, the user can pick one (`session.Answer`) or reply | |
| in words, such as "the red one". | |
| - **Speech.** Pass the recognised text to `session.Command`. For a fully offline app, use a local recogniser such as | |
| whisper.cpp. | |
| ## Tools | |
| | Tool | What it does | | |
| |---|---| | |
| | `grab`, `release` | Picks up or drops an object with the left, right or both hands | | |
| | `place` | Puts an object on a surface or in a container | | |
| | `move_by`, `rotate` | Moves an object by a given distance or turns it | | |
| | `set_state` | Switches on or off, opens, closes, locks or unlocks | | |
| | `press` | Presses a button | | |
| | `highlight` | Highlights objects | | |
| | `teleport` | Moves the player to another room | | |
| | `find_objects` | Searches every room for objects | | |
| | `ask_clarification` | Asks which object the user means | | |
| Unity refuses a call that does not fit the object, such as picking up a table, and tells the model why. What each | |
| object can do comes from its label and a catalogue of about 200 everyday objects. For a label outside the catalogue, | |
| turn on `SceneItem.overrideAffordances` and set the flags yourself; **Validate Scene** lists these labels. | |
| ## Components | |
| | Component | Purpose | Main settings | | |
| |---|---|---| | |
| | `SceneWorld` | Describes the scene to the model and carries out the tools | `rooms`, `playerRoom`, `ToolExecuted` | | |
| | `SceneItem` | Marks an object the agent may use | `itemId`, `label`, `color`, `room`, `overrideAffordances` | | |
| | `SceneAgent` | Sends commands to the model and runs its tool calls | `maxTurns`, `fallbackDecoding`, `guardSubstitutes` | | |
| | `LocalModelServer` | Starts and stops llama.cpp with the game | `modelPath`, `contextSize`, `preferGpu`, `capFrameRate` | | |
| | `SceneAgentSession` | Holds one conversation, including follow-up questions | `logDirectory` | | |
| | `SceneAgentGaze` | Resolves "this" and "that" | | | |
| | `SceneAgentChatBox` | On-screen chat for testing | `toggleKey`, `visibleLines` | | |
| | `SceneAgentScriptRunner` | Plays a text file of commands and writes a report | `script`, `outputDirectory` | | |
| `fallbackDecoding` checks every answer before it reaches your scene: a call to a tool or value that does not exist is | |
| turned into a refusal. `guardSubstitutes` stops the model from acting on a similar object when the one the user named | |
| is in another room. Leave both on. | |
| ## Performance | |
| - **Cap the frame rate.** An uncapped renderer on the same GPU can make the model several times slower. | |
| `LocalModelServer` caps rendering at 30 frames per second while the model runs on the GPU. In XR, set `capFrameRate` | |
| to 0 and let the headset set the frame rate. | |
| - **Keep rooms small.** Under about 60 interactive objects per room, multi-step commands fit the default 4,096-token | |
| context. For larger rooms, raise `LocalModelServer.contextSize`, which needs more GPU memory. | |
| ## Limitations | |
| - English only, and only the 11 tools. The agent cannot create or delete objects or change materials, lighting or | |
| physics. | |
| - It finds and fetches objects in other rooms ("bring me the book"), but when the lamp is in another room, "turn on | |
| the lamp" gets "there is no lamp here". The user should name the room or go there first. | |
| - Placed objects go to the centre of the target, so two objects placed on the same shelf overlap. Arrange them in | |
| `ToolExecuted` if this matters in your scene. | |
| - The model reads the scene as data, not through the camera. Keep object names and colours accurate. | |
| - Not intended for robots or safety-critical control. | |
| ## Troubleshooting | |
| | Problem | Fix | | |
| |---|---| | |
| | "llama-server not found" | Copy every file of the llama.cpp release into the model folder. | | |
| | "no model (.gguf) found" | Put the `.gguf` file in the model folder, or set `LocalModelServer.modelPath`. | | |
| | The model runs on the CPU | The llama.cpp build does not match your GPU, or GPU memory is full. See the log at `LocalModelServer.LogPath`. | | |
| | Every answer is slow | Cap the frame rate and check the GPU temperature. | | |
| | "the scene and request need N tokens" | The room has too many objects. Split it or raise `LocalModelServer.contextSize`. | | |
| | "object ... is not visible" | The object's `room` is not `SceneWorld.playerRoom`. | | |
| | A call on your object is refused | Its label is not in the catalogue. Use `overrideAffordances` or a more common name. | | |
| ## Licence | |
| The package is released under the [MIT License](LICENSE.md). The model is released under Apache-2.0, like its base | |
| model, Qwen3-4B. llama.cpp is released under the MIT License. | |