Maverick XR Agent for Unity

Control a Unity scene with typed or spoken English. This package runs Maverick-4B-Unity-XR-Agent on the player's own machine through llama.cpp: no account, no internet connection and no API key.

You choose which objects the agent may use. For each command the model reads those objects and calls one of 11 tools, which Unity then carries out: "put the red mug on the table", "turn on the lamp", "where is the stapler?". When a command could mean more than one object, it asks which one. When it cannot do something, it says so.

The article Offline Voice Control for Unity XR Scenes explains how the model was trained and tested.

Requirements

  • Unity 2022.3 LTS or later
  • Windows x64 (Linux and macOS are untested)
  • About 3 GB of free GPU memory. Without a GPU the model runs on the CPU, but slowly.
  • About 3 GB of disk space for the model and llama.cpp

Installation

  1. In Unity, open Window > Package Manager, click +, choose Add package from git URL and enter:

    https://huggingface.co/ErenAta00/Maverick-Unity.git
    
  2. Choose Tools > Scene Agent > Open Model Folder. This creates Assets/StreamingAssets/SceneAgent/.

  3. Copy two things into that folder:

    • llama.cpp for Windows from the llama.cpp releases: the CUDA 12 build for NVIDIA GPUs, the Vulkan build for other GPUs, or the CPU build. Copy every file in the archive, not only llama-server.exe.
    • The model Maverick-4B-Unity-XR-Agent-Q4_K_M.gguf from the model page.

Unity copies StreamingAssets into your build, so players do not need to install anything.

Quick start

  1. Choose Tools > Scene Agent > Set Up Scene. This adds a Scene Agent object with all components connected, including a chat box for testing.
  2. Give your objects plain names, such as Red Mug, Table or Floor Lamp.
  3. Select the objects the agent may use and choose Tools > Scene Agent > Make Selected Interactive. Each one gets a SceneItem with an id (mug_1), a label (mug), a colour (red), a room and the object it stands on.
  4. Choose Tools > Scene Agent > Validate Scene and fix anything it reports.
  5. Enter Play mode, type put the red mug on the table and press Enter. If the agent asks which mug you mean, click one. F1 hides the chat box.

Using it from code

using System.Collections.Generic;
using SceneAgent;
using UnityEngine;

public class VoiceCommands : MonoBehaviour
{
    public SceneAgentSession session;   // on the Scene Agent object

    void Start()
    {
        session.Line += text => Debug.Log(text);
        session.Finished += result =>
        {
            if (result.clarificationQuestion != null) ShowChoices(session.Candidates);
        };
        session.world.ToolExecuted += (tool, args, result) =>
        {
            if (tool == "teleport") MoveRigTo(session.world.playerRoom);
        };
    }

    // Typed text or the output of a speech recogniser.
    public void OnCommand(string text) => StartCoroutine(session.Command(text));

    // The user's answer to "Which one do you mean?".
    public void OnPick(string itemId) => StartCoroutine(session.Answer(itemId));

    void ShowChoices(List<string> itemIds) { /* your UI */ }
    void MoveRigTo(string room) { /* your camera rig */ }
}
  • Rooms. Give each SceneItem a room. The model sees only the player's room. It searches other rooms with find_objects and moves the player with teleport, which sets SceneWorld.playerRoom; move your camera rig in ToolExecuted.
  • "This" and "that". SceneAgentGaze points them at the object under the mouse, or at the centre of the view in XR. The objects need colliders.
  • Follow-up answers. When the agent asks which object you mean, the user can pick one (session.Answer) or reply in words, such as "the red one".
  • Speech. Pass the recognised text to session.Command. For a fully offline app, use a local recogniser such as whisper.cpp.

Tools

Tool What it does
grab, release Picks up or drops an object with the left, right or both hands
place Puts an object on a surface or in a container
move_by, rotate Moves an object by a given distance or turns it
set_state Switches on or off, opens, closes, locks or unlocks
press Presses a button
highlight Highlights objects
teleport Moves the player to another room
find_objects Searches every room for objects
ask_clarification Asks which object the user means

Unity refuses a call that does not fit the object, such as picking up a table, and tells the model why. What each object can do comes from its label and a catalogue of about 200 everyday objects. For a label outside the catalogue, turn on SceneItem.overrideAffordances and set the flags yourself; Validate Scene lists these labels.

Components

Component Purpose Main settings
SceneWorld Describes the scene to the model and carries out the tools rooms, playerRoom, ToolExecuted
SceneItem Marks an object the agent may use itemId, label, color, room, overrideAffordances
SceneAgent Sends commands to the model and runs its tool calls maxTurns, fallbackDecoding, guardSubstitutes
LocalModelServer Starts and stops llama.cpp with the game modelPath, contextSize, preferGpu, capFrameRate
SceneAgentSession Holds one conversation, including follow-up questions logDirectory
SceneAgentGaze Resolves "this" and "that"
SceneAgentChatBox On-screen chat for testing toggleKey, visibleLines
SceneAgentScriptRunner Plays a text file of commands and writes a report script, outputDirectory

fallbackDecoding checks every answer before it reaches your scene: a call to a tool or value that does not exist is turned into a refusal. guardSubstitutes stops the model from acting on a similar object when the one the user named is in another room. Leave both on.

Performance

  • Cap the frame rate. An uncapped renderer on the same GPU can make the model several times slower. LocalModelServer caps rendering at 30 frames per second while the model runs on the GPU. In XR, set capFrameRate to 0 and let the headset set the frame rate.
  • Keep rooms small. Under about 60 interactive objects per room, multi-step commands fit the default 4,096-token context. For larger rooms, raise LocalModelServer.contextSize, which needs more GPU memory.

Limitations

  • English only, and only the 11 tools. The agent cannot create or delete objects or change materials, lighting or physics.
  • It finds and fetches objects in other rooms ("bring me the book"), but when the lamp is in another room, "turn on the lamp" gets "there is no lamp here". The user should name the room or go there first.
  • Placed objects go to the centre of the target, so two objects placed on the same shelf overlap. Arrange them in ToolExecuted if this matters in your scene.
  • The model reads the scene as data, not through the camera. Keep object names and colours accurate.
  • Not intended for robots or safety-critical control.

Troubleshooting

Problem Fix
"llama-server not found" Copy every file of the llama.cpp release into the model folder.
"no model (.gguf) found" Put the .gguf file in the model folder, or set LocalModelServer.modelPath.
The model runs on the CPU The llama.cpp build does not match your GPU, or GPU memory is full. See the log at LocalModelServer.LogPath.
Every answer is slow Cap the frame rate and check the GPU temperature.
"the scene and request need N tokens" The room has too many objects. Split it or raise LocalModelServer.contextSize.
"object ... is not visible" The object's room is not SceneWorld.playerRoom.
A call on your object is refused Its label is not in the catalogue. Use overrideAffordances or a more common name.

Licence

The package is released under the MIT License. The model is released under Apache-2.0, like its base model, Qwen3-4B. llama.cpp is released under the MIT License.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Article mentioning ErenAta00/Maverick-Unity