Maverick-4B-Unity-XR-Agent

Maverick-4B-Unity-XR-Agent

Maverick-4B-Unity-XR-Agent turns spoken or typed English into tool calls for Unity 3D scenes. Say "put the red mug on the counter" and it returns the call your application executes. It can pick up, place, move and rotate objects, switch devices on and off, highlight objects, move the user between rooms and search other rooms. If a request could mean two objects, it asks which one. If it cannot do something, it says so.

The model is a fine-tune of Qwen3-4B. It is built for XR and VR applications such as training simulations, product showrooms, digital twins, accessibility tools and games.

The article Offline Voice Control for Unity XR Scenes explains how the model was trained and tested.

Results

We tested the model on 499 instructions that people wrote for the ALFRED benchmark. It chose the right action on the right object 83.8% of the time. Five other models, from 1.7B to 120B parameters, scored between 35.9% and 67.1% with the same scenes and tools. On object types that never appear in its training data, the model scored 91.7%. In a Unity test scene it carried out 440 commands with 97.3% accuracy and no runtime errors.

Benchmark summary

Model details

Base model Qwen3-4B
Method QLoRA supervised fine-tuning, merged into the base weights
Training data 20,091 English conversations
Language English
Format GGUF, Q4_K_M, 2.50 GB
Runtime llama.cpp (llama-server, OpenAI-compatible API)
GPU memory 3.06 GB
Context 4,096 tokens, enough for a room of about 90 objects
Tools grab, release, place, move_by, rotate, set_state, press, highlight, teleport, find_objects, ask_clarification
Licence Apache-2.0, the same as the base model

How it works

For each command, your application sends the model two things: a JSON description of the current room and the user's sentence. The room description lists every object's id, label, colour, state, support and position, as well as the rooms and what the user is holding and looking at.

The model answers with tool calls. Your application executes them and sends back the results, and the model confirms what it did in one sentence.

Examples from our Unity test scene:

Command Tool call Reply
pick up the hammer grab(hammer_1, right) Done. I picked up the hammer with your right hand.
move the chair 50 cm to the left move_by(chair_1, left, 50, cm) Done. I moved the chair left by 50 cm.
um could you uh close the window set_state(window_1, closed) Done. I closed the window.
pick up the toolbox (there are two) ask_clarification([toolbox_1, toolbox_2]) Which toolbox do you mean, the blue or the red one?
highlight that (looking at a box) highlight([cardboard_box_1]) Done. I highlighted the box.
where is the broom? find_objects(broom) The broom is in the storage, on the floor.
turn on the ladder – The ladder can't be turned on.
order more nails online – I can't do that; none of my tools can help with that.

Usage

Start the server:

llama-server -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M --jinja -c 4096 -ngl 99 --host 127.0.0.1 --port 8091

Send a command:

import json, urllib.request

SYSTEM = open("system_prompt.txt", encoding="utf-8").read()

scene = {
    "player": {"room": "kitchen", "gaze": None, "left_hand": None, "right_hand": None},
    "locations": ["kitchen", "living_room"],
    "visible_objects": [
        {"id": "mug_1", "label": "mug", "color": "red", "on": "table_1", "pos": [0.4, 0.8, 1.2]},
        {"id": "mug_2", "label": "mug", "color": "blue", "on": "counter_1", "pos": [2.0, 0.9, 0.3]},
        {"id": "table_1", "label": "table", "on": "floor", "pos": [0.5, 0.0, 1.2]},
        {"id": "counter_1", "label": "counter", "on": "floor", "pos": [2.0, 0.0, 0.3]},
    ],
}
body = {
    "messages": [
        {"role": "system", "content": SYSTEM},
        {"role": "user", "content": "Scene: " + json.dumps(scene) + "\nRequest: put the red mug on the counter"},
    ],
    "temperature": 0,
    "max_tokens": 192,
}
req = urllib.request.Request("http://127.0.0.1:8091/v1/chat/completions",
                             json.dumps(body).encode(), {"Content-Type": "application/json"})
print(json.load(urllib.request.urlopen(req))["choices"][0]["message"]["content"])

The request loop:

  1. Execute each tool call the model returns.
  2. Send each result back as a tool message: {"ok": true}, {"ok": false, "error": "..."}, or a list of matches for find_objects.
  3. Call the model again, until it replies in plain text or calls ask_clarification.

Use a temperature of 0.

System prompt

The model was trained with this prompt; use it unchanged. The repository also contains it as system_prompt.txt.

You are an assistant embedded in a 3D Unity scene. Turn the user's request into tool calls that use object ids exactly as listed in the scene; never invent an id or a tool. Only objects in the user's current room are listed. Match objects by meaning: a synonym or a more general word can name a listed object, and details the scene does not record (a colour when no colours are listed, or words like sliced, wooden, empty) do not rule an object out. If exactly one listed object fits, act on it. If several fit equally well, call ask_clarification with their ids. If none fits, use find_objects to look in the other rooms. place moves an object directly, so there is no need to grab it first. If a request is impossible or needs a tool you do not have, call no tool and briefly explain why. Answer questions about the scene in plain text.

# Tools
grab(object_id:str, hand:left|right|both)
release(hand:left|right|both)
place(object_id:str, target_id:str, relation?:on|in, orientation?:default|upright|horizontal|vertical|upside_down)
move_by(object_id:str, direction:left|right|forward|backward|up|down, distance:num, unit:m|cm)
rotate(object_id:str, degrees:num, axis?:yaw|pitch|roll)
set_state(object_id:str, state:on|off|open|closed|half_open|locked|unlocked)
press(object_id:str, hold_seconds?:num)
highlight(object_ids:list)
teleport(location_id:str)
find_objects(label:str, color?:str, room?:str)
ask_clarification(question:str, candidate_ids:list)

Reply with one <tool_call>{"name": ..., "arguments": {...}}</tool_call> block per call, or plain text.

Output check

Before executing a response, check it against the scene and the tool list:

Output Action
A valid call Execute it
An unknown tool, or an argument value outside the schema Replace it with a refusal
An unknown object id, or a malformed call Regenerate it with a grammar built from the scene

The check changes about 0.5% of responses and removes every invented tool and id. All results on this page were measured with it; without it, the ALFRED score is 83.6%.

Unity

The Unity package (Unity 2022.3 or later) starts the server, builds the scene description, runs the request loop and applies the output check. It also adds editor tools for setting up and validating a scene.

Install it from the Package Manager with Add package from git URL:

https://github.com/mcbu-xrlab/Maverick-Unity.git#v1.0.0

The same package is mirrored on Hugging Face at https://huggingface.co/ErenAta00/Maverick-Unity.git.

Then copy the GGUF file from this repository and llama.cpp's llama-server into Assets/StreamingAssets/SceneAgent/. The GitHub repository has a quick start and integration guides.

Evaluation

All test sets were frozen before training, and the released checkpoint was chosen on separate development sets.

ALFRED

ALFRED (Shridhar et al., CVPR 2020) is a public benchmark of household tasks that crowd workers described step by step. We turned its pick-up, put-down and lamp steps into single commands for grab, place and set_state:

  • Request: the worker's sentence.
  • Scene: the objects in the room at that step.
  • Expected answer: the object used in ALFRED's recorded demonstration.

The test has 499 steps from the validation splits. An answer counts as correct if the first action calls the right tool on the right object. None of ALFRED's text appears in the training data.

In about 5% of the items, the worker's wording names a different object from the one ALFRED recorded, for example "knife" for a butter knife. No model can therefore reach 100%.

ALFRED results by model

Model Parameters All Pick up Put down Switch on New object types ¹
Maverick-4B-Unity-XR-Agent 4B 83.8 88.5 73.0 96.0 91.7
Qwen3-1.7B 1.7B 35.9 49.0 9.0 63.6 23.3
Qwen3-4B 4B 57.1 41.0 58.5 86.9 55.0
Ministral 3 14B 14B 67.1 56.0 67.0 89.9 66.7
gpt-oss-120b 120B 57.1 57.5 46.0 78.8 53.3
Nemotron-3 Super 120B 58.9 60.0 46.0 82.8 55.0

Accuracy in %. Items: 499 in total, 200 pick up, 200 put down, 99 switch on, 60 new object types.

¹ Object types that never appear in the training data. Of these 60 items, 19 use the word "cup", which the training data contains only as a synonym for mug.

Error types (% of all items):

Where each model loses points on ALFRED

Model Invented tool or id Question instead of an action Text instead of an action Wrong object Wrong tool or extra calls
Maverick-4B-Unity-XR-Agent 0.0 0.4 3.4 9.0 3.4
Qwen3-4B 3.6 21.8 3.8 9.8 3.8
Ministral 3 14B 1.8 2.0 18.6 7.4 3.0
gpt-oss-120b 0.2 5.8 25.3 10.4 1.2
Nemotron-3 Super 0.4 14.0 15.8 9.6 1.2

The general-purpose models lose most of their points by asking a question or answering in text when one object clearly fits. Wrong-object errors are similar across all models (7–10%), and most of them come from the labelling mismatch described above. A manual review of this model's wrong-object cases found about 1% genuine errors.

Capability tests

We wrote 2,900 commands in 22 categories from templates. They use object types and sentence patterns that do not appear in the training data. Most categories come in pairs: with one matching object the model should act, and with two it should ask.

Results by capability group

Capability Commands Example Accuracy (%)
Object names: unseen types, synonyms, general terms 600 "Spotlight the house plant" 97.8
Spatial references: on, in, near, left, right, front, behind, order 750 "the pen on this side of the cart" 93.9
Details the scene does not record: colour, material, condition 400 "Set the gray box atop the desk" 99.5
Asking when two objects fit equally well 150 "Highlight the pink trash can" 100.0
Searching other rooms 100 "Do you know where the newspaper went?" 98.0
Reporting an object that does not exist 100 "Where might the sponge be?" 100.0
Declining an impossible action without trying it 100 "Activate the remote" 75.0
"this", "that", "it" and held objects 100 "Draw attention to this folder" 98.0
Verbs that mean a change of state 100 "Power off the stove knob" 99.0
Verbs with more than one meaning 150 "Put the crate upside down" 99.3
Speech-recognition errors 150 "spotlight the apple that's on the counter" 97.3
Out-of-scope requests, such as cooking or cleaning 100 "Clean the credit card" 98.0
Unsupported edits: resizing, recolouring, deleting ² 100 "Paint the key purple" 93.0
All 2,900 96.4

² These requests were kept out of the training data so that the features can be added later. Without the output check, this category scores 68%.

Unity

End to end: we ran 440 of the capability commands through the Unity package in a live scene. 97.3% were carried out correctly, with no runtime errors.

Package tests: the package was tested in a new project and in a built Windows game. It passed 54 of 55 checks covering:

  • the editor tools;
  • all 11 tools;
  • follow-up answers and unusual input;
  • several rooms and scene reloads;
  • process cleanup.

The one failure was a room with 170 objects, which does not fit in the 4,096-token context; the package reports a clear error.

Speed and memory

These figures were measured on a laptop RTX 3050 Ti (4 GB) under Windows 11, with Unity rendering the scene at 30 fps on the same GPU.

Measurement Value
GPU memory (llama.cpp) 3.06 GB
Shared system memory 96 MiB
First answer, model only, median 1.18 s
Per command, GPU cool 0.5–3.6 s
Per command, GPU at its thermal limit, median 5.9 s

Most commands take two model calls: one for the action and one for the confirmation. If Unity shares the GPU with the model, cap Unity's frame rate. Without a cap, commands took about five times longer.

Training

Data: 20,091 English conversations generated from templates and a catalogue of 209 object types. We held back 28 of these types for testing.

Area Share
Core scene actions and follow-up questions 38%
Spatial references 18%
Deciding to act, ask, search or decline 14%
Object names 14%
Details the scene does not record 7%
Verb meanings 6%
References and held objects 2%
Unsupported requests 1%

Data checks:

  • Every conversation was executed in a simulator and had to produce the intended result.
  • 501 conversations that overlapped with a test set were removed.
  • About 8% of the conversations contain simulated speech-recognition errors.

Setup:

Method QLoRA (Unsloth), rank 16, alpha 32, all attention and MLP projections
Schedule 1 epoch, 2,512 steps, batch size 8, learning rate 2e-4, cosine decay
Sequence length 1,536 tokens
Hardware 1 × NVIDIA T4
Export Merged into the de-quantised 4-bit base, converted to GGUF Q8_0, then quantised to Q4_K_M

Limitations

  • Other rooms: the model only searches other rooms for find and fetch requests. If the lamp is in another room, "turn on the lamp" gets "There is no lamp here".
  • Typing errors: a misspelled name such as "lapm" is treated as an unknown object.
  • Impossible requests: the model sometimes tries the action first and explains only after the engine rejects it.
  • Real instructions: about one in six still needs a retry or a clarification.
  • Scene questions: the model answers questions about the scene in text, but these answers were not evaluated.
  • Scope: only the 11 tools. The model cannot create or delete objects, and it does not handle materials, lighting or physics.
  • Language: English only.
  • Input: the model reads the scene description, not the camera image.
  • Test coverage: tested on one Windows laptop GPU only; not yet on XR headsets, Linux or macOS.

Intended use

Suitable for:

  • voice or text control of objects in Unity XR, VR, simulation, training, showroom, digital-twin and game scenes;
  • accessibility tools.

Not suitable for:

  • robots or safety-critical systems;
  • open-ended chat;
  • irreversible actions without a confirmation step.

Safety and privacy

  • Local: the server listens only on 127.0.0.1, and the model sends no data off the device.
  • Speech: for fully offline voice input, pair the model with an offline speech recogniser such as whisper.cpp.
  • Actions: the model can only request the 11 tools, and your application decides what each tool does.
  • Training data: contains no personal data.

Citation

@inproceedings{shridhar2020alfred,
  title     = {{ALFRED}: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks},
  author    = {Shridhar, Mohit and Thomason, Jesse and Gordon, Daniel and Bisk, Yonatan and Han, Winson and
               Mottaghi, Roozbeh and Zettlemoyer, Luke and Fox, Dieter},
  booktitle = {CVPR},
  year      = {2020}
}

@misc{qwen3technicalreport,
  title         = {Qwen3 Technical Report},
  author        = {Qwen Team},
  year          = {2025},
  eprint        = {2505.09388},
  archivePrefix = {arXiv}
}

Contact

Developed by Eren Ata at the Extended Reality Laboratory (XRLab), Manisa Celal Bayar University. For questions and bug reports, open a discussion in the Community tab.

Downloads last month
206
GGUF
Model size
4B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF

Finetuned
Qwen/Qwen3-4B
Finetuned
(1131)
this model

Papers for ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF

Article mentioning ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF