Instructions to use ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
Use Docker
docker model run hf.co/ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
- Ollama
How to use ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF with Ollama:
ollama run hf.co/ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF with Docker Model Runner:
docker model run hf.co/ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
- Lemonade
How to use ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Maverick-4B-Unity-XR-Agent-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Maverick-4B-Unity-XR-Agent
Maverick-4B-Unity-XR-Agent turns spoken or typed English into tool calls for Unity 3D scenes. Say "put the red mug on the counter" and it returns the call your application executes. It can pick up, place, move and rotate objects, switch devices on and off, highlight objects, move the user between rooms and search other rooms. If a request could mean two objects, it asks which one. If it cannot do something, it says so.
The model is a fine-tune of Qwen3-4B. It is built for XR and VR applications such as training simulations, product showrooms, digital twins, accessibility tools and games.
The article Offline Voice Control for Unity XR Scenes explains how the model was trained and tested.
Results
We tested the model on 499 instructions that people wrote for the ALFRED benchmark. It chose the right action on the right object 83.8% of the time. Five other models, from 1.7B to 120B parameters, scored between 35.9% and 67.1% with the same scenes and tools. On object types that never appear in its training data, the model scored 91.7%. In a Unity test scene it carried out 440 commands with 97.3% accuracy and no runtime errors.
Model details
| Base model | Qwen3-4B |
| Method | QLoRA supervised fine-tuning, merged into the base weights |
| Training data | 20,091 English conversations |
| Language | English |
| Format | GGUF, Q4_K_M, 2.50 GB |
| Runtime | llama.cpp (llama-server, OpenAI-compatible API) |
| GPU memory | 3.06 GB |
| Context | 4,096 tokens, enough for a room of about 90 objects |
| Tools | grab, release, place, move_by, rotate, set_state, press, highlight, teleport, find_objects, ask_clarification |
| Licence | Apache-2.0, the same as the base model |
How it works
For each command, your application sends the model two things: a JSON description of the current room and the user's sentence. The room description lists every object's id, label, colour, state, support and position, as well as the rooms and what the user is holding and looking at.
The model answers with tool calls. Your application executes them and sends back the results, and the model confirms what it did in one sentence.
Examples from our Unity test scene:
| Command | Tool call | Reply |
|---|---|---|
| pick up the hammer | grab(hammer_1, right) |
Done. I picked up the hammer with your right hand. |
| move the chair 50 cm to the left | move_by(chair_1, left, 50, cm) |
Done. I moved the chair left by 50 cm. |
| um could you uh close the window | set_state(window_1, closed) |
Done. I closed the window. |
| pick up the toolbox (there are two) | ask_clarification([toolbox_1, toolbox_2]) |
Which toolbox do you mean, the blue or the red one? |
| highlight that (looking at a box) | highlight([cardboard_box_1]) |
Done. I highlighted the box. |
| where is the broom? | find_objects(broom) |
The broom is in the storage, on the floor. |
| turn on the ladder | – | The ladder can't be turned on. |
| order more nails online | – | I can't do that; none of my tools can help with that. |
Usage
Start the server:
llama-server -hf ErenAta00/Maverick-4B-Unity-XR-Agent-GGUF:Q4_K_M --jinja -c 4096 -ngl 99 --host 127.0.0.1 --port 8091
Send a command:
import json, urllib.request
SYSTEM = open("system_prompt.txt", encoding="utf-8").read()
scene = {
"player": {"room": "kitchen", "gaze": None, "left_hand": None, "right_hand": None},
"locations": ["kitchen", "living_room"],
"visible_objects": [
{"id": "mug_1", "label": "mug", "color": "red", "on": "table_1", "pos": [0.4, 0.8, 1.2]},
{"id": "mug_2", "label": "mug", "color": "blue", "on": "counter_1", "pos": [2.0, 0.9, 0.3]},
{"id": "table_1", "label": "table", "on": "floor", "pos": [0.5, 0.0, 1.2]},
{"id": "counter_1", "label": "counter", "on": "floor", "pos": [2.0, 0.0, 0.3]},
],
}
body = {
"messages": [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": "Scene: " + json.dumps(scene) + "\nRequest: put the red mug on the counter"},
],
"temperature": 0,
"max_tokens": 192,
}
req = urllib.request.Request("http://127.0.0.1:8091/v1/chat/completions",
json.dumps(body).encode(), {"Content-Type": "application/json"})
print(json.load(urllib.request.urlopen(req))["choices"][0]["message"]["content"])
The request loop:
- Execute each tool call the model returns.
- Send each result back as a
toolmessage:{"ok": true},{"ok": false, "error": "..."}, or a list of matches forfind_objects. - Call the model again, until it replies in plain text or calls
ask_clarification.
Use a temperature of 0.
System prompt
The model was trained with this prompt; use it unchanged. The repository also contains it as system_prompt.txt.
You are an assistant embedded in a 3D Unity scene. Turn the user's request into tool calls that use object ids exactly as listed in the scene; never invent an id or a tool. Only objects in the user's current room are listed. Match objects by meaning: a synonym or a more general word can name a listed object, and details the scene does not record (a colour when no colours are listed, or words like sliced, wooden, empty) do not rule an object out. If exactly one listed object fits, act on it. If several fit equally well, call ask_clarification with their ids. If none fits, use find_objects to look in the other rooms. place moves an object directly, so there is no need to grab it first. If a request is impossible or needs a tool you do not have, call no tool and briefly explain why. Answer questions about the scene in plain text.
# Tools
grab(object_id:str, hand:left|right|both)
release(hand:left|right|both)
place(object_id:str, target_id:str, relation?:on|in, orientation?:default|upright|horizontal|vertical|upside_down)
move_by(object_id:str, direction:left|right|forward|backward|up|down, distance:num, unit:m|cm)
rotate(object_id:str, degrees:num, axis?:yaw|pitch|roll)
set_state(object_id:str, state:on|off|open|closed|half_open|locked|unlocked)
press(object_id:str, hold_seconds?:num)
highlight(object_ids:list)
teleport(location_id:str)
find_objects(label:str, color?:str, room?:str)
ask_clarification(question:str, candidate_ids:list)
Reply with one <tool_call>{"name": ..., "arguments": {...}}</tool_call> block per call, or plain text.
Output check
Before executing a response, check it against the scene and the tool list:
| Output | Action |
|---|---|
| A valid call | Execute it |
| An unknown tool, or an argument value outside the schema | Replace it with a refusal |
| An unknown object id, or a malformed call | Regenerate it with a grammar built from the scene |
The check changes about 0.5% of responses and removes every invented tool and id. All results on this page were measured with it; without it, the ALFRED score is 83.6%.
Unity
The Unity package (Unity 2022.3 or later) starts the server, builds the scene description, runs the request loop and applies the output check. It also adds editor tools for setting up and validating a scene.
Install it from the Package Manager with Add package from git URL:
https://github.com/mcbu-xrlab/Maverick-Unity.git#v1.0.0
The same package is mirrored on Hugging Face at https://huggingface.co/ErenAta00/Maverick-Unity.git.
Then copy the GGUF file from this repository and llama.cpp's llama-server into Assets/StreamingAssets/SceneAgent/.
The GitHub repository has a quick start and integration guides.
Evaluation
All test sets were frozen before training, and the released checkpoint was chosen on separate development sets.
ALFRED
ALFRED (Shridhar et al., CVPR 2020) is a public benchmark of household tasks
that crowd workers described step by step. We turned its pick-up, put-down and lamp steps into single commands for
grab, place and set_state:
- Request: the worker's sentence.
- Scene: the objects in the room at that step.
- Expected answer: the object used in ALFRED's recorded demonstration.
The test has 499 steps from the validation splits. An answer counts as correct if the first action calls the right tool on the right object. None of ALFRED's text appears in the training data.
In about 5% of the items, the worker's wording names a different object from the one ALFRED recorded, for example "knife" for a butter knife. No model can therefore reach 100%.
| Model | Parameters | All | Pick up | Put down | Switch on | New object types ¹ |
|---|---|---|---|---|---|---|
| Maverick-4B-Unity-XR-Agent | 4B | 83.8 | 88.5 | 73.0 | 96.0 | 91.7 |
| Qwen3-1.7B | 1.7B | 35.9 | 49.0 | 9.0 | 63.6 | 23.3 |
| Qwen3-4B | 4B | 57.1 | 41.0 | 58.5 | 86.9 | 55.0 |
| Ministral 3 14B | 14B | 67.1 | 56.0 | 67.0 | 89.9 | 66.7 |
| gpt-oss-120b | 120B | 57.1 | 57.5 | 46.0 | 78.8 | 53.3 |
| Nemotron-3 Super | 120B | 58.9 | 60.0 | 46.0 | 82.8 | 55.0 |
Accuracy in %. Items: 499 in total, 200 pick up, 200 put down, 99 switch on, 60 new object types.
¹ Object types that never appear in the training data. Of these 60 items, 19 use the word "cup", which the training data contains only as a synonym for mug.
Error types (% of all items):
| Model | Invented tool or id | Question instead of an action | Text instead of an action | Wrong object | Wrong tool or extra calls |
|---|---|---|---|---|---|
| Maverick-4B-Unity-XR-Agent | 0.0 | 0.4 | 3.4 | 9.0 | 3.4 |
| Qwen3-4B | 3.6 | 21.8 | 3.8 | 9.8 | 3.8 |
| Ministral 3 14B | 1.8 | 2.0 | 18.6 | 7.4 | 3.0 |
| gpt-oss-120b | 0.2 | 5.8 | 25.3 | 10.4 | 1.2 |
| Nemotron-3 Super | 0.4 | 14.0 | 15.8 | 9.6 | 1.2 |
The general-purpose models lose most of their points by asking a question or answering in text when one object clearly fits. Wrong-object errors are similar across all models (7–10%), and most of them come from the labelling mismatch described above. A manual review of this model's wrong-object cases found about 1% genuine errors.
Capability tests
We wrote 2,900 commands in 22 categories from templates. They use object types and sentence patterns that do not appear in the training data. Most categories come in pairs: with one matching object the model should act, and with two it should ask.
| Capability | Commands | Example | Accuracy (%) |
|---|---|---|---|
| Object names: unseen types, synonyms, general terms | 600 | "Spotlight the house plant" | 97.8 |
| Spatial references: on, in, near, left, right, front, behind, order | 750 | "the pen on this side of the cart" | 93.9 |
| Details the scene does not record: colour, material, condition | 400 | "Set the gray box atop the desk" | 99.5 |
| Asking when two objects fit equally well | 150 | "Highlight the pink trash can" | 100.0 |
| Searching other rooms | 100 | "Do you know where the newspaper went?" | 98.0 |
| Reporting an object that does not exist | 100 | "Where might the sponge be?" | 100.0 |
| Declining an impossible action without trying it | 100 | "Activate the remote" | 75.0 |
| "this", "that", "it" and held objects | 100 | "Draw attention to this folder" | 98.0 |
| Verbs that mean a change of state | 100 | "Power off the stove knob" | 99.0 |
| Verbs with more than one meaning | 150 | "Put the crate upside down" | 99.3 |
| Speech-recognition errors | 150 | "spotlight the apple that's on the counter" | 97.3 |
| Out-of-scope requests, such as cooking or cleaning | 100 | "Clean the credit card" | 98.0 |
| Unsupported edits: resizing, recolouring, deleting ² | 100 | "Paint the key purple" | 93.0 |
| All | 2,900 | 96.4 |
² These requests were kept out of the training data so that the features can be added later. Without the output check, this category scores 68%.
Unity
End to end: we ran 440 of the capability commands through the Unity package in a live scene. 97.3% were carried out correctly, with no runtime errors.
Package tests: the package was tested in a new project and in a built Windows game. It passed 54 of 55 checks covering:
- the editor tools;
- all 11 tools;
- follow-up answers and unusual input;
- several rooms and scene reloads;
- process cleanup.
The one failure was a room with 170 objects, which does not fit in the 4,096-token context; the package reports a clear error.
Speed and memory
These figures were measured on a laptop RTX 3050 Ti (4 GB) under Windows 11, with Unity rendering the scene at 30 fps on the same GPU.
| Measurement | Value |
|---|---|
| GPU memory (llama.cpp) | 3.06 GB |
| Shared system memory | 96 MiB |
| First answer, model only, median | 1.18 s |
| Per command, GPU cool | 0.5–3.6 s |
| Per command, GPU at its thermal limit, median | 5.9 s |
Most commands take two model calls: one for the action and one for the confirmation. If Unity shares the GPU with the model, cap Unity's frame rate. Without a cap, commands took about five times longer.
Training
Data: 20,091 English conversations generated from templates and a catalogue of 209 object types. We held back 28 of these types for testing.
| Area | Share |
|---|---|
| Core scene actions and follow-up questions | 38% |
| Spatial references | 18% |
| Deciding to act, ask, search or decline | 14% |
| Object names | 14% |
| Details the scene does not record | 7% |
| Verb meanings | 6% |
| References and held objects | 2% |
| Unsupported requests | 1% |
Data checks:
- Every conversation was executed in a simulator and had to produce the intended result.
- 501 conversations that overlapped with a test set were removed.
- About 8% of the conversations contain simulated speech-recognition errors.
Setup:
| Method | QLoRA (Unsloth), rank 16, alpha 32, all attention and MLP projections |
| Schedule | 1 epoch, 2,512 steps, batch size 8, learning rate 2e-4, cosine decay |
| Sequence length | 1,536 tokens |
| Hardware | 1 × NVIDIA T4 |
| Export | Merged into the de-quantised 4-bit base, converted to GGUF Q8_0, then quantised to Q4_K_M |
Limitations
- Other rooms: the model only searches other rooms for find and fetch requests. If the lamp is in another room, "turn on the lamp" gets "There is no lamp here".
- Typing errors: a misspelled name such as "lapm" is treated as an unknown object.
- Impossible requests: the model sometimes tries the action first and explains only after the engine rejects it.
- Real instructions: about one in six still needs a retry or a clarification.
- Scene questions: the model answers questions about the scene in text, but these answers were not evaluated.
- Scope: only the 11 tools. The model cannot create or delete objects, and it does not handle materials, lighting or physics.
- Language: English only.
- Input: the model reads the scene description, not the camera image.
- Test coverage: tested on one Windows laptop GPU only; not yet on XR headsets, Linux or macOS.
Intended use
Suitable for:
- voice or text control of objects in Unity XR, VR, simulation, training, showroom, digital-twin and game scenes;
- accessibility tools.
Not suitable for:
- robots or safety-critical systems;
- open-ended chat;
- irreversible actions without a confirmation step.
Safety and privacy
- Local: the server listens only on 127.0.0.1, and the model sends no data off the device.
- Speech: for fully offline voice input, pair the model with an offline speech recogniser such as whisper.cpp.
- Actions: the model can only request the 11 tools, and your application decides what each tool does.
- Training data: contains no personal data.
Citation
@inproceedings{shridhar2020alfred,
title = {{ALFRED}: A Benchmark for Interpreting Grounded Instructions for Everyday Tasks},
author = {Shridhar, Mohit and Thomason, Jesse and Gordon, Daniel and Bisk, Yonatan and Han, Winson and
Mottaghi, Roozbeh and Zettlemoyer, Luke and Fox, Dieter},
booktitle = {CVPR},
year = {2020}
}
@misc{qwen3technicalreport,
title = {Qwen3 Technical Report},
author = {Qwen Team},
year = {2025},
eprint = {2505.09388},
archivePrefix = {arXiv}
}
Contact
Developed by Eren Ata at the Extended Reality Laboratory (XRLab), Manisa Celal Bayar University. For questions and bug reports, open a discussion in the Community tab.
- Downloads last month
- 206
4-bit




