Maverick 1.0

An open source System 1 Decision Model for Flight.

Watch/play the flight simulator here  •  Source Code  •  Benchmark  •  Usage

Flight simulator environment with an objective to fly through 5 rings placed at random altitudes and rotations. Flying a Cessna 172, Maverick 1.0 finishes in 5.9 minutes, Jev and Qwen3 0.6B time out, and Laya crashes. Shown at 8× speed.

Flying a quadcopter drone, Maverick 1.0 finishes in 9.7 minutes, Jev finishes in 14.0 minutes, and Laya and Qwen3 0.6B time out. Shown at 13× speed.

The release of Jev, a System 1 model, has made quick, low-latency decision-making models popular. We present Maverick 1.0, a general-purpose flight controller that is able to fly both a Cessna 172 airplane and a quadcopter drone in a flight simulator, across a variety of objectives such as taking off, flying through rings (easy), flying through rings (hard), and flying a full circuit to a landing. The Cessna flies on the JSBSim physics engine. We used imitation learning (DAgger) and reinforcement learning (PPO) to train Qwen3 0.6B. Maverick 1.0 gets ultra-low latency on a GPU, faster than Jev, and still good-enough-to-fly latency on a CPU.

Benchmark

Model Successful flights Time per decision
Cessna 172 Drone
Maverick 1.0
Q8_0 · GPU
14 / 15 (93%)15 / 15 (100%)83 ms
Maverick 1.0
Q8_0 · CPU
13 / 15 (87%)15 / 15 (100%)626 ms
Jev
API
6 / 15 (40%)6 / 15 (40%)201 ms
Laya 0.4B
GPU
3 / 15 (20%)3 / 15 (20%)86 ms
Qwen3 0.6B (base)
Q8_0 · GPU
3 / 15 (20%)0 / 15 (0%)83 ms

Each model flew 15 flights in each aircraft: five objectives, three seeds each. The simulator runs in real time and never waits for the model, so a slow answer is like a pilot with slow reactions. A flight is successful when it completes its objective. Time per decision is the median time to answer one step, on a consumer GPU and a 6-core desktop CPU; Jev's includes the network.

  • The videos above show flying through rings in hard mode, which we never trained on.
  • Maverick 1.0 is robust to decision-making time. We trained it with answers arriving anywhere from 50 ms to 2.5 s late, so it flies almost as well on a CPU (626 ms per decision) as on a GPU (83 ms).

Training approach

We trained a small LoRA adapter, 1.7% of the model's size, on top of Qwen3 0.6B. This is how we trained it:

  1. Imitation learning (DAgger). Maverick first learns from a scripted autopilot. It flies on its own, the autopilot labels every situation Maverick gets into with the right answer, and Maverick retrains on those labels. Each label is the right answer for the moment Maverick's answer actually arrives, so it learns to answer ahead of its own delay.
  2. Reinforcement learning (PPO). Maverick flies thousands of practice flights. It earns reward for completing objectives and a little for progress toward them, and loses reward for crashes, hard landings and wasted time. The reward counts seconds of flight, not decisions, so it treats fast and slow models the same way.
  3. Robustness. More PPO in harder conditions: different airfields, heavier or weaker aircraft, gusts, noisy readings, and answers that arrive anywhere from 50 ms to 2.5 s late.

We then merged the adapter into Qwen3 0.6B and saved it as an 8-bit GGUF.

Example state passed in

At every step, the simulator passes the model the same kind of text prompt: the goal, the questions and the current state of the flight (numbers, no pixels). Maverick answers each question with one letter. It doesn't generate any text: we read the logits of the model's next token and pick the most likely option letter. The aircraft then holds those answers (a heading, a climb rate and a speed) until the next step.

The following is a state and questions passed to the model (aircraft: Cessna, objective: rings):

situation: flying to ring 3 of 4
airspeed_kt: 81
ground_speed_kt: 76
height_m: 150
vertical_speed_fpm: -80
pitch: nose 2 deg up
bank: wings level
heading_deg: 320
stall_warning: no
current_controls: turn=holding heading 316, climb=level, speed=80 kt
target: ring=3 of 4, distance_m=1500, direction=50 deg to the left, height_m=200, height_difference=50 m above you, climb_needed_fpm=260
wind: from 304 deg at 6.1 kt
Question Options Maverick 1.0 answer Qwen3 0.6B (base) answer
turn A. left 60° · B. left 20° · C. left 5° · D. straight · E. right 5° · F. right 20° · G. right 60° A. left 60° C. left 5°
climb A. climb fast · B. climb · C. level · D. flare · E. descend slowly · F. descend · G. descend fast B. climb D. flare
speed A. idle · B. 65 kt · C. 80 kt · D. 100 kt C. 80 kt C. 80 kt

Ring 3 is 50° to the left and 50 m higher. Maverick turns hard toward it and climbs. The base model barely turns and flares as if to land.

The full prompt for this step
<|im_start|>system
You are the decision maker in an environment. Read the goal, the questions and the current state, then answer every question with the letter of the option that best achieves the goal.<|im_end|>
<|im_start|>user
Goal: Take off, then fly through 4 rings in order. Each ring is a vertical circle of 60 m radius, passed by flying through it in its heading. No landing needed. The aircraft is a Cessna 172 airplane: it lifts off at about 55 kt after a takeoff roll, stalls below about 48 kt and lands on a runway at 55-80 kt.

Questions:
turn: Which way should the aircraft turn now? Each turn moves the heading it holds by that much.
  A. left 60°: a big turn left
  B. left 20°: turn left
  C. left 5°: a small correction left
  D. straight: hold the current heading
  E. right 5°: a small correction right
  F. right 20°: turn right
  G. right 60°: a big turn right
climb: Which climb rate or landing flare should the aircraft hold?
  A. climb fast: +1000 ft/min
  B. climb: +500 ft/min
  C. level: hold the current height
  D. flare: raise the nose to touch down gently at about -100 ft/min
  E. descend slowly: -250 ft/min
  F. descend: -500 ft/min
  G. descend fast: -1000 ft/min
speed: Which airspeed should the autothrottle hold?
  A. idle: no power: to flare, land or descend steeply
  B. 65 kt: approach and slow flight
  C. 80 kt: climbing
  D. 100 kt: cruise

State:
situation: flying to ring 3 of 4
airspeed_kt: 81
ground_speed_kt: 76
height_m: 150
vertical_speed_fpm: -80
pitch: nose 2 deg up
bank: wings level
heading_deg: 320
stall_warning: no
current_controls: turn=holding heading 316, climb=level, speed=80 kt
target: ring=3 of 4, distance_m=1500, direction=50 deg to the left, height_m=200, height_difference=50 m above you, climb_needed_fpm=260
wind: from 304 deg at 6.1 kt<|im_end|>
<|im_start|>assistant
<think>

</think>

turn: A
climb: B
speed: C

Usage

In our jym flight simulator environment (JSBSim physics engine)

jym is open source, and Maverick 1.0 is its default model.

git clone https://github.com/kashmoneygt/projectj && cd projectj/jym
uv run jym ui

Open http://127.0.0.1:8765 to watch Maverick race the autopilot, race other models, or fly yourself. The model downloads from this page the first time. To run it on a GPU: uv run jym ui --players maverick@gpu,reference.

With llama.cpp

Maverick 1.0 works with both of llama-server's endpoints: the usual completion endpoint, and the new decision model endpoint. Start the server, which downloads the model the first time:

llama-server -hf projectj/Flight-Autopilot-Decision-Model-Maverick1.0:Q8_0

Completion endpoint (/completion). This gives the exact answers Maverick was trained to give, and it is how jym runs it. Send the prompt up to a question's name, like the full prompt above up to turn:, and pick the most likely option letter for the next token:

import requests

prompt = open("prompt.txt").read().rstrip()  # the full prompt above, up to "turn:"
reply = requests.post("http://127.0.0.1:8080/completion", json={"prompt": prompt, "n_predict": 1, "n_probs": 10})
print(reply.json()["completion_probabilities"][0]["top_logprobs"][0]["token"])  # " A", so left 60°

Then add the letter and the next question's name, turn: A\nclimb:, and ask again. jym.logits builds these prompts.

Decision model endpoint (/v1/systemone, llama.cpp b11364 or newer). This answers all the questions in one request, with the System One API introduced with Jev, so a Jev client only needs a new base URL:

curl http://127.0.0.1:8080/v1/systemone -H "Content-Type: application/json" -d @request.json

The request holds the state, as a JSON object with its goal, and the questions. For the example above, Maverick answers left 60°, climb and 80 kt.

request.json for the example above
{
  "state": {
    "goal": "Take off, then fly through 4 rings in order. Each ring is a vertical circle of 60 m radius, passed by flying through it in its heading. No landing needed. The aircraft is a Cessna 172 airplane: it lifts off at about 55 kt after a takeoff roll, stalls below about 48 kt and lands on a runway at 55-80 kt.",
    "situation": "flying to ring 3 of 4",
    "airspeed_kt": 81,
    "ground_speed_kt": 76,
    "height_m": 150,
    "vertical_speed_fpm": -80,
    "pitch": "nose 2 deg up",
    "bank": "wings level",
    "heading_deg": 320,
    "stall_warning": false,
    "current_controls": {
      "turn": "holding heading 316",
      "climb": "level",
      "speed": "80 kt"
    },
    "target": {
      "ring": "3 of 4",
      "distance_m": 1500,
      "direction": "50 deg to the left",
      "height_m": 200,
      "height_difference": "50 m above you",
      "climb_needed_fpm": 260
    },
    "wind": "from 304 deg at 6.1 kt"
  },
  "questions": {
    "turn": {
      "type": "choice",
      "instructions": "Which way should the aircraft turn now? Each turn moves the heading it holds by that much.",
      "criteria": {
        "left 60°": "a big turn left",
        "left 20°": "turn left",
        "left 5°": "a small correction left",
        "straight": "hold the current heading",
        "right 5°": "a small correction right",
        "right 20°": "turn right",
        "right 60°": "a big turn right"
      }
    },
    "climb": {
      "type": "choice",
      "instructions": "Which climb rate or landing flare should the aircraft hold?",
      "criteria": {
        "climb fast": "+1000 ft/min",
        "climb": "+500 ft/min",
        "level": "hold the current height",
        "flare": "raise the nose to touch down gently at about -100 ft/min",
        "descend slowly": "-250 ft/min",
        "descend": "-500 ft/min",
        "descend fast": "-1000 ft/min"
      }
    },
    "speed": {
      "type": "choice",
      "instructions": "Which airspeed should the autothrottle hold?",
      "criteria": {
        "idle": "no power: to flare, land or descend steeply",
        "65 kt": "approach and slow flight",
        "80 kt": "climbing",
        "100 kt": "cruise"
      }
    }
  }
}

This endpoint answers each question on its own and reads the option letters a little differently from training. On 1,868 recorded decisions, it picks the same option as the completion endpoint 95% of the time, and when it differs, it picks a neighbouring option, such as climb instead of climb fast.

With transformers

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("Qwen/Qwen3-0.6B")
base = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-0.6B")
model = PeftModel.from_pretrained(base, "projectj/Flight-Autopilot-Decision-Model-Maverick1.0", subfolder="adapter")

with model.disable_adapter():
    ...  # the same model is the base Qwen3 0.6B here

Prompt it the same way, and read each answer from the logits of the option letters. The adapter is only 40 MB, so you can hot-swap it: keep one Qwen3 0.6B in memory and switch the flight controller on only when it has to fly. In the example above, the model answers A with the adapter and C without it.

With other flight simulators

To fly Maverick in another flight simulator, such as X-Plane, FlightGear or Microsoft Flight Simulator, build a bridge: read the aircraft's position, speed and attitude from the simulator's API, build the text state, and send the answers to the simulator's autopilot as heading, vertical speed and airspeed holds.

Files

File Description
Flight-Autopilot-Decision-Model-Maverick1.0-Q8_0.gguf Qwen3 0.6B with Maverick's adapter merged in, 8-bit, for llama.cpp and its /v1/systemone API (640 MB)
adapter/ Maverick's LoRA adapter (rank 16) for Qwen3 0.6B, in PEFT format (40 MB)

Limitations

  • It has only flown in simulation. A real aircraft needs a safety pilot and the platform's own safety features.
  • It reads numbers, not pixels. A camera-based setup has to turn images into the same state.

License and citation

Apache-2.0. Maverick 1.0 is an RL-trained fine-tune of Qwen3 0.6B (Apache-2.0).

@misc{maverick1.0,
  title  = {Maverick 1.0: an open source System 1 decision model for flight},
  author = {The projectj authors},
  year   = {2026},
  url    = {https://huggingface.co/projectj/Flight-Autopilot-Decision-Model-Maverick1.0}
}
Downloads last month
-
GGUF
Model size
0.6B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

8-bit

Video Preview
loading

Model tree for projectj/Flight-Autopilot-Decision-Model-Maverick1.0

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1332)
this model

Papers for projectj/Flight-Autopilot-Decision-Model-Maverick1.0