SFT-4B-LiteOS-MixedOpen

Qwen/Qwen3-VL-4B-Instruct supervised-finetuned on lite.osworld trajectories produced entirely by open-source models — no GPT data. Final checkpoint (epoch 3, 912 steps).

Serves as the init / reference model for HaoranLiu/DPO-LiteOS-SourceAblation.

Training data

Source HaoranLiu/LiteOS, desktop/use/train (train.synth)
Trajectories 1217, one per task
Sources Qwen3.8-27B (405) / Qwen3.5-27B (406) / EvoCUA-8B (406)
Task set tasks solved by both gpt and >=1 of these three, so neither arm of the source comparison gets a coverage advantage
Filter not exclude_reason and episode_return > 0.5
Render config scripts/configs/qwen3_vl/default/lite.osworld.yaml (4-image history, max_steps 30)

Hyperparameters

Epochs 3 (912 steps)
LR 5e-6 cosine to 1e-6, warmup fraction 0.1
Batch GLOBAL_BATCH_SIZE 4 trajectories/step, MBS 1
Adam 0.9 / 0.95, weight_decay 0.1, clip_grad 1.0
Precision bf16, seq_length 4096, seed 1234
Parallelism TP=2 PP=1 CP=1 -> DP=1 on 2xH100, OPTIM_CPU_OFFLOAD=1 (required)
Loss sft_loss, per-token

Results

lite.osworld eval, 332 tasks (369 minus exclude_reason), greedy, max_steps 30, group_size 1, mean over all 332.

model data steps mean succ succ_rate
HaoranLiu/SFT-4B-LiteOS (gpt-5.5 teacher) 1967 traj, perturb+synth 1472 0.3220 103 31.02%
epoch 1 1217 traj, open-source only 304 0.2043 64 19.28%
epoch 2 " 608 0.3213 101 30.42%
epoch 3 (this checkpoint) " 912 0.3439 108 32.53%

Two things to note when reading this:

It beats the single-GPT-teacher reference with 38% less data. 0.3439 vs 0.3220, +5 tasks, from 1217 trajectories instead of 1967 — and none of them from GPT. A smaller training set can only depress this arm, so the direction is not a data-volume artifact.

The epoch curve is steep and still rising. 0.2043 -> 0.3213 -> 0.3439 (+0.117, then +0.023). Epoch 1 is far below what the reference reaches, and it takes until epoch 3 to pass it. Training loss meanwhile falls 0.221 -> 0.144 -> 0.095, i.e. the model is memorising while eval is still improving. Open-source trajectories are longer (11.62 steps mean vs GPT's 8.36) and more meandering, so more passes are needed to extract the same competence. 3 epochs may not be the ceiling; it was capped to leave budget for DPO.

Caveats

  • Single seed, one sample per task at temperature 0. Repeated rollouts of one checkpoint move the mean by ~0.0035.
  • train.synth and the eval split are similar in construction, so some of the gain may be memorisation rather than capability. No base-model number was measured on this eval, so it is not established how much of the absolute score SFT contributes.
  • The reference (SFT-4B-LiteOS) used a different cohort mix (perturb+synth vs synth-only), so it is not a strict one-variable comparison. A matched GPT-only arm on this exact pool was not run.
Downloads last month
8
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HaoranLiu/SFT-4B-LiteOS-MixedOpen

Finetuned
(467)
this model