SFT-4B-LiteOS-MixedOpen
Qwen/Qwen3-VL-4B-Instruct supervised-finetuned on lite.osworld trajectories produced
entirely by open-source models — no GPT data. Final checkpoint (epoch 3, 912 steps).
Serves as the init / reference model for HaoranLiu/DPO-LiteOS-SourceAblation.
Training data
| Source | HaoranLiu/LiteOS, desktop/use/train (train.synth) |
| Trajectories | 1217, one per task |
| Sources | Qwen3.8-27B (405) / Qwen3.5-27B (406) / EvoCUA-8B (406) |
| Task set | tasks solved by both gpt and >=1 of these three, so neither arm of the source comparison gets a coverage advantage |
| Filter | not exclude_reason and episode_return > 0.5 |
| Render config | scripts/configs/qwen3_vl/default/lite.osworld.yaml (4-image history, max_steps 30) |
Hyperparameters
| Epochs | 3 (912 steps) |
| LR | 5e-6 cosine to 1e-6, warmup fraction 0.1 |
| Batch | GLOBAL_BATCH_SIZE 4 trajectories/step, MBS 1 |
| Adam | 0.9 / 0.95, weight_decay 0.1, clip_grad 1.0 |
| Precision | bf16, seq_length 4096, seed 1234 |
| Parallelism | TP=2 PP=1 CP=1 -> DP=1 on 2xH100, OPTIM_CPU_OFFLOAD=1 (required) |
| Loss | sft_loss, per-token |
Results
lite.osworld eval, 332 tasks (369 minus exclude_reason), greedy, max_steps 30,
group_size 1, mean over all 332.
| model | data | steps | mean | succ | succ_rate |
|---|---|---|---|---|---|
HaoranLiu/SFT-4B-LiteOS (gpt-5.5 teacher) |
1967 traj, perturb+synth | 1472 | 0.3220 | 103 | 31.02% |
| epoch 1 | 1217 traj, open-source only | 304 | 0.2043 | 64 | 19.28% |
| epoch 2 | " | 608 | 0.3213 | 101 | 30.42% |
| epoch 3 (this checkpoint) | " | 912 | 0.3439 | 108 | 32.53% |
Two things to note when reading this:
It beats the single-GPT-teacher reference with 38% less data. 0.3439 vs 0.3220, +5 tasks, from 1217 trajectories instead of 1967 — and none of them from GPT. A smaller training set can only depress this arm, so the direction is not a data-volume artifact.
The epoch curve is steep and still rising. 0.2043 -> 0.3213 -> 0.3439 (+0.117, then +0.023). Epoch 1 is far below what the reference reaches, and it takes until epoch 3 to pass it. Training loss meanwhile falls 0.221 -> 0.144 -> 0.095, i.e. the model is memorising while eval is still improving. Open-source trajectories are longer (11.62 steps mean vs GPT's 8.36) and more meandering, so more passes are needed to extract the same competence. 3 epochs may not be the ceiling; it was capped to leave budget for DPO.
Caveats
- Single seed, one sample per task at temperature 0. Repeated rollouts of one checkpoint move the mean by ~0.0035.
train.synthand the eval split are similar in construction, so some of the gain may be memorisation rather than capability. No base-model number was measured on this eval, so it is not established how much of the absolute score SFT contributes.- The reference (
SFT-4B-LiteOS) used a different cohort mix (perturb+synth vs synth-only), so it is not a strict one-variable comparison. A matched GPT-only arm on this exact pool was not run.
- Downloads last month
- 8
Model tree for HaoranLiu/SFT-4B-LiteOS-MixedOpen
Base model
Qwen/Qwen3-VL-4B-Instruct