Qwen3-4B-Instruct-2507 · IFStruct LoRA adapter

A LoRA adapter (rank 8, alpha 16, all linear layers) for Qwen/Qwen3-4B-Instruct-2507, trained with reinforcement learning (GRPO) to follow structured-output instructions: emit valid JSON/YAML that matches a requested schema, wrapper key, item count, code-block and no-commentary constraints.

IFStruct v1.0 result

Scored on the full 2,000-row test set of LiquidAI/ifstruct-v1.0 with the benchmark's own validator (byte-identical to Liquid4All/ifstruct).

Model pass@1 JSON YAML
this adapter 96.45 96.6 96.3
Qwen3-4B-Instruct-2507 (base; greedy, earlier run at max_tokens 4096) 76.15 76.4 75.9

Eval setup: greedy decoding (temperature 0), 1 sample per prompt, max_tokens 16000, no system prompt, the model's chat template (non-thinking model), no constrained decoding, vLLM 0.24. Full per-row outputs (prompt, response, validator errors) are in eval/.

Remaining failures (71/2000) are concentrated in prompts whose payload is long free text with heavy escaping (email threads, terminal sessions, dialogue). 18 of them are degenerate repetition loops inside a string value that run to the token cap.

Training

  • Method: GRPO (verl 0.9.0), CISPO policy loss, KL loss (low_var_kl, 0.01) to the base model, zero-variance group filtering, 16 prompts × 16 rollouts per step, max response 4096 tokens.
  • Reward: binary pass/fail from the IFStruct validator. No judge model.
  • Stage 1: fresh LoRA, lr 1e-4, rollout temperature 1.0 → 1.5, 105 steps on 2,406 synthetic prompts.
  • Stage 2 (this checkpoint): warm start from stage 1, lr 5e-5, rollout temperature 1.5, step 110 on 4,293 synthetic prompts (2 epochs).
  • Training prompts are synthetic. None of the 2,000 benchmark prompts or their entity types appear in the training data (checked by exact prompt and prompt-prefix match).

Disclosure: how the benchmark was used

This is not a blind held-out score. The benchmark influenced development in two ways:

  1. Data targeting. The synthetic data generator was steered toward the failure types the model showed on this benchmark (failure-mode mix matched to the previous checkpoint's benchmark failures). No benchmark rows were copied into training.
  2. Checkpoint selection. Checkpoints were compared on validation slices drawn from this benchmark (≈288 of the 2,000 rows) and on full-benchmark evaluations.

Expect a lower number on genuinely unseen structured-output distributions.

Usage

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base = "Qwen/Qwen3-4B-Instruct-2507"
tok = AutoTokenizer.from_pretrained(base)
model = AutoModelForCausalLM.from_pretrained(base, torch_dtype="auto", device_map="auto")
model = PeftModel.from_pretrained(model, "ichetandhembre/ifstruct-lora-adaptor")

With vLLM: vllm serve Qwen/Qwen3-4B-Instruct-2507 --enable-lora --max-lora-rank 8 --lora-modules ifstruct=ichetandhembre/ifstruct-lora-adaptor.

Downloads last month
15
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ichetandhembre/ifstruct-lora-adaptor

Adapter
(5796)
this model