Agent-Reflex-8B

Paper GitHub SFT Data RL Data

Agent-Reflex-8B is Qwen3-8B trained with Agent-Reflex, a framework that elicits proactive reflection in tool-using agents: noticing noisy tool feedback and user corrections, and recovering from them.

Training: Qwen3-8B → ReSFT (54,663 reflective trajectories, 5,127 steps) → ORPO (noisy agentic RL with skill-guided on-policy distillation). Skills are used only during training; the model runs skill-free.

Results

Model VitaBench τ²-Bench BFCL V4 ACEBench ToolSandbox Overall AgentNoiseBench
Qwen3-8B 11.4 26.2 40.4 54.0 69.1 40.2 12.3
EnvScaler-8B 15.8 37.9 47.6 60.0 80.5 48.4 17.6
Agent-Reflex-8B 17.5 54.9 48.9 76.0 81.8 55.8 21.4

Usage

Serve with vLLM (native function calling, thinking on), as in our evaluation:

vllm serve dongguanting/Agent-Reflex-8B --served-model-name Qwen/Qwen3-8B \
  --tensor-parallel-size 8 --max-model-len 40960 \
  --enable-auto-tool-choice --tool-call-parser hermes --reasoning-parser deepseek_r1

The full evaluation suite (τ²-Bench, BFCL V4, ACEBench, VitaBench, AgentNoiseBench, ToolSandbox) is in the GitHub repo.

Citation

@article{agentreflex2026,
  title  = {Agent-Reflex: Teaching Language Agents to Act Reflectively},
  author = {Anonymous},
  year   = {2026}
}
Downloads last month
16
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for dongguanting/Agent-Reflex-8B

Finetuned
Qwen/Qwen3-8B
Finetuned
(2170)
this model
Quantizations
2 models

Datasets used to train dongguanting/Agent-Reflex-8B

Collection including dongguanting/Agent-Reflex-8B