ES RL Model and Trajectories

This repository contains the step-20 Hugging Face export for qwen35-4b-math-grpo-dapo400-e15-seed1, the corresponding RL trajectories, and the latest evaluation outputs.

Contents

  • models/qwen35-4b-math-grpo-dapo400-e15-seed1-step20/: Hugging Face model export with 4 safetensors shards, tokenizer files, and generation/config metadata.
  • models/qwen35-4b-docvqa-grpo-four-gpu-from-step40-seed42-step120/: DocVQA Hugging Face model export with 4 safetensors shards, tokenizer files, and generation/config metadata.
  • trajectories/qwen35-4b-math-grpo-dapo400-e15-seed1/raw/: raw train and validation trajectory JSONL files.
  • trajectories/qwen35-4b-math-grpo-dapo400-e15-seed1/exports/attempt2_step001_020/: exported train/validation trajectory JSONL files for steps 1-20.
  • trajectories/qwen35-4b-docvqa-grpo-four-gpu-from-step40-seed42/trace2skill_collection_step80/: DocVQA trace2skill trajectory collection outputs, including RL rollouts and sample-only trajectories.
  • eval/qwen35-4b-math-grpo-dapo400-e15-seed1/step20_eval_50x4096_rerun_20260805/: evaluation summary and per-sample outputs.
  • eval/qwen35-4b-docvqa-grpo-four-gpu-from-step40-seed42/: DocVQA post-eval outputs, trace2skill sample-only eval outputs, and RL curve CSV.

Evaluation

Evaluation was run with 4 samples per item, temperature 1.0, top-p 1.0, top-k 40, presence penalty 2.0, and paper_react_cli mode.

Dataset Records Mean Score Max@4
DAPO100 400 0.6875 0.8300
AIME2026 120 0.5833 0.7667
DocVQA post-eval 400 0.4522 0.5147

DocVQA trace2skill sample-only evaluation used 8 samples per item on 50 evolve items and reached mean score 0.4338 with max@8 0.7585.

The full optimizer state, Megatron/FSDP training checkpoint files, tool workspaces, server logs, and cache directories are not included in this repo.

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading