Yuyao Ge
Add the model card and the overview figure
3a0fed1
|
Raw History Blame Contribute Delete
7.28 kB
---
license: mit
library_name: transformers
pipeline_tag: text-generation
base_model: Qwen/Qwen2.5-7B-Instruct
tags:
- agentic-rl
- llm-agent
- skill-library
- grpo
- search-qa
---
# SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles
[![arXiv](https://img.shields.io/badge/arXiv-2610.09832-b31b1b)](https://arxiv.org/abs/2610.09832)
[![PDF](https://img.shields.io/badge/PDF-full%20paper-white)](https://arxiv.org/pdf/2610.09832)
[![Code](https://img.shields.io/badge/Code-SkillForge-181717?logo=github)](https://github.com/YuyaoGe/SkillForge)
[![NeurIPS 2026](https://img.shields.io/badge/NeurIPS-2026-1783ff)](https://geyuyao.com/skillforge/)
[![Project page](https://img.shields.io/badge/project-SkillForge-b5613c)](https://geyuyao.com/skillforge/)
[![SkillForge collection](https://img.shields.io/badge/%F0%9F%A4%97-collection-ffc107)](https://huggingface.co/collections/YuyaoGe/skillforge)
[![License: MIT](https://img.shields.io/badge/license-MIT-f5de53)](https://huggingface.co/YuyaoGe/SkillForge_Search_7B/blob/main/README.md)
---
![SkillForge pipeline: seed skills are pre-retired under the base model, the survivors seed retirement-aware cold-start, then skills and the policy co-evolve through trial, active, stable and retired states.](assets/overview.png)
## Introduction
This is the **final SkillForge policy for Search-Augmented QA**: a full fine-tune of
`Qwen2.5-7B-Instruct` trained with GRPO while its skill library was simultaneously
forged β€” retired, stabilized, demoted and mutated β€” under the fitness-driven
lifecycle described in the paper.
Most memory-augmented agents keep the library **append-only**. A skill that was
correct at step 20 encodes a procedure the policy has outgrown by step 120, and it is
still being retrieved into the context. SkillForge instead scores each skill against
the policy's own rollouts and moves it between four states β€” `trial`, `active`,
`stable`, `retired` β€” so the library and the model co-evolve.
On Search-Augmented QA this reaches **48.7%** overall accuracy across 51,713 test
samples, the best of any method compared, against **46.8%** for SkillRL and **45.2%**
for ZeroSearch.
> **What is in this repository.** The policy weights and tokenizer only. The evolved
> skill library is *not* shipped here: the agent is the policy **plus** the retrieved
> skills injected into its context. Skill contents, fitness trajectories and
> retirement events are documented in SkillFurnace (Appendix C of the paper) and in
> the paper's case studies.
The same method, two more environments: [ALFWorld](https://huggingface.co/YuyaoGe/SkillForge_Alfworld_7B) and [WebShop](https://huggingface.co/YuyaoGe/SkillForge_Webshop_7B). All three sit in the
[SkillForge collection](https://huggingface.co/collections/YuyaoGe/skillforge), alongside the paper.
## Results
Search-Augmented QA, per dataset (%). \* = in-domain, \*\* = out-of-domain. The
`Overall` column is the sample-weighted micro-average over the same 51,713-sample
union for every method.
| Method | NQ\* | TriviaQA\*\* | PopQA\*\* | HotpotQA\* | 2Wiki\*\* | MuSiQue\*\* | Bamboogle\*\* | **Overall** |
|---|---|---|---|---|---|---|---|---|
| RAG | 27.4 | 58.2 | 17.8 | 25.8 | 23.2 | 9.4 | 16.8 | 29.4 |
| Search-R1 | 39.3 | 61.0 | 39.7 | 37.0 | 40.1 | 14.6 | 36.8 | 42.9 |
| ZeroSearch | 43.6 | 61.8 | **51.5** | 34.6 | 35.2 | 18.4 | 27.8 | 45.2 |
| EvolveR | 43.5 | 63.4 | 44.6 | 38.2 | 42.0 | 15.6 | 54.4 | 45.8 |
| SkillRL | 45.9 | 63.3 | 45.9 | 43.2 | 40.3 | 20.2 | 73.8 | 46.8 |
| **SkillForge** | **48.2** | **65.0** | 50.1 | **43.9** | 40.4 | **20.3** | **77.2** | **48.7** |
The largest single gains are out-of-domain: **+3.4** on Bamboogle and **+5.9** over
ZeroSearch on MuSiQue. ZeroSearch keeps PopQA, where it leads by 1.4.
### The library this model ended up with
| | Seed library | Final library |
|---|---|---|
| Total skills | 41 | **85** |
| General | 10 | β€” |
| Task-specific | 20 | β€” |
| Common-mistake | 11 | β€” |
The run saturates its skill cap (`S_max = 85`). This environment retires the
fewest skills of the three (59 events against ALFWorld's 138), and the two categories
that stand out here β€” redundant with the policy and contradicting the environment β€”
account for almost a quarter of them.
## How it was trained
1. **Pre-retirement.** The seed library, inherited from the SkillRL release, is scored
under 400 rollout episodes of the *base* model with skills injected. Anything whose
proto-fitness falls below `delta_pre = 0.3` is retired before training starts.
2. **Retirement-aware cold start.** The base model is fine-tuned with cross-entropy on
the rollouts that survived, with trajectories that leaned on a since-retired skill
filtered out.
3. **Skill-policy co-evolution.** GRPO takes over and, every 10 steps, the forging
cycle reads each skill's runtime fitness and promotes, demotes, retires or mutates
it. Mutation is LLM-guided (Kimi-K2.5 as teacher); at most 5 mutations and 3
retirements per cycle; retrieval is top-6 by task-type match.
| Hyperparameter | Value |
|---|---|
| Base model | `Qwen2.5-7B-Instruct` |
| Optimizer | GRPO, via `verl` |
| Learning rate | 1e-6 |
| Batch size / group size | 16 / 8 |
| Clip epsilon / KL beta | 0.2 / 0.001 |
| Training steps | 200 |
| Sampling temperature (train and eval) | 1.0 |
| Hardware | 64 x NVIDIA H200 |
## Quick Start
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "YuyaoGe/SkillForge_Search_7B"
tok = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="bfloat16", device_map="auto")
# Skills are retrieved (top-6, by task-type match) and injected into the context.
# The exact prompt template and the action space are in the paper.
messages = [
{"role": "system", "content": "<retrieved skills for this task type>"},
{"role": "user", "content": "<observation>\n\n> "},
]
ids = tok.apply_chat_template(messages, add_generation_prompt=True,
return_tensors="pt").to(model.device)
out = model.generate(ids, max_new_tokens=256, do_sample=False)
print(tok.decode(out[0][ids.shape[-1]:], skip_special_tokens=True))
```
The training loop, the skill lifecycle and the environment harnesses are in
[YuyaoGe/SkillForge](https://github.com/YuyaoGe/SkillForge) β€” but that is the code, not the runtime library: this
policy still needs its skills retrieved and injected at call time.
This is a research artifact: a 7B text policy for a text environment. It emits search
queries and answers in the Search-R1 action format; the retrieval backend is not part
of this repository.
## Citation
```bibtex
@misc{ge2026skillforgecoevolvingskillsagents,
title={SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles},
author={Yuyao Ge and Yiwei Wang and Yuchen He and Baolong Bi and Lingrui Mei and Jiayu Yao and Lizhe Chen and Shenghua Liu},
year={2026},
eprint={2610.09832},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2610.09832},
}
```
## Acknowledgments
Training runs on [`verl`](https://github.com/volcengine/verl) for the GRPO loop, with
seed skill libraries inherited from the SkillRL release.