AdithyaSK's picture
AdithyaSK HF Staff
Fix disconnected evaluation lines and GIF previews
a0bcdab verified
|
Raw History Blame Contribute Delete
8.96 kB
---
base_model: LiquidAI/LFM2.5-2.6B
base_model_relation: finetune
library_name: transformers
pipeline_tag: text-generation
language:
- en
license: other
tags:
- trl
- openenv
- harbor
- agent
- smoldataenvs
- grpo
- opencode
datasets:
- FineEnvs/SmolDataEnvs-harbor-train
thumbnail: https://huggingface.co/spaces/AdithyaSK/multi-harness-rl/resolve/main/app/public/og/og-image.png
model-index:
- name: LFM2.5-2.6B-opencode-RL (main, step 1000)
results:
- task:
type: text-generation
name: Agentic data analysis
dataset:
type: FineEnvs/SmolDataEnvs-harbor-test
name: SmolDataEnvs (250 fixed test tasks)
split: test
args:
harness: all four
metrics:
- type: pass_at_1
name: Pass@1 (%) — Overall
value: 52.3
- task:
type: text-generation
name: Agentic data analysis
dataset:
type: FineEnvs/SmolDataEnvs-harbor-test
name: SmolDataEnvs (250 fixed test tasks)
split: test
args:
harness: opencode
metrics:
- type: pass_at_1
name: Pass@1 (%) — OpenCode
value: 58.0
- task:
type: text-generation
name: Agentic data analysis
dataset:
type: FineEnvs/SmolDataEnvs-harbor-test
name: SmolDataEnvs (250 fixed test tasks)
split: test
args:
harness: claude-code
metrics:
- type: pass_at_1
name: Pass@1 (%) — Claude Code
value: 42.0
- task:
type: text-generation
name: Agentic data analysis
dataset:
type: FineEnvs/SmolDataEnvs-harbor-test
name: SmolDataEnvs (250 fixed test tasks)
split: test
args:
harness: codex
metrics:
- type: pass_at_1
name: Pass@1 (%) — Codex
value: 43.2
- task:
type: text-generation
name: Agentic data analysis
dataset:
type: FineEnvs/SmolDataEnvs-harbor-test
name: SmolDataEnvs (250 fixed test tasks)
split: test
args:
harness: mini-swe-agent
metrics:
- type: pass_at_1
name: Pass@1 (%) — Mini-SWE-Agent
value: 66.0
license_name: lfm1.0
license_link: LICENSE
---
[![The ultimate guide to multi-harness RL](https://huggingface.co/spaces/AdithyaSK/multi-harness-rl/resolve/main/app/public/og/og-image.png)](https://huggingface.co/spaces/AdithyaSK/multi-harness-rl)
# LFM2.5-2.6B-opencode-RL
A full fine-tune of [LiquidAI/LFM2.5-2.6B](https://huggingface.co/LiquidAI/LFM2.5-2.6B) for agentic data-analysis tasks, trained with asynchronous GRPO (TRL Async GRPO) using **OpenCode**. This release is **step 1,000** on **`main`**, scoring **52.3% pass@1** across four evaluation harnesses.
[Article](https://huggingface.co/spaces/AdithyaSK/multi-harness-rl) · [Collection](https://huggingface.co/collections/FineEnvs/multi-harness-rl-6abdfaaa8d74dacd481d5212) · [Evaluation tasks](https://huggingface.co/datasets/FineEnvs/SmolDataEnvs-harbor-test) · [Training dashboard](https://huggingface.co/spaces/FineEnvs/data-agent-training-comparison-trackio)
## Training and evaluation curves
[![Animated training, pass@1 and tool-use curves through step 1,000](https://huggingface.co/FineEnvs/LFM2.5-2.6B-opencode-RL/resolve/main/assets/curves.gif?v=2)](https://huggingface.co/spaces/AdithyaSK/multi-harness-rl)
Training curves use a trailing 50-step mean (at least 10 observations). Evaluation markers show measured checkpoints; hollow markers and dotted segments indicate incomplete coverage. Faint background lines show the full recorded trajectory; colored lines reveal the measured checkpoints. The first frame shows the completed chart before replaying, so previews also contain the full curves. Tool-call savings compare tasks solved by both the base model and checkpoint, with a different matched cohort at each checkpoint. The animation stops at this revision's step 1,000.
[Static chart](https://huggingface.co/FineEnvs/LFM2.5-2.6B-opencode-RL/resolve/main/assets/curves.png) · [Plotted data](https://huggingface.co/FineEnvs/LFM2.5-2.6B-opencode-RL/resolve/main/assets/curves.json) · [Interactive article](https://huggingface.co/spaces/AdithyaSK/multi-harness-rl)
## Training
TRL Async GRPO with **correctness plus a tool-call efficiency bonus**, initialized from the base model. All training rollouts use OpenCode. Rollouts run through OpenEnv × Harbor in E2B sandboxes.
| Setting | Value |
|---|---|
| Training task pool | 1,000 tasks: 400 medium / 600 hard |
| Checkpoint | 1,000 optimizer steps |
| Learning rate | 3e-6 |
| Rollouts per GRPO group / maximum staleness | 8 / 4 optimizer steps |
| Optimizer / precision | paged AdamW 8-bit / bfloat16 |
| Training harnesses | OpenCode |
| Sampling temperature / top-p | 0.8 / 1.0 |
| Per-call output budget, training / evaluation | 4,096 / 4,096 tokens |
The reward is `correctness × (1 + 0.1 × 15 / (15 + tool_calls))`, using verified native harness action counts. Incorrect answers receive zero; missing or unverified action counts receive no efficiency bonus. Reported pass@1 measures correctness only. The 1,000-task pool is not a claim that every task contributed an optimizer update.
## Evaluation
250 fixed SmolDataEnvs test tasks (33 easy, 118 medium, 99 hard), each evaluated under four harnesses: **1,000 graded task/harness cells**. Pass@1 uses the first graded attempt per cell; infrastructure retries do not turn it into pass@k. All cells below are graded.
| Harness | Correct / evaluated | Pass@1 |
|---|---:|---:|
| OpenCode | 145 / 250 | 58.0% |
| Claude Code | 105 / 250 | 42.0% |
| Codex | 108 / 250 | 43.2% |
| Mini-SWE-Agent | 165 / 250 | 66.0% |
| **Overall** | **523 / 1,000** | **52.3%** |
Harness versions: OpenCode 1.18.31, Claude Code 2.1.270, Codex 0.154.0, Mini-SWE-Agent 2.4.6. Full scores, including difficulty breakdowns, are in [eval_results.json](eval_results.json).
The [main branch](https://huggingface.co/FineEnvs/LFM2.5-2.6B-opencode-RL/tree/main) contains the final step-1,000 checkpoint (52.3%). The [step-900 branch](https://huggingface.co/FineEnvs/LFM2.5-2.6B-opencode-RL/tree/step-900) contains the best observed checkpoint (52.4%), used in the article's SFT-versus-RL comparison. **Best checkpoint selection used this test set**, not a separate validation set.
These are single-run results on a specific task set and harness versions. Harness mix, training exposure and compute differ across runs; the scores do not isolate a causal effect of the harness or objective.
## Load the checkpoint
The repository contains full saved model weights and the saved tokenizer/chat template, not a LoRA adapter. Training used Transformers 5.14.1; use a compatible Transformers release.
```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "FineEnvs/LFM2.5-2.6B-opencode-RL"
revision = "main"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForCausalLM.from_pretrained(
model_id, revision=revision, dtype=torch.bfloat16, device_map="auto"
)
```
Reproducing the task scores requires the agent harness and tools described in the [article](https://huggingface.co/spaces/AdithyaSK/multi-harness-rl); a plain chat prompt is not the same evaluation.
## Related releases
- [LFM2.5-2.6B-multiharness-RL](https://huggingface.co/FineEnvs/LFM2.5-2.6B-multiharness-RL)
- [LFM2.5-2.6B-opencode-RL](https://huggingface.co/FineEnvs/LFM2.5-2.6B-opencode-RL)
- [LFM2.5-2.6B-multiharness-SFT](https://huggingface.co/FineEnvs/LFM2.5-2.6B-multiharness-SFT)
- [LFM2.5-2.6B-opencode-SFT](https://huggingface.co/FineEnvs/LFM2.5-2.6B-opencode-SFT)
- [Qwen3.5-2B-multiharness-RL](https://huggingface.co/FineEnvs/Qwen3.5-2B-multiharness-RL)
- [Qwen3.5-2B-opencode-RL](https://huggingface.co/FineEnvs/Qwen3.5-2B-opencode-RL)
- [Qwen3.5-2B-opencode-standalone-RL](https://huggingface.co/FineEnvs/Qwen3.5-2B-opencode-standalone-RL)
All models, datasets, environments and the article are linked in the [multi-harness RL collection](https://huggingface.co/collections/FineEnvs/multi-harness-rl-6abdfaaa8d74dacd481d5212).
## License and provenance
Derived from [LiquidAI/LFM2.5-2.6B](https://huggingface.co/LiquidAI/LFM2.5-2.6B) under the LFM Open License v1.0. The base model's license is included unchanged in [LICENSE](LICENSE). FineEnvs modified the weights by reinforcement learning; this is not an official Liquid AI release. See [NOTICE](NOTICE) and [release_manifest.json](release_manifest.json) for modification notices, the pinned base revision, checkpoint identity and file checksums. Optimizer, scheduler, RNG and trainer state are excluded.
## Citation
For the experiment, methodology and interpretation, cite the [main article](https://huggingface.co/spaces/AdithyaSK/multi-harness-rl):
```bibtex
@misc{kolavi2026multiharnessrl,
author = {Adithya S Kolavi},
title = {The ultimate guide to multi-harness RL},
year = {2026},
url = {https://huggingface.co/spaces/AdithyaSK/multi-harness-rl}
}
```