AdithyaSK's picture
AdithyaSK HF Staff
Fix disconnected evaluation lines and GIF previews
a0bcdab verified
|
Raw History Blame Contribute Delete
8.96 kB
metadata
base_model: LiquidAI/LFM2.5-2.6B
base_model_relation: finetune
library_name: transformers
pipeline_tag: text-generation
language:
  - en
license: other
tags:
  - trl
  - openenv
  - harbor
  - agent
  - smoldataenvs
  - grpo
  - opencode
datasets:
  - FineEnvs/SmolDataEnvs-harbor-train
thumbnail: >-
  https://huggingface.co/spaces/AdithyaSK/multi-harness-rl/resolve/main/app/public/og/og-image.png
model-index:
  - name: LFM2.5-2.6B-opencode-RL (main, step 1000)
    results:
      - task:
          type: text-generation
          name: Agentic data analysis
        dataset:
          type: FineEnvs/SmolDataEnvs-harbor-test
          name: SmolDataEnvs (250 fixed test tasks)
          split: test
          args:
            harness: all four
        metrics:
          - type: pass_at_1
            name: Pass@1 (%) — Overall
            value: 52.3
      - task:
          type: text-generation
          name: Agentic data analysis
        dataset:
          type: FineEnvs/SmolDataEnvs-harbor-test
          name: SmolDataEnvs (250 fixed test tasks)
          split: test
          args:
            harness: opencode
        metrics:
          - type: pass_at_1
            name: Pass@1 (%) — OpenCode
            value: 58
      - task:
          type: text-generation
          name: Agentic data analysis
        dataset:
          type: FineEnvs/SmolDataEnvs-harbor-test
          name: SmolDataEnvs (250 fixed test tasks)
          split: test
          args:
            harness: claude-code
        metrics:
          - type: pass_at_1
            name: Pass@1 (%) — Claude Code
            value: 42
      - task:
          type: text-generation
          name: Agentic data analysis
        dataset:
          type: FineEnvs/SmolDataEnvs-harbor-test
          name: SmolDataEnvs (250 fixed test tasks)
          split: test
          args:
            harness: codex
        metrics:
          - type: pass_at_1
            name: Pass@1 (%) — Codex
            value: 43.2
      - task:
          type: text-generation
          name: Agentic data analysis
        dataset:
          type: FineEnvs/SmolDataEnvs-harbor-test
          name: SmolDataEnvs (250 fixed test tasks)
          split: test
          args:
            harness: mini-swe-agent
        metrics:
          - type: pass_at_1
            name: Pass@1 (%) — Mini-SWE-Agent
            value: 66
license_name: lfm1.0
license_link: LICENSE

The ultimate guide to multi-harness RL

LFM2.5-2.6B-opencode-RL

A full fine-tune of LiquidAI/LFM2.5-2.6B for agentic data-analysis tasks, trained with asynchronous GRPO (TRL Async GRPO) using OpenCode. This release is step 1,000 on main, scoring 52.3% pass@1 across four evaluation harnesses.

Article · Collection · Evaluation tasks · Training dashboard

Training and evaluation curves

Animated training, pass@1 and tool-use curves through step 1,000

Training curves use a trailing 50-step mean (at least 10 observations). Evaluation markers show measured checkpoints; hollow markers and dotted segments indicate incomplete coverage. Faint background lines show the full recorded trajectory; colored lines reveal the measured checkpoints. The first frame shows the completed chart before replaying, so previews also contain the full curves. Tool-call savings compare tasks solved by both the base model and checkpoint, with a different matched cohort at each checkpoint. The animation stops at this revision's step 1,000.

Static chart · Plotted data · Interactive article

Training

TRL Async GRPO with correctness plus a tool-call efficiency bonus, initialized from the base model. All training rollouts use OpenCode. Rollouts run through OpenEnv × Harbor in E2B sandboxes.

Setting Value
Training task pool 1,000 tasks: 400 medium / 600 hard
Checkpoint 1,000 optimizer steps
Learning rate 3e-6
Rollouts per GRPO group / maximum staleness 8 / 4 optimizer steps
Optimizer / precision paged AdamW 8-bit / bfloat16
Training harnesses OpenCode
Sampling temperature / top-p 0.8 / 1.0
Per-call output budget, training / evaluation 4,096 / 4,096 tokens

The reward is correctness × (1 + 0.1 × 15 / (15 + tool_calls)), using verified native harness action counts. Incorrect answers receive zero; missing or unverified action counts receive no efficiency bonus. Reported pass@1 measures correctness only. The 1,000-task pool is not a claim that every task contributed an optimizer update.

Evaluation

250 fixed SmolDataEnvs test tasks (33 easy, 118 medium, 99 hard), each evaluated under four harnesses: 1,000 graded task/harness cells. Pass@1 uses the first graded attempt per cell; infrastructure retries do not turn it into pass@k. All cells below are graded.

Harness Correct / evaluated Pass@1
OpenCode 145 / 250 58.0%
Claude Code 105 / 250 42.0%
Codex 108 / 250 43.2%
Mini-SWE-Agent 165 / 250 66.0%
Overall 523 / 1,000 52.3%

Harness versions: OpenCode 1.18.31, Claude Code 2.1.270, Codex 0.154.0, Mini-SWE-Agent 2.4.6. Full scores, including difficulty breakdowns, are in eval_results.json.

The main branch contains the final step-1,000 checkpoint (52.3%). The step-900 branch contains the best observed checkpoint (52.4%), used in the article's SFT-versus-RL comparison. Best checkpoint selection used this test set, not a separate validation set.

These are single-run results on a specific task set and harness versions. Harness mix, training exposure and compute differ across runs; the scores do not isolate a causal effect of the harness or objective.

Load the checkpoint

The repository contains full saved model weights and the saved tokenizer/chat template, not a LoRA adapter. Training used Transformers 5.14.1; use a compatible Transformers release.

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

model_id = "FineEnvs/LFM2.5-2.6B-opencode-RL"
revision = "main"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForCausalLM.from_pretrained(
    model_id, revision=revision, dtype=torch.bfloat16, device_map="auto"
)

Reproducing the task scores requires the agent harness and tools described in the article; a plain chat prompt is not the same evaluation.

Related releases

All models, datasets, environments and the article are linked in the multi-harness RL collection.

License and provenance

Derived from LiquidAI/LFM2.5-2.6B under the LFM Open License v1.0. The base model's license is included unchanged in LICENSE. FineEnvs modified the weights by reinforcement learning; this is not an official Liquid AI release. See NOTICE and release_manifest.json for modification notices, the pinned base revision, checkpoint identity and file checksums. Optimizer, scheduler, RNG and trainer state are excluded.

Citation

For the experiment, methodology and interpretation, cite the main article:

@misc{kolavi2026multiharnessrl,
  author = {Adithya S Kolavi},
  title = {The ultimate guide to multi-harness RL},
  year = {2026},
  url = {https://huggingface.co/spaces/AdithyaSK/multi-harness-rl}
}