Instructions to use FineEnvs/LFM2.5-2.6B-opencode-RL with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FineEnvs/LFM2.5-2.6B-opencode-RL with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="FineEnvs/LFM2.5-2.6B-opencode-RL") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("FineEnvs/LFM2.5-2.6B-opencode-RL") model = AutoModelForCausalLM.from_pretrained("FineEnvs/LFM2.5-2.6B-opencode-RL", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use FineEnvs/LFM2.5-2.6B-opencode-RL with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FineEnvs/LFM2.5-2.6B-opencode-RL" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FineEnvs/LFM2.5-2.6B-opencode-RL", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/FineEnvs/LFM2.5-2.6B-opencode-RL
- SGLang
How to use FineEnvs/LFM2.5-2.6B-opencode-RL with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "FineEnvs/LFM2.5-2.6B-opencode-RL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FineEnvs/LFM2.5-2.6B-opencode-RL", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "FineEnvs/LFM2.5-2.6B-opencode-RL" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FineEnvs/LFM2.5-2.6B-opencode-RL", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use FineEnvs/LFM2.5-2.6B-opencode-RL with Docker Model Runner:
docker model run hf.co/FineEnvs/LFM2.5-2.6B-opencode-RL
LFM2.5-2.6B-opencode-RL
A full fine-tune of LiquidAI/LFM2.5-2.6B for agentic data-analysis tasks, trained with asynchronous GRPO (TRL Async GRPO) using OpenCode. This release is step 1,000 on main, scoring 52.3% pass@1 across four evaluation harnesses.
Article · Collection · Evaluation tasks · Training dashboard
Training and evaluation curves
Training curves use a trailing 50-step mean (at least 10 observations). Evaluation markers show measured checkpoints; hollow markers and dotted segments indicate incomplete coverage. Faint background lines show the full recorded trajectory; colored lines reveal the measured checkpoints. The first frame shows the completed chart before replaying, so previews also contain the full curves. Tool-call savings compare tasks solved by both the base model and checkpoint, with a different matched cohort at each checkpoint. The animation stops at this revision's step 1,000.
Static chart · Plotted data · Interactive article
Training
TRL Async GRPO with correctness plus a tool-call efficiency bonus, initialized from the base model. All training rollouts use OpenCode. Rollouts run through OpenEnv × Harbor in E2B sandboxes.
| Setting | Value |
|---|---|
| Training task pool | 1,000 tasks: 400 medium / 600 hard |
| Checkpoint | 1,000 optimizer steps |
| Learning rate | 3e-6 |
| Rollouts per GRPO group / maximum staleness | 8 / 4 optimizer steps |
| Optimizer / precision | paged AdamW 8-bit / bfloat16 |
| Training harnesses | OpenCode |
| Sampling temperature / top-p | 0.8 / 1.0 |
| Per-call output budget, training / evaluation | 4,096 / 4,096 tokens |
The reward is correctness × (1 + 0.1 × 15 / (15 + tool_calls)), using verified native harness action counts. Incorrect answers receive zero; missing or unverified action counts receive no efficiency bonus. Reported pass@1 measures correctness only. The 1,000-task pool is not a claim that every task contributed an optimizer update.
Evaluation
250 fixed SmolDataEnvs test tasks (33 easy, 118 medium, 99 hard), each evaluated under four harnesses: 1,000 graded task/harness cells. Pass@1 uses the first graded attempt per cell; infrastructure retries do not turn it into pass@k. All cells below are graded.
| Harness | Correct / evaluated | Pass@1 |
|---|---|---|
| OpenCode | 145 / 250 | 58.0% |
| Claude Code | 105 / 250 | 42.0% |
| Codex | 108 / 250 | 43.2% |
| Mini-SWE-Agent | 165 / 250 | 66.0% |
| Overall | 523 / 1,000 | 52.3% |
Harness versions: OpenCode 1.18.31, Claude Code 2.1.270, Codex 0.154.0, Mini-SWE-Agent 2.4.6. Full scores, including difficulty breakdowns, are in eval_results.json.
The main branch contains the final step-1,000 checkpoint (52.3%). The step-900 branch contains the best observed checkpoint (52.4%), used in the article's SFT-versus-RL comparison. Best checkpoint selection used this test set, not a separate validation set.
These are single-run results on a specific task set and harness versions. Harness mix, training exposure and compute differ across runs; the scores do not isolate a causal effect of the harness or objective.
Load the checkpoint
The repository contains full saved model weights and the saved tokenizer/chat template, not a LoRA adapter. Training used Transformers 5.14.1; use a compatible Transformers release.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
model_id = "FineEnvs/LFM2.5-2.6B-opencode-RL"
revision = "main"
tokenizer = AutoTokenizer.from_pretrained(model_id, revision=revision)
model = AutoModelForCausalLM.from_pretrained(
model_id, revision=revision, dtype=torch.bfloat16, device_map="auto"
)
Reproducing the task scores requires the agent harness and tools described in the article; a plain chat prompt is not the same evaluation.
Related releases
- LFM2.5-2.6B-multiharness-RL
- LFM2.5-2.6B-opencode-RL
- LFM2.5-2.6B-multiharness-SFT
- LFM2.5-2.6B-opencode-SFT
- Qwen3.5-2B-multiharness-RL
- Qwen3.5-2B-opencode-RL
- Qwen3.5-2B-opencode-standalone-RL
All models, datasets, environments and the article are linked in the multi-harness RL collection.
License and provenance
Derived from LiquidAI/LFM2.5-2.6B under the LFM Open License v1.0. The base model's license is included unchanged in LICENSE. FineEnvs modified the weights by reinforcement learning; this is not an official Liquid AI release. See NOTICE and release_manifest.json for modification notices, the pinned base revision, checkpoint identity and file checksums. Optimizer, scheduler, RNG and trainer state are excluded.
Citation
For the experiment, methodology and interpretation, cite the main article:
@misc{kolavi2026multiharnessrl,
author = {Adithya S Kolavi},
title = {The ultimate guide to multi-harness RL},
year = {2026},
url = {https://huggingface.co/spaces/AdithyaSK/multi-harness-rl}
}
- Downloads last month
- 245
Model tree for FineEnvs/LFM2.5-2.6B-opencode-RL
Dataset used to train FineEnvs/LFM2.5-2.6B-opencode-RL
Collection including FineEnvs/LFM2.5-2.6B-opencode-RL
Evaluation results
- Pass@1 (%) — Overall on SmolDataEnvs (250 fixed test tasks)test set self-reported52.300
- Pass@1 (%) — OpenCode on SmolDataEnvs (250 fixed test tasks)test set self-reported58.000
- Pass@1 (%) — Claude Code on SmolDataEnvs (250 fixed test tasks)test set self-reported42.000
- Pass@1 (%) — Codex on SmolDataEnvs (250 fixed test tasks)test set self-reported43.200
- Pass@1 (%) — Mini-SWE-Agent on SmolDataEnvs (250 fixed test tasks)test set self-reported66.000

