sergiopaniego HF Staff
Add 07 PortSimEnv and the Simulation RL Environments article
e4d28dd verified |
Download README.md from FineEnvs/README: direct link, hf CLI and curl.
- Browser
- Download file 10.1 kB
-
https://huggingface.co/spaces/FineEnvs/README/resolve/main/README.md
- Command line
-
hf download hf://spaces/FineEnvs/README/README.md
-
curl -L -o README.md https://huggingface.co/spaces/FineEnvs/README/resolve/main/README.md
10.1 kB
| title: FineEnvs | |
| emoji: π€ | |
| colorFrom: yellow | |
| colorTo: purple | |
| sdk: static | |
| pinned: false | |
| license: mit | |
|  | |
| [](https://github.com/adithya-s-k/FineEnvs) | |
| [](https://huggingface.co/spaces/AdithyaSK/rl-environments-guide) | |
| [](https://huggingface.co/spaces/AdithyaSK/rl-environments-101-slides) | |
| # π€ FineEnvs: Open RL Environments | |
| FineEnvs is a home for **end-to-end RL environment recipes**, built to make it easier to **explore, reproduce, train, and evaluate agent systems**. | |
| Explore complete and reproducible environment projects from us and the community, including: | |
| * π **Open RL environments** | |
| * π§© **End-to-end environment recipes** | |
| * π» **Complete implementations** | |
| * π¦ **Models, datasets, and artifacts** | |
| * π§ͺ **Training and evaluation setups** | |
| * π **Demos and Spaces** | |
| * π **Tutorials and guides** | |
| All the reproducible code β environments, rollouts, training configs, notebooks, article and slide sources β lives in one repo: **[github.com/adithya-s-k/FineEnvs](https://github.com/adithya-s-k/FineEnvs)**. The artifacts those produce live here on the Hub. | |
| # FineEnvs Projects | |
| A growing collection of open projects, environments, resources, and artifacts. | |
| | Project | What it is | Explore | | |
| | :------------------------- | :----------------------------------------------------------------------------------------------------------------------- | :------------------------------------------------------------------------------ | | |
| | **FineEnvs Academy** | Articles, guides, tutorials, slides, and hands-on resources for learning how to build RL environments and agent systems. | [Explore β](https://huggingface.co/collections/FineEnvs/fineenvs-academy-6a942aed464a2eaa85047596) | | |
| | **Data Agent** | Training SLMs for data science with multi-harness RL environments. | [Explore β](https://huggingface.co/collections/FineEnvs/data-agent-6a942bca75a18973f8285023) | | |
| | **MiMo-V2.6-RL in Harbor** | All 7,780 of Xiaomi's MiMo-V2.6 RL environments as Harbor tasks, plus an explorer to browse them and run graded rollouts. | [Explore β](https://huggingface.co/collections/FineEnvs/mimo-v26-rl-in-harbor-6ab836dba014bd7580f8a3fe) Β· [Explorer β](https://huggingface.co/spaces/FineEnvs/MiMo-RL-Envs-Explorer) | | |
| | **Repo2RLEnv** | Verifiable coding and terminal RL environments in Harbor format, with per-task quality labels and provenance. | [Explore β](https://huggingface.co/collections/FineEnvs/repo2rlenv-verifiable-rl-environments-6aa82300d7494c050f50508d) | | |
| Each numbered project below is a self-contained recipe: an environment, a training run, and every artifact it produced. | |
| | # | Project | What it is | Explore | | |
| | :-- | :--- | :--- | :--- | | |
| | **00** | **RL Environments 101** | Three environments implemented six times over, one per framework. Same logic, six dialects. | [Source β](https://github.com/adithya-s-k/FineEnvs/tree/main/00-environments-101) Β· [Collection β](https://huggingface.co/collections/FineEnvs/rl-envs-101-6abca516b22e765e0b1aa0d1) | | |
| | **01** | **LaTeX OCR** | Qwen3-VL-2B trained to read rendered math into LaTeX, scored by a reward served from a live Space. | [Collection β](https://huggingface.co/collections/FineEnvs/latex-ocr-6aa7ed8498b3ffd222ad1c8c) | | |
| | **02** | **Watercolour** | Qwen3.5-35B-A3B trained to paint watercolours by writing p5.brush sketches, rewarded by taste rather than correctness. | [Collection β](https://huggingface.co/collections/FineEnvs/paint-with-code-6a955b79d63f67f1631d9be6) | | |
| | **03** | **GeoGuesser** | A multi-turn visual geolocation environment, and the 4B trained on it until it outscored `gpt-5.4-mini` and `claude-haiku-4.5`. | [Collection β](https://huggingface.co/collections/FineEnvs/geoguesser-env-6a969f8db267fe0e85fa1ab6) | | |
| | **04** | **SmolDataEnvs** | 5.5K+ data-analysis tasks for hill-climbing small models, graded deterministically with no LLM judge: plain prompts, verified SFT traces, and Harbor task suites. | [Collection β](https://huggingface.co/collections/FineEnvs/smoldataenvs-6ab4f2f6e09b7cb872ebc867) | | |
| | **05** | **SmolDataEnvs: Multi-harness RL** | One small model trained with GRPO inside four unmodified coding agents (OpenCode, Claude Code, Codex, Mini-SWE-Agent), with LFM2.5-2.6B and Qwen3.5-2B checkpoints. | [Article β](https://huggingface.co/spaces/FineEnvs/multi-harness-rl) Β· [Slides β](https://huggingface.co/spaces/FineEnvs/multi-harness-rl-slides) Β· [Collection β](https://huggingface.co/collections/FineEnvs/smoldataenvs-multi-harness-rl-6abdfaaa8d74dacd481d5212) | | |
| | **06** | **Multilingual** | Two OpenEnv servers: a million document pages in 22 languages with Sarvam Indic OCR Bench, and read speech in all 102 FLEURS languages. Gemma 4 trained on each to read and to hear Kannada. | [Collection β](https://huggingface.co/collections/FineEnvs/multilingual-multimodal-envs-6ac0c27c137f93e0799603e4) | | |
| | **07** | **PortSimEnv** | Re-plan a broken week of container-ship dockings at the Port of Barcelona, built from the port's real 2024 records and graded against a plan CP-SAT proves optimal. | [Article β](https://huggingface.co/spaces/FineEnvs/simulation-rl-environments) Β· [Play β](https://huggingface.co/spaces/FineEnvs/PortSimEnv) Β· [Collection β](https://huggingface.co/collections/FineEnvs/simulation-rl-envs) | | |
| # Articles & Talks | |
| | | What it covers | Read / Watch | | |
| | :--- | :--- | :--- | | |
| | π **The Ultimate Guide to RL Environments** | Building and scaling RL environments in the LLM era β how frameworks are built, how rewards are wired, how they scale to thousands of concurrent sessions. | [Read β](https://huggingface.co/spaces/AdithyaSK/rl-environments-guide) | | |
| | ποΈ **RL Environments 101** | From "what is an env?" to training your own: RL fundamentals β environment anatomy β OpenEnv β training with TRL. | [Watch β](https://huggingface.co/spaces/AdithyaSK/rl-environments-101-slides) | | |
| | π **Scaling RL for LLMs** | RL environments and RL training β what an environment is, how reward hacking happens, how to train against your own. AMD AI Dev Day. | [Watch β](https://huggingface.co/spaces/AdithyaSK/scaling-rl-for-llms-amd-ai-dev-day) | | |
| | π **Multi-Harness Training** | OpenEnv Γ Harbor β why an environment's failure model decides whether it can be trained against. | [Watch β](https://huggingface.co/spaces/AdithyaSK/multi-harness-training-slides) | | |
| | π§ **The Ultimate Guide to Multi-Harness RL** | Training small models on SmolDataEnvs with the same tasks and reward but a different tool loop each time (TRL, native OpenCode, Harbor), and what changes. | [Read β](https://huggingface.co/spaces/FineEnvs/multi-harness-rl) | | |
| | π€ **Training a Coding Agent Through a Harness You Did Not Write** | Multi-harness RL talk by Sergio Paniego Blanco: one model, four unmodified coding agents, GRPO. | [Watch β](https://huggingface.co/spaces/FineEnvs/multi-harness-rl-slides) | | |
| | π **How to turn a game into an RL environment** | The technical intuition, end to end: curating the data, designing the environment, shipping it with OpenEnv, and training a 4B against it with TRL. | [Read β](https://huggingface.co/spaces/FineEnvs/geoguesser-article) | | |
| | π’ **Simulation RL Environments** | Turning real-world work into RL environments. Part 1 builds PortSimEnv from the Port of Barcelona's 2024 records and grades every plan against a proven optimum. | [Read β](https://huggingface.co/spaces/FineEnvs/simulation-rl-environments) | | |
| # Environments | |
| Three reference environments, each implemented across six frameworks β `openenv`, `ors`, `nemo_gym`, `verifiers`, `skyrl_gym`, `gem`. Same logic, six dialects. [Source β](https://github.com/adithya-s-k/FineEnvs/tree/main/00-environments-101) Β· [RL Envs 101 collection β](https://huggingface.co/collections/FineEnvs/rl-envs-101-6abca516b22e765e0b1aa0d1) | |
| | Environment | Tools | OpenEnv | ORS | NeMo Gym | | |
| | :--- | :--: | :--- | :--- | :--- | | |
| | **Jupyter agent** β real code execution in an E2B sandbox | 4 | [Space](https://huggingface.co/spaces/AdithyaSK/jupyter-agent-openenv) | [Space](https://huggingface.co/spaces/AdithyaSK/jupyter-agent-ors) | [Space](https://huggingface.co/spaces/AdithyaSK/jupyter-agent-nemo-gym) | | |
| | **Wordle** β multi-turn, pure Python, no backend | 1 | [Space](https://huggingface.co/spaces/FineEnvs/wordle-openenv) | [Space](https://huggingface.co/spaces/FineEnvs/wordle-ors) | [Space](https://huggingface.co/spaces/FineEnvs/wordle-nemo-gym) | | |
| | **Desktop** β computer-use, vision-driven Linux desktop | 19 | [Space](https://huggingface.co/spaces/AdithyaSK/desktop-openenv) | [Space](https://huggingface.co/spaces/AdithyaSK/desktop-ors) | β | | |
| # Build your own | |
| Five agent skills turn a plain-English description into a runnable RL environment across four frameworks β works with Claude Code, Cursor, Codex, OpenCode, Gemini CLI and others. | |
| ```bash | |
| npx skills add adithya-s-k/FineEnvs | |
| ``` | |
| **We're looking for new end-to-end recipes** β a task, an environment, a training run, and honest results. [Contributing guide β](https://github.com/adithya-s-k/FineEnvs/blob/main/CONTRIBUTING.md) | |
| ## Citation | |
| ```bibtex | |
| @misc{fineenvs, | |
| author = {Kolavi, Adithya S}, | |
| title = {FineEnvs: Open Source RL Environments for LLM Agents}, | |
| year = {2026}, | |
| url = {https://github.com/adithya-s-k/FineEnvs} | |
| } | |
| ``` | |