Spaces:
Running
Download DESIGN.md from FineEnvs/RL-Explorer: direct link, hf CLI and curl.
- Browser
- Download file 24.4 kB
-
https://huggingface.co/spaces/FineEnvs/RL-Explorer/resolve/main/DESIGN.md
- Command line
-
hf download hf://spaces/FineEnvs/RL-Explorer/DESIGN.md
-
curl -L -o DESIGN.md https://huggingface.co/spaces/FineEnvs/RL-Explorer/resolve/main/DESIGN.md
RL Explorer: design
A Space for exploring the RL environments on the Hugging Face Hub, whatever their framework:
- Harbor task datasets (folders with
task.toml), indexed and run on HF Sandboxes; - OpenEnv Spaces, live: woken or restarted, their own app, a playground with rewards, their Task API, and an MCP server for coding agents; OpenEnv Γ Harbor Spaces with what they can run;
- custom environments: datasets whose tasks are rows (Verifiers, NeMo Gym, verl, MiMo's raw release, Harbor tasks packed into parquet, traces, any table of prompts), each read by a processor (CUSTOM_ENVS.md).
Browse an environment's tasks, read exactly what an agent gets and how it is graded, run an agent on a task in a
sandbox, watch it, and compare rollouts. The MiMo RL Environment Explorer did this for one release; MiMo is one
environment among the others here (its raw rows through the mimo processor, linked to its Harbor twins).
What a Harbor dataset is
A task is a folder with task.toml (settings, metadata), instruction.md (the prompt), environment/
(a Dockerfile, or a prebuilt docker_image in task.toml, sometimes a compose file), tests/ (the
verifier: test.sh writes /logs/verifier/reward.txt, one number, or reward.json, named scores) and
often solution/. A survey of the 30 most-downloaded harbor datasets (2026-10-03) found:
| Layout | Examples | Share |
|---|---|---|
tasks/<task>/ |
terminal-bench-2.1/3.0, MiMo, data-agent, repo2rlenv | ~60% |
tasks/<group>/.../<task>/ |
terminal-bench-science | few |
<task>/ at the repository root |
terminal-bench-2.0, harbor-mix, WildClawBench, dabstep | ~25% |
<benchmark>/tasks/<task>/ |
skilltrainbench-public, Lego-RL | few |
packed (tasks.parquet, .tar, .zip) |
TaskTrove, TerminalWorld, FACET | ~10%: read as rows (below) |
So a task is found by where its task.toml is, not by a fixed path. Most tasks ship only a Dockerfile;
a prebuilt docker_image is the minority (terminal-bench-2.1, LHTB, MiMo, apex). Metadata keys differ per
dataset, and some are answers (gold_answer, expected_output, solution): those are never shown.
Architecture
Two Spaces, one image, one private bucket (FineEnvs/rl-explorer-data, mounted at /data by both):
HF RL Explorer (public Space, app.main) Admin (private Space, app.admin_app, FineEnvs only)
catalog: Harbor datasets + OpenEnv Spaces overview, rollouts (with who ran them), environments
indexes and packs (/data/indexes, /data/packs) (pin, hide, feature, index), collections, indexes,
runner: OpenEnv's run_rollout + capture proxy settings (rollout switch, limits, agents, banner), audit
(/capture), HF Sandbox on the visitor's account β
store: /data/runs β writes /data/admin/settings.json (+ audit.jsonl)
β² re-reads settings.json (mtime) βββββββββββββββββββββββ reads /data/runs, builds indexes
The admin shares the bucket, not the process: pins, hidden environments, collections, rollout switches, rollouts taken off Community and requests to stop a running rollout all go through the settings file, which the explorer re-reads within seconds. The admin Space is private (only org members reach it) and every admin route checks membership again (public member list, or the memberships the Hub reported at sign-in).
- Catalog. Datasets tagged
harborand Spaces taggedopenenvorrl-environment(Docker, not articles), plus the collections admins curate (a dataset id, orspace:<id>), and the signed-in visitor's own datasets, private ones included. Private and gated datasets are read with the visitor's token; every request re-checks access. - Index and pack. On first open (or ahead of time with
python -m app.precache), one listing finds everytask.toml; the files a task page shows are fetched in parallel (one GET each, 48 at a time). A repository too big to list in 25 s (tasks that each ship a whole app, like microsoft/ProgramDistill's 4,375) is walked instead: a level is listed, the Hub's batchedpaths-infosays which of its folders hold atask.toml(500 a call), and only the files the index reads are looked up; the task list is published as soon as it's known (the page opens, and fills in as files arrive), and a task's other folders are listed when someone opens them in the file viewer. The index is one row per task (title, facets, grading, environment, network policy, whether it can run here); the pack holds every task's text, de-duplicated, and its file list, so task pages never wait on the Hub. The first open shows a step timeline with live counts. - Answers stay out. Answer-like metadata keys,
solution/, data files beside the grader (rubrics, expected outputs), and answers written into grader scripts (answer-named heredocs and literals) are withheld or masked. - Rollouts.
openenv.harbor.rollout.run_rollout, the same code as OpenEnv's Harbor UI. The agent (OpenCode, Terminus 2, mini-SWE-agent, Pi) runs in an HF Sandbox from the task's image, or, for a Dockerfile-only task, from the Dockerfile's base image with its RUN/COPY/ENV/WORKDIR replayed first (app/dockerfile.py; about 99% of the indexed Dockerfiles qualify).Sandbox.createis wrapped so the sandbox runs on the visitor's account. Tasks that ask for no internet are refused with the reason: HF Sandboxes can't enforce it. Models: Inference Providers through the router (:fastest) with the visitor's token, or an endpoint they bring (re-checked before use). - OpenEnv Spaces, live (
app/spaces_live.py). What a server offers comes from itself: its/openapi.json(routes: reset/step, the Task API, custom ones),/metadata,/schema,/list_environmentsandtools/liston/mcp; its web page from the README'sbase_path, then/web, then/. A sleeping Space is woken the way its Hub page lets any visitor (POST /spaces/<id>/start, at most once a minute); a broken one is restarted with the visitor's own token, which works only if they may write to it. The playground holds one WebSocket to the server's/wsper session (reset, step, state, and MCP calls as{"type": "mcp"}on the same episode: the HTTP routes are stateless), shows each observation (images, audio, chats) and reward, and plays a Task API task by its split and index. OpenEnv Γ Harbor servers show their engine, sandboxes, validated agents and served datasets, mapped back to the Hub (/data/org__nameisorg/name) so their tasks run on the visitor's account here. Other servers (NeMo Gym, ORS, any FastAPI app) are used through their own routes, read from their OpenAPI: a form per route from its request schema, called with the session's own cookie jar (NeMo Gym keeps an episode in a cookie; it stays on the explorer). Only plain paths the server published, GET or POST, path parameters checked; GET answers drop answer-like fields. OpenEnv servers' own extra routes aren't offered (geoguesser's return the answer's coordinates): they're used through reset/step and tools. - MCP for coding agents (
app/mcp_bridge.py,/mcp/<org>/<name>). OpenEnv's/mcpanswerstools/listandtools/callbut not MCP's handshake, so Claude Code, Codex, Cursor, VS Code, Gemini CLI and the rest can't use it directly. The bridge speaks Streamable HTTP: the Space's own tools in a session of their own, plusreset/step/statefor servers without tools (step's input is the action schema), pluslist_splits/list_tasks/get_taskfor the Task API, plus one tool per route for servers that aren't OpenEnv (NeMo Gym'sseed_session,guess,verify); images come back as image content. Sessions are capped per address, rate-limited, and end idle. - One contract for every format (
app/envs/contract.py, CUSTOM_ENVS.md). An environment is read by an adapter that answers five questions: its tasks (a summary, cards, each task as sections of blocks, its files), what they run in, how they can run (run options, each by a runner, with its own inputs), which other tasks are the same task (aliases: their rollouts are shared), and nothing more for MCP. So one environment page (/d/<org>/<name>,env.js) and one task page (/t/<org>/<name>/<ref>,task.js) serve Harbor folders, the MiMo release and any rows dataset; a raw MiMo row and its Harbor conversion link to each other and list each other's rollouts. Adapters:harbor.py(the index above),mimo.py,rows.py(with processors: rows from the Hub's dataset viewer, or straight from JSON Lines / JSON / parquet / CSV; a processor turns a row into a task; answers left out by name at any depth). The registry picks the surest adapter that confirms; linkers and aliasers relate tasks across adapters. - The MiMo release (
app/mimo,app/envs/mimo.py). The MiMo RL Environment Explorer's backend, kept close to its source: all 7,780 tasks with the brief, the agent's exact prompt, systems and their databases, workspace previews, Code tasks' repositories (snapshots in/data/repo-snapshots), the grader with its judge prompts, bands and music features, and the reward design across the release. Two runners: its Harbor conversion (the default) or its own harness (OpenCode in the task's image, Xiaomi's graders, a judge model the visitor picks), whose sandbox reaches models through/api/llm/<capability>and never holds a token. Its run page (events, live), compare view and community (leaderboard by domain, the tasks nobody has tried) work for every environment. - MCP for environments (
/mcp/d/<org>/<name>): describe, list and search tasks, read one in full as text, read its files, a random task; answers withheld as on the page. Any adapter, no code. - Visibility. Public rollouts appear in Community without who ran them, private ones only to their owner; only finished, graded rollouts are shared; a private dataset's rollouts stay private; admins can take one off.
Look
The MiMo explorer's design (its web/app.css, unchanged; web/rlx.css, fv.css and sp.css add the rest), with
the Hugging Face logo and the name HF RL Explorer. Explore opens with FineEnvs' banner (one link, closable for a
week), the numbers, two rows of trending environments with a pager, then one search, filters (kind: Harbor, OpenEnv,
Verifiers, NeMo Gym, other) and cards. Elsewhere the header's search lists trending environments before you type.
Files open in an editor-like viewer (folders, highlighting, line links, full view) that fits whatever column it's in.
Security
uv run pytest tests -q (no network). test_security.py: sessions can't be forged or outlive their expiry;
admin routes are org-only and absent from the public app; cross-site writes are refused; private rollouts and
datasets stay private, per token; public rollouts carry no user or endpoint; answers are withheld or masked;
task files stay inside their task; visitor endpoints must be public https; tokens never reach the store; headers.
test_live.py: only a Space's own hf.space host is ever called (redirects elsewhere aren't followed, base_path
can't point off it), private Spaces are refused, waking is limited, restarting needs write access, playground
sessions belong to who started them and are capped, the Task API browser drops answers, and the MCP bridge speaks
the protocol (handshake, notifications, errors as tool results, sessions, rate limit, no cross-origin pages).
test_envs.py: processors claim their formats, answers stay out (nested, packed solutions, MiMo's patches), archives
can't escape or blow up, filters can't inject, the direct reader pages right, and the walker never lists a task's
own folders. test_contract.py: a toy adapter registered from outside works end to end (pages, private keys stripped,
runs refused unless the chosen option can run, its fields checked, rollouts merged across aliases, MCP).
test_mimo.py: the MiMo adapter's tiles, facets and map, row refs, the default runner, judges from the pool only, no
token in a MiMo run record, links back from the Harbor twins, scrubbed public events, sandboxed raw files.
tests/ui-audit.mjs clicks through every page at four widths in both themes (needs network).
Never sent to a Space: the visitor's token. A Space's run_rollout is offered only when it has its own engine.
Next
A Space adapter (an OpenEnv server's Task API tasks as task pages, its playground as a run option); a folder-of-tasks
processor for repositories without task.toml; model token costs for Harbor rollouts; MiMo's repository snapshots
synced into this bucket.
Data layer
The catalog is built offline and read as an immutable SQLite snapshot, so no app process fetches the Hub's listing or builds indexes to answer a page, and the home page asks the server for one page of cards instead of downloading every environment (about 4 MB of JSON) to filter it in the browser.
HF scheduled Job, hourly (scripts/schedule_indexer.py) Explorer / admin Spaces (app/snapshot.py)
python -m app.indexer --store /data every 60 s: read snapshots/LATEST.v1.json
1 listing catalog._listing_fetch (every Space, ~2.5 min) new? copy the .db to local disk, check sha256,
2 indexes rebuild what changed (catalog's builders), budget open read-only + immutable, check schema,
3 build SQLite on local disk: envs, tasks, FTS, indexes counts, quick_check; swap (queries in flight
4 publish snapshots/catalog-<UTC>.db, then the pointer finish on the old one)
5 prune keep 5 (and whatever a pointer names) none / corrupt: build the same DB from the
β live catalog (local dev, first boot)
βββββ bucket FineEnvs/rl-explorer-data, mounted at /data βββββββββββ²
- Indexer (
app/indexer.py). One run fetches the listing and refuses it when it is much smaller than the last snapshot's (overall, datasets or Spaces under 80%: catalog skips a Hub tag listing that fails, and a half listing must not empty the site;--forceoverrides). It writeslisting.json.gzfor the app's own code paths, then rebuilds dataset indexes incrementally: a dataset whose index is current (sameINDEX_VERSION, HublastModifiedolder than the index's build) costs no Hub call; one touched around or after its build has its revision checked; changed revisions, older formats, indexes left partial and never-indexed Harbor datasets are rebuilt in that order (featured first, then by trending) with catalog's ownbuild_index, within a time and count budget. A build that fails isn't retried at the same revision for a day (indexer-state.json). Then it builds the snapshot on local disk (journal_mode=OFF,synchronous=OFF, rows streamed one dataset's index at a time, FTS rebuilt, indexes,ANALYZE,VACUUM,quick_check), copies it tosnapshots/catalog-<UTC ts>.db, checks it landed whole (size, and sha256 when the store can be re-read), and only then writes the pointersnapshots/LATEST.v<SCHEMA>.json({"db", "sha256", "size", "built_at", "counts", "schema"}) atomically. A run that fails anywhere before that publishes nothing. Old snapshots are pruned (the newest 5, and any a pointer of any schema names). Reruns are safe: current indexes are skipped and every run publishes a complete snapshot. - Snapshot (
app/snapshot.py).envs(one row per public environment: id, key, kind, framework, heading, brief, counters, dates, stage, MCP, OpenEnv version, badges and tags as JSON, indexed task count and revision, the facet values that don't depend on admin settings, and a lowercased search blob),envs_fts(FTS5, trigram: substring search, as the page always did),tasks(dataset, ref, title, brief, category, difficulty, group, how it runs, how it's graded, tags, a few extras),tasks_fts(FTS5, unicode61, token prefixes),meta. Only what the explorer already shows publicly: no private or gated dataset, no index a visitor built with their own token, no metadata values, no file texts. 7,225 environments and 51,498 tasks (the Harbor indexes and MiMo's release) made a 43 MB file, built in about a second; a run with the full listing took under five minutes. - Reading it. Each process checks the pointer on startup and every 60 s; SQLite never opens a file on the
bucket mount (the database is copied to
RLX_CACHE_DIR/snapshotsfirst), connections are read-only (mode=ro&immutable=1, a small pool per snapshot), and a swap retires the old snapshot when its last query is done. A pointer that names a missing, corrupt, truncated or mismatched database (sha256, schema, row counts, quick_check) is rejected and logged (snapshot.rejected), the current one kept, and the same file isn't fetched again for ten minutes. With no snapshot at all, the same database is built in the process fromcatalog.environments()and the indexes on disk, so local development and a first boot work without the Job. BumpingSCHEMApublishes under a new pointer name: an app of the old version keeps reading its own. - What admins change is applied per query. Hidden environments, pins and collections come from the admin settings on every request (an admin's change shows within seconds, between indexer runs), as do public rollout counts.
- Search API (
app/search_api.py, a router).GET /api/search?q=&kind=&collection=&f=&sort=&page=&size=returns a page of cards, the total and every facet's counts, with the page's exact semantics: the Kind taxonomy (OpenEnv Spaces including FineEnvs' curated servers, other Spaces including ORS, Harbor whenever an index found task folders, Verifiers, NeMo Gym, OpenEnv datasets, verl, other rows), pinned first in trending, FineEnvs' curated servers first among OpenEnv Spaces, each facet counted over what the other filters leave, ties in listing order.trending=24adds the trending row and header numbers to the same response (one request for the first paint);mine=1merges the signed-in visitor's own datasets, private ones included (a private response). Search words are data: each is a quoted FTS string (and checked again as a plain substring), so operators,NEAR,*, column filters and stray quotes can't change the query.GET /api/search/tasks?q=&env=searches every indexed task (each word a quoted prefix, best match first). Responses carryCache-Control(public 30 s, private formine), a weak ETag (304 on a repeat) andServer-Timing. A parity test compares every answer, order and facet with a plain-Python port of the page's old filtering. - Home page (
web/js/home.js): the first paint is one request of 12 KB on the wire (55 KB of JSON; it was 458 KB on the wire, 4 MB of JSON), and the first card shows in 283 ms locally instead of 467 ms (7.1 s to 5.2 s on DevTools' Fast 3G, where the data request itself went from 3.4 s to 0.65 s); search (debounced), filters, sort and paging ask the server, an older request is cancelled when a newer one starts, a slow answer shows placeholder cards of the real size, and the URL keeps the page's state as before. - Scheduling (
scripts/schedule_indexer.py, prints the Job;--applycreates it): the explorer Space's own image,@hourlywith concurrency off,cpu-upgrade, a 55-minute timeout, the bucket mounted read-write at/data(STORAGE_DIR=/data; written files upload on close, and a rename is one bucket batch, so temp-then-rename and the pointer-last order hold for readers on other mounts, which see changes within ~10β30 s),HF_TOKENas a Job secret (higher rate limits for public downloads only).--store hf://buckets/<org>/<name>publishes through the bucket API instead, with the same database-then-pointer order. - On the Spaces, once the Job runs: set
RLX_LISTING_FROM_STORE=1on both Spaces, so they read the Job'slisting.json.gz(a stat every 30 s, a read when it changes) instead of each fetching the Hub's whole listing every 15 minutes; a listing older than 3 h (the Job stopped) is fetched again as before./readyzreports the snapshot in use;GET /api/search/statussays when it last looked.GET /api/search/rank?key=gives one environment's trending rank among its kind (the Space page's "#599 of 6,818"), so no page loads the whole listing any more.
Production checklist (before the first deploy)
- Indexes and packs at INDEX_VERSION 13: rebuild (
uv run python -m app.precache --featured --top 40 --use-tokenor one indexer run) and sync toFineEnvs/rl-explorer-data; old-version files are rebuilt on demand, slowly. - MiMo repo snapshots (2,651) synced from
FineEnvs/mimo-explorer-runsinto the explorer's bucket. - Spaces:
FineEnvs/RL-Explorer(public; the sandbox reaches its/capture) and the private admin Space, both with the bucket mounted at/data,SESSION_SECRETand the OAuth app's secrets set. - Indexer Job:
uv run python scripts/schedule_indexer.py --applyonce the explorer Space has built; thenRLX_LISTING_FROM_STORE=1on both Spaces. - Smoke tests on the Space:
/readyz;scripts/check_judge_relay.py's path through the Space's/capture(a model-graded rollout, e.g. a MiMo General twin with Inkling as judge: reward recorded,judge_calls > 0);node tests/ui-audit.mjs https://<space>.
Admin source boundary
This public tooling directory contains only the Explorer. The dashboard application,
its assets, membership checks, tests and deployment tooling are maintained separately
in the private FineEnvs/RL-Explorer-admin Space. The Explorer reads shared settings
and moderation requests from the data bucket without importing the admin app.
Search indexing and social previews
The Space card uses the committed web/social/rl-explorer.png thumbnail, so a
sharing crawler does not have to start the application. HTML pages also publish
Open Graph and Twitter metadata with 1200Γ630 previews. Page-specific images fit
long names into the card; public task links retain their own canonical address
both in the initial HTML and after browser navigation.
/sitemap.xml links to shards of up to 10,000 public task URLs. Dataset tasks come
from local indexes or the immutable catalog snapshot, with MiMo entries deduplicated.
OpenEnv task coordinates come from advertised split counts on checked public
Spaces. Their sitemaps expand ranges one shard at a time, without downloading tasks
or creating millions of URLs in memory. These counts describe published entries,
not individually tested episodes or Google-indexed pages.
Space-check inventory refreshes outside request threads and never holds the reader lock during bucket I/O. Writes update the cached record immediately; concurrent refreshes preserve newer evidence. A cold process temporarily has no verified Spaces, and cached evidence still expires according to its original check time.
Task HTML includes public task text or structured input facts. Public row datasets and OpenEnv Task API records can be read anonymously on demand, with four concurrent reads and bounded caches. These reads do not build an index, wake a Space, run an episode, or forward a visitor's token. The existing answer-withholding rules apply. Invalid tasks return 404; temporary upstream failures return 503 with Retry-After. Private, gated and hidden content is excluded from sitemap generation.
Public read APIs may be fetched to render the interactive pages, but API responses
carry X-Robots-Tag: noindex. Rollouts, account pages and comparisons remain excluded.
Submit https://fineenvs-rl-explorer.hf.space/sitemap.xml for the matching URL-prefix
property in Google Search Console to monitor discovery and indexing. Robots.txt
also advertises it. Google chooses which pages to index; no application setting
can guarantee indexing or rankings. See Google's sitemap guidance:
https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap