RL-Explorer / DESIGN.md
AdithyaSK's picture
AdithyaSK HF Staff
Deploy HF RL Explorer
858027d verified
|
Raw History Blame Contribute Delete
24.4 kB

RL Explorer: design

A Space for exploring the RL environments on the Hugging Face Hub, whatever their framework:

  • Harbor task datasets (folders with task.toml), indexed and run on HF Sandboxes;
  • OpenEnv Spaces, live: woken or restarted, their own app, a playground with rewards, their Task API, and an MCP server for coding agents; OpenEnv Γ— Harbor Spaces with what they can run;
  • custom environments: datasets whose tasks are rows (Verifiers, NeMo Gym, verl, MiMo's raw release, Harbor tasks packed into parquet, traces, any table of prompts), each read by a processor (CUSTOM_ENVS.md).

Browse an environment's tasks, read exactly what an agent gets and how it is graded, run an agent on a task in a sandbox, watch it, and compare rollouts. The MiMo RL Environment Explorer did this for one release; MiMo is one environment among the others here (its raw rows through the mimo processor, linked to its Harbor twins).

What a Harbor dataset is

A task is a folder with task.toml (settings, metadata), instruction.md (the prompt), environment/ (a Dockerfile, or a prebuilt docker_image in task.toml, sometimes a compose file), tests/ (the verifier: test.sh writes /logs/verifier/reward.txt, one number, or reward.json, named scores) and often solution/. A survey of the 30 most-downloaded harbor datasets (2026-10-03) found:

Layout Examples Share
tasks/<task>/ terminal-bench-2.1/3.0, MiMo, data-agent, repo2rlenv ~60%
tasks/<group>/.../<task>/ terminal-bench-science few
<task>/ at the repository root terminal-bench-2.0, harbor-mix, WildClawBench, dabstep ~25%
<benchmark>/tasks/<task>/ skilltrainbench-public, Lego-RL few
packed (tasks.parquet, .tar, .zip) TaskTrove, TerminalWorld, FACET ~10%: read as rows (below)

So a task is found by where its task.toml is, not by a fixed path. Most tasks ship only a Dockerfile; a prebuilt docker_image is the minority (terminal-bench-2.1, LHTB, MiMo, apex). Metadata keys differ per dataset, and some are answers (gold_answer, expected_output, solution): those are never shown.

Architecture

Two Spaces, one image, one private bucket (FineEnvs/rl-explorer-data, mounted at /data by both):

HF RL Explorer (public Space, app.main)              Admin (private Space, app.admin_app, FineEnvs only)
  catalog: Harbor datasets + OpenEnv Spaces            overview, rollouts (with who ran them), environments
  indexes and packs (/data/indexes, /data/packs)       (pin, hide, feature, index), collections, indexes,
  runner: OpenEnv's run_rollout + capture proxy        settings (rollout switch, limits, agents, banner), audit
  (/capture), HF Sandbox on the visitor's account               β”‚
  store: /data/runs                                             β”‚ writes /data/admin/settings.json (+ audit.jsonl)
        β–²  re-reads settings.json (mtime) β—„β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ reads /data/runs, builds indexes

The admin shares the bucket, not the process: pins, hidden environments, collections, rollout switches, rollouts taken off Community and requests to stop a running rollout all go through the settings file, which the explorer re-reads within seconds. The admin Space is private (only org members reach it) and every admin route checks membership again (public member list, or the memberships the Hub reported at sign-in).

  • Catalog. Datasets tagged harbor and Spaces tagged openenv or rl-environment (Docker, not articles), plus the collections admins curate (a dataset id, or space:<id>), and the signed-in visitor's own datasets, private ones included. Private and gated datasets are read with the visitor's token; every request re-checks access.
  • Index and pack. On first open (or ahead of time with python -m app.precache), one listing finds every task.toml; the files a task page shows are fetched in parallel (one GET each, 48 at a time). A repository too big to list in 25 s (tasks that each ship a whole app, like microsoft/ProgramDistill's 4,375) is walked instead: a level is listed, the Hub's batched paths-info says which of its folders hold a task.toml (500 a call), and only the files the index reads are looked up; the task list is published as soon as it's known (the page opens, and fills in as files arrive), and a task's other folders are listed when someone opens them in the file viewer. The index is one row per task (title, facets, grading, environment, network policy, whether it can run here); the pack holds every task's text, de-duplicated, and its file list, so task pages never wait on the Hub. The first open shows a step timeline with live counts.
  • Answers stay out. Answer-like metadata keys, solution/, data files beside the grader (rubrics, expected outputs), and answers written into grader scripts (answer-named heredocs and literals) are withheld or masked.
  • Rollouts. openenv.harbor.rollout.run_rollout, the same code as OpenEnv's Harbor UI. The agent (OpenCode, Terminus 2, mini-SWE-agent, Pi) runs in an HF Sandbox from the task's image, or, for a Dockerfile-only task, from the Dockerfile's base image with its RUN/COPY/ENV/WORKDIR replayed first (app/dockerfile.py; about 99% of the indexed Dockerfiles qualify). Sandbox.create is wrapped so the sandbox runs on the visitor's account. Tasks that ask for no internet are refused with the reason: HF Sandboxes can't enforce it. Models: Inference Providers through the router (:fastest) with the visitor's token, or an endpoint they bring (re-checked before use).
  • OpenEnv Spaces, live (app/spaces_live.py). What a server offers comes from itself: its /openapi.json (routes: reset/step, the Task API, custom ones), /metadata, /schema, /list_environments and tools/list on /mcp; its web page from the README's base_path, then /web, then /. A sleeping Space is woken the way its Hub page lets any visitor (POST /spaces/<id>/start, at most once a minute); a broken one is restarted with the visitor's own token, which works only if they may write to it. The playground holds one WebSocket to the server's /ws per session (reset, step, state, and MCP calls as {"type": "mcp"} on the same episode: the HTTP routes are stateless), shows each observation (images, audio, chats) and reward, and plays a Task API task by its split and index. OpenEnv Γ— Harbor servers show their engine, sandboxes, validated agents and served datasets, mapped back to the Hub (/data/org__name is org/name) so their tasks run on the visitor's account here. Other servers (NeMo Gym, ORS, any FastAPI app) are used through their own routes, read from their OpenAPI: a form per route from its request schema, called with the session's own cookie jar (NeMo Gym keeps an episode in a cookie; it stays on the explorer). Only plain paths the server published, GET or POST, path parameters checked; GET answers drop answer-like fields. OpenEnv servers' own extra routes aren't offered (geoguesser's return the answer's coordinates): they're used through reset/step and tools.
  • MCP for coding agents (app/mcp_bridge.py, /mcp/<org>/<name>). OpenEnv's /mcp answers tools/list and tools/call but not MCP's handshake, so Claude Code, Codex, Cursor, VS Code, Gemini CLI and the rest can't use it directly. The bridge speaks Streamable HTTP: the Space's own tools in a session of their own, plus reset/step/ state for servers without tools (step's input is the action schema), plus list_splits/list_tasks/get_task for the Task API, plus one tool per route for servers that aren't OpenEnv (NeMo Gym's seed_session, guess, verify); images come back as image content. Sessions are capped per address, rate-limited, and end idle.
  • One contract for every format (app/envs/contract.py, CUSTOM_ENVS.md). An environment is read by an adapter that answers five questions: its tasks (a summary, cards, each task as sections of blocks, its files), what they run in, how they can run (run options, each by a runner, with its own inputs), which other tasks are the same task (aliases: their rollouts are shared), and nothing more for MCP. So one environment page (/d/<org>/<name>, env.js) and one task page (/t/<org>/<name>/<ref>, task.js) serve Harbor folders, the MiMo release and any rows dataset; a raw MiMo row and its Harbor conversion link to each other and list each other's rollouts. Adapters: harbor.py (the index above), mimo.py, rows.py (with processors: rows from the Hub's dataset viewer, or straight from JSON Lines / JSON / parquet / CSV; a processor turns a row into a task; answers left out by name at any depth). The registry picks the surest adapter that confirms; linkers and aliasers relate tasks across adapters.
  • The MiMo release (app/mimo, app/envs/mimo.py). The MiMo RL Environment Explorer's backend, kept close to its source: all 7,780 tasks with the brief, the agent's exact prompt, systems and their databases, workspace previews, Code tasks' repositories (snapshots in /data/repo-snapshots), the grader with its judge prompts, bands and music features, and the reward design across the release. Two runners: its Harbor conversion (the default) or its own harness (OpenCode in the task's image, Xiaomi's graders, a judge model the visitor picks), whose sandbox reaches models through /api/llm/<capability> and never holds a token. Its run page (events, live), compare view and community (leaderboard by domain, the tasks nobody has tried) work for every environment.
  • MCP for environments (/mcp/d/<org>/<name>): describe, list and search tasks, read one in full as text, read its files, a random task; answers withheld as on the page. Any adapter, no code.
  • Visibility. Public rollouts appear in Community without who ran them, private ones only to their owner; only finished, graded rollouts are shared; a private dataset's rollouts stay private; admins can take one off.

Look

The MiMo explorer's design (its web/app.css, unchanged; web/rlx.css, fv.css and sp.css add the rest), with the Hugging Face logo and the name HF RL Explorer. Explore opens with FineEnvs' banner (one link, closable for a week), the numbers, two rows of trending environments with a pager, then one search, filters (kind: Harbor, OpenEnv, Verifiers, NeMo Gym, other) and cards. Elsewhere the header's search lists trending environments before you type. Files open in an editor-like viewer (folders, highlighting, line links, full view) that fits whatever column it's in.

Security

uv run pytest tests -q (no network). test_security.py: sessions can't be forged or outlive their expiry; admin routes are org-only and absent from the public app; cross-site writes are refused; private rollouts and datasets stay private, per token; public rollouts carry no user or endpoint; answers are withheld or masked; task files stay inside their task; visitor endpoints must be public https; tokens never reach the store; headers. test_live.py: only a Space's own hf.space host is ever called (redirects elsewhere aren't followed, base_path can't point off it), private Spaces are refused, waking is limited, restarting needs write access, playground sessions belong to who started them and are capped, the Task API browser drops answers, and the MCP bridge speaks the protocol (handshake, notifications, errors as tool results, sessions, rate limit, no cross-origin pages). test_envs.py: processors claim their formats, answers stay out (nested, packed solutions, MiMo's patches), archives can't escape or blow up, filters can't inject, the direct reader pages right, and the walker never lists a task's own folders. test_contract.py: a toy adapter registered from outside works end to end (pages, private keys stripped, runs refused unless the chosen option can run, its fields checked, rollouts merged across aliases, MCP). test_mimo.py: the MiMo adapter's tiles, facets and map, row refs, the default runner, judges from the pool only, no token in a MiMo run record, links back from the Harbor twins, scrubbed public events, sandboxed raw files. tests/ui-audit.mjs clicks through every page at four widths in both themes (needs network).

Never sent to a Space: the visitor's token. A Space's run_rollout is offered only when it has its own engine.

Next

A Space adapter (an OpenEnv server's Task API tasks as task pages, its playground as a run option); a folder-of-tasks processor for repositories without task.toml; model token costs for Harbor rollouts; MiMo's repository snapshots synced into this bucket.

Data layer

The catalog is built offline and read as an immutable SQLite snapshot, so no app process fetches the Hub's listing or builds indexes to answer a page, and the home page asks the server for one page of cards instead of downloading every environment (about 4 MB of JSON) to filter it in the browser.

HF scheduled Job, hourly (scripts/schedule_indexer.py)          Explorer / admin Spaces (app/snapshot.py)
  python -m app.indexer --store /data                              every 60 s: read snapshots/LATEST.v1.json
   1 listing   catalog._listing_fetch (every Space, ~2.5 min)       new? copy the .db to local disk, check sha256,
   2 indexes   rebuild what changed (catalog's builders), budget        open read-only + immutable, check schema,
   3 build     SQLite on local disk: envs, tasks, FTS, indexes          counts, quick_check; swap (queries in flight
   4 publish   snapshots/catalog-<UTC>.db, then the pointer             finish on the old one)
   5 prune     keep 5 (and whatever a pointer names)                  none / corrupt: build the same DB from the
        β”‚                                                              live catalog (local dev, first boot)
        └──── bucket FineEnvs/rl-explorer-data, mounted at /data ──────────▲
  • Indexer (app/indexer.py). One run fetches the listing and refuses it when it is much smaller than the last snapshot's (overall, datasets or Spaces under 80%: catalog skips a Hub tag listing that fails, and a half listing must not empty the site; --force overrides). It writes listing.json.gz for the app's own code paths, then rebuilds dataset indexes incrementally: a dataset whose index is current (same INDEX_VERSION, Hub lastModified older than the index's build) costs no Hub call; one touched around or after its build has its revision checked; changed revisions, older formats, indexes left partial and never-indexed Harbor datasets are rebuilt in that order (featured first, then by trending) with catalog's own build_index, within a time and count budget. A build that fails isn't retried at the same revision for a day (indexer-state.json). Then it builds the snapshot on local disk (journal_mode=OFF, synchronous=OFF, rows streamed one dataset's index at a time, FTS rebuilt, indexes, ANALYZE, VACUUM, quick_check), copies it to snapshots/catalog-<UTC ts>.db, checks it landed whole (size, and sha256 when the store can be re-read), and only then writes the pointer snapshots/LATEST.v<SCHEMA>.json ({"db", "sha256", "size", "built_at", "counts", "schema"}) atomically. A run that fails anywhere before that publishes nothing. Old snapshots are pruned (the newest 5, and any a pointer of any schema names). Reruns are safe: current indexes are skipped and every run publishes a complete snapshot.
  • Snapshot (app/snapshot.py). envs (one row per public environment: id, key, kind, framework, heading, brief, counters, dates, stage, MCP, OpenEnv version, badges and tags as JSON, indexed task count and revision, the facet values that don't depend on admin settings, and a lowercased search blob), envs_fts (FTS5, trigram: substring search, as the page always did), tasks (dataset, ref, title, brief, category, difficulty, group, how it runs, how it's graded, tags, a few extras), tasks_fts (FTS5, unicode61, token prefixes), meta. Only what the explorer already shows publicly: no private or gated dataset, no index a visitor built with their own token, no metadata values, no file texts. 7,225 environments and 51,498 tasks (the Harbor indexes and MiMo's release) made a 43 MB file, built in about a second; a run with the full listing took under five minutes.
  • Reading it. Each process checks the pointer on startup and every 60 s; SQLite never opens a file on the bucket mount (the database is copied to RLX_CACHE_DIR/snapshots first), connections are read-only (mode=ro&immutable=1, a small pool per snapshot), and a swap retires the old snapshot when its last query is done. A pointer that names a missing, corrupt, truncated or mismatched database (sha256, schema, row counts, quick_check) is rejected and logged (snapshot.rejected), the current one kept, and the same file isn't fetched again for ten minutes. With no snapshot at all, the same database is built in the process from catalog.environments() and the indexes on disk, so local development and a first boot work without the Job. Bumping SCHEMA publishes under a new pointer name: an app of the old version keeps reading its own.
  • What admins change is applied per query. Hidden environments, pins and collections come from the admin settings on every request (an admin's change shows within seconds, between indexer runs), as do public rollout counts.
  • Search API (app/search_api.py, a router). GET /api/search?q=&kind=&collection=&f=&sort=&page=&size= returns a page of cards, the total and every facet's counts, with the page's exact semantics: the Kind taxonomy (OpenEnv Spaces including FineEnvs' curated servers, other Spaces including ORS, Harbor whenever an index found task folders, Verifiers, NeMo Gym, OpenEnv datasets, verl, other rows), pinned first in trending, FineEnvs' curated servers first among OpenEnv Spaces, each facet counted over what the other filters leave, ties in listing order. trending=24 adds the trending row and header numbers to the same response (one request for the first paint); mine=1 merges the signed-in visitor's own datasets, private ones included (a private response). Search words are data: each is a quoted FTS string (and checked again as a plain substring), so operators, NEAR, *, column filters and stray quotes can't change the query. GET /api/search/tasks?q=&env= searches every indexed task (each word a quoted prefix, best match first). Responses carry Cache-Control (public 30 s, private for mine), a weak ETag (304 on a repeat) and Server-Timing. A parity test compares every answer, order and facet with a plain-Python port of the page's old filtering.
  • Home page (web/js/home.js): the first paint is one request of 12 KB on the wire (55 KB of JSON; it was 458 KB on the wire, 4 MB of JSON), and the first card shows in 283 ms locally instead of 467 ms (7.1 s to 5.2 s on DevTools' Fast 3G, where the data request itself went from 3.4 s to 0.65 s); search (debounced), filters, sort and paging ask the server, an older request is cancelled when a newer one starts, a slow answer shows placeholder cards of the real size, and the URL keeps the page's state as before.
  • Scheduling (scripts/schedule_indexer.py, prints the Job; --apply creates it): the explorer Space's own image, @hourly with concurrency off, cpu-upgrade, a 55-minute timeout, the bucket mounted read-write at /data (STORAGE_DIR=/data; written files upload on close, and a rename is one bucket batch, so temp-then-rename and the pointer-last order hold for readers on other mounts, which see changes within ~10–30 s), HF_TOKEN as a Job secret (higher rate limits for public downloads only). --store hf://buckets/<org>/<name> publishes through the bucket API instead, with the same database-then-pointer order.
  • On the Spaces, once the Job runs: set RLX_LISTING_FROM_STORE=1 on both Spaces, so they read the Job's listing.json.gz (a stat every 30 s, a read when it changes) instead of each fetching the Hub's whole listing every 15 minutes; a listing older than 3 h (the Job stopped) is fetched again as before. /readyz reports the snapshot in use; GET /api/search/status says when it last looked. GET /api/search/rank?key= gives one environment's trending rank among its kind (the Space page's "#599 of 6,818"), so no page loads the whole listing any more.

Production checklist (before the first deploy)

  1. Indexes and packs at INDEX_VERSION 13: rebuild (uv run python -m app.precache --featured --top 40 --use-token or one indexer run) and sync to FineEnvs/rl-explorer-data; old-version files are rebuilt on demand, slowly.
  2. MiMo repo snapshots (2,651) synced from FineEnvs/mimo-explorer-runs into the explorer's bucket.
  3. Spaces: FineEnvs/RL-Explorer (public; the sandbox reaches its /capture) and the private admin Space, both with the bucket mounted at /data, SESSION_SECRET and the OAuth app's secrets set.
  4. Indexer Job: uv run python scripts/schedule_indexer.py --apply once the explorer Space has built; then RLX_LISTING_FROM_STORE=1 on both Spaces.
  5. Smoke tests on the Space: /readyz; scripts/check_judge_relay.py's path through the Space's /capture (a model-graded rollout, e.g. a MiMo General twin with Inkling as judge: reward recorded, judge_calls > 0); node tests/ui-audit.mjs https://<space>.

Admin source boundary

This public tooling directory contains only the Explorer. The dashboard application, its assets, membership checks, tests and deployment tooling are maintained separately in the private FineEnvs/RL-Explorer-admin Space. The Explorer reads shared settings and moderation requests from the data bucket without importing the admin app.

Search indexing and social previews

The Space card uses the committed web/social/rl-explorer.png thumbnail, so a sharing crawler does not have to start the application. HTML pages also publish Open Graph and Twitter metadata with 1200Γ—630 previews. Page-specific images fit long names into the card; public task links retain their own canonical address both in the initial HTML and after browser navigation.

/sitemap.xml links to shards of up to 10,000 public task URLs. Dataset tasks come from local indexes or the immutable catalog snapshot, with MiMo entries deduplicated. OpenEnv task coordinates come from advertised split counts on checked public Spaces. Their sitemaps expand ranges one shard at a time, without downloading tasks or creating millions of URLs in memory. These counts describe published entries, not individually tested episodes or Google-indexed pages.

Space-check inventory refreshes outside request threads and never holds the reader lock during bucket I/O. Writes update the cached record immediately; concurrent refreshes preserve newer evidence. A cold process temporarily has no verified Spaces, and cached evidence still expires according to its original check time.

Task HTML includes public task text or structured input facts. Public row datasets and OpenEnv Task API records can be read anonymously on demand, with four concurrent reads and bounded caches. These reads do not build an index, wake a Space, run an episode, or forward a visitor's token. The existing answer-withholding rules apply. Invalid tasks return 404; temporary upstream failures return 503 with Retry-After. Private, gated and hidden content is excluded from sitemap generation.

Public read APIs may be fetched to render the interactive pages, but API responses carry X-Robots-Tag: noindex. Rollouts, account pages and comparisons remain excluded. Submit https://fineenvs-rl-explorer.hf.space/sitemap.xml for the matching URL-prefix property in Google Search Console to monitor discovery and indexing. Robots.txt also advertises it. Google chooses which pages to index; no application setting can guarantee indexing or rankings. See Google's sitemap guidance: https://developers.google.com/search/docs/crawling-indexing/sitemaps/build-sitemap