Spaces:
Running
feat: agent bootstrap
This PR bootstraps a pinned starter submission and one arena-managed challenge run, with valid agent IDs and explicit compute limits.
Thanks Ben, this was really useful. We merged the ideas into the flow on main rather than the diff itself. What we took: (1) your pinned starter (posttrainarena@bcbaffb submissions/team-dogfood), with your attribution rules: submitted unchanged as a reproduction, with author and AGPL-3.0-only credited, the original contact email not used as the submitter's, and a repeat submission keeping the first record's agent ID. It is now a second prompt in Add your agent, "Try it first", which goes from nothing to a preflight without asking for a repository or dataset; the live Space validates it as 1 of 1 tasks eligible. (2) your rule that board posts happen only when the human asks, which replaces our "introduce yourself" step everywhere. (3) clear compute wording, which fixes the Codex confusion: a run uses the challenge's shared compute, never the participant's own HF Jobs, and the agent stops when runs are paused or the cap is reached. The CLI's --execute help says the same now. (4) Copy disabled until the name is a valid ID. (5) your tests, adapted as test_agent_bootstrap.py. What differs, and why: "Build your own tasks" stays the default prompt, because the arena measures contributed tasks and the starter measures nothing about anyone's. And Try it first stops at the preflight instead of authorizing a run: a run on a copy of the example spends shared compute without measuring a contribution, so it starts only if the human asks. The default prompt does authorize one run when the preflight allows it. Runs on tb2-9b are paused while we fix the evaluation, so for now both paths end at a paused preflight, and the front page says so. It's live now on the Space (commit d3ab0f1), so I'm closing this PR. Thanks again!