Spaces:
Sleeping
A newer version of the Gradio SDK is available: 6.22.0
Architecture
This document explains how the agent works, step by step, for someone who has never seen the codebase before.
What this agent does
The agent answers real-world questions that may include file attachments (PDF, image, Excel, audio, etc.). Each answer should be a short, exact value β a number, a name, or a comma-separated list.
For each question it:
- Reads the question text (and attachment, if any).
- Uses tools and reasoning to find the answer.
- Returns only the answer value (no extra explanation).
- Can submit batches to a scoring API, which checks exact match against expected answers.
Big picture
Think of the system as a small factory line:
Scoring API β Pipeline β CodeAgent + Tools β Verifier β Submit answer
β β β
questions optional extras format check
attachments (plan, hints) + one retry
You can run this factory in two ways:
| Entry point | When you use it |
|---|---|
app.py |
Gradio UI on Hugging Face Space β click one button to run the full batch and submit |
run_local.py |
Your laptop β same logic, good for development and faster iteration |
Both call the same core class: GaiaAgent.
End-to-end flow (one question)
Here is what happens when you ask the agent a single question.
1. Question arrives
GaiaAgent receives:
questionβ the question textfile_pathβ optional local path to a downloaded attachmenttask_idβ optional ID for logging and artifacts
If the question has an attachment, the file is downloaded first (file_resolver.py β GET /files/{task_id} on the scoring API).
2. Pipeline prepares context (agent/pipeline.py)
Before the language model runs, the pipeline may add extra context to the prompt:
Strategy hints (optional) β agent/retriever.py
If RETRIEVER_ENABLED=1, the agent reads past runs from memory/experience.jsonl and injects hints like βfor numeric questions, using execute_code worked before.β It never injects actual answers from past runs β only strategies.
Plan (optional) β agent/planner.py
If PIPELINE_DEPTH is standard or full, one extra LLM call produces a short JSON plan (2β4 steps). At minimal depth (default), planning is skipped.
Prompt assembly β agent/agent_runner.build_prompt()
The final prompt includes the question, attachment hints (βthis is a PDF, use read_pdfβ), and any plan/hints above.
3. Think mode is chosen (agent/think_mode.py)
Some models support extended internal reasoning. The agent turns that on or off per question:
- Off for trivial questions (βopposite of leftβ)
- On for calculations, attachments, long questions, YouTube/arXiv lookups
Controlled by THINK_MODE=auto|on|off.
4. CodeAgent runs (agent/agent_runner.py)
The heart of the system is a smolagents CodeAgent: a language model that writes Python code blocks to call tools, inspect results, and eventually call final_answer("...").
Why code? Many questions need several steps: search the web, read a file, compute a number, then answer. The agent loops:
LLM writes code β tools run β results go back β LLM writes more code β β¦ β final_answer
Default limit: 12 steps (AGENT_MAX_STEPS).
The system prompt enforces answer format rules: short values, no βFINAL ANSWER:β prefix inside final_answer(), lists formatted as a, b, c.
5. Tools (agent/tools.py + tools.py)
The agent has one toolbox shared by all questions:
| Category | Tools | Typical use |
|---|---|---|
| Web | DuckDuckGo search, visit webpage, Wikipedia, fetch URL as markdown | Facts, current info |
| Academic | arXiv search | Paper titles, authors |
| Files | read PDF, Excel, CSV, text; OCR on images; transcribe audio; describe image | Attachments |
| Video | YouTube transcript | Questions about video content |
| Compute | execute_code (local Python/bash), math helpers (add, multiply, β¦) |
Counting, parsing |
| Browser | Playwright (if BROWSER_ENABLED=1) |
Heavy JavaScript pages |
Tools return text the model reads on the next step. Nothing runs unless the model asks for it in code.
6. Verifier checks the answer (agent/verifier.py)
The raw model output is not sent directly. The verifier:
- Strips hidden reasoning blocks some models emit
- Extracts the value from
final_answer("...")or similar patterns - Normalizes via
scoring.normalize_answer()(trim quotes, remove βFINAL ANSWER:β prefix) - Checks format β empty, multi-line, or overly long answers fail
- Spot-checks URLs mentioned in the trace (optional HTTP fetch)
- Critic pass (standard/full depth only) β optional second LLM asks βis this supported by evidence?β
If verification fails, the pipeline retries once with the verifierβs issues in the prompt. On retry, think mode is forced on.
7. Optional extras (depth flags)
PIPELINE_DEPTH |
What changes |
|---|---|
minimal (default) |
Agent + verifier + one retry. Fastest. |
standard |
+ planner, numeric double-compute hint, optional critic |
full |
+ self-correction loop (up to 3 rounds), optional majority vote (VOTE_RUNS) |
8. Answer returned and logged
The pipeline writes optional artifacts under artifacts/{task_id}/:
notes.mdβ short summaryevidence.jsonβ answer, depth, think mode, etc.plan.mdβ if planning ran
The string returned to app.py or run_local.py is the normalized answer only.
9. Submission (batch runs)
For full evaluation, app.py or run_local.py:
GET /questionsβ fetch all tasks from the scoring API- Run the pipeline on each
POST /submitwith{username, agent_code, answers: [{task_id, submitted_answer}, β¦]}- Scoring API returns score (% correct)
Grading is exact string match (case-insensitive after normalization). Ground truth stays on the server β the agent never stores expected answers locally.
Language model routing (model_provider.py)
The same agent code runs locally and on the Space, but the LLM backend changes:
| Where | Default backend | How |
|---|---|---|
| Your machine | llama.cpp or Ollama | Local server, no cloud credits |
| HF Space | Groq (if GROQ_API_KEY set) |
Fast cloud API, free tier; throttled + auto-retry on 429 |
| HF Space fallback | Cerebras β Google Gemini | When Groq limits hit; add keys only β models are chosen automatically |
| HF Space fallback | Hugging Face Inference | Needs HF_TOKEN; monthly credits often run out (402 error) |
Set LLM_PROVIDER explicitly to override: ollama, llamacpp, groq, cerebras, google, or hf.
Provider rotation (PROVIDER_FALLBACK_ENABLED=1, default on): Cerebras models first, then Gemini, then Groq. Order via PROVIDER_FALLBACK_ORDER=cerebras,google,groq. Hard limits (TPD, 413, context) skip to the next slot immediately (default 5s between calls).
Project map
Final_Assignment_Template/
βββ agent/
β βββ gaia_agent.py β public API (GaiaAgent)
β βββ pipeline.py β orchestrates the full flow
β βββ agent_runner.py β CodeAgent + system prompt
β βββ tools.py β extended tools
β βββ verifier.py β answer extraction + checks
β βββ planner.py β optional planning (standard/full)
β βββ retriever.py β optional strategy hints
β βββ think_mode.py β per-question reasoning toggle
β βββ self_correction.py β full-depth correction loops
β βββ voting.py β full-depth majority vote
β βββ evolve.py β post-eval memory (SELF_EVOLVE=1)
β βββ code_interpreter.pyβ local Python/bash execution
β βββ memory/store.py β experience.jsonl persistence
β βββ artifacts/writer.pyβ plan/notes/evidence files
βββ tools.py β base tools (search, files, audio, β¦)
βββ model_provider.py β LLM backend selection
βββ provider_chain.py β cross-provider fallback (Groq β Cerebras β Gemini)
βββ groq_model.py β throttling, retry, Groq model chain
βββ scoring.py β normalize_answer
βββ file_resolver.py β download attachments from scoring API
βββ app.py β Gradio Space UI
βββ run_local.py β CLI runner + score mode
βββ eval/ β fixtures, course client, run_eval.py
Data policy
- Allowed: scoring API, local fixtures you write, your own run history for strategy hints
- Not used: bulk benchmark downloads, storing ground-truth answers for retrieval
Configuration cheat sheet
See .env.example for the full list. The most important knobs:
| Variable | Effect |
|---|---|
LLM_PROVIDER |
Which LLM backend to use |
OLLAMA_MODEL / GROQ_API_KEY |
Local or Space model access |
GROQ_MODEL |
Groq model id (Space default: Scout 17B) |
GROQ_FALLBACK_ENABLED |
Auto-switch Groq models after hard limits (default 1) |
GROQ_MODEL_FALLBACK_CHAIN |
Comma-separated Groq fallback model ids |
CEREBRAS_API_KEY / GOOGLE_API_KEY |
Cross-provider fallback after Groq (no model config needed) |
PROVIDER_FALLBACK_ORDER |
Provider order, e.g. cerebras,google,groq |
GROQ_MIN_REQUEST_INTERVAL |
Seconds between cloud API calls (default 5) |
GROQ_MAX_RETRIES |
Retries on 429 (default 5) |
HF_USERNAME |
Your HF username for API submit |
PIPELINE_DEPTH |
minimal / standard / full |
THINK_MODE |
auto / on / off |
AGENT_MAX_STEPS |
Max tool loops per question (default 12) |
RETRIEVER_ENABLED |
Use past strategy hints |
SELF_EVOLVE |
Record learning after graded runs |
Mental model for debugging
When something goes wrong, ask:
- Model errors (402, 404, Groq)? β Check
model_provider.pyand Space secrets - Wrong answer but agent ran? β Read Space logs or
artifacts/trace; verifier may have stripped a bad format - Attachment ignored? β Check file downloaded; prompt includes path and suffix hint
- Timeout / too slow? β Reduce
AGENT_MAX_STEPS, use faster model, orPIPELINE_DEPTH=minimal - Score 0% locally but API says higher? β Local summary uses submit response; ensure
HF_USERNAMEis set
Testing without burning LLM credits
pytest tests/ -q # unit tests, mocked LLM
python eval/run_eval.py --mode fixtures # hand-written Q&A pairs
python run_local.py --mode single # one cheap smoke question
Fixtures live in eval/fixtures/ β you control the expected answers.