--- title: GAIA Lite emoji: ๐Ÿค– colorFrom: green colorTo: blue sdk: gradio sdk_version: 5.25.2 app_file: app.py pinned: false hf_oauth: true # optional, default duration is 8 hours/480 minutes. Max duration is 30 days/43200 minutes. hf_oauth_expiration_minutes: 480 --- # ๐Ÿค– **GAIA Lite** ## ๐ŸŒŸ **Introduction** **GAIA Lite** is a small tool-calling agent for learning LangGraph and for the GAIA / Hugging Face Agents Course homework. It keeps a short tool list, a bounded thinkโ€“act loop, regex answer extraction, and a cached evaluation runner. ## ๐Ÿš€ **Key Features** - **๐Ÿ” Multi-Modal Search**: Web search, Wikipedia, and arXiv paper search - **๐Ÿ’ป Code Execution**: Support for Python, Bash, SQL, C, and Java - **๐Ÿ–ผ๏ธ Image Processing**: Analysis, transformation, OCR, and generation - **๐Ÿ“„ Document Processing**: PDF, CSV, Excel, and text file analysis - **๐Ÿ“ File Upload Support**: Handle multiple file types with drag-and-drop - **๐Ÿงฎ Mathematical Operations**: Complete set of mathematical tools - **๐Ÿ’ฌ Conversational Interface**: Natural chat-based interaction - **๐Ÿ“Š Evaluation System**: Automated benchmark testing and submission ## ๐Ÿ—๏ธ **Project Structure** ``` gaia-lite/ โ”œโ”€โ”€ app.py # Gradio Q&A chatbot โ”œโ”€โ”€ evaluation_app.py # GAIA eval + cached submit โ”œโ”€โ”€ agent.py # Tools + LangGraph loop โ”œโ”€โ”€ extract_answer.py # Parse/normalize FINAL ANSWER โ”œโ”€โ”€ files_util.py # Attachment download + prompt paths โ”œโ”€โ”€ code_interpreter.py # Python (and other) code execution โ”œโ”€โ”€ image_processing.py # Optional image helpers (not wired) โ”œโ”€โ”€ system_prompt.txt # Agent instructions โ”œโ”€โ”€ answers_cache.json # Local eval cache (gitignored) โ”œโ”€โ”€ requirements.txt โ””โ”€โ”€ README.md ``` ## ๐Ÿ› ๏ธ **Tools (kept small on purpose)** The agent binds six tools. Extra math/image-generation helpers were removed from the loop so the model spends less time picking the wrong tool. - `web_search` โ€” Tavily, up to 3 hits - `wiki_search` โ€” Wikipedia, up to 2 pages - `execute_python` โ€” calculation, pandas, dates, CSV/Excel via code - `download_file_from_url` โ€” save a remote file locally - `read_file` โ€” preview text, CSV, Excel, PDF - `extract_text_from_image` โ€” Tesseract OCR Vector-store retrieval is **off by default**. Pass `build_graph(use_retriever=True)` only if you explicitly want similar-question context. ## ๐ŸŽฏ **How to Use** ### **Q&A Chatbot Interface (app.py)** 1. **Start the Chatbot:** ```bash python app.py ``` 2. **Access the Interface:** - Open `http://localhost:7860` in your browser - Upload files (images, documents, CSV, etc.) if needed - Ask questions in natural language - Get comprehensive answers with tool usage 3. **Supported Interactions:** - **Text Questions**: "What is the capital of France?" - **Math Problems**: "Calculate the square root of 144" - **Code Requests**: "Write a Python function to sort a list" - **Image Analysis**: Upload an image and ask "What do you see?" - **Data Analysis**: Upload a CSV and ask "What are the trends?" - **Web Search**: "What are the latest AI developments?" ### **Evaluation Runner (evaluation_app.py)** 1. **Run the Evaluation:** ```bash python evaluation_app.py ``` 2. **Benchmark Testing:** - Log in with your Hugging Face account - Click "Run Evaluation & Submit All Answers" - Monitor progress as the agent processes GAIA benchmark questions - View results and scores automatically ## ๐Ÿ”ง **Technical Architecture** ### **LangGraph State Machine** ``` START โ†’ maybe_example โ†’ assistant โ‡„ tools โ†‘ โ†“ โ””โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”€โ”˜ ``` 1. **maybe_example**: no-op unless `use_retriever=True` 2. **assistant**: Groq Qwen with tools; after 8 tool rounds or a repeated identical call, it is forced to emit `FINAL ANSWER` 3. **tools**: `ToolNode` runs the selected Python functions 4. **extract_final_answer**: regex parse + light normalization before scoring submit Evaluation writes `answers_cache.json`, skips successful tasks on rerun, and submits from cache in a separate button. Attachments use `GET /files/{task_id}`, with an optional Hugging Face GAIA dataset fallback. ### **Vector Database Integration** - **Supabase Vector Store**: Stores GAIA benchmark Q&A pairs - **Semantic Search**: Finds similar questions for context - **HuggingFace Embeddings**: sentence-transformers/all-mpnet-base-v2 ### **Multi-Modal File Support** - **Images**: JPG, PNG, GIF, BMP, WebP - **Documents**: PDF, DOC, DOCX, TXT, MD - **Data**: CSV, Excel, JSON - **Code**: Python, Bash, SQL, C, Java ## โš™๏ธ **Installation & Setup** ### Hugging Face Space secrets (most common crash) If logs say `GROQ_API_KEY` is missing, open the Space โ†’ **Settings โ†’ Variables and secrets** and add at least: - `GROQ_API_KEY` โ€” from https://console.groq.com/keys - `TAVILY_API_KEY` โ€” for web search Then **Restart** the Space. Do not put keys in public git. ### Local setup ```bash git clone https://github.com/fisherman611/gaia-agent.git gaia-lite cd gaia-lite pip install -r requirements.txt ``` Create a `.env` file with your API keys: ```env SUPABASE_URL=your_supabase_url SUPABASE_SERVICE_ROLE_KEY=your_supabase_key GROQ_API_KEY=your_groq_api_key TAVILY_API_KEY=your_tavily_api_key HUGGINGFACEHUB_API_TOKEN=your_hf_token LANGSMITH_API_KEY=your_langsmith_key LANGSMITH_TRACING=true LANGSMITH_PROJECT=gaia-lite LANGSMITH_ENDPOINT=https://api.smith.langchain.com ``` ### **4. Database Setup (Supabase)** Execute this SQL in your Supabase database: ```sql -- Enable pgvector extension CREATE EXTENSION IF NOT EXISTS vector; -- Create match function for documents2 table CREATE OR REPLACE FUNCTION public.match_documents_2( query_embedding vector(768) ) RETURNS TABLE( id bigint, content text, metadata jsonb, embedding vector(768), similarity double precision ) LANGUAGE sql STABLE AS $$ SELECT id, content, metadata, embedding, 1 - (embedding <=> query_embedding) AS similarity FROM public.documents2 ORDER BY embedding <=> query_embedding LIMIT 10; $$; -- Grant permissions GRANT EXECUTE ON FUNCTION public.match_documents_2(vector) TO anon, authenticated; ``` ## ๐Ÿš€ **Running the Application** ### **Chatbot Interface** ```bash python app.py ``` Access at: `http://localhost:7860` ### **Evaluation Runner** ```bash python evaluation_app.py ``` Access at: `http://localhost:7860` ### **Live Demo** Upstream template Space: [fisherman611/gaia-agent](https://huggingface.co/spaces/fisherman611/gaia-agent) ## ๐Ÿ”— **Resources** - [GAIA Benchmark](https://huggingface.co/spaces/gaia-benchmark/leaderboard) - [Hugging Face Agents Course](https://huggingface.co/agents-course) - [LangGraph Documentation](https://langchain-ai.github.io/langgraph/) - [Supabase Vector Store](https://supabase.com/docs/guides/ai/vector-columns) ## ๐Ÿค **Contributing** Contributions are welcome! Areas for improvement: - **New Tools**: Add specialized tools for specific domains - **UI Enhancements**: Improve the chatbot interface - **Performance**: Optimize response times and accuracy - **Documentation**: Expand examples and use cases ## ๐Ÿ“„ **License** This project is licensed under the [MIT License](https://mit-license.org/).