DATAD2's picture
Upload 14 files
c73f03c verified
|
Raw History Blame Contribute Delete
7.49 kB
---
title: GAIA Lite
emoji: 🤖
colorFrom: green
colorTo: blue
sdk: gradio
sdk_version: 5.25.2
app_file: app.py
pinned: false
hf_oauth: true
# optional, default duration is 8 hours/480 minutes. Max duration is 30 days/43200 minutes.
hf_oauth_expiration_minutes: 480
---
# 🤖 **GAIA Lite**
## 🌟 **Introduction**
**GAIA Lite** is a small tool-calling agent for learning LangGraph and for the GAIA / Hugging Face Agents Course homework. It keeps a short tool list, a bounded think–act loop, regex answer extraction, and a cached evaluation runner.
## 🚀 **Key Features**
- **🔍 Multi-Modal Search**: Web search, Wikipedia, and arXiv paper search
- **💻 Code Execution**: Support for Python, Bash, SQL, C, and Java
- **🖼️ Image Processing**: Analysis, transformation, OCR, and generation
- **📄 Document Processing**: PDF, CSV, Excel, and text file analysis
- **📁 File Upload Support**: Handle multiple file types with drag-and-drop
- **🧮 Mathematical Operations**: Complete set of mathematical tools
- **💬 Conversational Interface**: Natural chat-based interaction
- **📊 Evaluation System**: Automated benchmark testing and submission
## 🏗️ **Project Structure**
```
gaia-lite/
├── app.py # Gradio Q&A chatbot
├── evaluation_app.py # GAIA eval + cached submit
├── agent.py # Tools + LangGraph loop
├── extract_answer.py # Parse/normalize FINAL ANSWER
├── files_util.py # Attachment download + prompt paths
├── code_interpreter.py # Python (and other) code execution
├── image_processing.py # Optional image helpers (not wired)
├── system_prompt.txt # Agent instructions
├── answers_cache.json # Local eval cache (gitignored)
├── requirements.txt
└── README.md
```
## 🛠️ **Tools (kept small on purpose)**
The agent binds six tools. Extra math/image-generation helpers were removed from the loop so the model spends less time picking the wrong tool.
- `web_search` — Tavily, up to 3 hits
- `wiki_search` — Wikipedia, up to 2 pages
- `execute_python` — calculation, pandas, dates, CSV/Excel via code
- `download_file_from_url` — save a remote file locally
- `read_file` — preview text, CSV, Excel, PDF
- `extract_text_from_image` — Tesseract OCR
Vector-store retrieval is **off by default**. Pass `build_graph(use_retriever=True)` only if you explicitly want similar-question context.
## 🎯 **How to Use**
### **Q&A Chatbot Interface (app.py)**
1. **Start the Chatbot:**
```bash
python app.py
```
2. **Access the Interface:**
- Open `http://localhost:7860` in your browser
- Upload files (images, documents, CSV, etc.) if needed
- Ask questions in natural language
- Get comprehensive answers with tool usage
3. **Supported Interactions:**
- **Text Questions**: "What is the capital of France?"
- **Math Problems**: "Calculate the square root of 144"
- **Code Requests**: "Write a Python function to sort a list"
- **Image Analysis**: Upload an image and ask "What do you see?"
- **Data Analysis**: Upload a CSV and ask "What are the trends?"
- **Web Search**: "What are the latest AI developments?"
### **Evaluation Runner (evaluation_app.py)**
1. **Run the Evaluation:**
```bash
python evaluation_app.py
```
2. **Benchmark Testing:**
- Log in with your Hugging Face account
- Click "Run Evaluation & Submit All Answers"
- Monitor progress as the agent processes GAIA benchmark questions
- View results and scores automatically
## 🔧 **Technical Architecture**
### **LangGraph State Machine**
```
START → maybe_example → assistant ⇄ tools
↑ ↓
└───────────┘
```
1. **maybe_example**: no-op unless `use_retriever=True`
2. **assistant**: Groq Qwen with tools; after 8 tool rounds or a repeated identical call, it is forced to emit `FINAL ANSWER`
3. **tools**: `ToolNode` runs the selected Python functions
4. **extract_final_answer**: regex parse + light normalization before scoring submit
Evaluation writes `answers_cache.json`, skips successful tasks on rerun, and submits from cache in a separate button. Attachments use `GET /files/{task_id}`, with an optional Hugging Face GAIA dataset fallback.
### **Vector Database Integration**
- **Supabase Vector Store**: Stores GAIA benchmark Q&A pairs
- **Semantic Search**: Finds similar questions for context
- **HuggingFace Embeddings**: sentence-transformers/all-mpnet-base-v2
### **Multi-Modal File Support**
- **Images**: JPG, PNG, GIF, BMP, WebP
- **Documents**: PDF, DOC, DOCX, TXT, MD
- **Data**: CSV, Excel, JSON
- **Code**: Python, Bash, SQL, C, Java
## ⚙️ **Installation & Setup**
### Hugging Face Space secrets (most common crash)
If logs say `GROQ_API_KEY` is missing, open the Space → **Settings → Variables and secrets** and add at least:
- `GROQ_API_KEY` — from https://console.groq.com/keys
- `TAVILY_API_KEY` — for web search
Then **Restart** the Space. Do not put keys in public git.
### Local setup
```bash
git clone https://github.com/fisherman611/gaia-agent.git gaia-lite
cd gaia-lite
pip install -r requirements.txt
```
Create a `.env` file with your API keys:
```env
SUPABASE_URL=your_supabase_url
SUPABASE_SERVICE_ROLE_KEY=your_supabase_key
GROQ_API_KEY=your_groq_api_key
TAVILY_API_KEY=your_tavily_api_key
HUGGINGFACEHUB_API_TOKEN=your_hf_token
LANGSMITH_API_KEY=your_langsmith_key
LANGSMITH_TRACING=true
LANGSMITH_PROJECT=gaia-lite
LANGSMITH_ENDPOINT=https://api.smith.langchain.com
```
### **4. Database Setup (Supabase)**
Execute this SQL in your Supabase database:
```sql
-- Enable pgvector extension
CREATE EXTENSION IF NOT EXISTS vector;
-- Create match function for documents2 table
CREATE OR REPLACE FUNCTION public.match_documents_2(
query_embedding vector(768)
)
RETURNS TABLE(
id bigint,
content text,
metadata jsonb,
embedding vector(768),
similarity double precision
)
LANGUAGE sql STABLE
AS $$
SELECT
id,
content,
metadata,
embedding,
1 - (embedding <=> query_embedding) AS similarity
FROM public.documents2
ORDER BY embedding <=> query_embedding
LIMIT 10;
$$;
-- Grant permissions
GRANT EXECUTE ON FUNCTION public.match_documents_2(vector) TO anon, authenticated;
```
## 🚀 **Running the Application**
### **Chatbot Interface**
```bash
python app.py
```
Access at: `http://localhost:7860`
### **Evaluation Runner**
```bash
python evaluation_app.py
```
Access at: `http://localhost:7860`
### **Live Demo**
Upstream template Space: [fisherman611/gaia-agent](https://huggingface.co/spaces/fisherman611/gaia-agent)
## 🔗 **Resources**
- [GAIA Benchmark](https://huggingface.co/spaces/gaia-benchmark/leaderboard)
- [Hugging Face Agents Course](https://huggingface.co/agents-course)
- [LangGraph Documentation](https://langchain-ai.github.io/langgraph/)
- [Supabase Vector Store](https://supabase.com/docs/guides/ai/vector-columns)
## 🤝 **Contributing**
Contributions are welcome! Areas for improvement:
- **New Tools**: Add specialized tools for specific domains
- **UI Enhancements**: Improve the chatbot interface
- **Performance**: Optimize response times and accuracy
- **Documentation**: Expand examples and use cases
## 📄 **License**
This project is licensed under the [MIT License](https://mit-license.org/).