File size: 7,294 Bytes
33516f7
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
# React Agent

A lightweight LangChain ReAct agent for querying the Future of Education podcast knowledge base.

## Overview

This agent implements the **ReAct pattern** (Reason β†’ Act/tool call β†’ Observe β†’ repeat) using `langchain.agents.create_agent`. It's intended for fast, tool-grounded Q&A over the local stores (Chroma + Neo4j).

**Features:**
- **Conversation memory** via LangGraph's SQLite checkpointer (survives process restart)
- **Multi-provider support** (OpenAI, Anthropic, Google, Fireworks)
- **Streaming** tool calls and responses

**When to use this vs deep_research_agent:**
- **React Agent**: Quick questions, single-topic lookups, testing tools
- **Deep Research Agent**: Complex multi-part research, comprehensive reports, cross-episode synthesis

## File Structure

| File | Purpose |
|------|---------|
| `graph.py` | Main agent using `langchain.agents.create_agent` |
| `checkpointer.py` | SqliteSaver helper (gitignored `.checkpoints/react_agent.sqlite`) |
| `prompts.py` | System prompt guiding agent behavior |
| `configuration.py` | Settings (model, max iterations, etc.) with model registry |
| `utils.py` | Helper functions for model creation and tool management |
| `tools.py` | Re-exports shared tools (Chroma, Neo4j) |
| `chat_cli.py` | Interactive terminal interface |
| `__init__.py` | Package exports |

## Architecture

```
User Query
    β”‚
    β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  ReAct Agent (create_agent)         β”‚
β”‚  β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”‚
β”‚  β”‚ 1. Reason: What do I need?  β”‚    β”‚
β”‚  β”‚ 2. Act: Call a tool         │◄───┼─── Tools:
β”‚  β”‚ 3. Observe: Read result     β”‚    β”‚    β€’ search_knowledge_base (Chroma)
β”‚  β”‚ 4. Repeat or Answer         β”‚    β”‚    β€’ query_knowledge_graph (Neo4j)
β”‚  β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β”‚    β€’ inspect_graph_schema
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
    β”‚
    β–Ό
  Answer
```

## Implementation Details

### Dynamic Schema Loading
The agent dynamically loads the Neo4j graph schema at startup (in `prompts.py`) to minimize tool calls. This allows the agent to write Cypher queries immediately without needing an initial `inspect_graph_schema` call, reducing latency and cost.

### Multi-Provider Support

Model creation is delegated to the shared factory in `src/llm/factory.py`, which supports multiple providers and validates the required API key env var for the chosen model.

Model names accept **stable aliases** (recommended) that are mapped to provider-specific IDs under the hood:
- **OpenAI**: `gpt-5`, `gpt-5-mini`
- **Anthropic**: `claude-sonnet-4-20250514` (alias), plus internal IDs like `claude-sonnet-4-5`
- **Google Gemini**: `gemini-flash-latest` (alias)

### Model Configuration

Models are configured via the `Configuration` class which supports:
- Model selection (default: `gemini-flash-latest`)
- Max tokens (default: 4000)
- Temperature (default: 0.0)
- Timeout (default: 30 seconds)
- Max iterations (default: 25)

Configuration can be overridden at runtime via `RunnableConfig`.

## Usage

### Interactive Chat

```bash
# Interactive chat (prints the full thread_id)
uv run python -m src.react_agent.chat_cli
```

### Replay a thread after restart

Checkpoints are stored at `.checkpoints/react_agent.sqlite` (gitignored; override with `REACT_AGENT_CHECKPOINT_PATH`). The CLI prints the full `thread_id`. Pass it back on the next process to continue the same conversation:

```bash
# Process 1 β€” start a chat and copy the printed thread_id
uv run python -m src.react_agent.chat_cli
# thread_id: 550e8400-e29b-41d4-a716-446655440000
# You: Remember that my favorite episode is about Two Hour Learning.
# ... type exit ...

# Process 2 β€” same thread_id continues the conversation
uv run python -m src.react_agent.chat_cli --thread-id 550e8400-e29b-41d4-a716-446655440000
# You: What did I say my favorite episode was?
```

You can also set `REACT_AGENT_THREAD_ID` instead of `--thread-id`.

### Programmatic Usage

```python
from src.react_agent.graph import react_agent, get_react_agent
from langchain_core.messages import HumanMessage

# Use default agent with durable SQLite conversation memory
# Pass a thread_id so the same conversation survives process restart
result = await react_agent.ainvoke(
    {"messages": [HumanMessage(content="What is Two Hour Learning?")]},
    config={"configurable": {"thread_id": "my-session-123"}}
)

# Follow-up questions in the same thread remember context
result = await react_agent.ainvoke(
    {"messages": [HumanMessage(content="Tell me more about that")]},
    config={"configurable": {"thread_id": "my-session-123"}}
)

# Create agent with custom config
from langchain_core.runnables import RunnableConfig

config = RunnableConfig(
    configurable={
        "model": "claude-sonnet-4-5",
        "max_tokens": 8000,
        "temperature": 0.1,
        "thread_id": "custom-thread"
    }
)
custom_agent = get_react_agent(config)
result = await custom_agent.ainvoke({
    "messages": [HumanMessage(content="What is Two Hour Learning?")]
}, config=config)
```

## Comparison with deep_research_agent

| Aspect | React Agent | Deep Research Agent |
|--------|-------------|---------------------|
| Framework | LangChain `create_agent` | LangGraph custom graph |
| Graph complexity | Single ReAct loop | Multi-node (clarify β†’ brief β†’ supervisor β†’ researchers β†’ report) |
| Use case | Quick Q&A | In-depth research |
| Tool access | Direct | Delegated via researchers |
| Output | Single response | Structured report |
| State | Minimal (messages + checkpointer memory) | Rich (brief, notes, iterations) |
| Memory | SqliteSaver checkpointer (disk, by thread_id) | Graph state |
| Model creation | Provider-specific classes | `init_chat_model` with configurable fields |

## Tools Available

The agent has access to these tools (re-exported from `deep_research_agent.tools`):

1. **search_knowledge_base(query)** - Semantic search over podcast transcripts and episode summaries in ChromaDB
2. **query_knowledge_graph(query)** - Run Cypher queries against the Neo4j knowledge graph
3. **inspect_graph_schema()** - Get the Neo4j schema (nodes, relationships) to help write Cypher

## Configuration

Default settings in `configuration.py`:
- Model: `gemini-flash-latest`
- Max iterations: 25 (safety limit)
- Max tokens: 4000
- Temperature: 0 (deterministic)
- Timeout: 30 seconds

Override via `RunnableConfig`:
```python
from langchain_core.runnables import RunnableConfig

config = RunnableConfig(
    configurable={
        "model": "gpt-5",
        "max_iterations": 5,
        "max_tokens": 8000,
        "temperature": 0.1
    }
)
result = await react_agent.ainvoke(input, config=config)
```

Or set environment variable:
```bash
export REACT_AGENT_DEFAULT_MODEL="claude-sonnet-4-20250514"
```

## Dependencies

Dependencies are managed in the repo’s `pyproject.toml`. The key runtime pieces are LangChain (+ provider integrations) and `python-dotenv`.