Spaces:
Paused
Paused
Commit ·
fc60784
1
Parent(s): ebd834c
fix: web UI history bug + wrong query after clarification + README update
Browse filesBug 1: app_web.py included the current user message in conversation
history sent to query understanding, causing the question to appear
twice in the prompt. Fixed by excluding the just-appended message.
Bug 2: After clarification, verify_and_respond received only the
original vague query instead of the combined (original + clarification).
Fixed to use the combined query, consistent with CLI behavior.
README: Updated "How It Works" flow to include query understanding step.
Fixed retrieval config (top_k: 5 -> 50, added max_distance: 0.55).
Added query_understanding config section.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
- README.md +17 -6
- app_web.py +10 -5
README.md
CHANGED
|
@@ -276,28 +276,34 @@ Changes you make in the sidebar (switching models, toggling web search) apply on
|
|
| 276 |
You ask a question
|
| 277 |
|
|
| 278 |
v
|
| 279 |
-
1.
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 280 |
-- Optionally search Semantic Scholar for academic papers
|
| 281 |
|
|
| 282 |
v
|
| 283 |
-
|
| 284 |
-- System prompt enforces strict citation rules
|
| 285 |
|
|
| 286 |
v
|
| 287 |
-
|
| 288 |
-- Check that cited claims overlap with source text
|
| 289 |
-- Ask the AI to audit its own answer against the sources
|
| 290 |
-- If errors found: correct and re-verify (up to 3 times)
|
| 291 |
-- If still failing: refuse to answer
|
| 292 |
|
|
| 293 |
v
|
| 294 |
-
|
| 295 |
```
|
| 296 |
|
| 297 |
**Key design principles:**
|
| 298 |
|
| 299 |
- **Your documents always come first.** Local sources are the primary authority. Web sources (if enabled) are supplementary and never override your documents.
|
| 300 |
- **No sources = no answer.** If the chatbot cannot find relevant passages in your documents or the web, it will say so rather than make something up.
|
|
|
|
| 301 |
- **Every claim gets a citation.** The answer format uses numbered endnotes (`[1]`, `[2]`, etc.) with a full reference list at the bottom.
|
| 302 |
- **Direct quotes are marked.** When the chatbot uses three or more consecutive words from a source, they appear in quotation marks.
|
| 303 |
|
|
@@ -346,14 +352,19 @@ embeddings:
|
|
| 346 |
retrieval:
|
| 347 |
chunk_size: 1000 # Characters per chunk (when splitting documents)
|
| 348 |
chunk_overlap: 100 # Overlap between chunks (preserves context)
|
| 349 |
-
top_k: 50 #
|
| 350 |
-
max_distance: 0.55 #
|
| 351 |
|
| 352 |
web_search:
|
| 353 |
enabled: true # true | false
|
| 354 |
backend: "semantic_scholar" # Search engine for academic papers
|
| 355 |
max_results: 5 # Papers to retrieve per query
|
| 356 |
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 357 |
verification:
|
| 358 |
enabled: true # Set to false to skip verification (faster but riskier)
|
| 359 |
max_iterations: 3 # Max correction attempts before refusing
|
|
|
|
| 276 |
You ask a question
|
| 277 |
|
|
| 278 |
v
|
| 279 |
+
1. UNDERSTAND -- The AI reformulates your question for better search
|
| 280 |
+
-- Expands abbreviations, resolves follow-up references
|
| 281 |
+
-- May ask a clarification question if genuinely ambiguous
|
| 282 |
+
|
|
| 283 |
+
v
|
| 284 |
+
2. RETRIEVE -- Search your local vector database for relevant passages
|
| 285 |
-- Optionally search Semantic Scholar for academic papers
|
| 286 |
|
|
| 287 |
v
|
| 288 |
+
3. GENERATE -- Send the question + retrieved passages to the AI model
|
| 289 |
-- System prompt enforces strict citation rules
|
| 290 |
|
|
| 291 |
v
|
| 292 |
+
4. VERIFY -- Scan for warning phrases ("based on my knowledge...")
|
| 293 |
-- Check that cited claims overlap with source text
|
| 294 |
-- Ask the AI to audit its own answer against the sources
|
| 295 |
-- If errors found: correct and re-verify (up to 3 times)
|
| 296 |
-- If still failing: refuse to answer
|
| 297 |
|
|
| 298 |
v
|
| 299 |
+
5. DISPLAY -- Show the verified answer with numbered references
|
| 300 |
```
|
| 301 |
|
| 302 |
**Key design principles:**
|
| 303 |
|
| 304 |
- **Your documents always come first.** Local sources are the primary authority. Web sources (if enabled) are supplementary and never override your documents.
|
| 305 |
- **No sources = no answer.** If the chatbot cannot find relevant passages in your documents or the web, it will say so rather than make something up.
|
| 306 |
+
- **Smart query understanding.** Before searching, the chatbot reformulates your question to improve retrieval accuracy -- expanding abbreviations, resolving references from prior conversation, and adding relevant keywords. It shows you what it searched for, so you can see how your question was interpreted.
|
| 307 |
- **Every claim gets a citation.** The answer format uses numbered endnotes (`[1]`, `[2]`, etc.) with a full reference list at the bottom.
|
| 308 |
- **Direct quotes are marked.** When the chatbot uses three or more consecutive words from a source, they appear in quotation marks.
|
| 309 |
|
|
|
|
| 352 |
retrieval:
|
| 353 |
chunk_size: 1000 # Characters per chunk (when splitting documents)
|
| 354 |
chunk_overlap: 100 # Overlap between chunks (preserves context)
|
| 355 |
+
top_k: 50 # Candidate pool cap for vector search
|
| 356 |
+
max_distance: 0.55 # Relevance threshold (0=identical, 1=unrelated)
|
| 357 |
|
| 358 |
web_search:
|
| 359 |
enabled: true # true | false
|
| 360 |
backend: "semantic_scholar" # Search engine for academic papers
|
| 361 |
max_results: 5 # Papers to retrieve per query
|
| 362 |
|
| 363 |
+
query_understanding:
|
| 364 |
+
enabled: true # Set to false to skip query reformulation
|
| 365 |
+
max_history: 6 # Conversation messages used for context
|
| 366 |
+
max_clarifications: 1 # Max clarification rounds before forcing a search
|
| 367 |
+
|
| 368 |
verification:
|
| 369 |
enabled: true # Set to false to skip verification (faster but riskier)
|
| 370 |
max_iterations: 3 # Max correction attempts before refusing
|
app_web.py
CHANGED
|
@@ -153,16 +153,20 @@ def render_chat():
|
|
| 153 |
max_history = qu_cfg.get("max_history", 6)
|
| 154 |
|
| 155 |
search_query = combined
|
|
|
|
| 156 |
|
| 157 |
if qu_enabled:
|
|
|
|
|
|
|
|
|
|
| 158 |
history = [
|
| 159 |
{"role": m["role"], "content": m["content"]}
|
| 160 |
-
for m in
|
| 161 |
]
|
| 162 |
try:
|
| 163 |
qu_result = understand_query(combined, cfg, history)
|
| 164 |
except Exception:
|
| 165 |
-
qu_result = {"action": "search", "search_query": combined, "original_query":
|
| 166 |
|
| 167 |
if qu_result["action"] == "clarify" and pending is None:
|
| 168 |
# Ask clarification — store original query, show question
|
|
@@ -181,7 +185,7 @@ def render_chat():
|
|
| 181 |
with st.chat_message("assistant"):
|
| 182 |
with st.status("Searching knowledge base...", expanded=True) as status:
|
| 183 |
# Show reformulated query if different
|
| 184 |
-
if search_query !=
|
| 185 |
status.update(label=f'Searching for: "{search_query}"...')
|
| 186 |
|
| 187 |
# Retrieval
|
|
@@ -194,8 +198,9 @@ def render_chat():
|
|
| 194 |
label=f"Found {n_local} local + {n_web} web sources. Generating response..."
|
| 195 |
)
|
| 196 |
|
| 197 |
-
# Verification and response generation (uses
|
| 198 |
-
|
|
|
|
| 199 |
|
| 200 |
# Update status based on verification outcome
|
| 201 |
if result.get("refused"):
|
|
|
|
| 153 |
max_history = qu_cfg.get("max_history", 6)
|
| 154 |
|
| 155 |
search_query = combined
|
| 156 |
+
response_query = combined # query passed to verify_and_respond
|
| 157 |
|
| 158 |
if qu_enabled:
|
| 159 |
+
# Exclude the just-appended user message to avoid sending
|
| 160 |
+
# the current question twice (once in history, once as query)
|
| 161 |
+
prior_messages = st.session_state.messages[:-1]
|
| 162 |
history = [
|
| 163 |
{"role": m["role"], "content": m["content"]}
|
| 164 |
+
for m in prior_messages[-max_history:]
|
| 165 |
]
|
| 166 |
try:
|
| 167 |
qu_result = understand_query(combined, cfg, history)
|
| 168 |
except Exception:
|
| 169 |
+
qu_result = {"action": "search", "search_query": combined, "original_query": combined}
|
| 170 |
|
| 171 |
if qu_result["action"] == "clarify" and pending is None:
|
| 172 |
# Ask clarification — store original query, show question
|
|
|
|
| 185 |
with st.chat_message("assistant"):
|
| 186 |
with st.status("Searching knowledge base...", expanded=True) as status:
|
| 187 |
# Show reformulated query if different
|
| 188 |
+
if search_query != response_query:
|
| 189 |
status.update(label=f'Searching for: "{search_query}"...')
|
| 190 |
|
| 191 |
# Retrieval
|
|
|
|
| 198 |
label=f"Found {n_local} local + {n_web} web sources. Generating response..."
|
| 199 |
)
|
| 200 |
|
| 201 |
+
# Verification and response generation (uses combined query
|
| 202 |
+
# so the LLM answers the clarified question, not just the original)
|
| 203 |
+
result = verify_and_respond(response_query, retrieval_result, cfg)
|
| 204 |
|
| 205 |
# Update status based on verification outcome
|
| 206 |
if result.get("refused"):
|