LIANJie-Jason Claude Opus 4.6 commited on
Commit
fc60784
·
1 Parent(s): ebd834c

fix: web UI history bug + wrong query after clarification + README update

Browse files

Bug 1: app_web.py included the current user message in conversation
history sent to query understanding, causing the question to appear
twice in the prompt. Fixed by excluding the just-appended message.

Bug 2: After clarification, verify_and_respond received only the
original vague query instead of the combined (original + clarification).
Fixed to use the combined query, consistent with CLI behavior.

README: Updated "How It Works" flow to include query understanding step.
Fixed retrieval config (top_k: 5 -> 50, added max_distance: 0.55).
Added query_understanding config section.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

Files changed (2) hide show
  1. README.md +17 -6
  2. app_web.py +10 -5
README.md CHANGED
@@ -276,28 +276,34 @@ Changes you make in the sidebar (switching models, toggling web search) apply on
276
  You ask a question
277
  |
278
  v
279
- 1. RETRIEVE -- Search your local vector database for relevant passages
 
 
 
 
 
280
  -- Optionally search Semantic Scholar for academic papers
281
  |
282
  v
283
- 2. GENERATE -- Send the question + retrieved passages to the AI model
284
  -- System prompt enforces strict citation rules
285
  |
286
  v
287
- 3. VERIFY -- Scan for warning phrases ("based on my knowledge...")
288
  -- Check that cited claims overlap with source text
289
  -- Ask the AI to audit its own answer against the sources
290
  -- If errors found: correct and re-verify (up to 3 times)
291
  -- If still failing: refuse to answer
292
  |
293
  v
294
- 4. DISPLAY -- Show the verified answer with numbered references
295
  ```
296
 
297
  **Key design principles:**
298
 
299
  - **Your documents always come first.** Local sources are the primary authority. Web sources (if enabled) are supplementary and never override your documents.
300
  - **No sources = no answer.** If the chatbot cannot find relevant passages in your documents or the web, it will say so rather than make something up.
 
301
  - **Every claim gets a citation.** The answer format uses numbered endnotes (`[1]`, `[2]`, etc.) with a full reference list at the bottom.
302
  - **Direct quotes are marked.** When the chatbot uses three or more consecutive words from a source, they appear in quotation marks.
303
 
@@ -346,14 +352,19 @@ embeddings:
346
  retrieval:
347
  chunk_size: 1000 # Characters per chunk (when splitting documents)
348
  chunk_overlap: 100 # Overlap between chunks (preserves context)
349
- top_k: 50 # Max chunks to consider per query
350
- max_distance: 0.55 # Only keep chunks with distance below this (0-1)
351
 
352
  web_search:
353
  enabled: true # true | false
354
  backend: "semantic_scholar" # Search engine for academic papers
355
  max_results: 5 # Papers to retrieve per query
356
 
 
 
 
 
 
357
  verification:
358
  enabled: true # Set to false to skip verification (faster but riskier)
359
  max_iterations: 3 # Max correction attempts before refusing
 
276
  You ask a question
277
  |
278
  v
279
+ 1. UNDERSTAND -- The AI reformulates your question for better search
280
+ -- Expands abbreviations, resolves follow-up references
281
+ -- May ask a clarification question if genuinely ambiguous
282
+ |
283
+ v
284
+ 2. RETRIEVE -- Search your local vector database for relevant passages
285
  -- Optionally search Semantic Scholar for academic papers
286
  |
287
  v
288
+ 3. GENERATE -- Send the question + retrieved passages to the AI model
289
  -- System prompt enforces strict citation rules
290
  |
291
  v
292
+ 4. VERIFY -- Scan for warning phrases ("based on my knowledge...")
293
  -- Check that cited claims overlap with source text
294
  -- Ask the AI to audit its own answer against the sources
295
  -- If errors found: correct and re-verify (up to 3 times)
296
  -- If still failing: refuse to answer
297
  |
298
  v
299
+ 5. DISPLAY -- Show the verified answer with numbered references
300
  ```
301
 
302
  **Key design principles:**
303
 
304
  - **Your documents always come first.** Local sources are the primary authority. Web sources (if enabled) are supplementary and never override your documents.
305
  - **No sources = no answer.** If the chatbot cannot find relevant passages in your documents or the web, it will say so rather than make something up.
306
+ - **Smart query understanding.** Before searching, the chatbot reformulates your question to improve retrieval accuracy -- expanding abbreviations, resolving references from prior conversation, and adding relevant keywords. It shows you what it searched for, so you can see how your question was interpreted.
307
  - **Every claim gets a citation.** The answer format uses numbered endnotes (`[1]`, `[2]`, etc.) with a full reference list at the bottom.
308
  - **Direct quotes are marked.** When the chatbot uses three or more consecutive words from a source, they appear in quotation marks.
309
 
 
352
  retrieval:
353
  chunk_size: 1000 # Characters per chunk (when splitting documents)
354
  chunk_overlap: 100 # Overlap between chunks (preserves context)
355
+ top_k: 50 # Candidate pool cap for vector search
356
+ max_distance: 0.55 # Relevance threshold (0=identical, 1=unrelated)
357
 
358
  web_search:
359
  enabled: true # true | false
360
  backend: "semantic_scholar" # Search engine for academic papers
361
  max_results: 5 # Papers to retrieve per query
362
 
363
+ query_understanding:
364
+ enabled: true # Set to false to skip query reformulation
365
+ max_history: 6 # Conversation messages used for context
366
+ max_clarifications: 1 # Max clarification rounds before forcing a search
367
+
368
  verification:
369
  enabled: true # Set to false to skip verification (faster but riskier)
370
  max_iterations: 3 # Max correction attempts before refusing
app_web.py CHANGED
@@ -153,16 +153,20 @@ def render_chat():
153
  max_history = qu_cfg.get("max_history", 6)
154
 
155
  search_query = combined
 
156
 
157
  if qu_enabled:
 
 
 
158
  history = [
159
  {"role": m["role"], "content": m["content"]}
160
- for m in st.session_state.messages[-max_history:]
161
  ]
162
  try:
163
  qu_result = understand_query(combined, cfg, history)
164
  except Exception:
165
- qu_result = {"action": "search", "search_query": combined, "original_query": original_query}
166
 
167
  if qu_result["action"] == "clarify" and pending is None:
168
  # Ask clarification — store original query, show question
@@ -181,7 +185,7 @@ def render_chat():
181
  with st.chat_message("assistant"):
182
  with st.status("Searching knowledge base...", expanded=True) as status:
183
  # Show reformulated query if different
184
- if search_query != original_query:
185
  status.update(label=f'Searching for: "{search_query}"...')
186
 
187
  # Retrieval
@@ -194,8 +198,9 @@ def render_chat():
194
  label=f"Found {n_local} local + {n_web} web sources. Generating response..."
195
  )
196
 
197
- # Verification and response generation (uses original query)
198
- result = verify_and_respond(original_query, retrieval_result, cfg)
 
199
 
200
  # Update status based on verification outcome
201
  if result.get("refused"):
 
153
  max_history = qu_cfg.get("max_history", 6)
154
 
155
  search_query = combined
156
+ response_query = combined # query passed to verify_and_respond
157
 
158
  if qu_enabled:
159
+ # Exclude the just-appended user message to avoid sending
160
+ # the current question twice (once in history, once as query)
161
+ prior_messages = st.session_state.messages[:-1]
162
  history = [
163
  {"role": m["role"], "content": m["content"]}
164
+ for m in prior_messages[-max_history:]
165
  ]
166
  try:
167
  qu_result = understand_query(combined, cfg, history)
168
  except Exception:
169
+ qu_result = {"action": "search", "search_query": combined, "original_query": combined}
170
 
171
  if qu_result["action"] == "clarify" and pending is None:
172
  # Ask clarification — store original query, show question
 
185
  with st.chat_message("assistant"):
186
  with st.status("Searching knowledge base...", expanded=True) as status:
187
  # Show reformulated query if different
188
+ if search_query != response_query:
189
  status.update(label=f'Searching for: "{search_query}"...')
190
 
191
  # Retrieval
 
198
  label=f"Found {n_local} local + {n_web} web sources. Generating response..."
199
  )
200
 
201
+ # Verification and response generation (uses combined query
202
+ # so the LLM answers the clarified question, not just the original)
203
+ result = verify_and_respond(response_query, retrieval_result, cfg)
204
 
205
  # Update status based on verification outcome
206
  if result.get("refused"):