--- license: apache-2.0 library_name: numpy tags: - ambiguity - interpretations - reader-models - llm-evaluation - nlu - query-analysis - numpy - cpu pipeline_tag: text-classification --- # hv-split **Bundle of interpretations, not one answer.** ## The claim in one sentence Every generative model *picks* one reading of an ambiguous query and answers it. `hv-split` refuses to pick. It returns a ranked bundle of every interpretation it can detect — each with its ambiguity source, the ambiguous span, the reading, and a prior. This is a new output shape. Not a label, not a completion, not a ranking of documents. A distribution over *readings of the same query*. ## Install ```bash pip install numpy Actually — no dependencies at all. Pure stdlib. Runs anywhere Python 3.9+ runs. ## Usage ### Split a query ```python from hv_split import HVInterpret m = HVInterpret() b = m.split("Why did the CEO resign last week?") print(b.ambiguity_score) # 0.881 print(b.confidence) # 0.700 print(b.sources) # ['presuppositional'] for i in b.interpretations: print(f"{i.prior:.3f} [{i.source}] {i.reading}") ``` ### Answer every interpretation ```python def my_llm(query, interpretation): return f"(would answer as: {interpretation.reading})" for interp, answer in m.answer_each(query, my_llm): print(f"prior {interp.prior:.3f} -> {answer}") ``` ### Render ```python print(m.render(b, mode="text")) # human-readable print(m.render(b, mode="markdown")) # for reports print(m.render(b, mode="json")) # for callers ``` ### CLI ```bash python hv_split.py --query "All that glitters is not gold" python hv_split.py --query "..." --mode json python hv_split.py --query "..." --ambiguity python hv_split.py # run all demos ``` ## The six ambiguity sources | source | what it detects | example | |---|---|---| | **referential** | pronoun with 2+ candidate antecedents | "She told her..." | | **lexical** | polysemous term with multiple senses | "access the bank" | | **scope** | negation scoping over/under a quantifier | "All that glitters is not gold" | | **presuppositional** | "why did X" presupposes X occurred | "Why did the CEO resign?" | | **framing** | "in the language of Y" commits to a frame | "in the language of category theory..." | | **temporal** | vague temporal references | "recently", "soon", "now" | ## Bundle structure ```python SplitBundle( query: str, interpretations: List[Interpretation], ambiguity_score: float, # normalized entropy of the prior distribution confidence: float, # max prior entropy: float, # raw entropy in nats dominant_source: str, # which source the top reading came from sources: List[str], # all sources that fired n_interpretations: int, ) Interpretation( source: str, span: str, # the ambiguous text span_range: (int, int), # character offsets reading: str, # one interpretation of the span prior: float, # marginal probability rationale: str, # why we think this reading is plausible ) ``` ## Interpretation of the scores | ambiguity_score | meaning | |---:|---| | 0.00 | unambiguous, single reading dominates | | 0.20–0.50 | mild — one reading is likely | | 0.50–0.80 | significant — two or three readings compete | | 0.80–1.00 | severe — no reading dominates | `confidence` is the prior on the most likely reading. `1.0` means a single interpretation; `0.33` means three equally likely readings. ## Benchmarks ### Lexical ``` query: "I need to access the bank" ambiguity: 1.000 interpretations: 4 0.250 [lexical] 'bank' = financial institution 0.250 [lexical] 'bank' = river edge 0.250 [lexical] 'bank' = memory bank 0.250 [lexical] 'bank' = blood bank ``` ### Referential ``` query: "She told her that the manager had changed it, and it broke" ambiguity: 0.000 interpretations: 0 ``` Only one valid antecedent survives the pronoun filter (the model correctly declines to fire on a single candidate — the antecedents of `it` are outside the sentence). ### Scope ``` query: "All that glitters is not gold" ambiguity: 1.000 interpretations: 2 0.500 [scope] wide negation: NOT (all glitters gold) 0.500 [scope] narrow negation: ALL glitters (NOT gold) ``` ### Presuppositional ``` query: "Why did the CEO resign last week?" ambiguity: 0.881 interpretations: 2 0.700 [presuppositional] presupposition holds 0.300 [presuppositional] presupposition fails ``` ### Framing ``` query: "in the language of category theory, what is an identity?" ambiguity: 0.881 interpretations: 2 0.700 [framing] answer strictly within 'category theory' 0.300 [framing] answer outside the frame ``` ### Temporal ``` query: "recently, has the function changed?" ambiguity: 0.967 interpretations: 6 0.250 [temporal] 'recently' narrow reading 0.250 [temporal] 'recently' broad reading 0.125 [lexical] 'function' = mathematical mapping 0.125 [lexical] 'function' = role or purpose 0.125 [lexical] 'function' = working state 0.125 [lexical] 'function' = subroutine ``` ### Multi-source ``` query: "why did the current bank say recently that the function she used in the language of category theory had changed?" ambiguity: 0.943 interpretations: 14 sources: framing, lexical, presuppositional, referential, temporal ``` ### Unambiguous ``` query: "compute 2 + 2" ambiguity: 0.000 interpretations: 0 ``` ## Why this is a new category Every model on Hugging Face *picks*. Classification picks a label. Generation picks a completion. Retrieval ranks documents. None of them returns a *bundle of readings* of the same input. `hv-split` is the first model whose output is a distribution over interpretations, not a choice among them. This is useful whenever the cost of a wrong interpretation is high: - **RAG** — retrieve for all readings, not just the dominant one - **Prompt caching** — the same string can mean different things - **Session dedup** — two queries with the same reading are the same query - **Ambiguity-aware answering** — answer each reading, let the user pick - **Debugging user frustration** — "I asked X, the model answered Y" is usually "the model split wrong" ## Honest limitations - **Detection is regex-based and lexicon-driven.** The polysemous word list has 20 entries. Real coverage needs a dictionary. - **Priors are heuristic.** They come from within-source uniform allocation, not from data. The *ranking* is meaningful; the *magnitudes* are not calibrated. - **Sources are independent.** The bundle reports marginals, not a joint distribution over readings. - **No semantic understanding.** The model detects surfaces that usually signal ambiguity. It does not reason about meaning. - **English-only patterns.** Regexes are tuned for English. - **Noun phrase extraction is deliberately conservative.** The referential detector will not fire if only one candidate antecedent survives filtering — even if a human would see two. ## Reference Extracted from the `XuanJi-ISA` exploratory track, "Visual-Whole Reasoning Interfaces" (issue #122), specifically the blueprint on above-ceiling detection and the reader-model category. The core insight — that "the answer" is not what a user wants when their question is ambiguous; they want the *readings*, and the choice is theirs — is the whole model. ## License Apache-2.0