|
Download README.md from zeechimp/hv-split: direct link, hf CLI and curl.
- Browser
- Download file 7.57 kB
-
https://huggingface.co/zeechimp/hv-split/resolve/main/README.md
- Command line
-
hf download hf://zeechimp/hv-split/README.md
-
curl -L -o README.md https://huggingface.co/zeechimp/hv-split/resolve/main/README.md
7.57 kB
| license: apache-2.0 | |
| library_name: numpy | |
| tags: | |
| - ambiguity | |
| - interpretations | |
| - reader-models | |
| - llm-evaluation | |
| - nlu | |
| - query-analysis | |
| - numpy | |
| - cpu | |
| pipeline_tag: text-classification | |
| # hv-split | |
| **Bundle of interpretations, not one answer.** | |
| ## The claim in one sentence | |
| Every generative model *picks* one reading of an ambiguous query and | |
| answers it. `hv-split` refuses to pick. It returns a ranked bundle of | |
| every interpretation it can detect β each with its ambiguity source, | |
| the ambiguous span, the reading, and a prior. | |
| This is a new output shape. Not a label, not a completion, not a | |
| ranking of documents. A distribution over *readings of the same query*. | |
| ## Install | |
| ```bash | |
| pip install numpy | |
| Actually β no dependencies at all. Pure stdlib. Runs anywhere Python 3.9+ | |
| runs. | |
| ## Usage | |
| ### Split a query | |
| ```python | |
| from hv_split import HVInterpret | |
| m = HVInterpret() | |
| b = m.split("Why did the CEO resign last week?") | |
| print(b.ambiguity_score) # 0.881 | |
| print(b.confidence) # 0.700 | |
| print(b.sources) # ['presuppositional'] | |
| for i in b.interpretations: | |
| print(f"{i.prior:.3f} [{i.source}] {i.reading}") | |
| ``` | |
| ### Answer every interpretation | |
| ```python | |
| def my_llm(query, interpretation): | |
| return f"(would answer as: {interpretation.reading})" | |
| for interp, answer in m.answer_each(query, my_llm): | |
| print(f"prior {interp.prior:.3f} -> {answer}") | |
| ``` | |
| ### Render | |
| ```python | |
| print(m.render(b, mode="text")) # human-readable | |
| print(m.render(b, mode="markdown")) # for reports | |
| print(m.render(b, mode="json")) # for callers | |
| ``` | |
| ### CLI | |
| ```bash | |
| python hv_split.py --query "All that glitters is not gold" | |
| python hv_split.py --query "..." --mode json | |
| python hv_split.py --query "..." --ambiguity | |
| python hv_split.py # run all demos | |
| ``` | |
| ## The six ambiguity sources | |
| | source | what it detects | example | | |
| |---|---|---| | |
| | **referential** | pronoun with 2+ candidate antecedents | "She told her..." | | |
| | **lexical** | polysemous term with multiple senses | "access the bank" | | |
| | **scope** | negation scoping over/under a quantifier | "All that glitters is not gold" | | |
| | **presuppositional** | "why did X" presupposes X occurred | "Why did the CEO resign?" | | |
| | **framing** | "in the language of Y" commits to a frame | "in the language of category theory..." | | |
| | **temporal** | vague temporal references | "recently", "soon", "now" | | |
| ## Bundle structure | |
| ```python | |
| SplitBundle( | |
| query: str, | |
| interpretations: List[Interpretation], | |
| ambiguity_score: float, # normalized entropy of the prior distribution | |
| confidence: float, # max prior | |
| entropy: float, # raw entropy in nats | |
| dominant_source: str, # which source the top reading came from | |
| sources: List[str], # all sources that fired | |
| n_interpretations: int, | |
| ) | |
| Interpretation( | |
| source: str, | |
| span: str, # the ambiguous text | |
| span_range: (int, int), # character offsets | |
| reading: str, # one interpretation of the span | |
| prior: float, # marginal probability | |
| rationale: str, # why we think this reading is plausible | |
| ) | |
| ``` | |
| ## Interpretation of the scores | |
| | ambiguity_score | meaning | | |
| |---:|---| | |
| | 0.00 | unambiguous, single reading dominates | | |
| | 0.20β0.50 | mild β one reading is likely | | |
| | 0.50β0.80 | significant β two or three readings compete | | |
| | 0.80β1.00 | severe β no reading dominates | | |
| `confidence` is the prior on the most likely reading. `1.0` means a | |
| single interpretation; `0.33` means three equally likely readings. | |
| ## Benchmarks | |
| ### Lexical | |
| ``` | |
| query: "I need to access the bank" | |
| ambiguity: 1.000 | |
| interpretations: 4 | |
| 0.250 [lexical] 'bank' = financial institution | |
| 0.250 [lexical] 'bank' = river edge | |
| 0.250 [lexical] 'bank' = memory bank | |
| 0.250 [lexical] 'bank' = blood bank | |
| ``` | |
| ### Referential | |
| ``` | |
| query: "She told her that the manager had changed it, and it broke" | |
| ambiguity: 0.000 | |
| interpretations: 0 | |
| ``` | |
| Only one valid antecedent survives the pronoun filter (the model | |
| correctly declines to fire on a single candidate β the antecedents of | |
| `it` are outside the sentence). | |
| ### Scope | |
| ``` | |
| query: "All that glitters is not gold" | |
| ambiguity: 1.000 | |
| interpretations: 2 | |
| 0.500 [scope] wide negation: NOT (all glitters gold) | |
| 0.500 [scope] narrow negation: ALL glitters (NOT gold) | |
| ``` | |
| ### Presuppositional | |
| ``` | |
| query: "Why did the CEO resign last week?" | |
| ambiguity: 0.881 | |
| interpretations: 2 | |
| 0.700 [presuppositional] presupposition holds | |
| 0.300 [presuppositional] presupposition fails | |
| ``` | |
| ### Framing | |
| ``` | |
| query: "in the language of category theory, what is an identity?" | |
| ambiguity: 0.881 | |
| interpretations: 2 | |
| 0.700 [framing] answer strictly within 'category theory' | |
| 0.300 [framing] answer outside the frame | |
| ``` | |
| ### Temporal | |
| ``` | |
| query: "recently, has the function changed?" | |
| ambiguity: 0.967 | |
| interpretations: 6 | |
| 0.250 [temporal] 'recently' narrow reading | |
| 0.250 [temporal] 'recently' broad reading | |
| 0.125 [lexical] 'function' = mathematical mapping | |
| 0.125 [lexical] 'function' = role or purpose | |
| 0.125 [lexical] 'function' = working state | |
| 0.125 [lexical] 'function' = subroutine | |
| ``` | |
| ### Multi-source | |
| ``` | |
| query: "why did the current bank say recently that the function | |
| she used in the language of category theory had changed?" | |
| ambiguity: 0.943 | |
| interpretations: 14 | |
| sources: framing, lexical, presuppositional, referential, temporal | |
| ``` | |
| ### Unambiguous | |
| ``` | |
| query: "compute 2 + 2" | |
| ambiguity: 0.000 | |
| interpretations: 0 | |
| ``` | |
| ## Why this is a new category | |
| Every model on Hugging Face *picks*. Classification picks a label. | |
| Generation picks a completion. Retrieval ranks documents. | |
| None of them returns a *bundle of readings* of the same input. | |
| `hv-split` is the first model whose output is a distribution over | |
| interpretations, not a choice among them. This is useful whenever the | |
| cost of a wrong interpretation is high: | |
| - **RAG** β retrieve for all readings, not just the dominant one | |
| - **Prompt caching** β the same string can mean different things | |
| - **Session dedup** β two queries with the same reading are the same query | |
| - **Ambiguity-aware answering** β answer each reading, let the user pick | |
| - **Debugging user frustration** β "I asked X, the model answered Y" is | |
| usually "the model split wrong" | |
| ## Honest limitations | |
| - **Detection is regex-based and lexicon-driven.** The polysemous word | |
| list has 20 entries. Real coverage needs a dictionary. | |
| - **Priors are heuristic.** They come from within-source uniform | |
| allocation, not from data. The *ranking* is meaningful; the | |
| *magnitudes* are not calibrated. | |
| - **Sources are independent.** The bundle reports marginals, not a | |
| joint distribution over readings. | |
| - **No semantic understanding.** The model detects surfaces that | |
| usually signal ambiguity. It does not reason about meaning. | |
| - **English-only patterns.** Regexes are tuned for English. | |
| - **Noun phrase extraction is deliberately conservative.** The | |
| referential detector will not fire if only one candidate antecedent | |
| survives filtering β even if a human would see two. | |
| ## Reference | |
| Extracted from the `XuanJi-ISA` exploratory track, "Visual-Whole | |
| Reasoning Interfaces" (issue #122), specifically the blueprint on | |
| above-ceiling detection and the reader-model category. | |
| The core insight β that "the answer" is not what a user wants when | |
| their question is ambiguous; they want the *readings*, and the choice | |
| is theirs β is the whole model. | |
| ## License | |
| Apache-2.0 |