Download README.md from zeechimp/hv-split: direct link, hf CLI and curl.
- Browser
- Download file 7.57 kB
-
https://huggingface.co/zeechimp/hv-split/resolve/main/README.md
- Command line
-
hf download hf://zeechimp/hv-split/README.md
-
curl -L -o README.md https://huggingface.co/zeechimp/hv-split/resolve/main/README.md
license: apache-2.0
library_name: numpy
tags:
- ambiguity
- interpretations
- reader-models
- llm-evaluation
- nlu
- query-analysis
- numpy
- cpu
pipeline_tag: text-classification
hv-split
Bundle of interpretations, not one answer.
The claim in one sentence
Every generative model picks one reading of an ambiguous query and
answers it. hv-split refuses to pick. It returns a ranked bundle of
every interpretation it can detect β each with its ambiguity source,
the ambiguous span, the reading, and a prior.
This is a new output shape. Not a label, not a completion, not a ranking of documents. A distribution over readings of the same query.
Install
pip install numpy
Actually β no dependencies at all. Pure stdlib. Runs anywhere Python 3.9+
runs.
## Usage
### Split a query
```python
from hv_split import HVInterpret
m = HVInterpret()
b = m.split("Why did the CEO resign last week?")
print(b.ambiguity_score) # 0.881
print(b.confidence) # 0.700
print(b.sources) # ['presuppositional']
for i in b.interpretations:
print(f"{i.prior:.3f} [{i.source}] {i.reading}")
Answer every interpretation
def my_llm(query, interpretation):
return f"(would answer as: {interpretation.reading})"
for interp, answer in m.answer_each(query, my_llm):
print(f"prior {interp.prior:.3f} -> {answer}")
Render
print(m.render(b, mode="text")) # human-readable
print(m.render(b, mode="markdown")) # for reports
print(m.render(b, mode="json")) # for callers
CLI
python hv_split.py --query "All that glitters is not gold"
python hv_split.py --query "..." --mode json
python hv_split.py --query "..." --ambiguity
python hv_split.py # run all demos
The six ambiguity sources
| source | what it detects | example |
|---|---|---|
| referential | pronoun with 2+ candidate antecedents | "She told her..." |
| lexical | polysemous term with multiple senses | "access the bank" |
| scope | negation scoping over/under a quantifier | "All that glitters is not gold" |
| presuppositional | "why did X" presupposes X occurred | "Why did the CEO resign?" |
| framing | "in the language of Y" commits to a frame | "in the language of category theory..." |
| temporal | vague temporal references | "recently", "soon", "now" |
Bundle structure
SplitBundle(
query: str,
interpretations: List[Interpretation],
ambiguity_score: float, # normalized entropy of the prior distribution
confidence: float, # max prior
entropy: float, # raw entropy in nats
dominant_source: str, # which source the top reading came from
sources: List[str], # all sources that fired
n_interpretations: int,
)
Interpretation(
source: str,
span: str, # the ambiguous text
span_range: (int, int), # character offsets
reading: str, # one interpretation of the span
prior: float, # marginal probability
rationale: str, # why we think this reading is plausible
)
Interpretation of the scores
| ambiguity_score | meaning |
|---|---|
| 0.00 | unambiguous, single reading dominates |
| 0.20β0.50 | mild β one reading is likely |
| 0.50β0.80 | significant β two or three readings compete |
| 0.80β1.00 | severe β no reading dominates |
confidence is the prior on the most likely reading. 1.0 means a
single interpretation; 0.33 means three equally likely readings.
Benchmarks
Lexical
query: "I need to access the bank"
ambiguity: 1.000
interpretations: 4
0.250 [lexical] 'bank' = financial institution
0.250 [lexical] 'bank' = river edge
0.250 [lexical] 'bank' = memory bank
0.250 [lexical] 'bank' = blood bank
Referential
query: "She told her that the manager had changed it, and it broke"
ambiguity: 0.000
interpretations: 0
Only one valid antecedent survives the pronoun filter (the model
correctly declines to fire on a single candidate β the antecedents of
it are outside the sentence).
Scope
query: "All that glitters is not gold"
ambiguity: 1.000
interpretations: 2
0.500 [scope] wide negation: NOT (all glitters gold)
0.500 [scope] narrow negation: ALL glitters (NOT gold)
Presuppositional
query: "Why did the CEO resign last week?"
ambiguity: 0.881
interpretations: 2
0.700 [presuppositional] presupposition holds
0.300 [presuppositional] presupposition fails
Framing
query: "in the language of category theory, what is an identity?"
ambiguity: 0.881
interpretations: 2
0.700 [framing] answer strictly within 'category theory'
0.300 [framing] answer outside the frame
Temporal
query: "recently, has the function changed?"
ambiguity: 0.967
interpretations: 6
0.250 [temporal] 'recently' narrow reading
0.250 [temporal] 'recently' broad reading
0.125 [lexical] 'function' = mathematical mapping
0.125 [lexical] 'function' = role or purpose
0.125 [lexical] 'function' = working state
0.125 [lexical] 'function' = subroutine
Multi-source
query: "why did the current bank say recently that the function
she used in the language of category theory had changed?"
ambiguity: 0.943
interpretations: 14
sources: framing, lexical, presuppositional, referential, temporal
Unambiguous
query: "compute 2 + 2"
ambiguity: 0.000
interpretations: 0
Why this is a new category
Every model on Hugging Face picks. Classification picks a label. Generation picks a completion. Retrieval ranks documents.
None of them returns a bundle of readings of the same input.
hv-split is the first model whose output is a distribution over
interpretations, not a choice among them. This is useful whenever the
cost of a wrong interpretation is high:
- RAG β retrieve for all readings, not just the dominant one
- Prompt caching β the same string can mean different things
- Session dedup β two queries with the same reading are the same query
- Ambiguity-aware answering β answer each reading, let the user pick
- Debugging user frustration β "I asked X, the model answered Y" is usually "the model split wrong"
Honest limitations
- Detection is regex-based and lexicon-driven. The polysemous word list has 20 entries. Real coverage needs a dictionary.
- Priors are heuristic. They come from within-source uniform allocation, not from data. The ranking is meaningful; the magnitudes are not calibrated.
- Sources are independent. The bundle reports marginals, not a joint distribution over readings.
- No semantic understanding. The model detects surfaces that usually signal ambiguity. It does not reason about meaning.
- English-only patterns. Regexes are tuned for English.
- Noun phrase extraction is deliberately conservative. The referential detector will not fire if only one candidate antecedent survives filtering β even if a human would see two.
Reference
Extracted from the XuanJi-ISA exploratory track, "Visual-Whole
Reasoning Interfaces" (issue #122), specifically the blueprint on
above-ceiling detection and the reader-model category.
The core insight β that "the answer" is not what a user wants when their question is ambiguous; they want the readings, and the choice is theirs β is the whole model.
License
Apache-2.0