hv-split / README.md
zeechimp's picture
Create README.md
a4b7de2 verified
|
Raw History Blame Contribute Delete
7.57 kB
---
license: apache-2.0
library_name: numpy
tags:
- ambiguity
- interpretations
- reader-models
- llm-evaluation
- nlu
- query-analysis
- numpy
- cpu
pipeline_tag: text-classification
---
# hv-split
**Bundle of interpretations, not one answer.**
## The claim in one sentence
Every generative model *picks* one reading of an ambiguous query and
answers it. `hv-split` refuses to pick. It returns a ranked bundle of
every interpretation it can detect β€” each with its ambiguity source,
the ambiguous span, the reading, and a prior.
This is a new output shape. Not a label, not a completion, not a
ranking of documents. A distribution over *readings of the same query*.
## Install
```bash
pip install numpy
Actually β€” no dependencies at all. Pure stdlib. Runs anywhere Python 3.9+
runs.
## Usage
### Split a query
```python
from hv_split import HVInterpret
m = HVInterpret()
b = m.split("Why did the CEO resign last week?")
print(b.ambiguity_score) # 0.881
print(b.confidence) # 0.700
print(b.sources) # ['presuppositional']
for i in b.interpretations:
print(f"{i.prior:.3f} [{i.source}] {i.reading}")
```
### Answer every interpretation
```python
def my_llm(query, interpretation):
return f"(would answer as: {interpretation.reading})"
for interp, answer in m.answer_each(query, my_llm):
print(f"prior {interp.prior:.3f} -> {answer}")
```
### Render
```python
print(m.render(b, mode="text")) # human-readable
print(m.render(b, mode="markdown")) # for reports
print(m.render(b, mode="json")) # for callers
```
### CLI
```bash
python hv_split.py --query "All that glitters is not gold"
python hv_split.py --query "..." --mode json
python hv_split.py --query "..." --ambiguity
python hv_split.py # run all demos
```
## The six ambiguity sources
| source | what it detects | example |
|---|---|---|
| **referential** | pronoun with 2+ candidate antecedents | "She told her..." |
| **lexical** | polysemous term with multiple senses | "access the bank" |
| **scope** | negation scoping over/under a quantifier | "All that glitters is not gold" |
| **presuppositional** | "why did X" presupposes X occurred | "Why did the CEO resign?" |
| **framing** | "in the language of Y" commits to a frame | "in the language of category theory..." |
| **temporal** | vague temporal references | "recently", "soon", "now" |
## Bundle structure
```python
SplitBundle(
query: str,
interpretations: List[Interpretation],
ambiguity_score: float, # normalized entropy of the prior distribution
confidence: float, # max prior
entropy: float, # raw entropy in nats
dominant_source: str, # which source the top reading came from
sources: List[str], # all sources that fired
n_interpretations: int,
)
Interpretation(
source: str,
span: str, # the ambiguous text
span_range: (int, int), # character offsets
reading: str, # one interpretation of the span
prior: float, # marginal probability
rationale: str, # why we think this reading is plausible
)
```
## Interpretation of the scores
| ambiguity_score | meaning |
|---:|---|
| 0.00 | unambiguous, single reading dominates |
| 0.20–0.50 | mild β€” one reading is likely |
| 0.50–0.80 | significant β€” two or three readings compete |
| 0.80–1.00 | severe β€” no reading dominates |
`confidence` is the prior on the most likely reading. `1.0` means a
single interpretation; `0.33` means three equally likely readings.
## Benchmarks
### Lexical
```
query: "I need to access the bank"
ambiguity: 1.000
interpretations: 4
0.250 [lexical] 'bank' = financial institution
0.250 [lexical] 'bank' = river edge
0.250 [lexical] 'bank' = memory bank
0.250 [lexical] 'bank' = blood bank
```
### Referential
```
query: "She told her that the manager had changed it, and it broke"
ambiguity: 0.000
interpretations: 0
```
Only one valid antecedent survives the pronoun filter (the model
correctly declines to fire on a single candidate β€” the antecedents of
`it` are outside the sentence).
### Scope
```
query: "All that glitters is not gold"
ambiguity: 1.000
interpretations: 2
0.500 [scope] wide negation: NOT (all glitters gold)
0.500 [scope] narrow negation: ALL glitters (NOT gold)
```
### Presuppositional
```
query: "Why did the CEO resign last week?"
ambiguity: 0.881
interpretations: 2
0.700 [presuppositional] presupposition holds
0.300 [presuppositional] presupposition fails
```
### Framing
```
query: "in the language of category theory, what is an identity?"
ambiguity: 0.881
interpretations: 2
0.700 [framing] answer strictly within 'category theory'
0.300 [framing] answer outside the frame
```
### Temporal
```
query: "recently, has the function changed?"
ambiguity: 0.967
interpretations: 6
0.250 [temporal] 'recently' narrow reading
0.250 [temporal] 'recently' broad reading
0.125 [lexical] 'function' = mathematical mapping
0.125 [lexical] 'function' = role or purpose
0.125 [lexical] 'function' = working state
0.125 [lexical] 'function' = subroutine
```
### Multi-source
```
query: "why did the current bank say recently that the function
she used in the language of category theory had changed?"
ambiguity: 0.943
interpretations: 14
sources: framing, lexical, presuppositional, referential, temporal
```
### Unambiguous
```
query: "compute 2 + 2"
ambiguity: 0.000
interpretations: 0
```
## Why this is a new category
Every model on Hugging Face *picks*. Classification picks a label.
Generation picks a completion. Retrieval ranks documents.
None of them returns a *bundle of readings* of the same input.
`hv-split` is the first model whose output is a distribution over
interpretations, not a choice among them. This is useful whenever the
cost of a wrong interpretation is high:
- **RAG** β€” retrieve for all readings, not just the dominant one
- **Prompt caching** β€” the same string can mean different things
- **Session dedup** β€” two queries with the same reading are the same query
- **Ambiguity-aware answering** β€” answer each reading, let the user pick
- **Debugging user frustration** β€” "I asked X, the model answered Y" is
usually "the model split wrong"
## Honest limitations
- **Detection is regex-based and lexicon-driven.** The polysemous word
list has 20 entries. Real coverage needs a dictionary.
- **Priors are heuristic.** They come from within-source uniform
allocation, not from data. The *ranking* is meaningful; the
*magnitudes* are not calibrated.
- **Sources are independent.** The bundle reports marginals, not a
joint distribution over readings.
- **No semantic understanding.** The model detects surfaces that
usually signal ambiguity. It does not reason about meaning.
- **English-only patterns.** Regexes are tuned for English.
- **Noun phrase extraction is deliberately conservative.** The
referential detector will not fire if only one candidate antecedent
survives filtering β€” even if a human would see two.
## Reference
Extracted from the `XuanJi-ISA` exploratory track, "Visual-Whole
Reasoning Interfaces" (issue #122), specifically the blueprint on
above-ceiling detection and the reader-model category.
The core insight β€” that "the answer" is not what a user wants when
their question is ambiguous; they want the *readings*, and the choice
is theirs β€” is the whole model.
## License
Apache-2.0