hv-split / README.md
zeechimp's picture
Create README.md
a4b7de2 verified
|
Raw History Blame Contribute Delete
7.57 kB
metadata
license: apache-2.0
library_name: numpy
tags:
  - ambiguity
  - interpretations
  - reader-models
  - llm-evaluation
  - nlu
  - query-analysis
  - numpy
  - cpu
pipeline_tag: text-classification

hv-split

Bundle of interpretations, not one answer.

The claim in one sentence

Every generative model picks one reading of an ambiguous query and answers it. hv-split refuses to pick. It returns a ranked bundle of every interpretation it can detect β€” each with its ambiguity source, the ambiguous span, the reading, and a prior.

This is a new output shape. Not a label, not a completion, not a ranking of documents. A distribution over readings of the same query.

Install

pip install numpy

Actually β€” no dependencies at all. Pure stdlib. Runs anywhere Python 3.9+
runs.

## Usage

### Split a query

```python
from hv_split import HVInterpret

m = HVInterpret()
b = m.split("Why did the CEO resign last week?")

print(b.ambiguity_score)   # 0.881
print(b.confidence)        # 0.700
print(b.sources)           # ['presuppositional']
for i in b.interpretations:
    print(f"{i.prior:.3f}  [{i.source}]  {i.reading}")

Answer every interpretation

def my_llm(query, interpretation):
    return f"(would answer as: {interpretation.reading})"

for interp, answer in m.answer_each(query, my_llm):
    print(f"prior {interp.prior:.3f}  ->  {answer}")

Render

print(m.render(b, mode="text"))       # human-readable
print(m.render(b, mode="markdown"))   # for reports
print(m.render(b, mode="json"))       # for callers

CLI

python hv_split.py --query "All that glitters is not gold"
python hv_split.py --query "..." --mode json
python hv_split.py --query "..." --ambiguity
python hv_split.py                    # run all demos

The six ambiguity sources

source what it detects example
referential pronoun with 2+ candidate antecedents "She told her..."
lexical polysemous term with multiple senses "access the bank"
scope negation scoping over/under a quantifier "All that glitters is not gold"
presuppositional "why did X" presupposes X occurred "Why did the CEO resign?"
framing "in the language of Y" commits to a frame "in the language of category theory..."
temporal vague temporal references "recently", "soon", "now"

Bundle structure

SplitBundle(
    query: str,
    interpretations: List[Interpretation],
    ambiguity_score: float,    # normalized entropy of the prior distribution
    confidence: float,          # max prior
    entropy: float,             # raw entropy in nats
    dominant_source: str,       # which source the top reading came from
    sources: List[str],         # all sources that fired
    n_interpretations: int,
)

Interpretation(
    source: str,
    span: str,                  # the ambiguous text
    span_range: (int, int),     # character offsets
    reading: str,               # one interpretation of the span
    prior: float,               # marginal probability
    rationale: str,             # why we think this reading is plausible
)

Interpretation of the scores

ambiguity_score meaning
0.00 unambiguous, single reading dominates
0.20–0.50 mild β€” one reading is likely
0.50–0.80 significant β€” two or three readings compete
0.80–1.00 severe β€” no reading dominates

confidence is the prior on the most likely reading. 1.0 means a single interpretation; 0.33 means three equally likely readings.

Benchmarks

Lexical

query: "I need to access the bank"
ambiguity: 1.000
interpretations: 4
  0.250  [lexical]  'bank' = financial institution
  0.250  [lexical]  'bank' = river edge
  0.250  [lexical]  'bank' = memory bank
  0.250  [lexical]  'bank' = blood bank

Referential

query: "She told her that the manager had changed it, and it broke"
ambiguity: 0.000
interpretations: 0

Only one valid antecedent survives the pronoun filter (the model correctly declines to fire on a single candidate β€” the antecedents of it are outside the sentence).

Scope

query: "All that glitters is not gold"
ambiguity: 1.000
interpretations: 2
  0.500  [scope]  wide negation: NOT (all glitters gold)
  0.500  [scope]  narrow negation: ALL glitters (NOT gold)

Presuppositional

query: "Why did the CEO resign last week?"
ambiguity: 0.881
interpretations: 2
  0.700  [presuppositional]  presupposition holds
  0.300  [presuppositional]  presupposition fails

Framing

query: "in the language of category theory, what is an identity?"
ambiguity: 0.881
interpretations: 2
  0.700  [framing]  answer strictly within 'category theory'
  0.300  [framing]  answer outside the frame

Temporal

query: "recently, has the function changed?"
ambiguity: 0.967
interpretations: 6
  0.250  [temporal]  'recently' narrow reading
  0.250  [temporal]  'recently' broad reading
  0.125  [lexical]   'function' = mathematical mapping
  0.125  [lexical]   'function' = role or purpose
  0.125  [lexical]   'function' = working state
  0.125  [lexical]   'function' = subroutine

Multi-source

query: "why did the current bank say recently that the function
        she used in the language of category theory had changed?"
ambiguity: 0.943
interpretations: 14
sources: framing, lexical, presuppositional, referential, temporal

Unambiguous

query: "compute 2 + 2"
ambiguity: 0.000
interpretations: 0

Why this is a new category

Every model on Hugging Face picks. Classification picks a label. Generation picks a completion. Retrieval ranks documents.

None of them returns a bundle of readings of the same input.

hv-split is the first model whose output is a distribution over interpretations, not a choice among them. This is useful whenever the cost of a wrong interpretation is high:

  • RAG β€” retrieve for all readings, not just the dominant one
  • Prompt caching β€” the same string can mean different things
  • Session dedup β€” two queries with the same reading are the same query
  • Ambiguity-aware answering β€” answer each reading, let the user pick
  • Debugging user frustration β€” "I asked X, the model answered Y" is usually "the model split wrong"

Honest limitations

  • Detection is regex-based and lexicon-driven. The polysemous word list has 20 entries. Real coverage needs a dictionary.
  • Priors are heuristic. They come from within-source uniform allocation, not from data. The ranking is meaningful; the magnitudes are not calibrated.
  • Sources are independent. The bundle reports marginals, not a joint distribution over readings.
  • No semantic understanding. The model detects surfaces that usually signal ambiguity. It does not reason about meaning.
  • English-only patterns. Regexes are tuned for English.
  • Noun phrase extraction is deliberately conservative. The referential detector will not fire if only one candidate antecedent survives filtering β€” even if a human would see two.

Reference

Extracted from the XuanJi-ISA exploratory track, "Visual-Whole Reasoning Interfaces" (issue #122), specifically the blueprint on above-ceiling detection and the reader-model category.

The core insight β€” that "the answer" is not what a user wants when their question is ambiguous; they want the readings, and the choice is theirs β€” is the whole model.

License

Apache-2.0