zeechimp commited on
Commit
a4b7de2
Β·
verified Β·
1 Parent(s): 5822f41

Create README.md

Browse files
Files changed (1) hide show
  1. README.md +263 -0
README.md ADDED
@@ -0,0 +1,263 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: apache-2.0
3
+ library_name: numpy
4
+ tags:
5
+ - ambiguity
6
+ - interpretations
7
+ - reader-models
8
+ - llm-evaluation
9
+ - nlu
10
+ - query-analysis
11
+ - numpy
12
+ - cpu
13
+ pipeline_tag: text-classification
14
+ ---
15
+
16
+ # hv-split
17
+
18
+ **Bundle of interpretations, not one answer.**
19
+
20
+ ## The claim in one sentence
21
+
22
+ Every generative model *picks* one reading of an ambiguous query and
23
+ answers it. `hv-split` refuses to pick. It returns a ranked bundle of
24
+ every interpretation it can detect β€” each with its ambiguity source,
25
+ the ambiguous span, the reading, and a prior.
26
+
27
+ This is a new output shape. Not a label, not a completion, not a
28
+ ranking of documents. A distribution over *readings of the same query*.
29
+
30
+ ## Install
31
+
32
+ ```bash
33
+ pip install numpy
34
+
35
+ Actually β€” no dependencies at all. Pure stdlib. Runs anywhere Python 3.9+
36
+ runs.
37
+
38
+ ## Usage
39
+
40
+ ### Split a query
41
+
42
+ ```python
43
+ from hv_split import HVInterpret
44
+
45
+ m = HVInterpret()
46
+ b = m.split("Why did the CEO resign last week?")
47
+
48
+ print(b.ambiguity_score) # 0.881
49
+ print(b.confidence) # 0.700
50
+ print(b.sources) # ['presuppositional']
51
+ for i in b.interpretations:
52
+ print(f"{i.prior:.3f} [{i.source}] {i.reading}")
53
+ ```
54
+
55
+ ### Answer every interpretation
56
+
57
+ ```python
58
+ def my_llm(query, interpretation):
59
+ return f"(would answer as: {interpretation.reading})"
60
+
61
+ for interp, answer in m.answer_each(query, my_llm):
62
+ print(f"prior {interp.prior:.3f} -> {answer}")
63
+ ```
64
+
65
+ ### Render
66
+
67
+ ```python
68
+ print(m.render(b, mode="text")) # human-readable
69
+ print(m.render(b, mode="markdown")) # for reports
70
+ print(m.render(b, mode="json")) # for callers
71
+ ```
72
+
73
+ ### CLI
74
+
75
+ ```bash
76
+ python hv_split.py --query "All that glitters is not gold"
77
+ python hv_split.py --query "..." --mode json
78
+ python hv_split.py --query "..." --ambiguity
79
+ python hv_split.py # run all demos
80
+ ```
81
+
82
+ ## The six ambiguity sources
83
+
84
+ | source | what it detects | example |
85
+ |---|---|---|
86
+ | **referential** | pronoun with 2+ candidate antecedents | "She told her..." |
87
+ | **lexical** | polysemous term with multiple senses | "access the bank" |
88
+ | **scope** | negation scoping over/under a quantifier | "All that glitters is not gold" |
89
+ | **presuppositional** | "why did X" presupposes X occurred | "Why did the CEO resign?" |
90
+ | **framing** | "in the language of Y" commits to a frame | "in the language of category theory..." |
91
+ | **temporal** | vague temporal references | "recently", "soon", "now" |
92
+
93
+ ## Bundle structure
94
+
95
+ ```python
96
+ SplitBundle(
97
+ query: str,
98
+ interpretations: List[Interpretation],
99
+ ambiguity_score: float, # normalized entropy of the prior distribution
100
+ confidence: float, # max prior
101
+ entropy: float, # raw entropy in nats
102
+ dominant_source: str, # which source the top reading came from
103
+ sources: List[str], # all sources that fired
104
+ n_interpretations: int,
105
+ )
106
+
107
+ Interpretation(
108
+ source: str,
109
+ span: str, # the ambiguous text
110
+ span_range: (int, int), # character offsets
111
+ reading: str, # one interpretation of the span
112
+ prior: float, # marginal probability
113
+ rationale: str, # why we think this reading is plausible
114
+ )
115
+ ```
116
+
117
+ ## Interpretation of the scores
118
+
119
+ | ambiguity_score | meaning |
120
+ |---:|---|
121
+ | 0.00 | unambiguous, single reading dominates |
122
+ | 0.20–0.50 | mild β€” one reading is likely |
123
+ | 0.50–0.80 | significant β€” two or three readings compete |
124
+ | 0.80–1.00 | severe β€” no reading dominates |
125
+
126
+ `confidence` is the prior on the most likely reading. `1.0` means a
127
+ single interpretation; `0.33` means three equally likely readings.
128
+
129
+ ## Benchmarks
130
+
131
+ ### Lexical
132
+
133
+ ```
134
+ query: "I need to access the bank"
135
+ ambiguity: 1.000
136
+ interpretations: 4
137
+ 0.250 [lexical] 'bank' = financial institution
138
+ 0.250 [lexical] 'bank' = river edge
139
+ 0.250 [lexical] 'bank' = memory bank
140
+ 0.250 [lexical] 'bank' = blood bank
141
+ ```
142
+
143
+ ### Referential
144
+
145
+ ```
146
+ query: "She told her that the manager had changed it, and it broke"
147
+ ambiguity: 0.000
148
+ interpretations: 0
149
+ ```
150
+
151
+ Only one valid antecedent survives the pronoun filter (the model
152
+ correctly declines to fire on a single candidate β€” the antecedents of
153
+ `it` are outside the sentence).
154
+
155
+ ### Scope
156
+
157
+ ```
158
+ query: "All that glitters is not gold"
159
+ ambiguity: 1.000
160
+ interpretations: 2
161
+ 0.500 [scope] wide negation: NOT (all glitters gold)
162
+ 0.500 [scope] narrow negation: ALL glitters (NOT gold)
163
+ ```
164
+
165
+ ### Presuppositional
166
+
167
+ ```
168
+ query: "Why did the CEO resign last week?"
169
+ ambiguity: 0.881
170
+ interpretations: 2
171
+ 0.700 [presuppositional] presupposition holds
172
+ 0.300 [presuppositional] presupposition fails
173
+ ```
174
+
175
+ ### Framing
176
+
177
+ ```
178
+ query: "in the language of category theory, what is an identity?"
179
+ ambiguity: 0.881
180
+ interpretations: 2
181
+ 0.700 [framing] answer strictly within 'category theory'
182
+ 0.300 [framing] answer outside the frame
183
+ ```
184
+
185
+ ### Temporal
186
+
187
+ ```
188
+ query: "recently, has the function changed?"
189
+ ambiguity: 0.967
190
+ interpretations: 6
191
+ 0.250 [temporal] 'recently' narrow reading
192
+ 0.250 [temporal] 'recently' broad reading
193
+ 0.125 [lexical] 'function' = mathematical mapping
194
+ 0.125 [lexical] 'function' = role or purpose
195
+ 0.125 [lexical] 'function' = working state
196
+ 0.125 [lexical] 'function' = subroutine
197
+ ```
198
+
199
+ ### Multi-source
200
+
201
+ ```
202
+ query: "why did the current bank say recently that the function
203
+ she used in the language of category theory had changed?"
204
+ ambiguity: 0.943
205
+ interpretations: 14
206
+ sources: framing, lexical, presuppositional, referential, temporal
207
+ ```
208
+
209
+ ### Unambiguous
210
+
211
+ ```
212
+ query: "compute 2 + 2"
213
+ ambiguity: 0.000
214
+ interpretations: 0
215
+ ```
216
+
217
+ ## Why this is a new category
218
+
219
+ Every model on Hugging Face *picks*. Classification picks a label.
220
+ Generation picks a completion. Retrieval ranks documents.
221
+
222
+ None of them returns a *bundle of readings* of the same input.
223
+
224
+ `hv-split` is the first model whose output is a distribution over
225
+ interpretations, not a choice among them. This is useful whenever the
226
+ cost of a wrong interpretation is high:
227
+
228
+ - **RAG** β€” retrieve for all readings, not just the dominant one
229
+ - **Prompt caching** β€” the same string can mean different things
230
+ - **Session dedup** β€” two queries with the same reading are the same query
231
+ - **Ambiguity-aware answering** β€” answer each reading, let the user pick
232
+ - **Debugging user frustration** β€” "I asked X, the model answered Y" is
233
+ usually "the model split wrong"
234
+
235
+ ## Honest limitations
236
+
237
+ - **Detection is regex-based and lexicon-driven.** The polysemous word
238
+ list has 20 entries. Real coverage needs a dictionary.
239
+ - **Priors are heuristic.** They come from within-source uniform
240
+ allocation, not from data. The *ranking* is meaningful; the
241
+ *magnitudes* are not calibrated.
242
+ - **Sources are independent.** The bundle reports marginals, not a
243
+ joint distribution over readings.
244
+ - **No semantic understanding.** The model detects surfaces that
245
+ usually signal ambiguity. It does not reason about meaning.
246
+ - **English-only patterns.** Regexes are tuned for English.
247
+ - **Noun phrase extraction is deliberately conservative.** The
248
+ referential detector will not fire if only one candidate antecedent
249
+ survives filtering β€” even if a human would see two.
250
+
251
+ ## Reference
252
+
253
+ Extracted from the `XuanJi-ISA` exploratory track, "Visual-Whole
254
+ Reasoning Interfaces" (issue #122), specifically the blueprint on
255
+ above-ceiling detection and the reader-model category.
256
+
257
+ The core insight β€” that "the answer" is not what a user wants when
258
+ their question is ambiguous; they want the *readings*, and the choice
259
+ is theirs β€” is the whole model.
260
+
261
+ ## License
262
+
263
+ Apache-2.0