hv-tail
Above-ceiling detector. Estimate whether the user is operating above the model's calibrated ceiling for this query type.
The claim in one sentence
A user who sees the answer, points at it, and is correct without
reading is operating above the mean. Modern LLMs are calibrated on the
mean. hv-tail is the first model on the Hub that detects the tail.
What it detects
| signal | what it captures |
|---|---|
| framing | "in the language of X", "from the perspective of Y" |
| precision | formal notation, technical vocabulary, backticks, math |
| undecidability | "is there", "can one prove", "well-defined", "GΓΆdel" |
| meta | "your reasoning", "are you sure", "what are your limits" |
| negative | "not the standard", "don't just", "beyond the obvious" |
| constraint | multiple simultaneous requirements |
| dialect | non-standard register signals |
| requery | history shows repeated refinements on one topic |
Each signal returns a value in [0, 1]. The raw score is
sum(signal Γ weight) / 3.0, capped at 1.0. Three strong weighted
signals saturate.
Install
pip install numpy
Actually β no dependencies at all. Pure stdlib. Runs anywhere Python 3.9+
runs.
## Usage
### Score a query
```python
from hv_tail import HVTail
m = HVTail()
r = m.score("what is a black hole")
# {'above_ceiling': False, 'below_ceiling': True,
# 'confidence': 1.000, 'raw_score': 0.000,
# 'recommendation': 'proceed', ...}
Score with history
r = m.score(
"then how does Hawking radiation preserve unitarity",
history=[
"what is a black hole",
"but why does the horizon have no local features",
"and does that mean information is destroyed",
],
)
# {'above_ceiling': False, 'confidence': 0.345,
# 'raw_score': 0.490, 'recommendation': 'ask', ...}
Verbose explanation
print(m.explain("your reasoning on that seems inconsistent..."))
Calibrate to your population
m.calibrate(
above_examples=[
"in the language of category theory, is there a functor...",
"your reasoning on that last answer seems inconsistent...",
],
below_examples=[
"what is a black hole",
"how do I make coffee",
],
)
The thresholds are fit from the labeled examples:
- Clean separation β
aboveandbelowhug the gap between the two clusters, with a proportional margin. - Overlap β thresholds maximize
TPR β FPR(above) andTNR β FNR(below).
CLI
# score a single query
python hv_tail.py --query "what is a black hole"
# verbose
python hv_tail.py --query "..." --explain
# score with history from a file
python hv_tail.py --query "..." --history history.json --explain
# JSON output
python hv_tail.py --query "..." --json
# no args = run all demos
python hv_tail.py
The formula
raw_score = min(1.0, sum(signal_i Γ weight_i) / 3.0)
where the eight signals and default weights are:
| signal | weight | what fires it |
|---|---|---|
| framing | 1.0 | explicit requests to reframe |
| precision | 1.2 | formal notation, technical terms |
| undecidability | 1.8 | questions that require "I don't know" |
| meta | 1.8 | questions about the model's own reasoning |
| negative | 1.1 | explicit rejection of the standard answer |
| constraint | 0.7 | multiple simultaneous requirements |
| dialect | 0.5 | non-standard register |
| requery | 1.8 | repeated refinements on one topic |
NORMALIZER = 3.0 means three strong weighted signals saturate the
score. This is a "noisy-OR-like" interpretation: independent weak
signals multiply their evidence.
The two weights that matter most:
- undecidability (1.8): a query that requires "I don't know" is the strongest signal that the mean-trained model is about to answer the wrong question.
- meta (1.8): a query about the model's own reasoning is a direct signal that the user is treating the model as an object, not an oracle.
Either one alone crosses the above-ceiling threshold. Both together saturate the score.
Recommended actions
| verdict | condition | action |
|---|---|---|
| proceed | raw_score β€ 0.20 |
the model can answer normally |
| ask | 0.20 < raw_score < 0.55 |
a clarifying question is needed |
| reframe | 0.55 β€ raw_score < 0.85 |
present multiple frames |
| refuse | raw_score β₯ 0.85 |
say "I don't know" |
The recommendation layer is what makes this model useful: not just a score, but an action. Every previous detector on the Hub returns a number. This one returns a decision.
Benchmarks
Mean-user query
query: "what is a black hole"
raw_score : 0.000
above_ceiling : False
below_ceiling : True
recommendation: proceed
All eight signals fire at zero. This is the reference point.
Tail-user query
query: "in the language of category theory, is there a functor
from the black hole information paradox to a broader
framework of unitarity that doesn't assume the standard
measurement postulate?"
raw_score : 1.000
above_ceiling : True
recommendation: refuse
Three signals fire (framing 0.33, precision 0.50, undecidability 1.00, negative 0.50). The score saturates. The model correctly identifies this as above the ceiling and recommends refusing rather than answering.
Borderline query
query: "explain this in the style of a Feynman diagram but not
the standard textbook version"
raw_score : 0.303
above_ceiling : False
below_ceiling : False
recommendation: ask
Framing 0.33 and negative 0.50, but nothing else. This is a user reaching toward the tail but still framing the question at the mean. The model says "ask" β a clarifying question would help.
Requery history
history:
[0] "what is a black hole"
[1] "but why does the horizon have no local features"
[2] "and does that mean information is destroyed"
query: "then how does Hawking radiation preserve unitarity"
raw_score : 0.490
above_ceiling : False
recommendation: ask
Requery fires at 0.67. Four consecutive refinements is a strong pattern; the model is on the edge of the ceiling.
Meta query
query: "your reasoning on that last answer seems inconsistent
with what you said earlier. can you explain your
calibration? are you sure the assumption you made is
defensible?"
raw_score : 0.631
above_ceiling : True
recommendation: reframe
Meta fires at 1.00. A user asking about the model's own reasoning is directly treating it as an object to inspect. This alone crosses the ceiling.
Batch comparison
| name | score | above | below | conf | recommendation |
|---|---|---|---|---|---|
| mean | 0.000 | False | True | 1.000 | proceed |
| borderline | 0.303 | False | False | 0.587 | ask |
| tail | 1.000 | True | False | 1.000 | refuse |
| requery | 0.490 | False | False | 0.345 | ask |
| meta | 0.631 | True | False | 0.179 | reframe |
Calibration behavior
The calibrate() method fits the two decision thresholds. It handles
both cases:
Narrow gap (above and below scores nearly touch):
above-threshold after : 0.0670
below-threshold after : 0.0223
gap : 0.0446
Wide gap (clusters separate cleanly):
above-threshold after : 0.5806
below-threshold after : 0.2500
gap : 0.3306
The margins are proportional to the gap. Narrow gaps produce tight thresholds; wide gaps produce comfortable ones.
Why this is a genuine new category
Every model on Hugging Face assumes the user is at the mean:
- text generation β assumes the reader wants text
- image captioning β assumes the image is the topic
- classification β assumes the label is the question
None of them model the user. hv-tail is the first model on the Hub
that classifies the person, not the content. Specifically: is this
person operating above the ceiling their tool was calibrated for?
This is the "genius fails the IQ test" problem, packaged as a
detector. It's the missing diagnostic in every LLM product's
infrastructure. Combined with hv-mode and hv-ttu, it forms the
first complete reader-model stack:
hv-tailβ who is asking?hv-modeβ what shape should the answer take?hv-ttuβ how long will it take to understand?
Honest limitations
- Signals are regex-based. They cover the common tail-user markers. A user operating above the ceiling without those markers will not be detected.
- Coefficients are heuristic. They are plausible and internally consistent. They are not fit to a labeled corpus. Calibrate on your own data.
- Threshold semantics are domain-specific.
raw_score β₯ 0.55means "above the mean," not "above GPT-4's mean" or "above Gemini's mean." Every model has its own calibration. - History assumes relevance. The
requerysignal assumes the caller passes history from the same conversation. Passing unrelated history will inflate the score. - No memory across sessions. The model has no persistent state. Each call is independent.
- No semantic understanding. It can detect that a user appears to be operating above the ceiling. It cannot verify that they actually are.
- Calibration is threshold-only. The signal weights are fixed. To re-weight signals for a specific population, edit the config directly.
Reference
Extracted from the XuanJi-ISA exploratory track, "Visual-Whole
Reasoning Interfaces" (issue #122), specifically the blueprint that
proposed reader models as an unmodeled axis in LLM delivery.
The core insight β that the mean is not the population, and that the tail is where the interesting users live β is the same one behind the entire reader-model category.
License
Apache-2.0