hv-ttu
Time-to-understanding predictor. Text in, predicted seconds out.
The claim in one sentence
Every LLM benchmark measures correctness. None measures cost to
the reader. hv-ttu is the missing metric: given text and a delivery
mode, predict how many seconds a reader will take to act on it.
What it predicts
TTU = time from "reader sees the artifact" to "reader acts on it."
Not included: network latency, generation time. Those are orthogonal. TTU is specifically the part that changes with the shape of the delivered answer.
Install
pip install numpy
Actually β no dependencies at all. Pure stdlib. Runs anywhere Python 3.9+
runs.
## Usage
### Quick score
```python
from hv_ttu import HVTTU
m = HVTTU()
seconds = m.ttu("A black hole is a region of spacetime...", mode="text")
print(f"{seconds:.1f}s") # e.g. "56.9s"
Compare all delivery modes
for mode, s in m.ttu_all_modes(text).items():
print(f"{mode:<14} {s:>8.1f}s")
Best mode + speedup
best, best_ttu = m.best_mode(text)
baseline = m.ttu(text, mode="text")
print(f"best: {best} ({baseline / best_ttu:.1f}x faster)")
# For a definition -> ("svg_tree", 7.24)
# For a procedure -> ("svg_strip", 5.57)
Reader / format mismatch
# A visual-first reader being handed raw text
m.ttu(text, mode="text", reader_mode="visual-first") # ~2.5x slower
# Same reader given the SVG tree
m.ttu(text, mode="svg_tree", reader_mode="visual-first") # ~1.0x
Content-type Γ mode fit
# A definition in pacman is wrong β no action to encode
m.ttu(definition_text, mode="pacman") # applies 5x content penalty
# A procedure in svg_strip is right β the sequence is the answer
m.ttu(procedure_text, mode="svg_strip") # applies 0.7x bonus
Breakdown
bd = m.ttu_breakdown(text, mode="text")
# {'scan': 0.5, 'read': 33.3, 'integrate': 7.6,
# 'verify': 10.5, 'act': 5.0,
# 'reader_fit': 1.0, 'content_fit': 1.0,
# 'subtotal': 56.9, 'total': 56.9}
Calibrate to your population
m.calibrate([
("the text they saw", "card_deck", 22.4), # actual seconds
("another text", "svg_tree", 8.9),
])
# The model fits a single global scale by least squares.
CLI
# score one text
echo "A black hole is..." | python hv_ttu.py
# compare modes
python hv_ttu.py --text "..." --all
# best mode
python hv_ttu.py --text "..." --best
# breakdown for a specific mode
python hv_ttu.py --text "..." --mode svg_tree --breakdown
# force a content type
python hv_ttu.py --text "..." --all --content-type procedure
# no args = run all demos
python hv_ttu.py
The formula
TTU(text, mode, reader, content_type) =
[ scan + read + integrate + verify + act ]
Γ reader_fit(reader, mode)
Γ content_fit(content_type, mode)
Γ scale
Every term is visible and tunable:
| term | what it measures | units |
|---|---|---|
scan |
time to find where to look | seconds |
read |
60 Β· effective_words / WPM[mode] | seconds |
integrate |
buffer burden β clauses, hedges, conditionals, abstractions | seconds |
verify |
trust checking β numbers, caveats, negations, citations | seconds |
act |
deciding what to do β 2s if there's a clear imperative, 5s otherwise | seconds |
effective_words = n_words Β· retention[mode] β a card deck doesn't
contain the same words as the paragraph it summarizes. This is the key
correction that makes card decks actually cheaper than walls of text.
Mode constants (defaults)
| mode | retention | WPM | scan (s) |
|---|---|---|---|
| text | 1.00 | 200 | 0.5 |
| card_deck | 0.40 | 300 | 0.4 |
| svg_tree | 0.15 | 450 | 1.2 |
| svg_strip | 0.20 | 400 | 1.0 |
| pacman | 0.05 | β | 2.0 |
| color_field | 0.05 | β | 0.8 |
| rhythm | 0.05 | β | 0.6 |
WPM values are within published reading-speed ranges. Retention factors are the fraction of content that survives the render β a card deck is ~40% of the original text, an SVG tree ~15%.
The fit penalties (why time isn't enough)
Speed alone would recommend pacman for every text β it's the fastest
to glance at. But rendering a definition as a Pac-Man screen produces a
scene that answers nothing. The scene is fast because it discards
everything; the reader can look at it and still not know what a black
hole is.
So the model has two multipliers applied to the raw time cost:
Reader fit β how well the format matches the reader's cognitive mode
| reader_mode | best format | penalty on wrong format |
|---|---|---|
| text-first | text | up to 2.5Γ |
| visual-first | svg_tree | up to 2.5Γ |
| music-first | rhythm | up to 3.0Γ |
| whole-pattern | svg_tree | up to 2.2Γ |
| mobile-native | card_deck | up to 2.0Γ |
| colors-first | color_field | up to 2.8Γ |
| designer | svg_tree | up to 2.4Γ |
Content fit β how well the format preserves the content's structure
| content_type | best format (bonus) | worst format (penalty) |
|---|---|---|
| definition | svg_tree (0.8Γ) | pacman (5.0Γ) |
| procedure | svg_strip (0.7Γ) | color_field (4.0Γ) |
| data | color_field (0.7Γ) | svg_strip (3.0Γ) |
| comparison | card_deck (0.9Γ) | rhythm (3.0Γ) |
| narrative | card_deck (0.9Γ) | color_field (3.0Γ) |
| code | text (1.0Γ) | rhythm (4.0Γ) |
| general | text/card_deck/svg_tree (1.0Γ) | pacman (2.0Γ) |
The two multipliers are orthogonal and multiply together. A visual-first
reader receiving a procedure as a Pac-Man screen pays
reader(1.10) Γ content(2.50) = 2.75Γ on top of base time. A
visual-first reader receiving the same procedure as svg_strip pays
reader(1.10) Γ content(0.70) = 0.77Γ β a bonus, because both fit.
Result: the model recommends the format that is both fast and semantically appropriate. Speed is one axis; fit is the other.
Benchmarks
Black hole definition (111 words, 6 sentences, 2 caveats)
Content type detected: definition.
| mode | TTU (s) | speedup | applicable? |
|---|---|---|---|
| text | 56.90 | 1.00Γ | yes |
| card_deck | 17.32 | 3.29Γ | yes |
| svg_tree | 7.24 | 7.86Γ | yes β best |
| pacman | 18.55 | 3.07Γ | not applicable |
| color_field | 18.45 | 3.08Γ | not applicable |
| rhythm | 17.85 | 3.19Γ | not applicable |
| svg_strip | 28.47 | 2.00Γ | not applicable |
Best applicable mode: svg_tree, 7.86Γ faster than text.
Flat-tire procedure (81 words, 6 steps, 1 warning)
Content type detected: procedure.
| mode | TTU (s) | speedup | applicable? |
|---|---|---|---|
| text | 40.90 | 1.00Γ | yes |
| card_deck | 12.04 | 3.40Γ | yes |
| svg_tree | 16.08 | 2.54Γ | yes |
| svg_strip | 5.57 | 7.34Γ | yes β best |
| pacman | 9.10 | 4.49Γ | yes (imperative present) |
| color_field | 24.16 | 1.69Γ | not applicable |
| rhythm | 17.52 | 2.33Γ | not applicable |
Best applicable mode: svg_strip, 7.34Γ faster than text.
Reader/format mismatch (black hole definition, text mode)
| reader_mode | TTU (s) | multiplier |
|---|---|---|
| text-first | 56.90 | 1.00Γ |
| visual-first | 142.25 | 2.50Γ |
| music-first | 170.70 | 3.00Γ |
| whole-pattern | 125.18 | 2.20Γ |
| mobile-native | 113.80 | 2.00Γ |
| colors-first | 159.32 | 2.80Γ |
| designer | 136.56 | 2.40Γ |
Calibration
The model predicts relative TTU correctly out of the box β a card deck is ~3β4Γ faster than a wall of text in any population. To fix the absolute scale for your population, provide observed triples:
m.calibrate([
(text_a, "card_deck", 22.4), # observed in real users, seconds
(text_b, "svg_tree", 8.9),
(text_c, "text", 61.2),
])
The model fits one global scale multiplier by least squares. Predictions
move to your units. Everything else stays intact.
When to use it
- Yes: measuring whether a delivery change actually helps. Ranking candidate outputs by comprehension cost. Reporting TTU in an LLM product's analytics. Any A/B test comparing text vs mode-matched output.
- Maybe: as a training signal for a model that generates low-TTU outputs. As a filter for a delivery router.
- No: as a ground-truth measure. This is a heuristic model, not a measurement. Validate against real users before making any claim stronger than "in relative terms, the card deck is faster."
Honest limitations
- Coefficients are not fit to a real user study. They are within published ranges and are internally consistent. They are not individually validated.
- Retention factors are estimates. A real card deck's word count depends on the summarizer.
- The compatibility table is heuristic. It encodes plausible penalties, not measured ones.
- The content-fit table is heuristic. Same caveat: plausible, not measured. A real study could re-fit both tables.
- Content classification is regex-based. It handles the common
cases. Mixed-type texts fall back to
general. - The model has no memory. It cannot use reader history, fatigue, or session state.
- Calibration is a single global scale. It cannot adjust per-mode or per-reader. Multi-scale calibration is future work.
Reference
Extracted from the XuanJi-ISA exploratory track, "Visual-Whole
Reasoning Interfaces" (issue #122), specifically the blueprint that
proposed TTU as the primary evaluation metric for reader-facing
delivery.
The core insight β that delivery shape is orthogonal to model capability, that comprehension cost is measurable, and that "fastest" is not "best" β is the same one behind the entire reader-model category.
License
Apache-2.0