hv-ttu

Time-to-understanding predictor. Text in, predicted seconds out.

The claim in one sentence

Every LLM benchmark measures correctness. None measures cost to the reader. hv-ttu is the missing metric: given text and a delivery mode, predict how many seconds a reader will take to act on it.

What it predicts

TTU = time from "reader sees the artifact" to "reader acts on it."

Not included: network latency, generation time. Those are orthogonal. TTU is specifically the part that changes with the shape of the delivered answer.

Install

pip install numpy

Actually β€” no dependencies at all. Pure stdlib. Runs anywhere Python 3.9+
runs.

## Usage

### Quick score

```python
from hv_ttu import HVTTU

m = HVTTU()
seconds = m.ttu("A black hole is a region of spacetime...", mode="text")
print(f"{seconds:.1f}s")   # e.g. "56.9s"

Compare all delivery modes

for mode, s in m.ttu_all_modes(text).items():
    print(f"{mode:<14} {s:>8.1f}s")

Best mode + speedup

best, best_ttu = m.best_mode(text)
baseline = m.ttu(text, mode="text")
print(f"best: {best} ({baseline / best_ttu:.1f}x faster)")
# For a definition -> ("svg_tree", 7.24)
# For a procedure  -> ("svg_strip", 5.57)

Reader / format mismatch

# A visual-first reader being handed raw text
m.ttu(text, mode="text", reader_mode="visual-first")   # ~2.5x slower
# Same reader given the SVG tree
m.ttu(text, mode="svg_tree", reader_mode="visual-first")  # ~1.0x

Content-type Γ— mode fit

# A definition in pacman is wrong β€” no action to encode
m.ttu(definition_text, mode="pacman")   # applies 5x content penalty
# A procedure in svg_strip is right β€” the sequence is the answer
m.ttu(procedure_text, mode="svg_strip")  # applies 0.7x bonus

Breakdown

bd = m.ttu_breakdown(text, mode="text")
# {'scan': 0.5, 'read': 33.3, 'integrate': 7.6,
#  'verify': 10.5, 'act': 5.0,
#  'reader_fit': 1.0, 'content_fit': 1.0,
#  'subtotal': 56.9, 'total': 56.9}

Calibrate to your population

m.calibrate([
    ("the text they saw",  "card_deck", 22.4),   # actual seconds
    ("another text",       "svg_tree",   8.9),
])
# The model fits a single global scale by least squares.

CLI

# score one text
echo "A black hole is..." | python hv_ttu.py

# compare modes
python hv_ttu.py --text "..." --all

# best mode
python hv_ttu.py --text "..." --best

# breakdown for a specific mode
python hv_ttu.py --text "..." --mode svg_tree --breakdown

# force a content type
python hv_ttu.py --text "..." --all --content-type procedure

# no args = run all demos
python hv_ttu.py

The formula

TTU(text, mode, reader, content_type) =
    [ scan + read + integrate + verify + act ]
    Γ— reader_fit(reader, mode)
    Γ— content_fit(content_type, mode)
    Γ— scale

Every term is visible and tunable:

term what it measures units
scan time to find where to look seconds
read 60 Β· effective_words / WPM[mode] seconds
integrate buffer burden β€” clauses, hedges, conditionals, abstractions seconds
verify trust checking β€” numbers, caveats, negations, citations seconds
act deciding what to do β€” 2s if there's a clear imperative, 5s otherwise seconds

effective_words = n_words Β· retention[mode] β€” a card deck doesn't contain the same words as the paragraph it summarizes. This is the key correction that makes card decks actually cheaper than walls of text.

Mode constants (defaults)

mode retention WPM scan (s)
text 1.00 200 0.5
card_deck 0.40 300 0.4
svg_tree 0.15 450 1.2
svg_strip 0.20 400 1.0
pacman 0.05 β€” 2.0
color_field 0.05 β€” 0.8
rhythm 0.05 β€” 0.6

WPM values are within published reading-speed ranges. Retention factors are the fraction of content that survives the render β€” a card deck is ~40% of the original text, an SVG tree ~15%.

The fit penalties (why time isn't enough)

Speed alone would recommend pacman for every text β€” it's the fastest to glance at. But rendering a definition as a Pac-Man screen produces a scene that answers nothing. The scene is fast because it discards everything; the reader can look at it and still not know what a black hole is.

So the model has two multipliers applied to the raw time cost:

Reader fit β€” how well the format matches the reader's cognitive mode

reader_mode best format penalty on wrong format
text-first text up to 2.5Γ—
visual-first svg_tree up to 2.5Γ—
music-first rhythm up to 3.0Γ—
whole-pattern svg_tree up to 2.2Γ—
mobile-native card_deck up to 2.0Γ—
colors-first color_field up to 2.8Γ—
designer svg_tree up to 2.4Γ—

Content fit β€” how well the format preserves the content's structure

content_type best format (bonus) worst format (penalty)
definition svg_tree (0.8Γ—) pacman (5.0Γ—)
procedure svg_strip (0.7Γ—) color_field (4.0Γ—)
data color_field (0.7Γ—) svg_strip (3.0Γ—)
comparison card_deck (0.9Γ—) rhythm (3.0Γ—)
narrative card_deck (0.9Γ—) color_field (3.0Γ—)
code text (1.0Γ—) rhythm (4.0Γ—)
general text/card_deck/svg_tree (1.0Γ—) pacman (2.0Γ—)

The two multipliers are orthogonal and multiply together. A visual-first reader receiving a procedure as a Pac-Man screen pays reader(1.10) Γ— content(2.50) = 2.75Γ— on top of base time. A visual-first reader receiving the same procedure as svg_strip pays reader(1.10) Γ— content(0.70) = 0.77Γ— β€” a bonus, because both fit.

Result: the model recommends the format that is both fast and semantically appropriate. Speed is one axis; fit is the other.

Benchmarks

Black hole definition (111 words, 6 sentences, 2 caveats)

Content type detected: definition.

mode TTU (s) speedup applicable?
text 56.90 1.00Γ— yes
card_deck 17.32 3.29Γ— yes
svg_tree 7.24 7.86Γ— yes β€” best
pacman 18.55 3.07Γ— not applicable
color_field 18.45 3.08Γ— not applicable
rhythm 17.85 3.19Γ— not applicable
svg_strip 28.47 2.00Γ— not applicable

Best applicable mode: svg_tree, 7.86Γ— faster than text.

Flat-tire procedure (81 words, 6 steps, 1 warning)

Content type detected: procedure.

mode TTU (s) speedup applicable?
text 40.90 1.00Γ— yes
card_deck 12.04 3.40Γ— yes
svg_tree 16.08 2.54Γ— yes
svg_strip 5.57 7.34Γ— yes β€” best
pacman 9.10 4.49Γ— yes (imperative present)
color_field 24.16 1.69Γ— not applicable
rhythm 17.52 2.33Γ— not applicable

Best applicable mode: svg_strip, 7.34Γ— faster than text.

Reader/format mismatch (black hole definition, text mode)

reader_mode TTU (s) multiplier
text-first 56.90 1.00Γ—
visual-first 142.25 2.50Γ—
music-first 170.70 3.00Γ—
whole-pattern 125.18 2.20Γ—
mobile-native 113.80 2.00Γ—
colors-first 159.32 2.80Γ—
designer 136.56 2.40Γ—

Calibration

The model predicts relative TTU correctly out of the box β€” a card deck is ~3–4Γ— faster than a wall of text in any population. To fix the absolute scale for your population, provide observed triples:

m.calibrate([
    (text_a, "card_deck", 22.4),   # observed in real users, seconds
    (text_b, "svg_tree",   8.9),
    (text_c, "text",      61.2),
])

The model fits one global scale multiplier by least squares. Predictions move to your units. Everything else stays intact.

When to use it

  • Yes: measuring whether a delivery change actually helps. Ranking candidate outputs by comprehension cost. Reporting TTU in an LLM product's analytics. Any A/B test comparing text vs mode-matched output.
  • Maybe: as a training signal for a model that generates low-TTU outputs. As a filter for a delivery router.
  • No: as a ground-truth measure. This is a heuristic model, not a measurement. Validate against real users before making any claim stronger than "in relative terms, the card deck is faster."

Honest limitations

  • Coefficients are not fit to a real user study. They are within published ranges and are internally consistent. They are not individually validated.
  • Retention factors are estimates. A real card deck's word count depends on the summarizer.
  • The compatibility table is heuristic. It encodes plausible penalties, not measured ones.
  • The content-fit table is heuristic. Same caveat: plausible, not measured. A real study could re-fit both tables.
  • Content classification is regex-based. It handles the common cases. Mixed-type texts fall back to general.
  • The model has no memory. It cannot use reader history, fatigue, or session state.
  • Calibration is a single global scale. It cannot adjust per-mode or per-reader. Multi-scale calibration is future work.

Reference

Extracted from the XuanJi-ISA exploratory track, "Visual-Whole Reasoning Interfaces" (issue #122), specifically the blueprint that proposed TTU as the primary evaluation metric for reader-facing delivery.

The core insight β€” that delivery shape is orthogonal to model capability, that comprehension cost is measurable, and that "fastest" is not "best" β€” is the same one behind the entire reader-model category.

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support