hv-tempo

Reading pace variation predictor. Given a text, output where the reader will slow down and where they'll speed up.

The claim in one sentence

hv-ttu predicts total comprehension time. hv-tempo predicts variation β€” the shape of the pace across a text.

What it produces

For each segment of the text (sentence or small group of sentences):

  • WPM β€” estimated reading speed for that span
  • duration_s β€” how long that span will take
  • slowdown β€” multiplicative factor vs the baseline
  • reasons β€” which features caused the slowdown
  • features β€” the raw surface signals

Plus aggregate stats: mean WPM, total duration, pace variance, slowest and fastest spans.

Install

pip install numpy

Actually β€” no dependencies. Pure stdlib.

## Usage

### Analyze a text

```python
from hv_tempo import HVTempo

m = HVTempo()
report = m.analyze("A black hole is a region of spacetime where gravity...")

print(report.mean_wpm)              # ~123
print(report.total_duration_s)      # ~54 s
print(report.pace_variance)         # 0.181
print(report.slowest.wpm)           # 117
print(report.slowest.reasons)       # ['long sentences', 'rare vocabulary', ...]

Just the numbers

m.wpm(text)              # mean WPM
m.duration(text)         # total seconds
m.slowest_spans(text, 5) # top-5 slowest segments
m.fastest_spans(text, 5) # top-5 fastest segments

CLI

python hv_tempo.py                          # run all demos
python hv_tempo.py --text "..."             # analyze a text
python hv_tempo.py --text "..." --wpm       # just the mean WPM
python hv_tempo.py --text "..." --duration  # just the duration
python hv_tempo.py --text "..." --slowest 5 # top 5 slowest spans
python hv_tempo.py --text "..." --json      # JSON output

The model

Per-span features:

feature what it captures
sentence_len_signed words per sentence above/below 15 (signed)
clauses_per_sentence commas + subordinators per sentence
rare_rate fraction of words outside a common-word list, length β‰₯ 7
abstract_rate fraction of words ending in -tion, -ness, -ity, etc.
digit_rate numbers per word
negation_rate negations per word
hedge_rate hedges per word ("may", "might", "typically")
conditional_rate conditionals per word ("if", "unless")
passive_rate "be + -ed" constructions per sentence
list_rate list markers per sentence (negative weight)

Slowdown is computed as a multiplicative factor:

slowdown = exp( Ξ£ weight_i Β· feature_i )
wpm = baseline_wpm / slowdown

Weights are signed: positive slows reading, negative speeds it up. Two features have negative weights: sentence_len_signed (for short sentences) and list_rate (lists always speed reading up).

Benchmarks

Sample texts

text words mean WPM duration pace variance
List-heavy 33 304 6.5 s n/a
Fiction (Hemingway) 51 270 11.3 s n/a
Technical (install) 83 180 27.6 s 0.055
Academic (black hole) 111 123 54.0 s 0.181

Reading the table:

  • List-heavy is fastest. Short clauses, list bonus.
  • Fiction is nearly as fast. Short sentences, common vocabulary.
  • Technical is medium. Numerals and rare vocabulary slow it down.
  • Academic is slowest. Long sentences, dense clauses, rare vocabulary.

Pace drivers per span

The model reports the top contributors per span. Examples:

  • Fiction: short sentences (8w avg) β€” faster, dense clauses (0.5/sent)
  • Academic: long sentences (24w avg), rare vocabulary (26%), dense clauses (1.0/sent)
  • Technical: rare vocabulary (24%), short sentences (12w avg) β€” faster, dense clauses (1.6/sent)
  • List-heavy: short sentences (6w avg) β€” faster, list structure (4 items) β€” faster, rare vocabulary (36%)

When to use it

  • Typography β€” where to adjust line length or font weight.
  • Editing β€” where to break long sentences.
  • Prose rhythm β€” testing whether a passage is uniform or varied.
  • Audiobook pacing β€” pre-planning where the narrator slows.
  • Ad copy β€” where the reader's eye drags.
  • Technical writing β€” identifying the load-bearing sentences.
  • Text difficulty β€” comparing two versions of the same content.

When not to use it

  • As ground truth. This is a heuristic model. It estimates relative variation, not absolute times. For absolute times, calibrate against real reading data.
  • For non-English text. The lexicons and patterns are English-specific.
  • For mathematical or code-heavy text. The surface features don't capture symbolic complexity.

Honest limitations

  • The weights are hand-tuned. They are plausible, internally consistent, and produce sensible rankings. They are not fit to a real user study.
  • The common-word list is ~800 words with a simple stemmer. Handles inflections (walked β†’ walk, friends β†’ friend) but not irregular forms.
  • The passive detector is regex-based. Catches be + -ed and a short list of irregular participles, but misses others.
  • The abstraction detector uses suffixes. Misses abstract words that don't fit suffix patterns (freedom, justice).
  • The baseline WPM is fixed. Real readers vary from 150 to 400 WPM. The model's output should be interpreted as a ratio to the baseline, not as an absolute time.
  • List detection requires bullets at line starts. A single stray "1." mid-sentence doesn't count as a list.
  • No calibration. Extending the model to fit real reading data is future work.

Reference

Part of the reader-model series. Companion to hv-ttu (total comprehension time) and hv-locality (feature-map locality).

hv-ttu answers how long will this take? hv-tempo answers where will it be slow?

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support