--- license: apache-2.0 library_name: numpy tags: - reading - reading-time - comprehension - memory - attention - discourse - heuristic - cpu pipeline_tag: text-classification language: - en --- # hv-reader **The reading experience, in one call.** Given a text, produce a `ReadingProfile`: pace, memory, passes, slip, and wall — the five axes that describe what it is like to read a text. ## The claim in one sentence The seven `hv-` models each describe one aspect of reading. `hv-reader` describes the *experience*: a five-axis signature that no single axis can produce. ## What it produces A `ReadingProfile` containing: - **pace** — mean reading speed, normalized to `[0, 1]` - **memory** — fraction of the text surviving a week - **passes** — how many reads are required to resolve the text - **slip** — probability of an attention lapse mid-read - **wall** — probability of reader refusal Plus two derived scalars: - **ttu_s** — total reading time in seconds - **hunger_final** — unresolved curiosity at the end And per-sentence detail: WPM, density, retention, fold load, slip, wall, wall type, reasons. ## Install ```bash pip install numpy Actually — no dependencies. Pure stdlib. ## Usage ### Demo ```bash python hv_reader.py Runs four synthetic samples — fiction, academic, overclaim, monotone — and prints the full profile for each. ### Analyze a text ```bash python hv_reader.py --text "The old man walked slowly to the boat." python hv_reader.py --text - < essay.txt python hv_reader.py --text "..." --profile # five-axis summary only python hv_reader.py --text "..." --json # full JSON ``` ### Python ```python from hv_reader import HVReader m = HVReader() profile = m.analyze(text) print(profile.pace) # 0.672 print(profile.memory) # 0.269 print(profile.passes) # 1.167 print(profile.slip) # 0.615 print(profile.wall) # 0.000 print(profile.summary) ``` ## The five axes ### pace — from `hv-tempo` Weighted sum of surface features (sentence length, clause density, rare vocabulary, numerals, negation, hedges, conditionals, passive voice). Normalized by `baseline_wpm = 220`. ``` pace = mean_wpm / 400 ``` A text at 400 WPM or faster has `pace = 1.0`. Fiction runs at ~0.67. Academic prose at ~0.30. ### memory — from `hv-forget` Per-sentence stability grows with density, rare-word ratio, and salience (numbers, proper nouns). Retention after 7 days follows Ebbinghaus decay: ``` stability = 1.0 + 5.0 × density + 3.0 × rare_ratio + 2.0 × salience retention = exp(−7 / stability) ``` Memory is the mean retention over all sentences. Dense, specific sentences last longer. Filler vanishes. ### passes — from `hv-fold` Novelty-based load: orphan definites, forward references, reentry markers. Fold grows with load: ``` passes = 1 + total_load / n_sentences ``` Fold of 1.0 = perfectly linear. Fold of 2.0+ = the text must be read twice to be understood. ### slip — from `hv-slip` Attention-lapse probability. Three signals: 1. **monotone** — sentences of similar length create drowsiness 2. **repetition** — content-word overlap with the previous sentence 3. **absence** — no numbers or proper nouns to anchor attention Slip is per-sentence, then averaged. A text of short similar sentences scores high. A text with varied rhythm and specific content scores low. ### wall — from `hv-wall` Highest per-sentence wall score. Three wall types, in priority order: 1. **contradiction** — the sentence contradicts an earlier claim 2. **overclaim** — certainty exceeds evidence 3. **gap** — a conclusion is drawn without prior support Wall is the *maximum*, not the mean. A single wall is a single wall. ## Benchmarks ### Four samples | sample | words | pace | memory | passes | slip | wall | TTU | |---|---:|---:|---:|---:|---:|---:|---:| | Fiction | 51 | 0.67 | 0.27 | 1.17 | 0.62 | 0.00 | 11.4 s | | Academic | 100 | 0.30 | 0.33 | 1.62 | 0.46 | 0.00 | 50.8 s | | Overclaim | 51 | 0.40 | 0.40 | 2.40 | 0.54 | 0.77 | 19.2 s | | Monotone | 30 | 0.71 | 0.19 | 1.67 | 0.87 | 0.00 | 6.3 s | **Reading the table:** - **Fiction** — fast pace, low memory, moderate slip. The reader finishes. - **Academic** — slow pace, moderate memory, moderate slip. The reader works but stays engaged. - **Overclaim** — mid pace, high memory, *high wall*. The reader stops. - **Monotone** — fast pace, low memory, *highest slip*. The reader's eye moves but the mind drifts. The five-axis profile separates all four, even where individual axes overlap. Fiction and monotone have similar pace. Academic and overclaim have similar memory. The *shape* of the profile differs. ### Why the profile is not the sum of its parts Two texts with identical pace can have different memory, slip, or wall. Two texts with identical wall can have very different slip. The profile is the *joint* description. Every single-axis model in the series collapses to one of the five. The profile is the artifact. ## When to use it - **Editing.** Find where the text breaks the reading experience. - **Prescription.** "Your essay has pace 0.72 and slip 0.81 — it reads fast but won't stick." - **Comparative analysis.** Rank texts by *reading difficulty shape*, not by a single readability score. - **Reader simulation.** Feed the profile into a decision model that predicts whether a reader finishes. - **Learning design.** Predict where students will zone out. - **Marketing copy.** Predict where attention drops before conversion. ## When *not* to use it - **As a reader study.** The axes are heuristic. The profile ranks; it does not measure. - **For non-English text.** All lexicons are English. - **For very short text.** Fewer than three sentences gives unstable per-axis values. - **As ground truth on comprehension.** `memory` predicts *what survives*, not *what is understood*. - **For poetry or experimental prose.** The models assume propositional content. ## Honest limitations - **The five axes are independent by construction.** A text can be slow and forgettable. A text can be fast and memorable. A text can be slow and unforgettable. The profile reports the joint shape, but the axes do not automatically produce a single scalar. Combining them into a "reader burden" or "expected completion" score is left to the caller. - **The pace model uses hand-tuned weights.** Not calibrated against eye-tracking or reading-time studies. - **The memory model uses Ebbinghaus decay with a specific half-life formula.** Real memory is content-dependent and reader-dependent. - **The slip model conflates monotony with attention lapse.** A reader can be lulled by a beautiful monotonous passage (Whitman) as easily as bored by filler. The model does not distinguish. - **The wall model's antonym table is hand-curated.** Contradictions that rely on entailment or world knowledge will not be detected. - **No calibration against real readers.** The demo samples are synthetic. The profile has not been validated against reader self-reports, completion rates, or comprehension tests. - **Fiction shows a slip score higher than intuition suggests.** Short sentences with similar lengths trigger the monotone detector even when the passage is engaging. The reason is the same as above: monotony and attention lapse are correlated but not identical. ## Reference Part of the reader-model series. Unifies seven component models into a single profile: | axis | component model | |---|---| | pace | `hv-tempo` | | memory | `hv-forget` | | passes | `hv-fold` | | slip | `hv-slip` | | wall | `hv-wall` | The component models remain independently useful. `hv-reader` is the joint signature. ## License Apache-2.0