|
Download README.md from zeechimp/hv-reader: direct link, hf CLI and curl.
- Browser
- Download file 7.81 kB
-
https://huggingface.co/zeechimp/hv-reader/resolve/main/README.md
- Command line
-
hf download hf://zeechimp/hv-reader/README.md
-
curl -L -o README.md https://huggingface.co/zeechimp/hv-reader/resolve/main/README.md
7.81 kB
| license: apache-2.0 | |
| library_name: numpy | |
| tags: | |
| - reading | |
| - reading-time | |
| - comprehension | |
| - memory | |
| - attention | |
| - discourse | |
| - heuristic | |
| - cpu | |
| pipeline_tag: text-classification | |
| language: | |
| - en | |
| # hv-reader | |
| **The reading experience, in one call.** | |
| Given a text, produce a `ReadingProfile`: pace, memory, passes, slip, and | |
| wall β the five axes that describe what it is like to read a text. | |
| ## The claim in one sentence | |
| The seven `hv-` models each describe one aspect of reading. `hv-reader` | |
| describes the *experience*: a five-axis signature that no single axis can | |
| produce. | |
| ## What it produces | |
| A `ReadingProfile` containing: | |
| - **pace** β mean reading speed, normalized to `[0, 1]` | |
| - **memory** β fraction of the text surviving a week | |
| - **passes** β how many reads are required to resolve the text | |
| - **slip** β probability of an attention lapse mid-read | |
| - **wall** β probability of reader refusal | |
| Plus two derived scalars: | |
| - **ttu_s** β total reading time in seconds | |
| - **hunger_final** β unresolved curiosity at the end | |
| And per-sentence detail: WPM, density, retention, fold load, slip, wall, | |
| wall type, reasons. | |
| ## Install | |
| ```bash | |
| pip install numpy | |
| Actually β no dependencies. Pure stdlib. | |
| ## Usage | |
| ### Demo | |
| ```bash | |
| python hv_reader.py | |
| Runs four synthetic samples β fiction, academic, overclaim, monotone β | |
| and prints the full profile for each. | |
| ### Analyze a text | |
| ```bash | |
| python hv_reader.py --text "The old man walked slowly to the boat." | |
| python hv_reader.py --text - < essay.txt | |
| python hv_reader.py --text "..." --profile # five-axis summary only | |
| python hv_reader.py --text "..." --json # full JSON | |
| ``` | |
| ### Python | |
| ```python | |
| from hv_reader import HVReader | |
| m = HVReader() | |
| profile = m.analyze(text) | |
| print(profile.pace) # 0.672 | |
| print(profile.memory) # 0.269 | |
| print(profile.passes) # 1.167 | |
| print(profile.slip) # 0.615 | |
| print(profile.wall) # 0.000 | |
| print(profile.summary) | |
| ``` | |
| ## The five axes | |
| ### pace β from `hv-tempo` | |
| Weighted sum of surface features (sentence length, clause density, rare | |
| vocabulary, numerals, negation, hedges, conditionals, passive voice). | |
| Normalized by `baseline_wpm = 220`. | |
| ``` | |
| pace = mean_wpm / 400 | |
| ``` | |
| A text at 400 WPM or faster has `pace = 1.0`. Fiction runs at ~0.67. | |
| Academic prose at ~0.30. | |
| ### memory β from `hv-forget` | |
| Per-sentence stability grows with density, rare-word ratio, and salience | |
| (numbers, proper nouns). Retention after 7 days follows Ebbinghaus decay: | |
| ``` | |
| stability = 1.0 + 5.0 Γ density + 3.0 Γ rare_ratio + 2.0 Γ salience | |
| retention = exp(β7 / stability) | |
| ``` | |
| Memory is the mean retention over all sentences. Dense, specific sentences | |
| last longer. Filler vanishes. | |
| ### passes β from `hv-fold` | |
| Novelty-based load: orphan definites, forward references, reentry markers. | |
| Fold grows with load: | |
| ``` | |
| passes = 1 + total_load / n_sentences | |
| ``` | |
| Fold of 1.0 = perfectly linear. Fold of 2.0+ = the text must be read twice | |
| to be understood. | |
| ### slip β from `hv-slip` | |
| Attention-lapse probability. Three signals: | |
| 1. **monotone** β sentences of similar length create drowsiness | |
| 2. **repetition** β content-word overlap with the previous sentence | |
| 3. **absence** β no numbers or proper nouns to anchor attention | |
| Slip is per-sentence, then averaged. A text of short similar sentences | |
| scores high. A text with varied rhythm and specific content scores low. | |
| ### wall β from `hv-wall` | |
| Highest per-sentence wall score. Three wall types, in priority order: | |
| 1. **contradiction** β the sentence contradicts an earlier claim | |
| 2. **overclaim** β certainty exceeds evidence | |
| 3. **gap** β a conclusion is drawn without prior support | |
| Wall is the *maximum*, not the mean. A single wall is a single wall. | |
| ## Benchmarks | |
| ### Four samples | |
| | sample | words | pace | memory | passes | slip | wall | TTU | | |
| |---|---:|---:|---:|---:|---:|---:|---:| | |
| | Fiction | 51 | 0.67 | 0.27 | 1.17 | 0.62 | 0.00 | 11.4 s | | |
| | Academic | 100 | 0.30 | 0.33 | 1.62 | 0.46 | 0.00 | 50.8 s | | |
| | Overclaim | 51 | 0.40 | 0.40 | 2.40 | 0.54 | 0.77 | 19.2 s | | |
| | Monotone | 30 | 0.71 | 0.19 | 1.67 | 0.87 | 0.00 | 6.3 s | | |
| **Reading the table:** | |
| - **Fiction** β fast pace, low memory, moderate slip. The reader finishes. | |
| - **Academic** β slow pace, moderate memory, moderate slip. The reader | |
| works but stays engaged. | |
| - **Overclaim** β mid pace, high memory, *high wall*. The reader stops. | |
| - **Monotone** β fast pace, low memory, *highest slip*. The reader's eye | |
| moves but the mind drifts. | |
| The five-axis profile separates all four, even where individual axes | |
| overlap. Fiction and monotone have similar pace. Academic and overclaim | |
| have similar memory. The *shape* of the profile differs. | |
| ### Why the profile is not the sum of its parts | |
| Two texts with identical pace can have different memory, slip, or wall. | |
| Two texts with identical wall can have very different slip. The profile is | |
| the *joint* description. Every single-axis model in the series collapses | |
| to one of the five. The profile is the artifact. | |
| ## When to use it | |
| - **Editing.** Find where the text breaks the reading experience. | |
| - **Prescription.** "Your essay has pace 0.72 and slip 0.81 β it reads | |
| fast but won't stick." | |
| - **Comparative analysis.** Rank texts by *reading difficulty shape*, not | |
| by a single readability score. | |
| - **Reader simulation.** Feed the profile into a decision model that | |
| predicts whether a reader finishes. | |
| - **Learning design.** Predict where students will zone out. | |
| - **Marketing copy.** Predict where attention drops before conversion. | |
| ## When *not* to use it | |
| - **As a reader study.** The axes are heuristic. The profile ranks; it | |
| does not measure. | |
| - **For non-English text.** All lexicons are English. | |
| - **For very short text.** Fewer than three sentences gives unstable | |
| per-axis values. | |
| - **As ground truth on comprehension.** `memory` predicts *what survives*, | |
| not *what is understood*. | |
| - **For poetry or experimental prose.** The models assume propositional | |
| content. | |
| ## Honest limitations | |
| - **The five axes are independent by construction.** A text can be slow | |
| and forgettable. A text can be fast and memorable. A text can be slow | |
| and unforgettable. The profile reports the joint shape, but the axes | |
| do not automatically produce a single scalar. Combining them into a | |
| "reader burden" or "expected completion" score is left to the caller. | |
| - **The pace model uses hand-tuned weights.** Not calibrated against | |
| eye-tracking or reading-time studies. | |
| - **The memory model uses Ebbinghaus decay with a specific half-life | |
| formula.** Real memory is content-dependent and reader-dependent. | |
| - **The slip model conflates monotony with attention lapse.** A reader | |
| can be lulled by a beautiful monotonous passage (Whitman) as easily as | |
| bored by filler. The model does not distinguish. | |
| - **The wall model's antonym table is hand-curated.** Contradictions that | |
| rely on entailment or world knowledge will not be detected. | |
| - **No calibration against real readers.** The demo samples are | |
| synthetic. The profile has not been validated against reader | |
| self-reports, completion rates, or comprehension tests. | |
| - **Fiction shows a slip score higher than intuition suggests.** Short | |
| sentences with similar lengths trigger the monotone detector even when | |
| the passage is engaging. The reason is the same as above: monotony and | |
| attention lapse are correlated but not identical. | |
| ## Reference | |
| Part of the reader-model series. Unifies seven component models into a | |
| single profile: | |
| | axis | component model | | |
| |---|---| | |
| | pace | `hv-tempo` | | |
| | memory | `hv-forget` | | |
| | passes | `hv-fold` | | |
| | slip | `hv-slip` | | |
| | wall | `hv-wall` | | |
| The component models remain independently useful. `hv-reader` is the | |
| joint signature. | |
| ## License | |
| Apache-2.0 |