hv-reader / README.md
zeechimp's picture
Create README.md
273450b verified
|
Raw History Blame Contribute Delete
7.81 kB
---
license: apache-2.0
library_name: numpy
tags:
- reading
- reading-time
- comprehension
- memory
- attention
- discourse
- heuristic
- cpu
pipeline_tag: text-classification
language:
- en
---
# hv-reader
**The reading experience, in one call.**
Given a text, produce a `ReadingProfile`: pace, memory, passes, slip, and
wall β€” the five axes that describe what it is like to read a text.
## The claim in one sentence
The seven `hv-` models each describe one aspect of reading. `hv-reader`
describes the *experience*: a five-axis signature that no single axis can
produce.
## What it produces
A `ReadingProfile` containing:
- **pace** β€” mean reading speed, normalized to `[0, 1]`
- **memory** β€” fraction of the text surviving a week
- **passes** β€” how many reads are required to resolve the text
- **slip** β€” probability of an attention lapse mid-read
- **wall** β€” probability of reader refusal
Plus two derived scalars:
- **ttu_s** β€” total reading time in seconds
- **hunger_final** β€” unresolved curiosity at the end
And per-sentence detail: WPM, density, retention, fold load, slip, wall,
wall type, reasons.
## Install
```bash
pip install numpy
Actually β€” no dependencies. Pure stdlib.
## Usage
### Demo
```bash
python hv_reader.py
Runs four synthetic samples β€” fiction, academic, overclaim, monotone β€”
and prints the full profile for each.
### Analyze a text
```bash
python hv_reader.py --text "The old man walked slowly to the boat."
python hv_reader.py --text - < essay.txt
python hv_reader.py --text "..." --profile # five-axis summary only
python hv_reader.py --text "..." --json # full JSON
```
### Python
```python
from hv_reader import HVReader
m = HVReader()
profile = m.analyze(text)
print(profile.pace) # 0.672
print(profile.memory) # 0.269
print(profile.passes) # 1.167
print(profile.slip) # 0.615
print(profile.wall) # 0.000
print(profile.summary)
```
## The five axes
### pace β€” from `hv-tempo`
Weighted sum of surface features (sentence length, clause density, rare
vocabulary, numerals, negation, hedges, conditionals, passive voice).
Normalized by `baseline_wpm = 220`.
```
pace = mean_wpm / 400
```
A text at 400 WPM or faster has `pace = 1.0`. Fiction runs at ~0.67.
Academic prose at ~0.30.
### memory β€” from `hv-forget`
Per-sentence stability grows with density, rare-word ratio, and salience
(numbers, proper nouns). Retention after 7 days follows Ebbinghaus decay:
```
stability = 1.0 + 5.0 Γ— density + 3.0 Γ— rare_ratio + 2.0 Γ— salience
retention = exp(βˆ’7 / stability)
```
Memory is the mean retention over all sentences. Dense, specific sentences
last longer. Filler vanishes.
### passes β€” from `hv-fold`
Novelty-based load: orphan definites, forward references, reentry markers.
Fold grows with load:
```
passes = 1 + total_load / n_sentences
```
Fold of 1.0 = perfectly linear. Fold of 2.0+ = the text must be read twice
to be understood.
### slip β€” from `hv-slip`
Attention-lapse probability. Three signals:
1. **monotone** β€” sentences of similar length create drowsiness
2. **repetition** β€” content-word overlap with the previous sentence
3. **absence** β€” no numbers or proper nouns to anchor attention
Slip is per-sentence, then averaged. A text of short similar sentences
scores high. A text with varied rhythm and specific content scores low.
### wall β€” from `hv-wall`
Highest per-sentence wall score. Three wall types, in priority order:
1. **contradiction** β€” the sentence contradicts an earlier claim
2. **overclaim** β€” certainty exceeds evidence
3. **gap** β€” a conclusion is drawn without prior support
Wall is the *maximum*, not the mean. A single wall is a single wall.
## Benchmarks
### Four samples
| sample | words | pace | memory | passes | slip | wall | TTU |
|---|---:|---:|---:|---:|---:|---:|---:|
| Fiction | 51 | 0.67 | 0.27 | 1.17 | 0.62 | 0.00 | 11.4 s |
| Academic | 100 | 0.30 | 0.33 | 1.62 | 0.46 | 0.00 | 50.8 s |
| Overclaim | 51 | 0.40 | 0.40 | 2.40 | 0.54 | 0.77 | 19.2 s |
| Monotone | 30 | 0.71 | 0.19 | 1.67 | 0.87 | 0.00 | 6.3 s |
**Reading the table:**
- **Fiction** β€” fast pace, low memory, moderate slip. The reader finishes.
- **Academic** β€” slow pace, moderate memory, moderate slip. The reader
works but stays engaged.
- **Overclaim** β€” mid pace, high memory, *high wall*. The reader stops.
- **Monotone** β€” fast pace, low memory, *highest slip*. The reader's eye
moves but the mind drifts.
The five-axis profile separates all four, even where individual axes
overlap. Fiction and monotone have similar pace. Academic and overclaim
have similar memory. The *shape* of the profile differs.
### Why the profile is not the sum of its parts
Two texts with identical pace can have different memory, slip, or wall.
Two texts with identical wall can have very different slip. The profile is
the *joint* description. Every single-axis model in the series collapses
to one of the five. The profile is the artifact.
## When to use it
- **Editing.** Find where the text breaks the reading experience.
- **Prescription.** "Your essay has pace 0.72 and slip 0.81 β€” it reads
fast but won't stick."
- **Comparative analysis.** Rank texts by *reading difficulty shape*, not
by a single readability score.
- **Reader simulation.** Feed the profile into a decision model that
predicts whether a reader finishes.
- **Learning design.** Predict where students will zone out.
- **Marketing copy.** Predict where attention drops before conversion.
## When *not* to use it
- **As a reader study.** The axes are heuristic. The profile ranks; it
does not measure.
- **For non-English text.** All lexicons are English.
- **For very short text.** Fewer than three sentences gives unstable
per-axis values.
- **As ground truth on comprehension.** `memory` predicts *what survives*,
not *what is understood*.
- **For poetry or experimental prose.** The models assume propositional
content.
## Honest limitations
- **The five axes are independent by construction.** A text can be slow
and forgettable. A text can be fast and memorable. A text can be slow
and unforgettable. The profile reports the joint shape, but the axes
do not automatically produce a single scalar. Combining them into a
"reader burden" or "expected completion" score is left to the caller.
- **The pace model uses hand-tuned weights.** Not calibrated against
eye-tracking or reading-time studies.
- **The memory model uses Ebbinghaus decay with a specific half-life
formula.** Real memory is content-dependent and reader-dependent.
- **The slip model conflates monotony with attention lapse.** A reader
can be lulled by a beautiful monotonous passage (Whitman) as easily as
bored by filler. The model does not distinguish.
- **The wall model's antonym table is hand-curated.** Contradictions that
rely on entailment or world knowledge will not be detected.
- **No calibration against real readers.** The demo samples are
synthetic. The profile has not been validated against reader
self-reports, completion rates, or comprehension tests.
- **Fiction shows a slip score higher than intuition suggests.** Short
sentences with similar lengths trigger the monotone detector even when
the passage is engaging. The reason is the same as above: monotony and
attention lapse are correlated but not identical.
## Reference
Part of the reader-model series. Unifies seven component models into a
single profile:
| axis | component model |
|---|---|
| pace | `hv-tempo` |
| memory | `hv-forget` |
| passes | `hv-fold` |
| slip | `hv-slip` |
| wall | `hv-wall` |
The component models remain independently useful. `hv-reader` is the
joint signature.
## License
Apache-2.0