hv-fold
Minimal reading passes predictor. Given a text, output the minimum number of times it must be read to be fully understood.
The claim in one sentence
hv-tempo predicts where the reader slows down. hv-fold predicts
how many times they must come back.
What it produces
For each sentence:
- load β the reference cost the sentence imposes on the reader
- n_refs β definite references in this sentence
- role β
demon|reentry|anchor|dependent|neutral - reasons β which features caused the load
Plus aggregates:
- fold β scalar, β₯ 1.0 (
1.0= perfectly linear) - fold_tier β
linear|two-pass|three-pass|recursive - reading_plan β one-line instruction
- anchor_sentences β sentences later material depends on
- reentry_points β sentences that explicitly demand backtracking
Install
pip install numpy
Actually β no dependencies. Pure stdlib.
## Usage
```python
from hv_fold import HVFold
m = HVFold()
report = m.analyze("This. The previous sentence is not the one you "
"should attend to. The point is elsewhere.")
print(report.fold) # 2.5
print(report.fold_tier) # 'recursive'
print(report.reading_plan) # 'Read recursively. Earlier material...'
### Just the numbers
```python
m.fold(text) # the scalar
m.reading_plan(text) # the plan string
m.reentry_points(text) # indices where the reader must jump back
CLI
python hv_fold.py # run all demos
python hv_fold.py --text "..." # full report
python hv_fold.py --text "..." --fold # just the scalar
python hv_fold.py --text "..." --tier # just the tier
python hv_fold.py --text "..." --plan # just the plan
python hv_fold.py --text "..." --json # JSON output
python hv_fold.py --text "..." --exempt-first # opening-sentence exemption
The model
Per-sentence load accumulates from five sources:
| source | weight | captures |
|---|---|---|
| standalone demonstrative | 2.0 | orphan anaphor ("This." "That.") |
| definite NP, no antecedent | 0.5 | "the dog" with no dog introduced |
| forward reference | 0.5 Γ delay | definite NP whose head appears only later |
| resolved reference | 0.1 | head appeared earlier; small memory cost |
| reentry marker | 1.0 | "above", "previous", "noted", "as discussed" |
unique referent (sun, moon, morning, β¦) |
0.0 | no cost β the referent is universal |
fold = 1 + Ξ£(load) / n_sentences
Tiers
| fold | tier | plan |
|---|---|---|
| < 1.15 | linear |
Read once, front to back. |
| < 1.75 | two-pass |
Re-scan for definite references without antecedents. |
| < 2.50 | three-pass |
Once for gist, once for coherence, once for detail. |
| β₯ 2.50 | recursive |
Earlier material only makes sense after later material. |
Benchmarks
| text | sentences | load | fold | tier |
|---|---|---|---|---|
| Linear (declarative) | 4 | 0.00 | 1.000 | linear |
| Narrative (orphan NPs) | 3 | 1.00 | 1.333 | two-pass |
| Academic (reentry-heavy) | 3 | 3.50 | 2.167 | three-pass |
| Meta (self-referential) | 3 | 4.50 | 2.500 | recursive |
Reading the table:
- Linear β pure declaratives with unique referents. Fold = 1. Read once.
- Narrative β orphan definites ("the room", "the chair") with no prior introduction. Small backtracking cost.
- Academic β reentry markers ("as noted above", "as discussed below") plus orphan definites. Coherence demands a second pass.
- Meta β self-referential. "This." has no antecedent by construction; the text explicitly tells the reader to look back.
When to use it
- Editing β find texts that demand re-reading.
- Typography β decide where to place footnotes or asides.
- Academic writing β see which paragraphs force backtracking.
- Translation QA β compare fold across versions of the same text.
- Literary analysis β rank chapters by reading difficulty curve.
When not to use it
- As an absolute count.
fold = 2.3does not mean "read exactly 2.3 times." It means "this text demands more backtracking than a fold-1.5 text and less than a fold-3.0 text." The tiers are calibrated, not derived. - For non-English text. All lexicons are English.
- For code or symbolic text. The model has no notion of identifier binding, scoping, or macro expansion.
Honest limitations
- The weights are hand-tuned. They are plausible, internally consistent, and produce sensible rankings. They are not fit to a real reading study.
- True forward references are not detected. The model finds definite NPs whose head noun first appears later, but it cannot detect cataphora involving pronouns ("When she arrived, Maryβ¦").
- Unique referents are a curated list. "Sun", "moon", "morning", and fifteen others are exempt. "Sky", "sea", "government" are not β the boundary is a judgment call, not a measurement.
- Reentry markers are detected lexically. "As noted" triggers a penalty; "as I have already made clear on multiple occasions" does not.
- No calibration against real reading data. Extending the model to fit eye-tracking or self-reported re-reading is future work.
Reference
Part of the reader-model series. Companion to hv-tempo (pace variation),
hv-core (argument robustness), hv-ttu (total comprehension time), and
hv-locality (feature-map locality).
hv-tempo answers where the reader slows down.
hv-fold answers how many times they come back.
License
Apache-2.0