hv-fold

Minimal reading passes predictor. Given a text, output the minimum number of times it must be read to be fully understood.

The claim in one sentence

hv-tempo predicts where the reader slows down. hv-fold predicts how many times they must come back.

What it produces

For each sentence:

  • load β€” the reference cost the sentence imposes on the reader
  • n_refs β€” definite references in this sentence
  • role β€” demon | reentry | anchor | dependent | neutral
  • reasons β€” which features caused the load

Plus aggregates:

  • fold β€” scalar, β‰₯ 1.0 (1.0 = perfectly linear)
  • fold_tier β€” linear | two-pass | three-pass | recursive
  • reading_plan β€” one-line instruction
  • anchor_sentences β€” sentences later material depends on
  • reentry_points β€” sentences that explicitly demand backtracking

Install

pip install numpy

Actually β€” no dependencies. Pure stdlib.

## Usage

```python
from hv_fold import HVFold

m = HVFold()
report = m.analyze("This. The previous sentence is not the one you "
                   "should attend to. The point is elsewhere.")

print(report.fold)          # 2.5
print(report.fold_tier)     # 'recursive'
print(report.reading_plan)  # 'Read recursively. Earlier material...'

### Just the numbers

```python
m.fold(text)             # the scalar
m.reading_plan(text)     # the plan string
m.reentry_points(text)   # indices where the reader must jump back

CLI

python hv_fold.py                          # run all demos
python hv_fold.py --text "..."             # full report
python hv_fold.py --text "..." --fold      # just the scalar
python hv_fold.py --text "..." --tier      # just the tier
python hv_fold.py --text "..." --plan      # just the plan
python hv_fold.py --text "..." --json      # JSON output
python hv_fold.py --text "..." --exempt-first   # opening-sentence exemption

The model

Per-sentence load accumulates from five sources:

source weight captures
standalone demonstrative 2.0 orphan anaphor ("This." "That.")
definite NP, no antecedent 0.5 "the dog" with no dog introduced
forward reference 0.5 Γ— delay definite NP whose head appears only later
resolved reference 0.1 head appeared earlier; small memory cost
reentry marker 1.0 "above", "previous", "noted", "as discussed"
unique referent (sun, moon, morning, …) 0.0 no cost β€” the referent is universal
fold = 1 + Ξ£(load) / n_sentences

Tiers

fold tier plan
< 1.15 linear Read once, front to back.
< 1.75 two-pass Re-scan for definite references without antecedents.
< 2.50 three-pass Once for gist, once for coherence, once for detail.
β‰₯ 2.50 recursive Earlier material only makes sense after later material.

Benchmarks

text sentences load fold tier
Linear (declarative) 4 0.00 1.000 linear
Narrative (orphan NPs) 3 1.00 1.333 two-pass
Academic (reentry-heavy) 3 3.50 2.167 three-pass
Meta (self-referential) 3 4.50 2.500 recursive

Reading the table:

  • Linear β€” pure declaratives with unique referents. Fold = 1. Read once.
  • Narrative β€” orphan definites ("the room", "the chair") with no prior introduction. Small backtracking cost.
  • Academic β€” reentry markers ("as noted above", "as discussed below") plus orphan definites. Coherence demands a second pass.
  • Meta β€” self-referential. "This." has no antecedent by construction; the text explicitly tells the reader to look back.

When to use it

  • Editing β€” find texts that demand re-reading.
  • Typography β€” decide where to place footnotes or asides.
  • Academic writing β€” see which paragraphs force backtracking.
  • Translation QA β€” compare fold across versions of the same text.
  • Literary analysis β€” rank chapters by reading difficulty curve.

When not to use it

  • As an absolute count. fold = 2.3 does not mean "read exactly 2.3 times." It means "this text demands more backtracking than a fold-1.5 text and less than a fold-3.0 text." The tiers are calibrated, not derived.
  • For non-English text. All lexicons are English.
  • For code or symbolic text. The model has no notion of identifier binding, scoping, or macro expansion.

Honest limitations

  • The weights are hand-tuned. They are plausible, internally consistent, and produce sensible rankings. They are not fit to a real reading study.
  • True forward references are not detected. The model finds definite NPs whose head noun first appears later, but it cannot detect cataphora involving pronouns ("When she arrived, Mary…").
  • Unique referents are a curated list. "Sun", "moon", "morning", and fifteen others are exempt. "Sky", "sea", "government" are not β€” the boundary is a judgment call, not a measurement.
  • Reentry markers are detected lexically. "As noted" triggers a penalty; "as I have already made clear on multiple occasions" does not.
  • No calibration against real reading data. Extending the model to fit eye-tracking or self-reported re-reading is future work.

Reference

Part of the reader-model series. Companion to hv-tempo (pace variation), hv-core (argument robustness), hv-ttu (total comprehension time), and hv-locality (feature-map locality).

hv-tempo answers where the reader slows down. hv-fold answers how many times they come back.

License

Apache-2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support