hv-manifold

The shape of a corpus in style space.

Given a corpus of texts, embed each text into a style vector, then compute the geometry of the resulting point cloud: its intrinsic dimension, its curvature, its connectivity, and its boundary.

The claim in one sentence

Every other hv- model reads a text. hv-manifold reads a corpus: a single geometric signature that describes the shape of an entire collection in style space.

What it produces

A ManifoldReport containing:

  • pr_dim β€” participation-ratio dimension (continuous intrinsic dim)
  • pca_dim β€” eigenvalues > 1 (Kaiser rule)
  • curvature β€” clustered (positive) or sprawling (negative)
  • beta_1_at_connection β€” cycle rank when the corpus first connects
  • n_clusters β€” connected components at the median edge length
  • cluster_sizes β€” size of each cluster
  • per_cluster_dim β€” per-cluster intrinsic dimension
  • boundary_points β€” outliers on the edge of the manifold
  • log_volume β€” log-determinant spread of the covariance

Plus per-text detail: curvature, local density, 2D projection, and component label.

Plus a global style label: bimodal, clustered, connected, uniform-clustered, core-periphery, fragmented, or loosely-clustered.

Install

pip install numpy

Actually β€” no dependencies. Pure stdlib.

## Usage

### Demo

```bash
python hv_manifold.py

Runs a 30-text synthetic corpus in five groups β€” casual, academic,
technical, poetic, repetitive β€” and prints the full geometric report.

### Analyze a directory

```bash
python hv_manifold.py --dir ./corpus
python hv_manifold.py --dir ./corpus --glob "*.txt" --limit 100
python hv_manifold.py --dir ./corpus --json
python hv_manifold.py --dir ./corpus --save-to ./model

### Python

```python
from hv_manifold import HVManifold

m = HVManifold()
report = m.analyze(texts, labels=labels)

print(report.pr_dim)            # 8.474
print(report.pca_dim)           # 5
print(report.curvature_mean)    # -0.400
print(report.n_clusters)        # 2
print(report.style)             # "bimodal"
print(report.summary)

The ten style axes

Each text is embedded into a 10-dimensional style vector, then standardized across the corpus.

axis what it measures
lexical_diversity types / tokens
rare_rate fraction of long, uncommon words
hedge_rate may, might, perhaps, probably
certainty_rate definitely, certainly, proven
negation_rate not, never, without
passive_rate passive constructions per sentence
clause_rate commas + subordinators per word
mean_sentence_length words / sentence
drift 1 βˆ’ Jaccard between first and last sentence
composition_depth max nesting depth (subordinators, brackets)

The manifold is the shape of these 10 vectors after standardization.

The geometry

pr_dim β€” participation-ratio dimension

Continuous intrinsic dimension. Reflects how many eigenvalues carry non-trivial variance:

pr_dim = (Σλ)² / Σλ²

A corpus spread evenly across 10 axes has pr_dim β‰ˆ 10. A corpus collapsing onto 2 axes has pr_dim β‰ˆ 2. The demo corpus reports 8.47.

pca_dim β€” Kaiser rule

Count of eigenvalues greater than 1 after standardization. The demo reports 5.

curvature β€” clustered or sprawling

Per-text local density (mean k-NN distance) versus the median:

curvature_i = (median βˆ’ density_i) / median
  • Positive mean β€” clustered. Texts have close neighbours; the corpus has centres of gravity.
  • Negative mean β€” sprawling. Texts sit on the edge of the manifold; the corpus has no centre.
  • High std β€” mixed. Some regions dense, others sparse.

The demo reports mean = βˆ’0.400, std = 1.045.

beta_1_at_connection β€” cycle rank

Cycle rank of the graph at the moment it first becomes connected, built by union-find over edges sorted by distance. Zero means the corpus connects as a tree. Positive means there are redundant paths.

n_clusters and cluster_sizes

Connected components at the median edge length. A robust threshold that does not require a bandwidth parameter.

The demo reports 2 clusters (sizes 29, 1) β€” one giant component and one outlier.

per_cluster_dim

Intrinsic dimension computed inside each cluster. Cluster 0 has 7.31; cluster 1 has 0.00 (a singleton has no dimension).

boundary_points

Texts whose local density exceeds boundary_factor Γ— median. These are the texts on the edge of the manifold β€” the outliers, the experiments, the ones that do not belong to any cluster.

log_volume

Log-determinant of the covariance eigenvalues. A monotone measure of spread. The demo reports βˆ’0.879.

The style label

A single string summarizing the manifold's shape:

label condition
connected one component, near-zero curvature
uniform-clustered one component, curvature > 0.10
core-periphery one component, curvature < βˆ’0.10
bimodal exactly two components
clustered three to five components
loosely-clustered six to n/4 components
fragmented more than max(6, n/4) components

The demo corpus is bimodal: 29 texts in one mass, one text alone.

Benchmarks

The demo corpus

30 synthetic texts across five groups:

group count character
C β€” casual 6 short, hedged, everyday
A β€” academic 6 long, subordinated, hedged
T β€” technical 6 numeric, dense, clause-heavy
P β€” poetic 6 long sentences, high drift, low clause rate
R β€” repetitive 6 short sentences, high lexical repetition

Reported geometry:

metric value
PR dim 8.474
PCA dim 5
curvature mean βˆ’0.400
curvature std 1.045
clusters @ median 2
cluster sizes 29, 1
local dims 7.31, 0.00
β₁ @ connect 0
connectivity Ξ΅ 5.620
boundary points 8
log-volume βˆ’0.879
style bimodal

Reading the report:

  • PR dim 8.47 β€” the corpus spreads across nearly the full 10-dimensional style space. It is not low-dimensional.
  • PCA dim 5 β€” but only 5 axes carry above-unit variance. Five style directions explain the corpus.
  • Curvature βˆ’0.40 β€” sprawling. The corpus does not have a centre.
  • 2 clusters (29, 1) β€” one text is geometrically isolated at the median edge length. That text is P β€” the poetic outlier the classifier flagged.
  • 8 boundary points β€” the two groups at the edges (C and A) contribute most of them. Casual and academic texts are the extremes of this corpus in style space.
  • Style: bimodal β€” the honest label. This corpus is not one thing.

Why the geometry is not the sum of its parts

A corpus can be high-dimensional (pr_dim large) but low-rank (pca_dim small). A corpus can be clustered (positive curvature) but one single component (n_clusters = 1). The manifold is the joint description of dimension, curvature, connectivity, and boundary. Every single-axis summary collapses one of these. The report is the artifact.

When to use it

  • Corpus diagnostics. "Is my training corpus one thing or five?"
  • Outlier detection. boundary_points names the texts on the edge.
  • Cluster triage. Before running k-means, ask how many clusters the geometry actually supports.
  • Style drift monitoring. Track pr_dim and curvature over time as a corpus grows.
  • Dataset cards. Report the manifold as a fingerprint of the corpus.
  • Retrieval sanity checks. If a corpus is bimodal, one retrieval index may not serve it.

When not to use it

  • For fewer than ~15 texts. The Jacobi eigensolver and k-NN density are unstable on tiny corpora.
  • As a ground-truth clustering. The clusters are at the median edge length β€” a robust default, not the answer.
  • For non-English corpora. The style lexicons are English.
  • For very long texts (whole books). Style vectors saturate; the manifold flattens.
  • For corpora with fewer than 3 features. PR dim and curvature both degenerate.
  • As a replacement for reading. The manifold tells you the shape. It does not tell you what the corpus says.

Honest limitations

  • The style axes are hand-designed. Not learned. They are one reasonable basis, not the only one.
  • The manifold is basis-dependent. Swap the 10 features and the geometry changes. hv-manifold reports the shape in this basis.
  • PCA dim depends on standardization. The Kaiser rule (eigen > 1) assumes unit-variance features. Correlated features can hide here.
  • Curvature is a k-NN density proxy. It is not Riemannian curvature. It is a signal of clustered vs. sprawling, nothing more.
  • β₁ at connection uses union-find only. It is a cycle-rank proxy, not a full persistent-homology computation. Loops that close after the first connection are not counted.
  • The style label is thresholded. bimodal and clustered sit on either side of an integer. Small corpora will flip between them.
  • Boundary points are density outliers, not semantic ones. A text can be a boundary point and still be completely on-topic.
  • No calibration against human judgment. The demo is synthetic. The labels have not been validated against expert corpus annotation.
  • The demo corpus is designed to be separable. Its geometry looks clean because the five groups are genuinely different. Real corpora will report connected and loosely-clustered more often.

Reference

Part of the corpus-model series. Sits above the text-level models:

model reads
hv-tempo one text's pace
hv-forget one text's memory
hv-fold one text's passes
hv-slip one text's slip
hv-wall one text's wall
hv-reader one text's reading profile
hv-manifold a whole corpus's shape

hv-manifold is orthogonal to the reading models. It does not score any single text. It describes how a collection of texts sits together in style space.

License

Apache-2.0

Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support