Paragraph Boundaries Are Not White Space:Compression Depth as the Signature of Hierarchical Structure
Abstract
Standard positional encodings represent position as a one-dimensional reading-order coordinate, but reading order alone does not determine hierarchical textual structure. We use a hierarchical rotary positional encoding (hRoPE) that represents paragraph, sentence, and token indices as separate channels, hold the token sequence fixed, intervene on the paragraph coordinate p1, and measure cross-paragraph attention with a token-distance-exact estimator. Attention is compressed relative to a token-distance-matched baseline in every corpus, but compression alone is not diagnostic of true structure: an architecturally identical channel with density-matched random labels is compressed too, more shallowly. What distinguishes real structure is the depth of compression, which is greater and corpus-dependent while the control's is not. Comparing eight corpus-only quantities across three constructs (lexical persistence, paragraph length, embedding-based coherence), none fully reproduces the cross-corpus ordering of depth, though embedding-based coherence comes closest. Compression depth, not its location, is the reproducible signature of genuine paragraph structure in our setting.
Community
Does a Transformer actually "see" paragraph structure, or just token distance? We use hierarchical RoPE (separate paragraph / sentence / token channels) and intervene on the paragraph coordinate while holding the token sequence fixed.
Key finding: cross-paragraph attention is compressed in every corpus, but so is a control channel with random labels. What separates real structure from noise is the depth of compression, which is deeper and corpus-dependent for genuine paragraphs.
📝 Blog post: Your LLM Has a Curved Space of Paragraphs
💻 Code: https://github.com/ShuyangenFrance/hrope
Happy to discuss!
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Content-Based Addressing for Long Context (2026)
- REIGN: Refurbished Embeddings with Integrated Guidance Networks for Efficient Context-Length Scaling (2026)
- Messages, Not Tokens: Grounded Coresets for Faithful VLM Compression (2026)
- Directional Contextual Representations for Dependency Relations: Why Cross-Direction Pairing Fails (2026)
- Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models (2026)
- When Is Noise Response Universal? Tokenization as the Hidden Variable in Language Models (2026)
- BindCLIP: One Balanced Coupling For Compositional Vision Language Scoring (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper