hv-mode

Cognitive delivery router. LLM text in, reader-mode-matched output out.

The claim in one sentence

LLM output is storage, not delivery. Users need their cognitive mode served, not the writer's mode. This model takes any LLM response and renders it as the shape the reader can actually use.

Seven modes

mode mime who it serves
text text/plain text-first (the current default)
card_deck text/html mobile-native, casual readers
svg_tree image/svg+xml whole-pattern, structural
svg_strip image/svg+xml procedural, step-by-step
pacman text/html action-first, non-readers
color_field image/svg+xml colors-first, data readers
rhythm image/svg+xml music-first, temporal

Every mode returns a self-contained string β€” SVG, HTML, or plain text. No external assets. No fonts to load. No JavaScript frameworks.

Install

pip install numpy

That's it. `hv-mode` has no other dependencies.

## Usage

### One-shot β€” pick the best mode automatically

```python
from hv_mode import HVMode

m = HVMode()
out = m.respond(
    llm_text="A black hole is a region of spacetime...",
    query="show me a picture of what a black hole is",
)
print(out["mode"])       # e.g. "svg_tree"
print(out["mime"])       # "image/svg+xml"
open("out.svg", "w").write(out["content"])

Content negotiation β€” client declares Accept

out = m.respond(llm_text=text, accept="image/svg+xml")
# -> {"mode": "svg_tree", "content": "<svg>...</svg>", ...}

out = m.respond(llm_text=text, accept="text/html")
# -> {"mode": "card_deck", "content": "<!DOCTYPE html>...", ...}

Force a specific mode

for mode in ["card_deck", "svg_strip", "pacman", "color_field", "rhythm"]:
    content, mime = m.render(m.extract(text), mode)
    ext = "html" if "html" in mime else "svg"
    open(f"out_{mode}.{ext}", "w").write(content)

Interactive mode menu (recommended demo)

html = m.mode_menu(llm_text, query)
open("mode_menu.html", "w").write(html)
# open in a browser: tabs to switch modes, score bars show detection

CLI

# read text from stdin, write best-mode render to stdout
echo "A black hole is..." | python hv_mode.py --query "show me a picture"

# force a mode
python hv_mode.py --text "..." --mode pacman --out out.html

# generate the interactive menu
python hv_mode.py --text "..." --menu --outdir ./out

# no args = run all demos
python hv_mode.py

API

Method Description
extract(text) Text β†’ structured Meaning
detect(query, user_history=None, meaning=None) Query β†’ distribution over 7 modes
best_mode(query, user_history=None, meaning=None) Argmax of detect
render(meaning, mode, title="") Meaning + mode β†’ (content, mime)
respond(text, query, accept, mode_override) One-shot
mode_menu(text, query) Interactive HTML with all 7 variants
observe(mode) Record a user's mode acceptance
save_pretrained(dir) / from_pretrained(dir) Persist config + history

The algorithm

Three stages, no training:

LLM text ──► [1] extract ──► Meaning ──► [2] route ──► mode ──► [3] serialize ──► output
                     β–²                        β–²
                     β”‚                        β”‚
              heuristics:                detection:
              β€’ sentence split           β€’ modality mentions
              β€’ content classify         β€’ register (formal/informal)
              β€’ entity detection         β€’ query length
              β€’ sequence extraction      β€’ user history (accumulated)
              β€’ number + caveat parse    β€’ meaning availability

Stage 1 β€” extract. Split into sentences; classify into one of DEFINITION / PROCEDURE / COMPARISON / DATA / NARRATIVE / CODE / GENERAL; pull entities, numbers + units, caveats, and β€” if procedural β€” an ordered step list.

Entity extraction uses three signals, in order of confidence:

  1. Explicit multi-word phrases (highest): "black hole", "event horizon", "general relativity", "hawking radiation", etc.
  2. Suffix patterns: " ", e.g. "X horizon", "X radiation", "X theory", "X algorithm".
  3. Proper nouns (lowest): capitalised words not at sentence start.

This avoids the two common failure modes: capturing sentence-starters ("The", "Although", "Note") and missing multi-word lowercase concepts.

Stage 2 β€” route. Combine five signals into a probability distribution over the seven modes:

  1. modality keywords in the query
  2. register (formal vs informal)
  3. query length
  4. accumulated user history
  5. mode availability β€” color_field is downweighted when there are no numbers, rhythm and svg_strip when there's no sequence, svg_tree and pacman when there are no entities

The Accept header overrides the auto-detection when present.

Stage 3 β€” serialize. Each mode is an independent serializer that takes Meaning and returns a string. text mode returns the raw input unchanged β€” it's a fallback, not a transformation.

Modes in detail

text β€” the fallback

Returns the original text, byte-for-byte. Nothing is transformed. This is the mode that 100% of LLM users currently receive by default.

card_deck β€” swipeable mobile deck

Each card holds one idea: summary, steps, key concepts, numbers, caveats. Horizontal scroll with CSS scroll-snap. Arrow keys navigate on desktop. Fits one card per screen β€” no vertical scrolling.

svg_tree β€” concept graph as SVG

Nodes are entities from the extractor. Edges are placeholder (complete graph). Title from the first sentence. Renders as a single inline SVG.

svg_strip β€” state sequence (Rubik's-cube format)

Procedures render as a strip of numbered boxes with arrows between them. Each box holds one step, word-wrapped to fit. This is the same format used in speedcubing tutorials, which is the canonical "procedural answer as a picture" shape.

pacman β€” one-screen action scene

An HTML canvas scene with an agent (you), a goal (first entity), obstacles (entities 2–3), threats (entities 4–5), and a dashed path. This is the "chimp-readable" format: everything visible at once, no text required.

color_field β€” numbers as a colored grid

Numeric values render as colored rectangles, ordered by magnitude from cool (low) to warm (high). Only renders when numbers are present.

rhythm β€” hits on a circular necklace

Sequence steps map to evenly-spaced hits on a 12-beat circle. Without a sequence, falls back to the first few numbers. Circular representation of a rhythmic pattern.

Benchmarks

Mode detection on the black hole sample

Query: "show me a picture of what a black hole is"

mode score
svg_tree 0.581
card_deck 0.389
text 0.009
svg_strip 0.005
rhythm 0.005
pacman 0.002
color_field 0.001

color_field is correctly near-zero because the sample text contains no numbers.

Content negotiation

Accept header routed mode
text/plain text
text/html card_deck
image/svg+xml svg_tree
image/* svg_tree
(none) text

Concept extraction on the black hole sample

Extracted entities:

['black hole', 'event horizon', 'hawking radiation',
 'general relativity', 'spacetime']

Five real concepts. The first three come from explicit phrase matches; "general relativity" from the phrase list; "spacetime" from suffix pattern matching.

Why this is different

Every existing "render LLM output better" tool either:

  • adds markdown formatting (still text)
  • generates an image from the whole response (loses structure)
  • wraps in a chat UI (same shape, prettier frame)

hv-mode changes the shape of the answer based on who is reading, without retraining the model. It's a delivery layer, not a capability layer. The LLM keeps producing text; the reader gets a picture, a card deck, a state sequence, or a color field.

When to use it

  • Yes: any LLM product where the reader is not the writer. Any chat interface where users complain the answers are "walls of text." Any mobile context where long text is hostile. Any context where the user's cognitive mode is known or inferable.
  • Maybe: streaming responses β€” the current implementation works on complete text, not token streams.
  • No: tasks where the LLM's raw text is the deliverable (code, legal contracts, poetry).

Reference

Extracted from the XuanJi-ISA exploratory track, "Visual-Whole Reasoning Interfaces" (issue #122) and the fifteen blueprints on human-first LLM delivery that followed it.

The core insight β€” that delivery shape is orthogonal to model capability β€” is the same one behind the "any user who sees a picture of the answer, points at it, and is correct" test. If a user can point at the answer without reading a paragraph, the delivery succeeded.

License

Apache-2.0


---

## πŸ“„ `requirements.txt`

```text
numpy>=1.24

πŸ“„ .gitattributes

*.npz filter=lfs diff=lfs merge=lfs -text
*.bin filter=lfs diff=lfs merge=lfs -text
*.safetensors filter=lfs diff=lfs merge=lfs -text

πŸ“„ LICENSE

Apache License
Version 2.0, January 2004
http://www.apache.org/licenses/

Copyright 2026 Sylv Q (zeechimp)

Licensed under the Apache License, Version 2.0 (the "License");
you may not use this file except in compliance with the License.
You may obtain a copy of the License at

    http://www.apache.org/licenses/LICENSE-2.0

Unless required by applicable law or agreed to in writing, software
distributed under the License is distributed on an "AS IS" BASIS,
WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
See the License for the specific language governing permissions and
limitations under the License.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support