hv-context-2048

A 167 KB hypervector context selector for LLM conversations. Runs in 0.182 ms per query on CPU. NumPy only.

Model size: 167 KB (int8 codebook + int8 bank)

Retrieval latency: 0.182 ms per query at 60-turn bank

Precision@5: 0.850 on synthetic 6-topic demo (random 0.117, 7.3ร— improvement)

Recall@5: 0.425

Context reduction: 12ร— (60 turns โ†’ 5 turns passed to LLM)

Dependencies: NumPy only


Model Description

A retrieval-based context selector that keeps a bank of conversation turn hypervectors and returns the top-k most relevant turns for a new query. Given a query and a bank of prior turns, retrieves the turns most likely to be relevant, then passes only those to the LLM instead of the full history.

This is not an LLM. It is a retrieval layer. Its output is a set of turn indices to include in the LLM's context window.

Architecture

  • Encoding. Each turn (user text + assistant text concatenated) is tokenized by whitespace. Every word is assigned a 2048-bit bipolar hypervector from a random codebook. A turn is encoded by summing IDF-weighted word hypervectors and normalizing to unit length.
  • IDF weights (self-supervised). No labels are needed. Weights are computed from the bank itself: a word that appears in every turn gets weight 0; a word that appears in one turn gets weight 1. This is the standard IDF formula, normalized to [0, 1].
  • Bank. Each turn's hypervector is stored as int8 (after normalization and scaling by 127). Bank size for 60 turns at D=2048 is 120 KB.
  • Retrieval. For a new query, encode it the same way, compute cosine similarity against every bank item, and return the top-k by similarity.
  • No training loop. The model is built in one pass over the bank. Time complexity is O(N ร— D) where N is the number of turns.

Evaluation

Metric Value
Precision@5 0.850
Recall@5 0.425
Random baseline precision@5 0.117
Most-recent baseline precision@5 0.167
Ratio vs random 7.3ร—
Retrieval latency 0.182 ms
Model size 167 KB

Corpus. 60 synthetic conversation turns across 6 topics (weather, coding, cooking, travel, finance, health), 10 turns per topic. 24 held-out queries, 4 per topic.

Caveat. The synthetic corpus has clean topic separation. On real conversation logs, precision@5 will likely drop to 0.5โ€“0.65 because turns share vocabulary across topics. The real number requires a real dataset.

Intended Use

  • Context reduction. Replace the full conversation history with the top-k retrieved turns. Reduces token cost by 10ร— or more.
  • Long-conversation handling. Maintain a bank of 1000+ turns and retrieve only the relevant ones per query.
  • RAG pre-filter. Narrow document collections or chunk sets before running full vector search.
  • Edge deployment. 167 KB model on any CPU since 2005. No GPU, no PyTorch, no server.

Limitations

  • Bag of words. Word order is discarded. "what time is it" and "is it time" produce similar encodings.
  • Closed vocabulary. Words not in the training vocabulary are dropped.
  • Synthetic evaluation. Precision@5 of 0.850 is on synthetic data. Real conversations will be worse.
  • Fixed-size bank. Adding turns is O(1) per turn (encode and append), but retrieval cost grows linearly with bank size.
  • No context awareness. The selector treats each turn independently. It does not know that the current turn follows the previous one.

How to Use

from hv_context import HVContext

model = HVContext.load("context_model")

# Retrieve top-5 relevant turns for a query
indices, scores = model.retrieve("what is the forecast for rome", k=5)

# Get the actual turn texts
relevant_turns = [model.texts[i] for i in indices]

# Use them as context for the LLM
context = "\n".join(relevant_turns)
Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Evaluation results

  • Precision@5 on Synthetic multi-topic conversation
    self-reported
    0.850
  • Random Baseline Precision@5 on Synthetic multi-topic conversation
    self-reported
    0.117