hv-temporal-context-2048

A 189 KB hypervector context selector for LLM conversations with a two-clock temporal signal. Runs in 0.2 ms per query on CPU. NumPy only.

Model size: 189 KB (int8 codebook + int8 bank + int8 surprises)

Precision@5: 0.967 on synthetic 6-topic demo (position-conditional)

Precision@5: 0.858 (pure similarity baseline)

Random baseline precision@5: 0.117

Context reduction: 12× (60 turns → 5 turns passed to LLM)

Dependencies: NumPy only


Model Description

A retrieval-based context selector that keeps a bank of conversation turn hypervectors and returns the top-k most relevant turns for a new query. Turns that are temporally distinctive — surprising relative to the running conversation average — are ranked higher when a topic hint is available.

This is not an LLM. It is a retrieval layer. Its output is a set of turn indices to include in the LLM's context window.

Architecture

The model uses three components.

1. Word-level hypervector encoding. Each turn is tokenized by whitespace. Every word maps to a 2048-bit bipolar hypervector from a random codebook. A turn is encoded by summing IDF-weighted word hypervectors and normalizing to unit length. IDF weights are self-supervised: no labels required.

2. Two-clock surprise. Two leaky integrators with different time constants run over the bank where h_t is the hypervector of turn t. S(t) is a "surprise" vector: it is large when recent turns deviate from the long-term average. The default parameters are α_f = 0.30, α_s = 0.90, found by grid search over the synthetic corpus.

3. Position-conditional retrieval. Given a query and a topic hint, the model looks up the position where that topic last appeared in the conversation and uses the surprise vector at that position to rerank the similarity candidates.

Evaluation

Metric Value
Precision@5, position-conditional 0.967
Precision@5, pure similarity 0.858
Random baseline 0.117
Retrieval latency 0.2 ms
Model size 189 KB

Corpus. 60 synthetic conversation turns across 6 topics (weather, coding, cooking, travel, finance, health), 10 turns per topic. 24 held-out queries, 4 per topic.

Caveat. The corpus has clean topic separation. On real conversation logs, both numbers will drop. The gap between them (0.11) is more likely to survive than the absolute values.

Intended Use

  • Context reduction. Replace the full conversation history with the top-k retrieved turns. Reduces token cost by 10× or more.
  • Topic-shift detection. The raw surprise magnitude ||S(t)|| is a scalar signal that peaks at topic boundaries. It can be used directly as a "the user changed the subject" detector.
  • Long-conversation handling. Maintain a bank of 1000+ turns and retrieve only the relevant ones per query.
  • Edge deployment. 189 KB model on any CPU since 2005.

Limitations

  • Position-conditional retrieval needs a topic hint. The 0.967 number assumes the model knows which topic the query belongs to (for example, from an upstream router). Without a hint, the pure-similarity path gives 0.858.
  • Bag of words. Word order is discarded. Word order changes will produce nearly identical encodings.
  • Closed vocabulary. Words not in the training vocabulary are dropped.
  • Synthetic evaluation. Both precision@5 numbers are on synthetic data.
  • The two-clock adds 0.11 on synthetic data. Whether that gap survives on real conversations is untested.

How to Use

from hv_temporal_context import HVTemporalContext

model = HVTemporalContext.load("hv-temporal-context-2048")

# Pure similarity (no topic hint needed)
indices, scores = model.retrieve("what is the forecast for rome", k=5)

# Position-conditional (uses a topic hint)
indices, scores = model.retrieve_temporal(
    "what is the forecast for rome", topic_hint="weather", k=5)

# Topic-shift score for a single turn
surprise = model.surprise_at(position=42)
Downloads last month
173
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results

  • Precision@5 (position-conditional) on Synthetic multi-topic conversation
    self-reported
    0.967
  • Precision@5 (pure similarity) on Synthetic multi-topic conversation
    self-reported
    0.858
  • Random Baseline on Synthetic multi-topic conversation
    self-reported
    0.117