Instructions to use PollardWeights/FlyBrain-Pollard-CNSv1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use PollardWeights/FlyBrain-Pollard-CNSv1 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S # Run inference directly in the terminal: llama cli -hf PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S # Run inference directly in the terminal: llama cli -hf PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S # Run inference directly in the terminal: ./llama-cli -hf PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S # Run inference directly in the terminal: ./build/bin/llama-cli -hf PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S
Use Docker
docker model run hf.co/PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S
- LM Studio
- Jan
- Ollama
How to use PollardWeights/FlyBrain-Pollard-CNSv1 with Ollama:
ollama run hf.co/PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S
- Unsloth Desktop
- Pi
How to use PollardWeights/FlyBrain-Pollard-CNSv1 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use PollardWeights/FlyBrain-Pollard-CNSv1 with Docker Model Runner:
docker model run hf.co/PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S
- Lemonade
How to use PollardWeights/FlyBrain-Pollard-CNSv1 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S
Run and chat with the model
lemonade run user.FlyBrain-Pollard-CNSv1-IQ3_S
List all available models
lemonade list
- Hermes Agent
How to use PollardWeights/FlyBrain-Pollard-CNSv1 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use PollardWeights/FlyBrain-Pollard-CNSv1 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "PollardWeights/FlyBrain-Pollard-CNSv1:IQ3_S" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
FlyBrain-Pollard-CNSv1
A fruit-fly connectome used as a language model's memory. The artifact here is the brain, not a model: 71 MB of connectome, synapse signs, trained adapters and gate. It attaches to a frozen stock Qwen2.5-0.5B-Instruct and is never merged into it.
A transformer's memory is its KV cache: it grows with every token, and past the window the beginning is gone. This does not grow. The state is 8,552 slots x 328 = 11.2 MB, constant whether the document is a thousand tokens or a million.
Measured
A six-letter string is stated once, buried under filler, and asked about far beyond the attention window. The model never sees more than 128 tokens at a time, so it can never see the fact and the question together. Words are generated fresh for every sample and never reused -- there is nothing to memorise, so the only way to score is to store and retrieve.
| FLOOR -- no brain, fact outside the window | 0.0% |
| RECALL -- with the brain | 100.0% |
| CONTROL -- a word the document never contained | 0.0% |
| CONTROL -- brain read a different document | 0.0% |
| live state, at any document length | 11.2 MB |
128 samples through the reference harness, reconfirmed at 64 samples through pollard-brainverify.
Per-token, decoded from the stored bits: ['100%', '100%', '100%', '100%'].
Both controls are the point. A brain that scores on the first row and also on the last two is leaking, not remembering.
Which file should I download?
This repo ships the brain and the backbone it attaches to, quantized by Pollard. The .pt is the
memory; the .gguf is the model. You need both.
| rung | size | what the allocator actually did | KL at that width |
|---|---|---|---|
| Q6_K | 506 MB | 24L @ q6_K | 0.0123 |
| IQ4_XS | 369 MB | 1L @ q5_K, 23L @ iq4_xs | 0.2058 |
| IQ3_S | 338 MB | 22L @ iq3_s, 2L @ iq2_s | 2.4001 |
None of these is a flat preset. Pollard measures each layer's sensitivity against the model's own calibration data and spends bits where they change the output, which is why IQ4_XS keeps one layer at 5 bits and IQ3_S drops two to 2 bits.
Read the KL column as a warning, not a score. It is the divergence from f16 when the WHOLE model is cast to that width -- the measured noise curve, which is what the allocator plans against. A mixed build does better than its own headline number, because the layers that would have cost the most were not cast that low. What the column tells you is where the cliff is: this model is fine at 4-bit and falls apart below 3, and 2-bit is unusable at KL ~14.
That cliff is model-specific and worth knowing: Qwen2.5-0.5B degrades far faster than larger models. The Qwen2-VL-2B in the sibling repo measures KL 0.61 at 3-bit where this one measures 2.40. Small models have less redundancy to spend.
Recommended: Q6_K unless you are tight on memory. IQ4_XS is the value pick. Take IQ3_S only if 338 MB versus 369 MB genuinely matters to you -- it is a real quality step down on a model this size.
Use it
pip install 'pollard-weights[flybrain]'
from pollard_flybrain import FlyBrain, load_backbone
from transformers import AutoTokenizer
M = "Qwen/Qwen2.5-0.5B-Instruct"
tok = AutoTokenizer.from_pretrained(M)
model = load_backbone(M, device="cpu") # frozen, never modified
brain = FlyBrain.load("FlyBrain-Pollard-CNSv1.pt").bind(model, tok)
brain.feed(open("long_document.txt").read()) # any length, 128 tokens at a time
print(brain.recall(" Question: what is the secret word? Answer: The secret word is"))
brain.save_state("session.flystate") # 11.2 MB, and it does not grow
The prompt is part of the experiment. A token is filed under the words immediately before it, so
a query has to reproduce that context. The document says "The secret word is ", so the query ends
with those same words. Ask "what is the secret word?" on its own and a brain that measures 100%
measures 46% -- the memory is intact, the question arrives at the wrong address. Verify with
pollard-brainverify, which has the correct construction built in.
Train one for YOUR model
A brain is fitted to one backbone: the address matrix has that model's hidden size and the token
codes come from its output embedding. bind() refuses a mismatch rather than returning confident
nonsense. So this file works with Qwen2.5-0.5B-Instruct; for anything else, train your own. It is
quick -- 100% exact recall by step 26 from a cold start:
pollard-connectome --list # or --human for the H01 human cortex graph
pollard-flybrain --train 900 --model <your-hf-id> \
--connectome graph.feather --signs signs.npy \
--probes corpus.txt --brain MyModel-FlyBrain.pt
pollard-brainverify --brain MyModel-FlyBrain.pt --model <your-hf-id> --filler corpus.txt
Honest scope
- It remembers; it does not reason. Language and reasoning come from the backbone. 8,552 slots are a memory, not a mind.
- One brain per backbone. They do not transfer between models.
- PyTorch path. The brain injects memory tokens into the residual stream through
inputs_embedsand readshidden_states, so it runs under transformers. It does not attach to a GGUF, EXL3 or MLX build today. - Single-fact recall is solved; multi-fact selection is not. One fact per document is 100%; several facts in one document is ~72-80% and is the open problem.
- The wiring is not doing the work on this task. A degree-preserving shuffle of the connectome -- same neurons, same degrees, same weights, 99.94% of edges rewired -- reaches 100% at the same step. The connectome supplies the topology and sparsity; on exact single-fact recall it is not measurably better than a degree-matched random graph. An earlier version of this card claimed a +9.2 point advantage; that was measured on a different architecture and does not hold here.
Connectome
MaleCNS v1.0 (FlyEM/Janelia), CC-BY. The full volume is 188,778 neurons and 26,028,386 synapses; this uses the associative-memory core -- mushroom body and central complex -- pruned to 8,552 neurons and 300,880 synapses.
- Downloads last month
- 433
3-bit
4-bit
6-bit