--- license: other license_name: cc-by-4.0-weights-non-commercial-data license_link: https://huggingface.co/datasets/dichspace/darulm language: - ru library_name: transformers pipeline_tag: text-generation datasets: - dichspace/darulm tags: - connectome - reservoir-computing - echo-state-network - fruit-fly - drosophila - malecns - russian - custom_code --- # nanofly-decoder-ru ![MaleCNS soma cloud coloured by this model's neuron states](connectome.png) 43,993 neurons at their measured MaleCNS v1.0 coordinates, coloured by this checkpoint's state at one tick while continuing the prompt *«Сегодня утром»*. Orange excited, blue inhibited, grey at rest. Frontal view; the optic lobes flank the central brain. A Russian language model whose recurrent layer is the measured wiring of a fruit fly. The connectome is a frozen [echo state network](https://en.wikipedia.org/wiki/Echo_state_network) reservoir — **no synapse is trained**. Only the input projection, per-neuron gain/bias/leak, one global scale and the readout learn. > **Non-commercial.** The training data ([DaruLM](https://huggingface.co/datasets/dichspace/darulm)) permits scientific, non-commercial use only. That restriction travels with these weights. > **Unfiltered.** No toxicity or profanity filtering at any stage. It emits Russian obscenity unprompted. Do not put it in front of users without a filter. - **Code:** [github.com/igorktech/nanoFLY](https://github.com/igorktech/nanoFLY) - **English sibling:** [`igorktech/nanofly-decoder-en`](https://huggingface.co/igorktech/nanofly-decoder-en) ## Usage ```python import torch from transformers import AutoTokenizer, AutoModelForCausalLM repo = "igorktech/nanofly-decoder-ru" tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True) model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).eval() model = model.to("cuda" if torch.cuda.is_available() else "cpu") ids = tok("Сегодня утром", return_tensors="pt").input_ids ids = torch.cat([torch.tensor([[model.config.bos_token_id]]), ids], dim=1).to(model.device) out = model.generate(ids, max_new_tokens=80, do_sample=True, top_k=50, temperature=0.7) print(tok.decode(out[0], skip_special_tokens=True)) ``` - **Sample, do not decode greedily** — greedy falls into repetition loops within a sentence or two. - Prepend BOS: every training example started with it. - Beam search and assisted generation are unsupported (stateful model). - ~960 forward passes/s on an RTX 5080, ~10/s on a laptop CPU. ## Architecture | | | |---|---| | connectome | MaleCNS v1.0 central brain — `cb_sensory`, `cb_intrinsic`, `visual_projection`, `descending_neuron`, `ascending_neuron` | | neurons / edges | 49,393 / 9,055,280 signed (623,728 dropped: modulatory or unknown transmitter) | | edge weight | sign of the presynaptic transmitter × synapse count, rows normalised to unit absolute weight. ACh +1; GABA, Glu, His −1; others 0 | | token input | 11,434 sensory-facing neurons, 8-slot delay line (slot *j* gets token *t−j*). No attention, no positional encoding | | held out | the 2,635 ORNs stay out of the token input, so the encoder-decoder variant can start from these weights | | dynamics | `x ← (1−a)·x + a·tanh(ρ·g·(Wx) + u + b)`, 2 ticks per token; `a` learned per neuron (init 0.5), `ρ` learned global (init 1.0 → **5.34**) | | readout | all 49,393 states → `Linear(49393→256)` → `LayerNorm` → `Linear(256→4096)` | | trainable | **17.82M** — readout 13.69M, input projection 2.93M, embedding 1.05M, per-neuron scalars 0.15M | ## Training | | | |---|---| | data | DaruLM — Pikabu, Lenta, Gazeta shards; 178,148 documents / 1,852 held out; 81,634,628 tokens; BPE vocab 4,096 | | mixture | **2 : 1 : 1 by token count**. `--mix` samples per document and the sources differ in length (349 / 446 / 1,516 tokens), so per-document weights are 8.7 : 3.4 : 1 | | objective | next-token cross entropy, truncated BPTT over 32-token windows, state carried across windows | | optimiser | AdamW — body 2e-3 (no decay), readout 5e-4 (decay 0.01), warmup 200 then cosine to 10%, clip 1.0 | | schedule | 2 epochs, 42,081 updates, batch 128 | | hardware | 1 × RTX 5080, 3.51 h at ~13,000 tok/s | ## Evaluation | | val loss | ppl | bits/char | |---|---|---|---| | **this model** | **3.738** | 42.0 | 1.77 | | English sibling, for scale | 1.933 | 6.9 | 0.92 | **Perplexities across different tokenizers are not comparable** — this model's vocabulary is 4× larger and 3.04 characters per token. Bits per character is the fair axis, and there the gap is under 2×, not 6×. The corpora also differ in difficulty: open-domain web Russian against a deliberately closed and repetitive TinyStories. Validation fell 98.1 → 42.0 over 23 evaluations and was still improving at the end; the checkpoint is undertrained. No shuffled-wiring control has been run for this model (the English one has: 1.933 real vs 1.979 degree-matched shuffle). Samples, top-k 50, temperature 0.7, prompt in bold: > **По данным синоптиков**, в городе Мой биологи в регионе было обнаружено в одном городе и блинском городе Уфе. Об этом сообщает пресс-служба столичных регионах страны. > **Вчера вечером я** решил подробно настроить на сайтах: — Чувак, которые я вам не сижу на пикабу сижу, что я хочу поделиться с =) Morphology, short-range agreement and register are learned — the first is recognisably newswire down to the "Об этом сообщает пресс-служба" formula, the second recognisably a Pikabu post. Meaning is not. ## Limitations - 17.8M trainable parameters over 163M token-steps of web Russian. Fluent-looking Russian that does not mean anything. - Greedy decoding degenerates into loops. Sampling is required. - 8-token delay line plus a short leaky recurrent memory; it cannot hold a subject across a sentence. - Unfiltered Pikabu, Lenta and Gazeta: obscenity, the biases of that data, and a mid-2010s news skew. DaruLM flags itself `not-for-all-audiences` and notes its domain splits are noisy. - A `tanh` rate neuron is not a spiking model: no spikes, no synaptic delays, no neuromodulation — modulatory edges are removed outright. - Central brain only; the optic lobes and ventral nerve cord of the 166,700-neuron CNS are absent. - Synapse count is a proxy for strength, and rows are normalised. Neither is measured physiology. ## Credits - **Connectome:** MaleCNS v1.0 — FlyEM / HHMI Janelia, University of Cambridge, MRC LMB, Google Research. CC BY 4.0. The published buffers derive from that release; keep the attribution when redistributing. - **Data:** [DaruLM](https://huggingface.co/datasets/dichspace/darulm) by dichspace, from corpora collected by Ilya Gusev. **Scientific, non-commercial use only** — the same restriction applies to these weights. - **Transmitter signs:** Shiu et al., *Nature* 2024. - **Connectome as reservoir:** Costi, Hadjiivanov, Dold, Hale, Izzo, 2025. - **Prior art:** [`ngxson/fly-llm-hf`](https://huggingface.co/ngxson/fly-llm-hf), whose graph subset this reproduces. The licence tag is `other`, not `cc-by-4.0`: the connectome would allow CC BY, the training data does not permit commercial use, and the stricter term governs. Modeling code Apache-2.0.