Parallax 8B lambda: mid-training checkpoints

These are research checkpoints from lambda, a training run of Parallax that is still in progress. Parallax is a mixture-of-experts language model trained on consumer GPUs spread over several countries. The hosts have no direct connections to each other and synchronize over the public internet.

  • Base model. It is pretrained on a mixture of web, document, math, code, knowledge-focused and book text. It is not instruction-tuned, chat-tuned or safety-tuned, and it will continue text rather than follow instructions.
  • Mid-training. Each export is a snapshot of a run that has not finished. Later exports are expected to differ, and quality between exports is not monotonic.
  • Research artifact. It is published so the training can be followed and inspected. It is not intended for production use.

Live training dashboard: parallax.chutes.ai.

Model

Parameters ~7.8B total, ~1.25B active per token
Layers 64: 18 gated delta-rule (GDN2) recurrent, 8 sliding-window attention (window 2048), 6 sparse attention, 32 mixture-of-experts
Experts 4096 routed experts (128 per MoE layer), 12 routed + 1 shared expert per token
Expert weights ternary (-1, 0, +1) with per-row scales; at most two nonzero pairs in every group of eight
Width d_model 1152; ReLU² expert activation
Tokenizer Llama 3 tokenizer (vocabulary padded to 128,384); tied input/output embeddings
Training context 4096 tokens
Output logit scale learnable, capped at 3.0

Training

  • Data: the same data mix as the earlier Parallax runs epsilon and kappa, not FineWeb-Edu. The main phase draws from seven sources totalling about 1.09T tokens: two general pretraining mixtures of web, document, math and code text (about 78% of the tokens), an additional web set, two math sets, a knowledge-focused set and a books set. At about 899.6B tokens the data switches to two further sources for the rest of the run; the learning rate does not change at that point. The run is planned as one pass of about 973B tokens.
  • Optimization: the dense trunk is synchronized with a decoupled DiLoCo scheme over the internet. Routed experts are trained through low-rank adapters on the GPUs that use them, and the adapter updates are folded into full-precision masters that publish new ternary versions. The tech report describes the details.
  • Learning rate: linear warm-up from 0 to 6e-4 over the first 19.66B tokens, then 6e-4 until 400B tokens, 3e-4 until 700B and 1.5e-4 through the end of training. There is no final decay.
  • Batch: 10 gradient-accumulation micro-batches per step (about 8.5M tokens per fleet step) for about the first 100B tokens, then 20 (about 17M tokens per step) after a switch made while the run continued. The fleet is 26 machines with 208 RTX 5090 GPUs.
  • Output logits: the learnable output logit scale is capped at 3.0 in training and inference (see Files and format).

Exports

One folder per export under exports/, named by training tokens (rounded down to whole billions) and fleet step. A new export is added about once an hour while the run continues. Published exports are never modified or removed.

Export Tokens Step Size val_mix (nats/token) 0-shot macro 5-shot macro Exported (UTC)
162B-tokens_step15667 (latest) 162.4B 15667 2.46 GB - - - 2026-10-05 06:45
157B-tokens_step15384 157.9B 15384 2.46 GB - - - 2026-10-05 06:13
155B-tokens_step15179 155.1B 15179 2.46 GB - - - 2026-10-05 05:46
150B-tokens_step14880 150.6B 14880 2.46 GB - - - 2026-10-05 05:12
147B-tokens_step14693 147.8B 14693 2.46 GB - - - 2026-10-05 04:45
143B-tokens_step14118 143.4B 14118 2.47 GB - - - 2026-10-05 04:13
140B-tokens_step13940 140.5B 13940 2.47 GB - - - 2026-10-05 03:44
135B-tokens_step13714 135.8B 13714 2.47 GB - - - 2026-10-05 03:14
132B-tokens_step13524 132.5B 13524 2.47 GB - - - 2026-10-05 02:48
127B-tokens_step13242 127.7B 13242 2.48 GB - - - 2026-10-05 02:12
123B-tokens_step13016 123.6B 13016 2.48 GB - - - 2026-10-05 01:43
119B-tokens_step12790 119.5B 12790 2.48 GB - - - 2026-10-05 01:13
115B-tokens_step12532 115.4B 12532 2.48 GB - - - 2026-10-05 00:38
111B-tokens_step12302 111.4B 12302 2.48 GB - - - 2026-10-05 00:10
107B-tokens_step12081 107.2B 12081 2.48 GB - - - 2026-10-04 23:41
103B-tokens_step11842 103.2B 11842 2.48 GB - - - 2026-10-04 23:14
99B-tokens_step11533 99.2B 11533 2.48 GB - - - 2026-10-04 22:39
95B-tokens_step11160 96.0B 11160 2.47 GB - - - 2026-10-04 22:12
92B-tokens_step10766 92.6B 10766 2.47 GB - - - 2026-10-04 21:43
89B-tokens_step10388 89.3B 10388 2.47 GB - - - 2026-10-04 21:12
85B-tokens_step9986 86.0B 9986 2.47 GB - - - 2026-10-04 20:41
82B-tokens_step9591 82.7B 9591 2.47 GB - - - 2026-10-04 20:09
79B-tokens_step9195 79.3B 9195 2.47 GB - - - 2026-10-04 19:39
76B-tokens_step8835 76.1B 8835 2.48 GB - - - 2026-10-04 19:13
72B-tokens_step8443 72.8B 8443 2.48 GB - - - 2026-10-04 18:40
69B-tokens_step8039 69.5B 8039 2.48 GB - - - 2026-10-04 18:11
66B-tokens_step7652 66.1B 7652 2.48 GB - - - 2026-10-04 17:42
62B-tokens_step7275 62.9B 7275 2.48 GB - - - 2026-10-04 17:10
59B-tokens_step6711 59.6B 6711 2.48 GB - - - 2026-10-04 16:40
56B-tokens_step6444 56.3B 6444 2.48 GB - - - 2026-10-04 16:09
53B-tokens_step6096 53.0B 6096 2.48 GB - - - 2026-10-04 15:40
49B-tokens_step5723 49.8B 5723 2.48 GB - - - 2026-10-04 15:12
46B-tokens_step5320 46.4B 5320 2.48 GB - - - 2026-10-04 14:40
43B-tokens_step4946 43.1B 4946 2.48 GB - - - 2026-10-04 14:10
39B-tokens_step4549 39.8B 4549 2.48 GB - - - 2026-10-04 13:38
36B-tokens_step4166 36.5B 4166 2.48 GB - - - 2026-10-04 13:15
33B-tokens_step3768 33.2B 3768 2.48 GB - - - 2026-10-04 12:40
29B-tokens_step3391 29.8B 3391 2.49 GB - - - 2026-10-04 12:16
26B-tokens_step3002 26.5B 3002 2.49 GB - - - 2026-10-04 11:43
23B-tokens_step2604 23.1B 2604 2.50 GB - - - 2026-10-04 11:15
19B-tokens_step2198 19.9B 2198 2.50 GB - - - 2026-10-04 10:40
16B-tokens_step1817 16.7B 1817 2.50 GB - - - 2026-10-04 10:11
13B-tokens_step1419 13.3B 1419 2.50 GB - - - 2026-10-04 09:42
6B-tokens_step635 6.6B 635 2.50 GB - - - 2026-10-04 08:41
  • val_mix: mean cross-entropy (nats per token, lower is better) on a fixed held-out validation set stratified over the nine sources of the data mix (2048 windows of 4096 tokens), the same set used for kappa. It is not comparable with the FineWeb-Edu validation number published for theta.
  • 0-shot macro: mean over 11 tasks (ARC-Challenge, ARC-Easy, BoolQ, COPA, HellaSwag, LAMBADA, OpenBookQA, PIQA, SciQ, SIQA, WinoGrande). 5-shot macro: the same tasks without LAMBADA (10 tasks). Per task, acc_norm is used for ARC, HellaSwag, OpenBookQA, PIQA and SciQ, and acc for the rest. Values are percentages.
  • The scores come from the project's own scorer, which runs the native ternary experts. They track progress within this run. Compare them with numbers from other evaluation harnesses with care.
  • Benchmark prompts are scored with every run of two or more newlines collapsed to a single newline (including the few-shot separator).
  • A - means the export has not been scored yet. The table fills in as scores arrive.

Latest export

exports/162B-tokens_step15667: 162.4B training tokens, step 15667.

hf download chutesai/parallax-8b-lambda --include "exports/162B-tokens_step15667/*" --local-dir parallax-8b-lambda
cd parallax-8b-lambda/exports/162B-tokens_step15667
tar -xf packed_experts.tar
sha256sum -c --quiet SHA256SUMS   # every file of the original export, byte for byte

Files and format

The files are in Parallax's native compact export format (inference only, no optimizer state). They are byte-identical to the export the training system produced:

File Contents
manifest.json export manifest: tensor inventory, per-file sha256 digests, token clock
model_config.json model configuration
coverage.json, layouts.json tensor coverage and expert frame layouts
indexer.bundle sparse-attention indexer weights
relay_pack/ trunk (non-expert) weights in bf16, with their own manifest
packed_experts.tar the 4096 routed experts (packed_experts/*.t24p, packed ternary codes and scales)
SHA256SUMS sha256 of every file of the original export
export_info.json step, tokens, time, sizes and digests of this export

The only change from the original export is packaging. The 4096 expert files are stored in one uncompressed tar to keep the repository's file count manageable. Extract it and check SHA256SUMS as shown above. Every upload was checked against the training system's own digests before and after it was published.

Logit-scale bound. The model's output logit scale is bounded: the forward pass uses exp(min(logit_scale_log, log 3.0)). The stored trunk tensor is the raw training parameter (it can sit slightly above the bound, e.g. from bf16 rounding), and each export records the bound in manifest.json (logit_scale: max, raw, effective) and in coverage.json (inference_policy.logit_scale_max). A loader must apply the recorded bound; the files themselves are left byte-identical.

Running it

This is a base model only. It is not chat- or instruction-tuned, so it does plain text completion: give it the start of a text and it continues it. It will not follow instructions or hold a conversation.

Standard transformers cannot load this format. Use our llama.cpp fork, https://github.com/chutesai/llama.cpp, which adds a dedicated runtime, llama-parallax, for CPU (x86-64, ARM64) and Apple GPUs (Metal):

git clone https://github.com/chutesai/llama.cpp && cd llama.cpp
cmake -S . -B build -DCMAKE_BUILD_TYPE=Release -DGGML_METAL=ON   # -DGGML_METAL=OFF for CPU only
cmake --build build --target llama-parallax -j 8
python tools/parallax/run.py --binary build/bin/llama-parallax --model parallax-t9.gguf \
  --tokenizer tokenizer.json --backend metal --experts lut9 \
  --prompt 'The capital of France is' --predict 64

The runtime reads GGUF files converted from these exports. Ready-made GGUF files are not published yet; tools/parallax/README.md in the fork covers conversion, options and the tokenizer.

Tech report

The Parallax tech report: https://parallax.chutes.ai/tech-report.pdf. It is AI-generated from the project's measurements and logs, and it is a living document that changes as the run progresses. It covers the predecessor run (kappa); lambda differs as described above.

Limitations

This is an early base model. It can produce incorrect, biased or nonsensical output. It has no alignment or safety tuning, and it is small and far from converged.

License

MIT.

Table updated 2026-10-05 06:54 UTC.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support