appvoid commited on
Commit
ea8b40b
·
verified ·
1 Parent(s): 127ee04

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +2 -76
README.md CHANGED
@@ -3,80 +3,6 @@ library_name: transformers
3
  pipeline_tag: text-generation
4
  tags: [bet, byte-level, recurrent, looped]
5
  ---
6
- # Cortex — SparkBET-9M
7
 
8
- Target repository: **appvoid/cortex**. The current training model has **9,353,876 parameters**: width 324, FFN 864, one prelude block, six physical recurrent body blocks, one coda block, 6 query heads / 2 KV heads, head dimension 54, rank-16 phase LoRA, two loop-level Hyper-Connection lanes, Deep-Delta body residuals, QK normalization, and continuous phase/stride conditioning. The byte vocabulary has 259 IDs (0–255 plus PAD/BOS/EOS), context is 1,024 IDs, and the maximum recurrent budget is L8.
9
-
10
- ## Run
11
-
12
- Import the notebook into Kaggle (two T4 GPUs) or a Modal notebook (A10/A10G or RTX PRO 6000), enable network access, provide `HF_TOKEN`, then **Run All**. Kaggle Secrets and environment variables are both supported. The same FP16 autocast + GradScaler precision policy is used on every supported CUDA family; larger-memory GPUs spend memory on larger physical microbatches before activation checkpointing is enabled.
13
-
14
- The trainer runs continuously until interrupted. A fresh run warms up once and then keeps a nonzero constant learning rate. Interrupting the training cell requests a clean stop and waits for a completed optimizer update, local full-state save, Hugging Face full-state upload, metrics upload, and inference export. A hard kill can recover from the most recent durable checkpoint.
15
-
16
- All project source remains embedded in readable notebook cells and is materialized by Run All. This includes data streams, the Cortex curriculum, model, trainer, full-state recovery, Hugging Face export, tokenizer/model wrappers, provenance files, and tests. Existing codec utility modules remain packaged for project continuity, but the current training registry contains **no image or audio datasets**.
17
-
18
- ## Active training data
19
-
20
- The active external language sources are:
21
-
22
- - `appvoid/rewrite6`
23
- - Ultra-FineWeb-L3 English multi-style
24
- - Ultra-FineWeb-L3 English QA
25
- - DCLM baseline 1.0
26
- - FineWeb-Edu score 2
27
- - FineMath 4+
28
- - `appvoid/no-prompt-oasst`
29
- - `appvoid/no-prompt-openhermes`
30
- - SmolLM-Corpus Cosmopedia v2
31
- - SmolLM-Corpus FineWeb-Edu-dedup
32
-
33
- The procedural Cortex curriculum remains active as its own weighted source. Dataset revisions and shard lists are pinned in `dataset_manifest.json`; each source keeps deterministic cursor/epoch/rejection state in full checkpoints.
34
-
35
- ## Objective and recurrent compute
36
-
37
- Every successful optimizer update trains the exact L8 trajectory plus one exact auxiliary trajectory, keeping the same objective:
38
-
39
- `CE(L8) + 0.20 * CE(Lr), r in {1,...,7}`.
40
-
41
- The auxiliary depth follows a deterministic, repeating **sustained progressive-data curriculum**, not a rapid 28-update depth rotation. With `aux_stage_base_updates=128` the L8+L1 stage receives 128 committed optimizer updates and batch draws, L8+L2 receives 256, then L3=384, L4=512, L5=640, L6=768, and L7=896; one curriculum spans 3,584 updates before repeating. Every update draws another global mixer batch (which can contain repeated finite-dataset samples after epochs, not guaranteed unique records). Hence L7 receives 7× the optimizer updates and example presentations allocated to L1, and **L8 remains supervised in every update**. Tune only `aux_stage_base_updates` to scale the entire schedule; the checkpoint persists this value and the curriculum origin. Depth is independent of microbatch count, DDP partition and overflow retries, and is reconstructed from the committed checkpoint step on resume. Each exact-budget trajectory receives its own continuous phase/stride coordinates; auxiliary losses are not snapshots from L8. Full BPTT is retained; activation checkpointing remains a GPU-memory fallback. More exposure encourages but does not guarantee monotonic validation performance.
42
-
43
- The exact previous SparkBET notebook fingerprint and exact original pre-schedule fingerprint are explicitly approved for *verified auxiliary-schedule migration*. Migrated checkpoints begin their new data stages at L1 using their committed update as `aux_curriculum_origin_step`, while preserving model, optimizer, scaler, and data cursors. A subsequent resume restores the saved origin and stage budget instead of restarting L1. Other fingerprints are rejected rather than silently restarting. Set `allow_verified_schedule_migration=False` to forbid even these approved migrations. The new objective fingerprint is stored on the next checkpoint. BET2/LeWorldModel experimental code is included, but `BET2_ABLATION_ENABLED=False` by default; standard training does not create or train the sidecar.
44
-
45
- ## Hardware policy
46
-
47
- Run All identifies T4, A10/A10G, and RTX PRO 6000 explicitly. It first probes the largest divisible physical microbatch without activation checkpointing. If none fits with safety headroom, it retries with checkpointing. This keeps model/context/precision fixed while using larger-memory GPUs for larger physical batches. TF32 remains disabled so the numerical precision policy does not silently change when a run moves between supported GPU families.
48
-
49
- ## Checkpoints, resume, and Hugging Face
50
-
51
- Complete local checkpoint directories contain `training.pt`, `metadata.json`, and `COMPLETE`. They preserve model, AdamW, GradScaler, per-rank RNG, exact mixer/data cursors, update counters, and training contract. Local retention keeps three complete checkpoints.
52
-
53
- Hugging Face full-state checkpoints are uploaded to a SparkBET-specific namespace inside **appvoid/cortex** so they cannot be confused with earlier architectures. Hub head retention keeps two complete SparkBET checkpoints; repository history is not rewritten. At startup the trainer compares verified local and Hub candidates newest-first and can recover across sessions. Upload failures preserve local state and are reported rather than deleting a completed checkpoint.
54
-
55
- TensorBoard event files are uploaded independently to `runs/`; the model page exposes them under the repository TensorBoard view. A standalone Transformers export is uploaded periodically and on a clean stop. The export contains SafeTensors weights, exact configuration, byte tokenizer, model card, dataset manifest, and custom model code for `trust_remote_code=True`.
56
-
57
- ## Hugging Face authentication
58
-
59
- `HF_TOKEN` is read from Kaggle Secrets first on Kaggle and otherwise from the environment. It is never printed. The token needs read access to any private training source you use and write access to `appvoid/cortex`.
60
-
61
- ## Inference
62
-
63
- The Hub export supports:
64
-
65
- ```python
66
- from transformers import AutoTokenizer, AutoModelForCausalLM
67
- import torch
68
-
69
- repo = "appvoid/cortex"
70
- tok = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
71
- model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True).cuda().eval()
72
- ids = torch.tensor([[257] + tok.encode("The next step is", add_special_tokens=False)], device="cuda")
73
- with torch.inference_mode(), torch.autocast("cuda", dtype=torch.float16):
74
- result = model.generate(ids, max_new_tokens=200, do_sample=False, use_cache=False)
75
- print(tok.decode(result[0], skip_special_tokens=True))
76
- ```
77
-
78
- The current model does not implement a recurrent KV cache, so generation recomputes the retained context for each new byte.
79
-
80
- ## Validation
81
-
82
- The notebook runs source-integrity checks, architecture fingerprint/count checks, semantic/curriculum tests, data-stream tests, checkpoint recovery tests, Hugging Face export structure tests, and DDP equivalence checks before training. GPU capacity is validated with a real full-context L8+L7 backward/Adam preflight before the first committed update.
 
3
  pipeline_tag: text-generation
4
  tags: [bet, byte-level, recurrent, looped]
5
  ---
6
+ # Cortex
7
 
8
+ Byte-level language model trained on lots of data with lots of loops.