File size: 2,767 Bytes
6302710
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
# Training a small LLM from scratch — an operating guide

For the next agent session that has to build, train, publish and evaluate a ~100M-parameter language model
on rented, capped, per-week GPU. It is the distilled version of one real project: 106,194,240 params,
999,817,216 tokens, 2×Tesla T4 on Kaggle, storage on the Hugging Face Hub, eight academic benchmarks, and a
public report — plus 54 written-up failure records spanning E-001…E-056 that are the actual source of these rules.

Read in order once; then use `06-checklists.md` as the working page.

| File | What it covers |
|---|---|
| `01-before-you-start.md` | The constitution: what to fix in writing, and why memory files are the real state |
| `02-data.md` | Mix design, tokenisation, sharding, the contamination audit, publishing reproducibly |
| `03-preflight.md` | The evidence ladder: free rehearsal → cheap smoke → measurement probe → run |
| `04-the-run.md` | Checkpoints, resume, session sizing, quota arithmetic, what to do when it breaks |
| `05-publish-and-evaluate.md` | Model card, clean-room load, benchmark protocol, honest reporting |
| `06-checklists.md` | Copy-paste gates for each phase, and the recurring failure patterns |
| `07-platform-notes.md` | Kaggle and HF mechanics measured rather than believed, with the numbers |

## The five sentences that matter most

1. **Every claim about a tool is a hypothesis until you have executed it** — the expensive mistakes in this
   project were all "documented behaviour" that the installed version did not have.
2. **Verify bytes, never listings.** A commit that returned, a file that is listed, and bytes that
   download are three different facts, and only the third authorises deleting your copy.
3. **Design for interruption; it is not an exception.** Every interval must end on a verified, pointed-at
   checkpoint on shared storage, and the only resume test that counts deletes the local disk first.
4. **Free compute is a strategy, not a fallback.** A CPU rehearsal that costs nine minutes buys back a
   five-hour session; run everything that does not need the accelerator there, forever.
5. **Write the report before the results exist.** Frozen choices, pre-registered targets and a contamination
   statement composed after you can see the scores are a different kind of document.

## What "done" means

A public model repo that a stranger loads with `from_pretrained` from an empty cache; a dataset repo whose
manifest, hashes and exclusion masks let someone rebuild that exact mix; a benchmark table where every row
names its metric, its shot count and its split, and can be re-run from published config; and a report that
says what failed. Model quality is not the deliverable — a trustworthy number is.