iskhare commited on
Commit
c6d35c2
·
verified ·
1 Parent(s): 9d1be78

Back up decoder-aligned joint MsTok step 28000 with FP64 generation evaluation

Browse files
Files changed (24) hide show
  1. iclr-debug-decoder-135b-20260916/decoder-mstok/checkpoint-28000-README.md +15 -0
  2. iclr-debug-decoder-135b-20260916/decoder-mstok/config.yaml +236 -0
  3. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/controller.py +54 -0
  4. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/evaluator.py +187 -0
  5. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/manifest.json +25 -0
  6. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/quality-summary.json +87 -0
  7. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-0.json +0 -0
  8. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-0.samples.json +0 -0
  9. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-1.json +0 -0
  10. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-1.samples.json +0 -0
  11. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-2.json +0 -0
  12. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-2.samples.json +0 -0
  13. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-3.json +0 -0
  14. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-3.samples.json +0 -0
  15. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-4.json +0 -0
  16. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-4.samples.json +0 -0
  17. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/report.md +44 -0
  18. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/summarize-quality.py +52 -0
  19. iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/summary.json +41 -0
  20. iclr-debug-decoder-135b-20260916/decoder-mstok/milestone-iter-28000.pt +3 -0
  21. iclr-debug-decoder-135b-20260916/decoder-mstok/teacher-provenance.json +6 -0
  22. iclr-debug-decoder-135b-20260916/git-provenance.json +3 -0
  23. iclr-debug-decoder-135b-20260916/manifest.json +525 -0
  24. iclr-debug-decoder-135b-20260916/provenance/hf-backup-step-28000-manifest.json +116 -0
iclr-debug-decoder-135b-20260916/decoder-mstok/checkpoint-28000-README.md ADDED
@@ -0,0 +1,15 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Decoder-aligned joint MsTok: step 28,000
2
+
3
+ Full resumable checkpoint from `iclr-debug-decoder-135b-20260916`, trained jointly from scratch on full OWT.
4
+
5
+ - 28,000 optimizer updates; 22,020,096,000 input positions including padding.
6
+ - Context 256; 16 square scales; global batch 3,072; eight H100 workers.
7
+ - Checkpoint contains online codec, generator, EMA teacher, semantic projector, optimizer, LR scheduler, per-rank RNG states, config, and actual token counter.
8
+ - Training source commit: see `git-provenance.json`. Resolved config and immutable run manifest are included.
9
+ - Generation evaluation: 640 samples across five seeds, temperature 1 without truncation, supplied true level-zero code, scored with FP64 GPT-2 Large and TF32 disabled.
10
+ - Mean generation PPL 111.484 +/- 5.392 standard error; mean token entropy 4.381 nats. 65.3% of samples contain Unicode replacement characters. This is an intermediate research checkpoint with substantial generation defects.
11
+ - Entropy >= 4.5 retains 182/640 samples; pooled generation PPL 203.126. See `evaluation/step-28000-fp64/` for full results and saved generations.
12
+
13
+ The full checkpoint is `milestone-iter-28000.pt`; component weights can be exported with `utils.iclr_training.export_training_checkpoint` from the matching source checkout. Local absolute paths in saved configuration record the original run environment and are not portable defaults.
14
+
15
+ See `../provenance/` at the run root for the SHA-256 inventory. The backup does not include training data or credentials. Training continues independently of this snapshot.
iclr-debug-decoder-135b-20260916/decoder-mstok/config.yaml ADDED
@@ -0,0 +1,236 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ task: mstok-next-concept
2
+ experiment: iclr-debug-decoder-135b-20260916-decoder-mstok
3
+ experiment_dir: /home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok
4
+ dataset:
5
+ train_source: /home/ubuntu/data/full_owt/train_gpt2.bin
6
+ validate_source: /home/ubuntu/data/full_owt/valid_gpt2.bin
7
+ pad_token_id: 50257
8
+ pre_tokenizer:
9
+ name: hf
10
+ tokenizer: hf
11
+ model_id: gpt2
12
+ special_tokens:
13
+ eos_token: <|endoftext|>
14
+ pad_token: <|pad|>
15
+ codec:
16
+ n_layers: 6
17
+ context_length: 256
18
+ embed_dim: 384
19
+ in_vocab_size: 50304
20
+ vqvae_vocab_size:
21
+ - 16384
22
+ - 16384
23
+ - 16384
24
+ - 16384
25
+ - 16384
26
+ - 16384
27
+ - 16384
28
+ - 16384
29
+ - 16384
30
+ - 16384
31
+ - 16384
32
+ - 16384
33
+ - 16384
34
+ - 16384
35
+ - 16384
36
+ - 16384
37
+ compression_factor: 1
38
+ dropout: 0.1
39
+ pre_quant_groupnorm: 4
40
+ pre_quant_dropout: 0.2
41
+ vector_quantizer_config:
42
+ clss: multiscale_residual_vector_quantizer
43
+ decay: 0.99
44
+ epsilon: 1.0e-05
45
+ commitment_cost: 0.25
46
+ learned_l1_sampling: true
47
+ learned_all_sampling: false
48
+ quant_resi:
49
+ enabled: true
50
+ ratio: 0.5
51
+ share_mode: 0
52
+ learnable_ratio: false
53
+ levels:
54
+ use_manual_levels: true
55
+ manual_levels:
56
+ - 1
57
+ - 4
58
+ - 9
59
+ - 16
60
+ - 25
61
+ - 36
62
+ - 49
63
+ - 64
64
+ - 81
65
+ - 100
66
+ - 121
67
+ - 144
68
+ - 169
69
+ - 196
70
+ - 225
71
+ - 256
72
+ aux:
73
+ fine_drop:
74
+ prob: 0.5
75
+ min_keep: 1
76
+ generator:
77
+ n_layer: 12
78
+ n_head: 12
79
+ bias: true
80
+ dropout: 0.1
81
+ n_embd: 768
82
+ context_length: 256
83
+ vocab_size:
84
+ - 16384
85
+ - 16384
86
+ - 16384
87
+ - 16384
88
+ - 16384
89
+ - 16384
90
+ - 16384
91
+ - 16384
92
+ - 16384
93
+ - 16384
94
+ - 16384
95
+ - 16384
96
+ - 16384
97
+ - 16384
98
+ - 16384
99
+ attn_pattern: block_diagonal
100
+ use_positional_encoding: true
101
+ use_level_encoding: true
102
+ prefix_len: 0
103
+ use_rope: true
104
+ rope_base: 10000
105
+ use_qk_norm: true
106
+ post_upsample_conv:
107
+ enabled: true
108
+ kernel_size: 3
109
+ use_relu2: true
110
+ use_flex_attention: true
111
+ shared_output_head: true
112
+ shared_head_per_level_bias: false
113
+ shared_head_per_level_scale: false
114
+ shared_head_adapter_rank: 0
115
+ objectives:
116
+ codec_weight: 1.0
117
+ ncp_weight: 1.0
118
+ soft_assignment_temperature: 1.0
119
+ prediction_temperature: 1.0
120
+ mstok_weight: 0.25
121
+ residual_weight: 1.0
122
+ reconstruction_weight: 1.0
123
+ teacher_ema_decay: 0.999
124
+ teacher_ema_warmup_steps: 1000
125
+ mstok_warmup_steps: 500
126
+ optimization:
127
+ codec_lr: 0.001
128
+ codec_min_lr: 0.0001
129
+ codec_warmup_iters: 0
130
+ codec_lr_decay_iters: 171662
131
+ generator_lr: 0.0005
132
+ generator_min_lr: 1.0e-05
133
+ generator_warmup_iters: 300
134
+ generator_lr_decay_iters: 171662
135
+ beta_1: 0.9
136
+ codec_beta_2: 0.99
137
+ generator_beta_2: 0.99
138
+ weight_decay: 0.1
139
+ max_grad_norm: 1.0
140
+ level_loss_alpha: 1.0
141
+ grad_accumulation_steps: 12
142
+ corruption:
143
+ mode: per_level
144
+ per_level_probs:
145
+ - 0.85
146
+ - 0.8321428571
147
+ - 0.8142857143
148
+ - 0.7964285714
149
+ - 0.7785714286
150
+ - 0.7607142857
151
+ - 0.7428571429
152
+ - 0.725
153
+ - 0.7071428571
154
+ - 0.6892857143
155
+ - 0.6714285714
156
+ - 0.6535714286
157
+ - 0.6357142857
158
+ - 0.6178571429
159
+ - 0.6
160
+ skip_level0: true
161
+ training:
162
+ log_interval: 10
163
+ eval_interval: 1000
164
+ checkpoint_interval: 1000
165
+ eval_batch_size: 4
166
+ val_iters: 25
167
+ keep_last: 3
168
+ seed: 55
169
+ codec_initialization_seed: 42
170
+ generator_initialization_seed: 55
171
+ total_iters: 171662
172
+ batch_size: 32
173
+ expected_world_size: 8
174
+ expected_global_batch_size: 3072
175
+ resume_checkpoint: null
176
+ milestone_steps:
177
+ - 1272
178
+ - 6358
179
+ - 17167
180
+ - 34333
181
+ - 68665
182
+ - 102997
183
+ - 137330
184
+ - 171662
185
+ generation_eval_steps:
186
+ - 17167
187
+ - 34333
188
+ - 68665
189
+ - 102997
190
+ - 137330
191
+ - 171662
192
+ torch_compile:
193
+ enable: true
194
+ scope: loss
195
+ mode: default
196
+ dynamic: false
197
+ fullgraph: true
198
+ backend: inductor
199
+ mixed_precision:
200
+ enable: true
201
+ wandb:
202
+ entity: mstok
203
+ project: iclr-debug
204
+ group: full-owt-ctx256-16sq-135b-v1
205
+ enable: true
206
+ id: iclr-debug-decoder-135b-20260916-decoder-mstok
207
+ resume: allow
208
+ gradients_and_params:
209
+ enable: false
210
+ log: all
211
+ log_freq: 1000
212
+ export:
213
+ vqvae_config_template: /home/ubuntu/mstok-runs/iclr-debug-joint-20260916/config/repro-ctx256/vqvae.yaml
214
+ ncp_config_template: /home/ubuntu/mstok-runs/iclr-debug-joint-20260916/config/repro-ctx256/ncp-sharedhead.yaml
215
+ semantic:
216
+ enabled: true
217
+ mode: eostok
218
+ weight: 0.5
219
+ warmup_steps: 500
220
+ teacher_id: FacebookAI/roberta-base
221
+ teacher_revision: e2da8e2f811d1448a5b465c236feacd80ffbac7b
222
+ tokenizer_revision: 607a30d783dfa663caf39e06633721c8d4cfcd7e
223
+ teacher_dim: 768
224
+ feature_layer: 6
225
+ projector_dim: 2048
226
+ projector_seed: 56
227
+ alignment_site: decoder
228
+ iclr_debug:
229
+ stage: decoder-mstok
230
+ target_positions: 135000000000
231
+ lr_horizon_positions: 135000000000
232
+ stop_step: 171662
233
+ pilot: false
234
+ tokenizer_checkpoint: null
235
+ tokenizer_steps: null
236
+ budget_unit: input positions including padding; log non-padding tokens separately
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/controller.py ADDED
@@ -0,0 +1,54 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import datetime
2
+ import json
3
+ import math
4
+ import os
5
+ from pathlib import Path
6
+ import subprocess
7
+ import sys
8
+
9
+ REPO = Path('/home/ubuntu/mstok-runs/iclr-debug-joint-20260916')
10
+ STAGE = Path('/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok')
11
+ OUT = STAGE / 'evaluation/step-28000-fp64'
12
+ EXPORT = STAGE / 'exports/step-28000'
13
+ ENV = dict(os.environ, CUDA_VISIBLE_DEVICES='0', OMP_NUM_THREADS='1',
14
+ TOKENIZERS_PARALLELISM='false', PYTORCH_ALLOC_CONF='expandable_segments:True')
15
+
16
+ def status(state, **extra):
17
+ payload = dict(state=state, updated_utc=datetime.datetime.now(datetime.timezone.utc).isoformat(),
18
+ controller_pid=os.getpid(), **extra)
19
+ temporary = OUT / 'manual-eval-status.tmp'
20
+ temporary.write_text(json.dumps(payload, indent=2) + '\n')
21
+ temporary.replace(OUT / 'manual-eval-status.json')
22
+ print(json.dumps(payload), flush=True)
23
+
24
+ try:
25
+ for seed in range(5):
26
+ result = OUT / f'random-seed-{seed}.json'
27
+ if not result.exists():
28
+ status('evaluating', seed=seed)
29
+ code = "import runpy,torch; torch.cuda.set_per_process_memory_fraction(0.15); runpy.run_path('/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/evaluator.py',run_name='__main__')"
30
+ command = [sys.executable, '-u', '-c', code, '--ckpt', str(EXPORT/'ncp.pt'),
31
+ '--codec', str(EXPORT/'vqvae.pt'), '--validation-source',
32
+ '/home/ubuntu/data/full_owt/valid_gpt2.bin', '--ref-model', 'gpt2-large',
33
+ '--ref-dtype', 'float64', '--seed', str(seed), '--n-samples', '128',
34
+ '--batch-size', '4', '--temperature', '1', '--top-k', '0', '--top-p', '1',
35
+ '--save-all-generations', '--out', str(result)]
36
+ with result.with_suffix('.log').open('a') as log:
37
+ subprocess.run(command, cwd=REPO, env=ENV, stdout=log, stderr=subprocess.STDOUT,
38
+ check=True, timeout=1800)
39
+ row = json.loads(result.read_text())
40
+ assert row['checkpoint_step'] == 28000 and row['seed'] == seed
41
+ assert row['reference_dtype'] == 'float64' and not row['tf32_allowed']
42
+ assert row['protocol']['n_levels'] == 16 and row['protocol']['n_provided_levels'] == 1
43
+ assert row['requested_samples'] == row['scored_samples'] == 128
44
+ assert row['scored_tokens'] > 0
45
+ assert math.isfinite(row['mean_ppl']) and math.isfinite(row['mean_entropy_nats'])
46
+ status('seed-complete', seed=seed, mean_ppl=row['mean_ppl'], entropy=row['mean_entropy_nats'])
47
+ with (OUT/'summarize.log').open('w') as log:
48
+ subprocess.run([sys.executable, 'evaluation/summarize_gen_ppl.py', '--input-glob',
49
+ str(OUT/'random-seed-[0-4].json'), '--out', str(OUT/'summary.json')],
50
+ cwd=REPO, env=ENV, stdout=log, stderr=subprocess.STDOUT, check=True)
51
+ status('complete', summary=str(OUT/'summary.json'))
52
+ except BaseException as exc:
53
+ status('failed', error=str(exc))
54
+ raise
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/evaluator.py ADDED
@@ -0,0 +1,187 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ """Evaluate random-sampling generation PPL from a completed NCP checkpoint."""
2
+
3
+ import argparse
4
+ import json
5
+ import math
6
+ import sys
7
+ from collections import Counter
8
+ from pathlib import Path
9
+
10
+ REPO_ROOT = Path("/home/ubuntu/mstok-runs/iclr-debug-joint-20260916")
11
+ if str(REPO_ROOT) not in sys.path:
12
+ sys.path.insert(0, str(REPO_ROOT))
13
+
14
+ import hydra
15
+ import torch
16
+ from omegaconf import OmegaConf
17
+ from transformers import AutoModelForCausalLM, AutoTokenizer
18
+
19
+ from data.serialized_dataset import SerializedDataset
20
+ from evaluation.gen_ppl import generation_perplexity
21
+ from utils.misc import set_manual_seed
22
+ from utils.registry import trainers
23
+
24
+
25
+ def main():
26
+ parser = argparse.ArgumentParser()
27
+ parser.add_argument("--ckpt", required=True)
28
+ parser.add_argument("--codec")
29
+ parser.add_argument("--validation-source")
30
+ parser.add_argument("--ref-model", default="gpt2-large")
31
+ parser.add_argument("--ref-dtype", choices=("bfloat16", "float32", "float64"), default="bfloat16",
32
+ help="Legacy default retained; new ICLR runs explicitly use float32")
33
+ parser.add_argument("--seed", type=int, default=0)
34
+ parser.add_argument("--n-samples", type=int, default=128)
35
+ parser.add_argument("--batch-size", type=int, default=8)
36
+ parser.add_argument("--temperature", type=float, default=1.0)
37
+ parser.add_argument("--top-k", type=int, default=0)
38
+ parser.add_argument("--top-p", type=float, default=1.0)
39
+ parser.add_argument("--out", required=True)
40
+ parser.add_argument("--save-all-generations", action="store_true")
41
+ args = parser.parse_args()
42
+
43
+ device = "cuda"
44
+ checkpoint = torch.load(args.ckpt, map_location="cpu", weights_only=False)
45
+ cfg = checkpoint["args"]
46
+ OmegaConf.set_struct(cfg, False)
47
+ if args.codec:
48
+ cfg.tokenizer.vqvae.checkpoint_path = args.codec
49
+ if args.validation_source:
50
+ cfg.dataset.validate_source = args.validation_source
51
+ cfg.optimization.token_corruption_prob = 0.0
52
+ if getattr(cfg.optimization, "stochastic_encode", None) is not None:
53
+ cfg.optimization.stochastic_encode.enabled = False
54
+ else:
55
+ cfg.optimization.stochastic_encode = {"enabled": False}
56
+
57
+ set_manual_seed(args.seed)
58
+ # Seed setup enables TF32 by default, so enforce evaluation precision after it.
59
+ if args.ref_dtype in ("float32", "float64"):
60
+ torch.backends.cuda.matmul.allow_tf32 = False
61
+ torch.backends.cudnn.allow_tf32 = False
62
+ Trainer = hydra.utils.get_class(trainers[cfg.task])
63
+ trainer = Trainer(cfg, device)
64
+ trainer.model.load_state_dict(checkpoint["model"])
65
+ model = trainer.model.to(device).eval()
66
+ tokenizer = trainer.tokenizer
67
+ levels = list(model.levels)
68
+ prefix_len = int(cfg.model.prefix_len)
69
+
70
+ dataset = SerializedDataset(
71
+ cfg.dataset.validate_source,
72
+ int(cfg.model.context_length),
73
+ cfg.task,
74
+ tokenizer.pad_token_id,
75
+ )
76
+ full_sequences, _ = dataset.get_rand_batch(args.n_samples)
77
+ prefix_sequences = full_sequences[:, :prefix_len]
78
+ target_sequences = full_sequences[:, prefix_len:]
79
+
80
+ prompts = []
81
+ generations = []
82
+ with torch.no_grad():
83
+ for start in range(0, args.n_samples, args.batch_size):
84
+ end = min(start + args.batch_size, args.n_samples)
85
+ prefix = prefix_sequences[start:end].to(device)
86
+ target_idx = tokenizer.encode_pre_tokenized_idx(
87
+ target_sequences[start:end].to(device)
88
+ )
89
+ level0 = target_idx[:, : levels[0]]
90
+ generated_idx = model.generate(
91
+ level0,
92
+ prefix,
93
+ temperature=args.temperature,
94
+ top_k=args.top_k,
95
+ top_p=args.top_p,
96
+ )
97
+ reconstructions = tokenizer.decode_multiscales(
98
+ generated_idx, skip_special_tokens=True
99
+ )
100
+ generations.extend(
101
+ reconstructions[row][-1] for row in range(len(reconstructions))
102
+ )
103
+ prompts.extend(
104
+ tokenizer.pre_tokenizer.decode_batch(
105
+ prefix.cpu().tolist(), skip_special_tokens=True
106
+ )
107
+ )
108
+
109
+ if args.save_all_generations:
110
+ # Preserve every generated sample even if scoring fails or some are empty.
111
+ samples_path = Path(args.out).with_suffix(".samples.json")
112
+ samples_path.parent.mkdir(parents=True, exist_ok=True)
113
+ samples_path.write_text(json.dumps(dict(seed=args.seed, checkpoint=args.ckpt,
114
+ prompts=prompts, generations=generations), indent=2) + "\n")
115
+ kept = [(prompt, text) for prompt, text in zip(prompts, generations) if text.strip()]
116
+ prompts = [prompt for prompt, _ in kept]
117
+ generations = [text for _, text in kept]
118
+
119
+ entropies = []
120
+ for text in generations:
121
+ ids = tokenizer.pre_tokenizer.encode(text)
122
+ if ids:
123
+ counts = Counter(ids)
124
+ entropies.append(
125
+ -sum((count / len(ids)) * math.log(count / len(ids)) for count in counts.values())
126
+ )
127
+
128
+ del model, trainer, tokenizer
129
+ torch.cuda.empty_cache()
130
+
131
+ ref_tokenizer = AutoTokenizer.from_pretrained(args.ref_model)
132
+ ref_model = AutoModelForCausalLM.from_pretrained(
133
+ args.ref_model, torch_dtype=getattr(torch, args.ref_dtype),
134
+ **({"attn_implementation": "eager"} if args.ref_dtype == "float64" else {})
135
+ ).to(device).eval()
136
+ assert all(p.dtype == getattr(torch, args.ref_dtype) for p in ref_model.parameters())
137
+ result = generation_perplexity(
138
+ ref_model,
139
+ ref_tokenizer,
140
+ prompts,
141
+ generations,
142
+ score_only_generated=True,
143
+ batch_size=args.batch_size,
144
+ device=device,
145
+ )
146
+
147
+ assert result.nll_sum.dtype == getattr(torch, args.ref_dtype)
148
+ assert bool(torch.isfinite(result.ppl).all()) and bool((result.token_count > 0).all())
149
+ payload = {
150
+ "checkpoint": str(Path(args.ckpt).resolve()),
151
+ "checkpoint_step": int(checkpoint["step"]),
152
+ "codec": str(Path(cfg.tokenizer.vqvae.checkpoint_path).resolve()),
153
+ "reference_model": args.ref_model,
154
+ "reference_dtype": args.ref_dtype,
155
+ "reference_attention_implementation": ref_model.config._attn_implementation,
156
+ "per_sample_nll_sum": result.nll_sum.tolist(),
157
+ "per_sample_token_count": result.token_count.tolist(),
158
+ "per_sample_ppl": result.ppl.tolist(),
159
+ "per_sample_entropy_nats": entropies,
160
+ "tf32_allowed": torch.backends.cuda.matmul.allow_tf32,
161
+ "protocol": {
162
+ "sampling": "random" if args.top_k == 0 and args.top_p == 1.0 else "truncated",
163
+ "temperature": args.temperature,
164
+ "top_k": args.top_k,
165
+ "top_p": args.top_p,
166
+ "n_provided_levels": 1,
167
+ "n_levels": len(levels),
168
+ "document_aware_validation_sampling": dataset.doc_offsets is not None,
169
+ },
170
+ "seed": args.seed,
171
+ "requested_samples": args.n_samples,
172
+ "scored_samples": len(generations),
173
+ "skipped_empty_samples": args.n_samples - len(generations),
174
+ "scored_tokens": int(result.token_count.sum()),
175
+ "mean_ppl": float(result.mean_ppl),
176
+ "median_ppl": float(result.ppl.median()),
177
+ "mean_entropy_nats": sum(entropies) / len(entropies),
178
+ "sample_generations": generations if args.save_all_generations else generations[:8],
179
+ }
180
+ output_path = Path(args.out)
181
+ output_path.parent.mkdir(parents=True, exist_ok=True)
182
+ output_path.write_text(json.dumps(payload, indent=2) + "\n")
183
+ print(json.dumps({key: value for key, value in payload.items() if key != "sample_generations"}, indent=2))
184
+
185
+
186
+ if __name__ == "__main__":
187
+ main()
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/manifest.json ADDED
@@ -0,0 +1,25 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "checkpoint_step": 28000,
3
+ "input_positions": 22020096000,
4
+ "source_checkpoint": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/checkpoint-iter-28000.pt",
5
+ "retained_checkpoint": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/source-checkpoint.pt",
6
+ "checkpoint_sha256": "d9eb43b22050308d6a75176a2978dadc8b0eeba38d905afa73521385e4a373c0",
7
+ "evaluator_sha256": "e2c87f66f1001853f58b9e78a646fc4e089fc894d05fa6c215a8baf493a84055",
8
+ "checkout": "/home/ubuntu/mstok-runs/iclr-debug-joint-20260916",
9
+ "reference_dtype": "float64",
10
+ "reference_attention": "eager",
11
+ "tf32_allowed": false,
12
+ "memory_fraction": 0.15,
13
+ "gpu": 0,
14
+ "batch_size": 4,
15
+ "seeds": [
16
+ 0,
17
+ 1,
18
+ 2,
19
+ 3,
20
+ 4
21
+ ],
22
+ "samples_per_seed": 128,
23
+ "sampling": "temperature-1-untruncated-supplied-true-level0",
24
+ "concurrent_with_training": true
25
+ }
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/quality-summary.json ADDED
@@ -0,0 +1,87 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "checkpoint_step": 28000,
3
+ "reference_dtype": "float64",
4
+ "tf32_allowed": false,
5
+ "unfiltered": {
6
+ "samples": 640,
7
+ "scored_tokens": 160778,
8
+ "pooled_token_weighted_ppl": 110.93052440167119,
9
+ "mean_entropy_nats": 4.381471981070202,
10
+ "fraction_samples_with_replacement_character": 0.653125,
11
+ "adjacent_repeated_word_fraction": 0.03943009422883846
12
+ },
13
+ "entropy_ge_4p5": {
14
+ "samples": 182,
15
+ "scored_tokens": 45589,
16
+ "pooled_token_weighted_ppl": 203.12568343002383,
17
+ "mean_entropy_nats": 4.5968029107321735,
18
+ "fraction_samples_with_replacement_character": 0.6208791208791209,
19
+ "adjacent_repeated_word_fraction": 0.03950939465147932
20
+ },
21
+ "acceptance_rate": 0.284375,
22
+ "comparison_step17167": {
23
+ "checkpoint_step": 17167,
24
+ "reference_dtype": "float64",
25
+ "threshold": 4.5,
26
+ "entropy_definition": "Empirical GPT-2 unigram entropy in nats within each generated sample",
27
+ "selection": "Filter existing 640 generations; no new generation; all accepted samples pooled",
28
+ "accepted": {
29
+ "samples": 115,
30
+ "scored_tokens": 28860,
31
+ "pooled_token_weighted_ppl": 189.63357736496337,
32
+ "mean_entropy_nats": 4.57744085314838,
33
+ "median_sample_ppl": 171.1510152167491,
34
+ "adjacent_repeated_word_fraction": 0.04295705405426942,
35
+ "samples_with_replacement_character": 63,
36
+ "fraction_samples_with_replacement_character": 0.5478260869565217
37
+ },
38
+ "unfiltered": {
39
+ "samples": 640,
40
+ "scored_tokens": 160898,
41
+ "pooled_token_weighted_ppl": 109.72590001293689,
42
+ "mean_entropy_nats": 4.3665410349745,
43
+ "median_sample_ppl": 99.25832710923365,
44
+ "adjacent_repeated_word_fraction": 0.04761292563004243,
45
+ "samples_with_replacement_character": 431,
46
+ "fraction_samples_with_replacement_character": 0.6734375
47
+ },
48
+ "acceptance_rate": 0.1796875,
49
+ "per_seed": [
50
+ {
51
+ "seed": 0,
52
+ "accepted": 26,
53
+ "proposed": 128,
54
+ "ppl": 198.79796628571316
55
+ },
56
+ {
57
+ "seed": 1,
58
+ "accepted": 21,
59
+ "proposed": 128,
60
+ "ppl": 171.26448994329286
61
+ },
62
+ {
63
+ "seed": 2,
64
+ "accepted": 25,
65
+ "proposed": 128,
66
+ "ppl": 201.94449304567118
67
+ },
68
+ {
69
+ "seed": 3,
70
+ "accepted": 21,
71
+ "proposed": 128,
72
+ "ppl": 201.29939871172186
73
+ },
74
+ {
75
+ "seed": 4,
76
+ "accepted": 22,
77
+ "proposed": 128,
78
+ "ppl": 174.01773580284262
79
+ }
80
+ ],
81
+ "mean_of_seed_ppls": 189.46481675784833,
82
+ "ppl_definition": "exp(sum saved FP64 reference NLL / sum scored tokens)",
83
+ "conditioning": "supplied true level-zero code",
84
+ "fp64_source": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-17167-fp64",
85
+ "entropy_source": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-17167-diversity-audit/step_17167-samples.json"
86
+ }
87
+ }
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-0.json ADDED
The diff for this file is too large to render. See raw diff
 
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-0.samples.json ADDED
The diff for this file is too large to render. See raw diff
 
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-1.json ADDED
The diff for this file is too large to render. See raw diff
 
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-1.samples.json ADDED
The diff for this file is too large to render. See raw diff
 
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-2.json ADDED
The diff for this file is too large to render. See raw diff
 
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-2.samples.json ADDED
The diff for this file is too large to render. See raw diff
 
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-3.json ADDED
The diff for this file is too large to render. See raw diff
 
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-3.samples.json ADDED
The diff for this file is too large to render. See raw diff
 
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-4.json ADDED
The diff for this file is too large to render. See raw diff
 
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-4.samples.json ADDED
The diff for this file is too large to render. See raw diff
 
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/report.md ADDED
@@ -0,0 +1,44 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ # Step 28,000 generation evaluation
2
+
3
+ FP64 GPT-2 Large generation PPL: 111.484 +/- 5.392 standard error across five seeds.
4
+
5
+ 640 generations, temperature 1 without truncation, supplied true level-zero code. TF32 disabled.
6
+
7
+ ## Quality measures
8
+
9
+ | Metric | All samples | Entropy >= 4.5 |
10
+ |---|---:|---:|
11
+ | Samples | 640.000 | 182.000 |
12
+ | Pooled token-weighted PPL | 110.931 | 203.126 |
13
+ | Mean entropy (nats) | 4.381 | 4.597 |
14
+ | Samples with replacement character (%) | 65.312 | 62.088 |
15
+ | Adjacent repeated words (%) | 3.943 | 3.951 |
16
+
17
+ ## First three seed-0 samples
18
+
19
+ Samples shown in original order, without selecting for quality.
20
+
21
+ ### Seed 0, index 0 — PPL 126.29, entropy 3.747
22
+
23
+ 64 - of. 2017 28 - February. - 13 (10 -: 8 2019 - - 5 July 2 1 2015 - - 5 5 - - 15 1015 2009 - 528 - 29 29 2014 - September 11 8. 14 -1141. 10. March 2016 - 9x0 2014 January - 6 2018 10 25 May07 - 24E 2017 2 35 - 61 pm 2019 - 4 -10 - February 7 -. June 2017 - 2017 8 8 412. 2017 - the 5 February, 2013 - 3.8, 5, - 3 2 7 - 2017 16 3.10 5 - April 2 -.. 660 - February 4an 2015. - 7 17 17 -10 - 12, July. - 8 C 8 2016 - 5 18a 9 8 December 1025 - September 1023,th the - November October17 October 2016 4.7 - the September -12 - May 14 --30 miles - 17 March - - March March 7 4 the17 6 416. 12 - December - 3 - 3 3 2015 -. -16 16 4 16 2016 -20 to to - 4 4 2017 - - the - 1201 - June 8 - 2017 6 other06 2018 June 12 AM 19 19 19 2015- 102 - 29
24
+
25
+ ### Seed 0, index 1 — PPL 119.24, entropy 4.375
26
+
27
+ K about that.
28
+
29
+ That’s true, they’re not wrong, that they get. But that’s the issue to itself.”
30
+
31
+ We need you to this reason, Dallas quarterback. Ryan Smith second shot, and then that that means and with to hindsight success.
32
+
33
+ HURWRY: All all. There’s no other reason to say his reason reason are probably thinks he should be be on the issue.
34
+
35
+ “If think this the run have guy once, a guy still still throw it stuff out there it’s just just going to be be the challenge I’m running or or I� ave running running on that day — if that Portland�s case, its�s true. That used�s not the reason for it because things I do not think. But it’s not that who everything, I can’t have I not not to stop stop or get better a real challenge.”
36
+
37
+ But then think this is, guess that we a got the offense, as whole game can do.
38
+
39
+
40
+ STHY? Well, it would be mean, if you don� MIDt, be game game. You I
41
+
42
+ ### Seed 0, index 2 — PPL 97.97, entropy 4.454
43
+
44
+ to take me very different time; but at the same time else. When I older parents would get to play out, all every time, I do that, something my life is much better than any else, anything around it, gets gets crazy for it, now it's now. There's the moment. If you just come out. It works out. If just did it, and it's always the main thing. want.. ItIt like naming, "I'd always thank for me." "�That's, 'Oh shit. How Can to it, right now is when you came out right when I was, then you really really come superb out. That could not get things great. So I was allowed to get for the work out. — "I've been in a but we was have a a sort of a done". Yeah, I just felt like something that because I was this that I got to get back. You could last was to see what I was getting in the process. So I'm not okay. At the point, want want to know that you want to let people do you, it's not obvious that you find, something else heal something is a negotiation. It's a. cool. Every person will get trust who someone
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/summarize-quality.py ADDED
@@ -0,0 +1,52 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ import json
2
+ import math
3
+ from pathlib import Path
4
+ import re
5
+ import statistics
6
+
7
+ OUT = Path(__file__).resolve().parent
8
+ rows=[]
9
+ for seed in range(5):
10
+ result=json.loads((OUT/f'random-seed-{seed}.json').read_text())
11
+ assert result['reference_dtype']=='float64' and not result['tf32_allowed']
12
+ assert result['checkpoint_step']==28000 and result['scored_samples']==128
13
+ for i,text in enumerate(result['sample_generations']):
14
+ words=re.findall(r'\b\w+\b',text.lower())
15
+ row=dict(seed=seed,index=i,text=text,entropy_nats=result['per_sample_entropy_nats'][i],
16
+ nll_sum=result['per_sample_nll_sum'][i],scored_tokens=result['per_sample_token_count'][i],
17
+ ppl=result['per_sample_ppl'][i],has_replacement_character=int('\ufffd' in text),
18
+ adjacent_repeated_word_fraction=sum(a==b for a,b in zip(words,words[1:]))/max(1,len(words)-1))
19
+ assert math.isfinite(row['ppl']) and row['scored_tokens']>0
20
+ rows.append(row)
21
+
22
+ def metrics(group):
23
+ if not group:
24
+ return dict(samples=0)
25
+ return dict(samples=len(group),scored_tokens=sum(r['scored_tokens'] for r in group),
26
+ pooled_token_weighted_ppl=math.exp(sum(r['nll_sum'] for r in group)/sum(r['scored_tokens'] for r in group)),
27
+ mean_entropy_nats=statistics.mean(r['entropy_nats'] for r in group),
28
+ fraction_samples_with_replacement_character=statistics.mean(r['has_replacement_character'] for r in group),
29
+ adjacent_repeated_word_fraction=statistics.mean(r['adjacent_repeated_word_fraction'] for r in group))
30
+
31
+ accepted=[r for r in rows if r['entropy_nats']>=4.5]
32
+ quality=dict(checkpoint_step=28000,reference_dtype='float64',tf32_allowed=False,
33
+ unfiltered=metrics(rows),entropy_ge_4p5=metrics(accepted),acceptance_rate=len(accepted)/len(rows),
34
+ comparison_step17167=json.loads((OUT.parent/'step-17167-entropy-ge-4p5/summary.json').read_text()))
35
+ (OUT/'quality-summary.json').write_text(json.dumps(quality,indent=2)+'\n')
36
+ (OUT/'accepted-entropy-ge-4p5.json').write_text(json.dumps(accepted,indent=2)+'\n')
37
+ summary=json.loads((OUT/'summary.json').read_text())
38
+ summary.update(reference_dtype='float64',tf32_allowed=False,quality_summary=str(OUT/'quality-summary.json'),
39
+ generation_tf32_allowed=False,total_requested_samples=640,total_skipped_empty_samples=0)
40
+ (OUT/'summary.json').write_text(json.dumps(summary,indent=2)+'\n')
41
+ lines=['# Step 28,000 generation evaluation','',
42
+ f"FP64 GPT-2 Large generation PPL: {summary['mean_gen_ppl']:.3f} +/- {summary['gen_ppl_se']:.3f} standard error across five seeds.",
43
+ '', '640 generations, temperature 1 without truncation, supplied true level-zero code. TF32 disabled.',
44
+ '', '## Quality measures','', '| Metric | All samples | Entropy >= 4.5 |','|---|---:|---:|']
45
+ for label,key,scale in [('Samples','samples',1),('Pooled token-weighted PPL','pooled_token_weighted_ppl',1),('Mean entropy (nats)','mean_entropy_nats',1),('Samples with replacement character (%)','fraction_samples_with_replacement_character',100),('Adjacent repeated words (%)','adjacent_repeated_word_fraction',100)]:
46
+ lines.append(f"| {label} | {quality['unfiltered'][key]*scale:.3f} | {quality['entropy_ge_4p5'].get(key,float('nan'))*scale:.3f} |")
47
+ lines+=['','## First three seed-0 samples','', 'Samples shown in original order, without selecting for quality.']
48
+ for row in rows[:3]:
49
+ lines+=['',f"### Seed {row['seed']}, index {row['index']} — PPL {row['ppl']:.2f}, entropy {row['entropy_nats']:.3f}",'',row['text']]
50
+ (OUT/'report.md').write_text('\n'.join(lines)+'\n')
51
+ print(json.dumps(summary,indent=2))
52
+ print(json.dumps({k:v for k,v in quality.items() if k!='comparison_step17167'},indent=2))
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/summary.json ADDED
@@ -0,0 +1,41 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "protocol": {
3
+ "sampling": "random",
4
+ "temperature": 1.0,
5
+ "top_k": 0,
6
+ "top_p": 1.0,
7
+ "n_provided_levels": 1,
8
+ "n_levels": 16,
9
+ "document_aware_validation_sampling": true
10
+ },
11
+ "reference_model": "gpt2-large",
12
+ "checkpoint": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/exports/step-28000/ncp.pt",
13
+ "checkpoint_step": 28000,
14
+ "seeds": [
15
+ 0,
16
+ 1,
17
+ 2,
18
+ 3,
19
+ 4
20
+ ],
21
+ "samples_per_seed": 128,
22
+ "total_scored_samples": 640,
23
+ "total_scored_tokens": 160778,
24
+ "per_seed_mean_ppl": {
25
+ "0": 124.87831485116675,
26
+ "1": 114.46117735800877,
27
+ "2": 119.77330723462428,
28
+ "3": 102.35317455545147,
29
+ "4": 95.95460863688895
30
+ },
31
+ "mean_gen_ppl": 111.48411652722804,
32
+ "gen_ppl_se": 5.392206594348971,
33
+ "mean_of_seed_median_ppl": 105.71566681551471,
34
+ "mean_entropy_nats": 4.381471981070202,
35
+ "reference_dtype": "float64",
36
+ "tf32_allowed": false,
37
+ "quality_summary": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/quality-summary.json",
38
+ "generation_tf32_allowed": false,
39
+ "total_requested_samples": 640,
40
+ "total_skipped_empty_samples": 0
41
+ }
iclr-debug-decoder-135b-20260916/decoder-mstok/milestone-iter-28000.pt ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d9eb43b22050308d6a75176a2978dadc8b0eeba38d905afa73521385e4a373c0
3
+ size 4879455521
iclr-debug-decoder-135b-20260916/decoder-mstok/teacher-provenance.json ADDED
@@ -0,0 +1,6 @@
 
 
 
 
 
 
 
1
+ {
2
+ "teacher_id": "FacebookAI/roberta-base",
3
+ "teacher_revision": "e2da8e2f811d1448a5b465c236feacd80ffbac7b",
4
+ "tokenizer_revision": "607a30d783dfa663caf39e06633721c8d4cfcd7e",
5
+ "token_map_sha256": "0d5aa4e0a98157722f03cf9a0c7a1f17178c370dd728896090077ea9efe0b593"
6
+ }
iclr-debug-decoder-135b-20260916/git-provenance.json ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ {
2
+ "commit": "1b5226d4bbdd34653cb5fdd58d436355fb76ac8c"
3
+ }
iclr-debug-decoder-135b-20260916/manifest.json ADDED
@@ -0,0 +1,525 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "format": "iclr-debug-production-v1",
3
+ "variant": "decoder-mstok",
4
+ "configs": {
5
+ "decoder-mstok": {
6
+ "task": "mstok-next-concept",
7
+ "experiment": "iclr-debug-decoder-135b-20260916-decoder-mstok",
8
+ "experiment_dir": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok",
9
+ "dataset": {
10
+ "train_source": "/home/ubuntu/data/full_owt/train_gpt2.bin",
11
+ "validate_source": "/home/ubuntu/data/full_owt/valid_gpt2.bin",
12
+ "pad_token_id": 50257
13
+ },
14
+ "pre_tokenizer": {
15
+ "name": "hf",
16
+ "tokenizer": "hf",
17
+ "model_id": "gpt2",
18
+ "special_tokens": {
19
+ "eos_token": "<|endoftext|>",
20
+ "pad_token": "<|pad|>"
21
+ }
22
+ },
23
+ "codec": {
24
+ "n_layers": 6,
25
+ "context_length": 256,
26
+ "embed_dim": 384,
27
+ "in_vocab_size": 50304,
28
+ "vqvae_vocab_size": [
29
+ 16384,
30
+ 16384,
31
+ 16384,
32
+ 16384,
33
+ 16384,
34
+ 16384,
35
+ 16384,
36
+ 16384,
37
+ 16384,
38
+ 16384,
39
+ 16384,
40
+ 16384,
41
+ 16384,
42
+ 16384,
43
+ 16384,
44
+ 16384
45
+ ],
46
+ "compression_factor": 1,
47
+ "dropout": 0.1,
48
+ "pre_quant_groupnorm": 4,
49
+ "pre_quant_dropout": 0.2,
50
+ "vector_quantizer_config": {
51
+ "clss": "multiscale_residual_vector_quantizer",
52
+ "decay": 0.99,
53
+ "epsilon": 1e-05,
54
+ "commitment_cost": 0.25,
55
+ "learned_l1_sampling": true,
56
+ "learned_all_sampling": false,
57
+ "quant_resi": {
58
+ "enabled": true,
59
+ "ratio": 0.5,
60
+ "share_mode": 0,
61
+ "learnable_ratio": false
62
+ },
63
+ "levels": {
64
+ "use_manual_levels": true,
65
+ "manual_levels": [
66
+ 1,
67
+ 4,
68
+ 9,
69
+ 16,
70
+ 25,
71
+ 36,
72
+ 49,
73
+ 64,
74
+ 81,
75
+ 100,
76
+ 121,
77
+ 144,
78
+ 169,
79
+ 196,
80
+ 225,
81
+ 256
82
+ ]
83
+ },
84
+ "aux": {
85
+ "fine_drop": {
86
+ "prob": 0.5,
87
+ "min_keep": 1
88
+ }
89
+ }
90
+ }
91
+ },
92
+ "generator": {
93
+ "n_layer": 12,
94
+ "n_head": 12,
95
+ "bias": true,
96
+ "dropout": 0.1,
97
+ "n_embd": 768,
98
+ "context_length": 256,
99
+ "vocab_size": [
100
+ 16384,
101
+ 16384,
102
+ 16384,
103
+ 16384,
104
+ 16384,
105
+ 16384,
106
+ 16384,
107
+ 16384,
108
+ 16384,
109
+ 16384,
110
+ 16384,
111
+ 16384,
112
+ 16384,
113
+ 16384,
114
+ 16384
115
+ ],
116
+ "attn_pattern": "block_diagonal",
117
+ "use_positional_encoding": true,
118
+ "use_level_encoding": true,
119
+ "prefix_len": 0,
120
+ "use_rope": true,
121
+ "rope_base": 10000,
122
+ "use_qk_norm": true,
123
+ "post_upsample_conv": {
124
+ "enabled": true,
125
+ "kernel_size": 3
126
+ },
127
+ "use_relu2": true,
128
+ "use_flex_attention": true,
129
+ "shared_output_head": true,
130
+ "shared_head_per_level_bias": false,
131
+ "shared_head_per_level_scale": false,
132
+ "shared_head_adapter_rank": 0
133
+ },
134
+ "objectives": {
135
+ "codec_weight": 1.0,
136
+ "ncp_weight": 1.0,
137
+ "soft_assignment_temperature": 1.0,
138
+ "prediction_temperature": 1.0,
139
+ "mstok_weight": 0.25,
140
+ "residual_weight": 1.0,
141
+ "reconstruction_weight": 1.0,
142
+ "teacher_ema_decay": 0.999,
143
+ "teacher_ema_warmup_steps": 1000,
144
+ "mstok_warmup_steps": 500
145
+ },
146
+ "optimization": {
147
+ "codec_lr": 0.001,
148
+ "codec_min_lr": 0.0001,
149
+ "codec_warmup_iters": 0,
150
+ "codec_lr_decay_iters": 171662,
151
+ "generator_lr": 0.0005,
152
+ "generator_min_lr": 1e-05,
153
+ "generator_warmup_iters": 300,
154
+ "generator_lr_decay_iters": 171662,
155
+ "beta_1": 0.9,
156
+ "codec_beta_2": 0.99,
157
+ "generator_beta_2": 0.99,
158
+ "weight_decay": 0.1,
159
+ "max_grad_norm": 1.0,
160
+ "level_loss_alpha": 1.0,
161
+ "grad_accumulation_steps": 12,
162
+ "corruption": {
163
+ "mode": "per_level",
164
+ "per_level_probs": [
165
+ 0.85,
166
+ 0.8321428571,
167
+ 0.8142857143,
168
+ 0.7964285714,
169
+ 0.7785714286,
170
+ 0.7607142857,
171
+ 0.7428571429,
172
+ 0.725,
173
+ 0.7071428571,
174
+ 0.6892857143,
175
+ 0.6714285714,
176
+ 0.6535714286,
177
+ 0.6357142857,
178
+ 0.6178571429,
179
+ 0.6
180
+ ],
181
+ "skip_level0": true
182
+ }
183
+ },
184
+ "training": {
185
+ "log_interval": 10,
186
+ "eval_interval": 1000,
187
+ "checkpoint_interval": 1000,
188
+ "eval_batch_size": 4,
189
+ "val_iters": 25,
190
+ "keep_last": 3,
191
+ "seed": 55,
192
+ "codec_initialization_seed": 42,
193
+ "generator_initialization_seed": 55,
194
+ "total_iters": 171662,
195
+ "batch_size": 32,
196
+ "expected_world_size": 8,
197
+ "expected_global_batch_size": 3072,
198
+ "resume_checkpoint": null,
199
+ "milestone_steps": [
200
+ 1272,
201
+ 6358,
202
+ 17167,
203
+ 34333,
204
+ 68665,
205
+ 102997,
206
+ 137330,
207
+ 171662
208
+ ],
209
+ "generation_eval_steps": [
210
+ 17167,
211
+ 34333,
212
+ 68665,
213
+ 102997,
214
+ 137330,
215
+ 171662
216
+ ]
217
+ },
218
+ "torch_compile": {
219
+ "enable": true,
220
+ "scope": "loss",
221
+ "mode": "default",
222
+ "dynamic": false,
223
+ "fullgraph": true,
224
+ "backend": "inductor"
225
+ },
226
+ "mixed_precision": {
227
+ "enable": true
228
+ },
229
+ "wandb": {
230
+ "entity": "mstok",
231
+ "project": "iclr-debug",
232
+ "group": "full-owt-ctx256-16sq-135b-v1",
233
+ "enable": true,
234
+ "id": "iclr-debug-decoder-135b-20260916-decoder-mstok",
235
+ "resume": "allow",
236
+ "gradients_and_params": {
237
+ "enable": false,
238
+ "log": "all",
239
+ "log_freq": 1000
240
+ }
241
+ },
242
+ "export": {
243
+ "vqvae_config_template": "/home/ubuntu/mstok-runs/iclr-debug-joint-20260916/config/repro-ctx256/vqvae.yaml",
244
+ "ncp_config_template": "/home/ubuntu/mstok-runs/iclr-debug-joint-20260916/config/repro-ctx256/ncp-sharedhead.yaml"
245
+ },
246
+ "semantic": {
247
+ "enabled": true,
248
+ "mode": "eostok",
249
+ "weight": 0.5,
250
+ "warmup_steps": 500,
251
+ "teacher_id": "FacebookAI/roberta-base",
252
+ "teacher_revision": "e2da8e2f811d1448a5b465c236feacd80ffbac7b",
253
+ "tokenizer_revision": "607a30d783dfa663caf39e06633721c8d4cfcd7e",
254
+ "teacher_dim": 768,
255
+ "feature_layer": 6,
256
+ "projector_dim": 2048,
257
+ "projector_seed": 56,
258
+ "alignment_site": "decoder"
259
+ },
260
+ "iclr_debug": {
261
+ "stage": "decoder-mstok",
262
+ "target_positions": 135000000000,
263
+ "lr_horizon_positions": 135000000000,
264
+ "stop_step": 171662,
265
+ "pilot": false,
266
+ "tokenizer_checkpoint": null,
267
+ "tokenizer_steps": null,
268
+ "budget_unit": "input positions including padding; log non-padding tokens separately"
269
+ }
270
+ }
271
+ },
272
+ "sources": {
273
+ "config/repro-ctx256/alignment-generator.yaml": "f62f36d078a3a373a19ea716a3ebc7eda5ae2fb7d670c5b935c734826bb76895",
274
+ "config/repro-ctx256/alignment-joint.yaml": "bb511950d20335e4237e406464ab26760a7fd9c9c9472c0507addd88d33d5178",
275
+ "config/repro-ctx256/alignment-tokenizer.yaml": "2d32955296f7d15990f284328c1cf8f52071abdba28033a65e272bb49cf5dd56",
276
+ "config/repro-ctx256/alignment.yaml": "3a486fa2ed28eca4fc406232669399b0625d274376ac43271f4ccf187b1e9077",
277
+ "config/repro-ctx256/joint.yaml": "8b4889111687a52e2d7d71995bda45a7f01f2777674c88191f338dca4c1732d2",
278
+ "config/repro-ctx256/mstok-semantic-eostok-matched10ep.yaml": "ae8a72c0c9f1cb5d2e301d83a1cdc08985a1f895bf627a5f75b1b3ce627f5e58",
279
+ "config/repro-ctx256/mstok-semantic-eostok.yaml": "2c9c48b21325c0e479990701b28116d78f611101724846747d173e52145fe15a",
280
+ "config/repro-ctx256/mstok-semantic-gear-matched10ep.yaml": "164cb8ac39fa407bb337f559ad8a9c62a34708fb900b8829d3ccf182961f2d0c",
281
+ "config/repro-ctx256/mstok-semantic-gear.yaml": "8cb8c998551daef0bf7fd9ec77b13b34c06662028d8820767e1432b5a9e8be54",
282
+ "config/repro-ctx256/mstok-semantic.yaml": "e0b00dbfa3eb838103f657c19bde20e36375c7fbe692623452ea093003f1ac98",
283
+ "config/repro-ctx256/mstok-w1-pilot.yaml": "21ee22eeccaaea2563c92ccc84f6df3871ea5da6a5ce980656088ea577dc1438",
284
+ "config/repro-ctx256/mstok-w1.yaml": "b6bf6079114b3c997fab89e85a53318bb2ec5dae55678513599d84152eab701b",
285
+ "config/repro-ctx256/mstok.yaml": "d9bb216f4eeb58857d72be399d5f8847d197ce70d0e504879f0788045483dc87",
286
+ "config/repro-ctx256/ncp-sharedhead.yaml": "86c9efbcb2b4f74b7253aa3bc4393de4a6a05e20ea30da573d1ca094ec600645",
287
+ "config/repro-ctx256/ncp.yaml": "93663654c4ce5ad0931db052edf507fdcb7c0ff9fc9722f4018820deee5c5acf",
288
+ "config/repro-ctx256/substitution-generator.yaml": "064f55887b9b2b8f3b0b52c094d4a193510dc1148ad7a5c7f53e317971f87c7c",
289
+ "config/repro-ctx256/substitution-joint.yaml": "b1576f4ecfe6fc6029a0cd868271b80a1d988b7f3667c6b1a8d24a6a4015df3b",
290
+ "config/repro-ctx256/substitution-tokenizer.yaml": "e29be5046f0fafbfe4c06973be4cd036a004e845348911cc0e6e05a7fca5d425",
291
+ "config/repro-ctx256/substitution.yaml": "ebba1c7e664de0e84913bdfd38e794837781c42c1b48cde90135ff9417856a95",
292
+ "config/repro-ctx256/vqvae.yaml": "05fde46e57c3e6b411e8a6afcac251f178e4b8a0856a0b99332bb284fb4710ce",
293
+ "data/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
294
+ "data/serialized_dataset.py": "5622c2d94236cd1b9b4ca18e8d5c272a86b54d193d99048fb9b7cd2687e8dbc6",
295
+ "data/tinysentences/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
296
+ "data/tinysentences/clean.py": "918a9a91aae484a5e06d6d212c2c0974946a4032e061c62a7b65851735bfc8dd",
297
+ "data/tinysentences/clean_parsed.py": "009a087f967e171bc6098484dbfd58635302900f6b07629629898f1d00fa70a3",
298
+ "data/tinysentences/generate_meta.py": "662714f1370e05567d412add168a7c3a05f4b1c9adb047fad374900bc1936a7e",
299
+ "data/tinysentences/generate_metadata_char.py": "afd3336ea03ff0653ccaf41c121d9b93debf7a380db6cd881b62b308ee6ee68c",
300
+ "data/tinysentences/generate_training_data_word.py": "82e2edcb7ea4db0609ccf2b047a0b207e7b8fcff4fc0b4ceee521ff278e8a8e8",
301
+ "data/tinysentences/parse_summaries.py": "2317159364dbbe7206f252d2939cab5d375395e0a536be0a1f60c5df88b31702",
302
+ "data/tinysentences/prepare.py": "35263b287f93337fda4207ed1991e9e8a820ab75139411bb7dd5f0a8a25bcbd9",
303
+ "data/tinysentences/sentences_only.py": "5e4d72bf612df27dad46cf606a40402606d897043ae24274dc2dfa88ec0b3e79",
304
+ "data/tinysentences/stats.py": "2665e9e6fac4a8f6c3e5bd8b07cbfd941ba6359a43981d89bdef58c136fdd3a9",
305
+ "data/tinysentences/summarize/merge_summaries_char.py": "bce94f4d9a824c2d25e45448023d59b7fbc42aba22b33efbc067cdbba93e1934",
306
+ "data/tinysentences/summarize/merge_summaries_word.py": "f6079110332cc38b2dd3571c30bedcca97098c39498f45eec1ef6e60cce7ac81",
307
+ "data/tinysentences/summarize/summarize_gemini.py": "282a438efce0dee351fd9833a59c56e4480bed4314dae40844201794ada3bd96",
308
+ "data/tinysentences/summarize/summarize_gpt.py": "fc08cbc37d6161e5b0057a1ede43272d96f7274524917d6c4d42c22f505c6e09",
309
+ "data/tinysentences/summarize/summarize_llama.py": "7f3f8230652a5b88ea2856b273d497cd5b633580b430c909e4f996cd0232832d",
310
+ "data/tinysentences/summarize/summarize_qwen.py": "b07e3c226ab0fdf5b21b025263adfa7c2625d640f0fac3c787bc4a2c9e32107f",
311
+ "data/tinysentences/summarize/summarize_together.py": "8c26ea052f2a64c2653d39f1ea5c13983dfbc9547f21cd7d7ea9d9b39033bffe",
312
+ "data/tinysentences/test.py": "9749244da77c078f550a78751664b45bb92067682f3aa27c2338419f1186c844",
313
+ "data/tinysentences/test_regex.py": "46652e189a0d5f5c0c7a2c90b857082004bf6727ef1d3ad42b631ac7bb0aa88f",
314
+ "data/tinysentences/tinysentences_char_dataset.py": "152ab9c30aec669a9db14df9e90412597870105eff9cfb737b876aad2323a206",
315
+ "data/tinysentences/tinysentences_dataset.py": "f90251d21871f6a7ad5ebd2ad0a6e59cd0b02ca245f7e46f93e6d970bf657bc7",
316
+ "data/tinysentences/tinysentences_vae.py": "30f7adde5413348faf5066448706747fe26a7fedd510ffea720aa0bf8283f6f8",
317
+ "data/tinysentences/to_csv.py": "c06c74648ed95b4494835c631c94fe1a00a9ff47a272f5e38546a7f94ff9be8b",
318
+ "data/tinysentences/tokenize/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
319
+ "data/tinysentences/tokenize/common.py": "e877c74df2a48d15991492b126cd0c5fe26731d61f3aa36d83ebf2cb9bda34a0",
320
+ "data/tinysentences/tokenize/tokenizer.py": "c16abc50b92a163f6978542b36b03c9a14aa00153c73117c4d52084d5de25c2a",
321
+ "data/tinysentences/tokenize/train_bpe.py": "fa8e55978dd5ca90117200648b468fb22822bf6981f15f0394c815e72eae1bdb",
322
+ "data/tinystories/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
323
+ "data/tinystories/download.py": "5afce37e912f6965b6722f0bc028076a655fea78304f1a213e4862dbe223cdcb",
324
+ "data/tinystories/stats.py": "216650fe19b78e8431c4512ff87a5c84af64d9e56ab8fdaf3952ca035725ea1b",
325
+ "data/wikitext/parse_wikitext.py": "ab583e0328f82afc3ba242410e8076be42ef1a16a51d66f3eb46a713cca591d0",
326
+ "evaluation/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
327
+ "evaluation/eval_random_gen_ppl_checkpoint.py": "221ea1917dc2d3bc19be39da0e0b8cca49295d0400a05da276cd50956876da8d",
328
+ "evaluation/gen_ppl.py": "e6853b14d27f93407505eb7607a1bd9fed49a7add1511ceca2b4a3076d533859",
329
+ "evaluation/gen_ppl_ncp.py": "96d68adee372e5186d0470659171eb03829e3d37f611e8f5917b45fb61e3a4ad",
330
+ "evaluation/gen_ppl_ntp.py": "b2638622bc041fa88e9d538be4229b3723ebec19c5b99fd6e5b73db4776dc952",
331
+ "evaluation/guess.py": "c27803225dd21b2deff1e7e494b9e3666180b93776340164ce6e953dacc50656",
332
+ "evaluation/ppl.py": "f3b581cb4849b0b2aa518729b8338838c7055697c3a26028dedb819c335ac838",
333
+ "evaluation/run_genppl_multiseed.py": "f8c9ac982f4caa4335b5f4d1a22d70f56dc11a706106a3dfb31d344eb8c93c2e",
334
+ "evaluation/run_genppl_release_ts.py": "d1bc5c00089118d9cf6c05c77459329ad40c2f8bc347b5aeedda8f5318de8796",
335
+ "evaluation/run_genppl_sweep.py": "ade4f39e373125848fdd32938e8d7bf86a0ec5108cca9bd75f0a966d092a798e",
336
+ "evaluation/show_8k_samples.py": "30ad825d5a24375e178750dfaaac45bc486bc3587f2e7d6d305d53bf37f8d226",
337
+ "evaluation/show_low_sample.py": "d696f1e54ca67e87565dce5864e0e6e2ad4194108cfb1c9c1e06eebd97008da6",
338
+ "evaluation/summarize_gen_ppl.py": "307f3d0d85884879295d5268aaf7f399f654b1b330aaa5728d6824ad24c6d15a",
339
+ "evaluation/sweep_gen_ppl_checkpoints.py": "3a9141e27d84b49c5493194dfc5a70bfa0a915cdf5ceb7a46afcf27cc6f33acb",
340
+ "evaluation/sweep_gen_ppl_ncp.py": "d91c2959cfed51d224ccb1c0d0f2c3b9a35c4027cc9ab8c65aa7603d587de761",
341
+ "evaluation/sweep_gen_ppl_ncp_scaling.py": "05249c1e00c5eb05042b6ea5b1ca008ecb8f0e4a77537ee755a27c153c25a17c",
342
+ "evaluation/sweep_gen_ppl_ncp_yolo.py": "4b26ef25b27af2f7fcbcb5ff3c312d47d2a4746ba759155d28802e978a681755",
343
+ "evaluation/sweep_gen_ppl_ntp.py": "14a04074741ee44580ec6557870d42534dd172d26c5be3e943ed5baf1b727303",
344
+ "evaluation/sweep_random_llama_yolo.py": "4062b54289fcd9407d15feb1ffd241b8722fd717c700350c22472fecec5ae46c",
345
+ "evaluation/val_oracle_ppl.py": "1a23f3c2897b6a6a3d239ed533979de35352d303bae094f4006bef4156f935c0",
346
+ "evaluation/val_vqvae_ppl.py": "4c293c502d3faf982d4ad66566b7a037faa7d28191d929de1621cb67f5ae8b5f",
347
+ "models/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
348
+ "models/common.py": "99aa0da356963a282936413d738ee1da72a3e3f277fecc28eddb553051582a4b",
349
+ "models/masked_lm.py": "80279614d74efc7b4c25306789d74e2f71045c54e4e4f485ee365f664705d7f6",
350
+ "models/next_concept.py": "743667aac3bdceb7bb7476a580700185b2638ddd9575e6357db57acc7d699339",
351
+ "models/next_token.py": "d0d292b9f5c032a351cbca1a8f0a6d41c3062890cb3e3f5b67d4e23a7cffeb77",
352
+ "models/quant.py": "b7b73abd4e55e10454cb680dea82124c38c222d37236c4c76b5c7d21123ac1b1",
353
+ "models/semantic_input.py": "b06658d4bc641397cbb973255eadfd2d6087fde244d801d3c70d29e3ce841bf2",
354
+ "models/vae.py": "02e2213a8878920374baf27dd10dfc03847dce866557c97e816c427d230c9ded",
355
+ "models/vqvae.py": "8eb3503eb1d990121b7d6138a4f8216aecfc72246d4a92bbecf98025c8a3d504",
356
+ "scripts/benchmark_iclr_debug.py": "dba86e8d578eec2188c633a479ccf673d58f3e5580b9074869cae1a352b8a043",
357
+ "scripts/benchmark_mstok_compile.py": "aad909da6f05ba25ac3d8f6b3cf450103bed55d0b93b1490764c0ef54f35b606",
358
+ "scripts/eval_owt_ckpts_gpt2_large.py": "4977a2d614698f60032fb6fd7fef614eb503d5b3883f499146a4fc048214a509",
359
+ "scripts/evaluate_mstok_sidecar.py": "9ba33e865fa0fad34bd2ab375d7fe38f37c441e2d96348f0e389e29dd9643455",
360
+ "scripts/evaluations/sweep_gen_ppl_shared_vocab_ncp.py": "2311ac71f9beba282426faa21390a5898c7d183a95f529755d6e1d6e8fc24050",
361
+ "scripts/neurips/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
362
+ "scripts/neurips/eval_gen_ppl_ncp_owt.py": "b92eff1ce198eb4ac8c88da7113b41d095c12acad92db541685175993db3535c",
363
+ "scripts/neurips/sweep_gen_ppl_owt_gpt2.py": "02adb1629cfa40ff51a75152e4da9c9ee00ead42b7ecda9e78cf7101e6826a67",
364
+ "scripts/preflight_owtsmall_mstok_ctx256.py": "a4f80921435b6cd136d089ec39bfa5442359d0dffbeefaaa1019178b8364657a",
365
+ "scripts/prepare_full_owt.py": "b1cb6b74397625c2b0643c4ee744ca3c89dfa6bbb7eeb7e9681e3a36f91a2c24",
366
+ "scripts/r1_lab/diagnose_at_40k.py": "b1582dcc29ccebc9255eaa77a7d56a7dbdef589ec5d5a9b0af150dede98874ab",
367
+ "scripts/r1_lab/inspect_summaries.py": "4154e8a6c1b3f68df5c623a57a1a7632b594c46c5f6d1d85cb08c476d6aa9cea",
368
+ "scripts/r1_lab/overnight_pipeline.py": "53f4e03ce1a5fb43fbca3ae455288f45ae8350b0a16f3d5318763011de6a8f96",
369
+ "scripts/r1_lab/prime_rope_qknorm_score_cache.py": "a791cdec03b415bfccdfd8d89c7e659aff8a2a5696d16ab7b4350ff608d0391b",
370
+ "scripts/r1_lab/sweep_gen_ppl_r1lab.py": "c7c0acf9aa6f07bb0a985cf0d5e5692c696e658ebb0470b7c92d9332565f9c06",
371
+ "scripts/r1_lab/sweep_gen_ppl_r1lab_v15.py": "524aa8680aa1b9fea2689509848c19e5a74a3ad72e78b0388bbca7acd43be22e",
372
+ "scripts/r1_lab/sweep_gen_ppl_rope_qknorm.py": "7822df057355b0a3a71ea5145742107c6e30f906859096237ff61d765fe36d50",
373
+ "scripts/r1_lab/train_variant.py": "9d0a2091784609ac9728226afd82f56fd4c1ac2dc08e89e1b17b60d34bb3d55e",
374
+ "scripts/run_alignment.py": "6312cd9d0976899763589fdcbd258c7681f6e9e96d7396cf35ff022d4695b862",
375
+ "scripts/run_iclr_debug.py": "a48a14067af036a0eba46d8de69922c615c05f6bbbab11587dca292ea57682ce",
376
+ "scripts/run_mstok_semantic.py": "8a7a38a50dd6694950bfa0e7b79718ec653c7e70be99fe13cbf2ee05c28f46a0",
377
+ "scripts/run_mstok_w1_pilot.py": "35224e462d0d11eb3e4326b0daa8bad7c99e1e30cdc6383310f8e2c8b5bb708f",
378
+ "scripts/run_substitution.py": "2cd22eaa58dab8b5231984eee340a3a2bdd0b88cc83333f7133447471e7f73f1",
379
+ "scripts/smoke_test_masked_l0.py": "d24f1270b3219a31212b68bcbb786f649c7c72f96e53f602d87ad6c17f3f73e5",
380
+ "scripts/test0_diagnostics.py": "ed389c666beef015ccac004f314a5a1428def303b46ac0e442f0236337344549",
381
+ "scripts/test0_r1_similarity.py": "4a486878bebd418a3b5454ddbf02585726ccbacc2700fdf591aa524766f9c9c9",
382
+ "scripts/test0_r1_uniqueness.py": "3ee59d4397d5b268151dbb8ca779c7f9bf8461d129d97fcd1a517c48ec2a69a1",
383
+ "scripts/test0_residual_norms.py": "3c4f68e5960e84796dc0bca4fd4583bf3b598552654f9b565fd5fcabd420f9e1",
384
+ "scripts/test_gen_equiv.py": "ae0e9c513422ee58bdc050ed6218ee54fc555dfb15ea8968c9ffc62a7a7ea99a",
385
+ "scripts/training_budget.py": "bb94c19c6dfbdd542946b10476b8d6958b8a189abd141cf62288fdb5fa9a2d38",
386
+ "scripts/vqvae_utilization.py": "5f90327f1507a5042b2f77b119fe4fc0df4e2131e5bbc21069dffd1fadc059c5",
387
+ "scripts/xsum/codec_ceiling.py": "ddb751e9152d5c88966e588d89cdceee4c5d91c4139aaedf369f65a45ecf2276",
388
+ "scripts/xsum/eval_merge.py": "6e313d6ed25227d0418b7422d58e09ceda4fc6dfae7d0fd663417b306241135c",
389
+ "scripts/xsum/eval_q0.py": "82fe04dccce28d7fc9c6ec8b4c9c91404a49527f9437913f44f51a6853ba9abc",
390
+ "scripts/xsum/eval_shard.py": "07192f6ba6219da1e8123ceeb99748c86da85322859ed5fa25cebcc3816f7ea8",
391
+ "scripts/xsum/eval_shard_cfg.py": "95ac4a8e2e3a73a21fa4303c62df80c086f23bcfe5ab34eb07ae90b585052ecb",
392
+ "scripts/xsum/eval_test.py": "7f92d578c31893e788e5fa7d9eef3098ba5602b3642873ce2ce031342c8be868",
393
+ "scripts/xsum/eval_uncond_rep.py": "ba4740b29bc091d866e2dfd91881e9a7d5c46a1bfa993d13eb63c0467397ec28",
394
+ "scripts/xsum/make_warmstart.py": "78793d6ef11fbf658c2837ef12cb843fc4817add8ba73c822b5e408901de3318",
395
+ "scripts/xsum/mbr_fast.py": "3db6e21d2dcce2e3bcf6f22474535d7eba0232b071b4271d2b172f787aed5c85",
396
+ "scripts/xsum/mbr_select.py": "107a061014c62395b7f5fd986fba9e55dd738054e2474022005cb7e9e0c9b8fb",
397
+ "scripts/xsum/mbr_shard_cfg.py": "3e7162375028c391276149eff288e9de40c874a3108ac53de0248198f4d08791",
398
+ "scripts/xsum/pool_append.py": "e50252f48c4d8d905be0de3eb56ca1d720f8f8545c6fdcb58ed27f0360c4e2ad",
399
+ "scripts/xsum/prep_xsum_le48.py": "3688ed1abc49d18f75d34f8382ff68a5afbd753b36cefdc0652d0bffb6a5f347",
400
+ "scripts/xsum/prep_xsum_mix30.py": "4bd47e52637c4b6bca5fd4956c1787b506a7f51b8a5bb0b68f19f728c4d48afb",
401
+ "scripts/xsum/prep_xsum_mix_split.py": "ec2fe74f08068ae8fda4ed59a86537a1d3f48276d205099bbfe16f8428ca36d5",
402
+ "scripts/xsum/prep_xsum_pass_uncond.py": "82ea50248d811e5182e63d0b4dfd6f995dd2f47eb963b5bb1c5d09bc533db7c1",
403
+ "scripts/xsum/prep_xsum_test_all.py": "7389fbdf5612ba524801c5b04109f810b139152ca50101b60acbb9b1e767b256",
404
+ "scripts/xsum/score_checkpoints.py": "41b412c06cd92cbd4caa6b586da7ddc7e913607a47c1a3aac0ee0a36a20a36be",
405
+ "tokenization/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
406
+ "tokenization/bpe_tokenizer.py": "0c8af84868a6aa40f981768f8ab49d28e501d55cce444edb1ad315de53c3445a",
407
+ "tokenization/build_word_vocab.py": "778f9666f4e7c642dc471773ee81c92c26975db37de1ebfffb605cae211351b5",
408
+ "tokenization/byte_tokenizer.py": "eb567a1772ae69738884ed9b94330aa39387c49e54d76e4b9752e67674fad810",
409
+ "tokenization/compression_ratio.py": "ed07ab1cc38d6176362631ce655bc974dd3623b87693f51279d5047e6bc6890f",
410
+ "tokenization/hf_tokenizer.py": "2ca8fae097f9c640bcb52b4788c37de8e0d9bb2fa971e30a5abb39f90ff97d4c",
411
+ "tokenization/tokenize.py": "f8ab54338613b6bd84513fdda85be173fd32274c35d97ef39a5c17d584f8108f",
412
+ "tokenization/tokenize_doc_level.py": "a06c81cf088fb81d2317ed1c3a1417a0e48ae0ea04962824f13f5d3beddc8902",
413
+ "tokenization/tokenize_vqvae.py": "14730eec3da87f6528e5fbd1166d5da79e10ac1700f7cd0be73bc40716bdebfc",
414
+ "tokenization/tokenizer.py": "e527d47e4ed8634c673d48a872a07c9cc8828d30988331de2d4bd3b7c21ddcb7",
415
+ "tokenization/train_bpe.py": "bf4e2a29eeba874024ed95f32fe720ae2320f72cb978361692db73e71be98683",
416
+ "tokenization/vqvae_tokenizer.py": "d7fa9d595865aa3e0b0d185bfd04721eb297fe5f657c4fca4ce69fc96aca308f",
417
+ "tokenization/word_tokenizer.py": "a405621bf35ed6cbb29abadd9d1d651e0783e6801fc7645b35e799d5de381ac2",
418
+ "train_mstok.py": "077b7c99c12daad2bbd986918d0f487411779de05a472af71124098da399b79d",
419
+ "trainer/alignment_trainer.py": "54f8ee7cca4d407223f2ce5146e6fc561ad748e20f2c702beb12ae040cd290c4",
420
+ "trainer/corruption.py": "f0f24c94652e6e9cf8e4ef30ae1a302e3232d64e765c931696d454db96e8b57d",
421
+ "trainer/joint_trainer.py": "8c8ebb55d21b6e223be1f7595c51f1b9257ccac40f8cff28260c1088656e2e49",
422
+ "trainer/mstok_trainer.py": "90712e396472d5dac9679b62af4f0b62afa6390a839d4d9bd9dae57406fe76a4",
423
+ "trainer/ncp_trainer.py": "73fc23f78f403dc39800b924dfb18ad42ffeab847fed30e931183c3d3acaee68",
424
+ "trainer/ntp_trainer.py": "5076d7aa29fac70869e1e309a15af6c4d3a91d6de8e00ac113a586841a9866e1",
425
+ "trainer/onpolicy.py": "6ff1fe4103d71001afe2d73e574d20c7242381e420929db14020141464ec6821",
426
+ "trainer/plain_two_stage.py": "4c5973ed87100b1b478aa1df0bb16cd3d1eb5844dffe5894003dd67c463f6c9a",
427
+ "trainer/scheduled_sampling.py": "f53515e23c2cec1231265f21a341ff77d11fa8a6129e0e382d540cc85571f5ed",
428
+ "trainer/semantic_mstok_trainer.py": "ebc02e80d279a445c99d8f09505cdfaf7632f1a661d6880f09cc1bd0fa90ef30",
429
+ "trainer/semantic_teacher.py": "7983b5da86d25954ed8267eee03d144c421305967705e0a242f4f025794ceeae",
430
+ "trainer/substitution_trainer.py": "33c7833b2e8dfabbb9ee56b258708055d3194b8d7453a1b8d93f59b44ff91465",
431
+ "utils/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
432
+ "utils/arg_utils.py": "e9eb7971f6273e9630b56fd87a03725b29183a8c80abb014c4c84f8b1bdb6445",
433
+ "utils/benchmark.py": "ccf436ef72a4e626c7f671d8e63a683057bf502a3852e5a7142410d1dc3e5501",
434
+ "utils/data.py": "78cce11e777813755cd171ccaa3247f89ab99ac4079e0157e55c72ce8a5a7fe5",
435
+ "utils/dist.py": "b2849d3ce03a8b1737e3dd61216e76e4a9e82d63ad1ea74cddf72e2285eaca64",
436
+ "utils/experiment_config.py": "616a70fefaac99ab0d856a1b6c36787a28fb9af8a4c55fc6d4765012fbaaeb38",
437
+ "utils/iclr_training.py": "10c6857fa2319952fe0df18067730b42e3ff755244e2b059f6c95882f8d27c8e",
438
+ "utils/logging.py": "901e67515cc65a35f575f05d16a3b9ef4a5b2cb0ea252cd7abbc80ecce9e88e8",
439
+ "utils/lr.py": "4d1ea2c25340afc4bdde8b582dab36f174e0b2965fa7bafa2a7539019dbe2b7e",
440
+ "utils/misc.py": "be4a323116a0956a77960d1baf613f268057d6a36023663d2af9f8e4728cbe77",
441
+ "utils/plot_levels.py": "916b39885ced191960c26c334a3b16f9c99751f6c29c46bd69d14fee414965f3",
442
+ "utils/registry.py": "f7a44c6d18433100873d5df462b9f08a8bdeb11cf3d92226334c7e5ece318318",
443
+ "utils/train.py": "43eef988000f3a5e20856653d725337a0fe0dbe276f0cf80e1b200292abe812d",
444
+ "train_iclr_debug.py": "e0a0adf4a30ef55925f9017de24d434299749332e58a507a6424265f7a5f796e",
445
+ "config/iclr-debug/README.md": "1c757d2ff37026d08c22324ec44cfa760b0599af62c8ad9b728ada471a8f2905",
446
+ "config/iclr-debug/benchmark-16sq.yaml": "18c9cef7032f9d3806d8b0c8b20fe0fc7d1ea98ad1d2440e24a67ee873c97f9b",
447
+ "config/iclr-debug/full-owt-audit.json": "fefa394834f6e9f65944d3b514633c29061071ebe6ca60fce21a6edef9748c1e",
448
+ "config/iclr-debug/training-defaults.yaml": "3b144b8d53f96fe035e36edca96aac456e6cf21202373a467a3cd875fe86b474",
449
+ "config/iclr-debug/NCM-BUDGET-AUDIT.md": "e734caeab7edb612de7ff508a8c0ec52445530ab7a35895b5b793d6a7b752e4b",
450
+ "config/iclr-debug/TRAINING.md": "91555457c5cf6af078075673ce96cb440bc8e22aeae5211c252d94f92ea87ac7",
451
+ "config/iclr-debug/BENCHMARK.md": "2781ea506212e12a0640a4b0d2a2c881355ec85756ca4014bccb1b1aedccdf59",
452
+ "config/iclr-debug/HF-MODEL-CARD.md": "de1fe176ebde262eacdf32af7d3f8c6497a764dd98e5a9094dad26de4aade38e"
453
+ },
454
+ "dataset": {
455
+ "dataset": "hazyresearch/ncm-tokenized-datasets",
456
+ "revision": "951e04517fd53d57521d8e09abd6511a662787d0",
457
+ "dtype": "little-endian uint16",
458
+ "eos_id": 50256,
459
+ "offset_units": "token positions; [0, each EOS position + 1] with final sentinel",
460
+ "sampling": "length-weighted within-segment windows; pad short segments with 50257",
461
+ "caveat": "EOS segments are not independently verified source-document identities",
462
+ "splits": {
463
+ "train": {
464
+ "bytes": 18070847356,
465
+ "tokens": 9035423678,
466
+ "sha256": "b24b78a07f56974fd0c9855dc16f25b836d2cb172be23eb5d3c650cd6d73f56a",
467
+ "min_token_id": 0,
468
+ "max_token_id": 50256,
469
+ "segments": 8009770,
470
+ "trailing_tokens": 0,
471
+ "eos_only_segments": 0,
472
+ "length_percentiles": {
473
+ "min": 132.0,
474
+ "p25": 416.0,
475
+ "p50": 718.0,
476
+ "p75": 1249.0,
477
+ "p99": 7610.0,
478
+ "max": 131288.0
479
+ },
480
+ "segments_shorter_than_256": 673011,
481
+ "expected_padding_fraction_at_256": 1.5955357567300495e-05,
482
+ "offsets_status": "created",
483
+ "offsets_sha256": "7e84564946b8cf8190dbdc1e86bdf2fe8cfd9235a517328c80b4df038ea0028b",
484
+ "offset_entries": 8009771
485
+ },
486
+ "valid": {
487
+ "bytes": 9186834,
488
+ "tokens": 4593417,
489
+ "sha256": "1f2ee17e08327a39d4e2417948d02c7b998d626d2dcb2661197a6e208f58266f",
490
+ "min_token_id": 0,
491
+ "max_token_id": 50256,
492
+ "segments": 3999,
493
+ "trailing_tokens": 0,
494
+ "eos_only_segments": 0,
495
+ "length_percentiles": {
496
+ "min": 143.0,
497
+ "p25": 433.0,
498
+ "p50": 734.0,
499
+ "p75": 1292.5,
500
+ "p99": 7122.719999999997,
501
+ "max": 24868.0
502
+ },
503
+ "segments_shorter_than_256": 363,
504
+ "expected_padding_fraction_at_256": 1.6064002171981108e-05,
505
+ "offsets_status": "created",
506
+ "offsets_sha256": "09a748ce1dcec7de36c77c48e39ccbf5c0293cc073689fa7c5f48bdfb9bb9445",
507
+ "offset_entries": 4000
508
+ }
509
+ },
510
+ "overlap": {
511
+ "exact_train_segment_matches": 0,
512
+ "unique_validation_segments_in_train": 0,
513
+ "unique_validation_segments": 3999,
514
+ "near_duplicates_checked": false
515
+ }
516
+ },
517
+ "hf_repo": null,
518
+ "generation_evaluation": true,
519
+ "versions": {
520
+ "torch": "2.13.0+cu126",
521
+ "transformers": "4.44.2",
522
+ "numpy": "1.26.4",
523
+ "hydra-core": "1.3.6"
524
+ }
525
+ }
iclr-debug-decoder-135b-20260916/provenance/hf-backup-step-28000-manifest.json ADDED
@@ -0,0 +1,116 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "repository": "iskhare/iclr-debug",
3
+ "checkpoint_step": 28000,
4
+ "files": [
5
+ {
6
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/milestone-iter-28000.pt",
7
+ "bytes": 4879455521,
8
+ "sha256": "d9eb43b22050308d6a75176a2978dadc8b0eeba38d905afa73521385e4a373c0"
9
+ },
10
+ {
11
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/manifest.json",
12
+ "bytes": 28971,
13
+ "sha256": "53dc5b666685e03efc016e4311b89ab94bb6babe710405c28ae58356406582d0"
14
+ },
15
+ {
16
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/git-provenance.json",
17
+ "bytes": 59,
18
+ "sha256": "96f601db3a051fd136086ca5f54c87ce2a72449e528d826a3b882541953e3fa2"
19
+ },
20
+ {
21
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/config.yaml",
22
+ "bytes": 4827,
23
+ "sha256": "073bd5a69f1d0a39019efffd79f4b7bd645e0c1792edead7e31b524515ef134f"
24
+ },
25
+ {
26
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/teacher-provenance.json",
27
+ "bytes": 270,
28
+ "sha256": "96e0063afe4211b1109509690a697fd8a7366941988c894f1027e1ead45938cf"
29
+ },
30
+ {
31
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/summary.json",
32
+ "bytes": 1174,
33
+ "sha256": "2bee9a7e46cdf0402e56360d2dadf306d3347f835432b6685241bcab04f712ba"
34
+ },
35
+ {
36
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/quality-summary.json",
37
+ "bytes": 2945,
38
+ "sha256": "3441e3ff7d5cce050e74ff6daf146514e0fc26621e1675d005e8a177a75e1d1c"
39
+ },
40
+ {
41
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/report.md",
42
+ "bytes": 3444,
43
+ "sha256": "4f15ced1ff8b7aeb09895f1e01ab27b27865d4bbd692d15874ccada09991e652"
44
+ },
45
+ {
46
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/manifest.json",
47
+ "bytes": 916,
48
+ "sha256": "993caaf29b702767a24bbddf23e27fb961b652fdcf5ce5771de39cd6f0dd1568"
49
+ },
50
+ {
51
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/evaluator.py",
52
+ "bytes": 7709,
53
+ "sha256": "e2c87f66f1001853f58b9e78a646fc4e089fc894d05fa6c215a8baf493a84055"
54
+ },
55
+ {
56
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/controller.py",
57
+ "bytes": 3094,
58
+ "sha256": "1de01fb6ca96049e786f86c61d07dae9a73897824e73c5a8be1aa47f2e41edd7"
59
+ },
60
+ {
61
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/summarize-quality.py",
62
+ "bytes": 3757,
63
+ "sha256": "60ba459f27cde464bd4ed8bd2f5518e5b97d120d315042788d9ad37fbfd059d1"
64
+ },
65
+ {
66
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-0.json",
67
+ "bytes": 139308,
68
+ "sha256": "60889f234d7ce2324f2f97782dc936971ad62ba8764fe0cc2902bc3ffc15aab6"
69
+ },
70
+ {
71
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-0.samples.json",
72
+ "bytes": 129390,
73
+ "sha256": "ccb0e33aba18b71c5c1d77c8c957f146604fc9b68f244e9db03508f60c188861"
74
+ },
75
+ {
76
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-1.json",
77
+ "bytes": 141566,
78
+ "sha256": "40b7bd89e42f6bc5b7cb12022373ffcdf743e21d3308a56ccddbb58d3d48fc4f"
79
+ },
80
+ {
81
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-1.samples.json",
82
+ "bytes": 131645,
83
+ "sha256": "e5f7af200bdbf470e43cdca1a689c245c2927846e51bb8ddd45fe0e7c35dafcd"
84
+ },
85
+ {
86
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-2.json",
87
+ "bytes": 143228,
88
+ "sha256": "77fac889717c95593b412a885ce8fde84fa05a56445e94cf951346f3e6fcb2f3"
89
+ },
90
+ {
91
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-2.samples.json",
92
+ "bytes": 133321,
93
+ "sha256": "8da15a810f2766becb09dc5ed56de844c61c0470eeb13ac376f362bd420a168a"
94
+ },
95
+ {
96
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-3.json",
97
+ "bytes": 141349,
98
+ "sha256": "f63cf95471df8e843388ff0c53335db775e17874ee3dab00ec80485c969fc2f6"
99
+ },
100
+ {
101
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-3.samples.json",
102
+ "bytes": 131426,
103
+ "sha256": "02fb5a585fde4f9188e8d9be168ed17b72f51d4586eb5d9e909dc2286f2295f7"
104
+ },
105
+ {
106
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-4.json",
107
+ "bytes": 140771,
108
+ "sha256": "f413441516002fd050c00d2283d3d9ac6a86cc90414d49fe514da110d8969d5f"
109
+ },
110
+ {
111
+ "path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-4.samples.json",
112
+ "bytes": 130875,
113
+ "sha256": "0a4d08f610465b45b68e31cdc448f092bb18fe447f5556ec5a1cea114f2338ab"
114
+ }
115
+ ]
116
+ }