Back up decoder-aligned joint MsTok step 28000 with FP64 generation evaluation
Browse files- iclr-debug-decoder-135b-20260916/decoder-mstok/checkpoint-28000-README.md +15 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/config.yaml +236 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/controller.py +54 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/evaluator.py +187 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/manifest.json +25 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/quality-summary.json +87 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-0.json +0 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-0.samples.json +0 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-1.json +0 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-1.samples.json +0 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-2.json +0 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-2.samples.json +0 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-3.json +0 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-3.samples.json +0 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-4.json +0 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-4.samples.json +0 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/report.md +44 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/summarize-quality.py +52 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/summary.json +41 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/milestone-iter-28000.pt +3 -0
- iclr-debug-decoder-135b-20260916/decoder-mstok/teacher-provenance.json +6 -0
- iclr-debug-decoder-135b-20260916/git-provenance.json +3 -0
- iclr-debug-decoder-135b-20260916/manifest.json +525 -0
- iclr-debug-decoder-135b-20260916/provenance/hf-backup-step-28000-manifest.json +116 -0
iclr-debug-decoder-135b-20260916/decoder-mstok/checkpoint-28000-README.md
ADDED
|
@@ -0,0 +1,15 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Decoder-aligned joint MsTok: step 28,000
|
| 2 |
+
|
| 3 |
+
Full resumable checkpoint from `iclr-debug-decoder-135b-20260916`, trained jointly from scratch on full OWT.
|
| 4 |
+
|
| 5 |
+
- 28,000 optimizer updates; 22,020,096,000 input positions including padding.
|
| 6 |
+
- Context 256; 16 square scales; global batch 3,072; eight H100 workers.
|
| 7 |
+
- Checkpoint contains online codec, generator, EMA teacher, semantic projector, optimizer, LR scheduler, per-rank RNG states, config, and actual token counter.
|
| 8 |
+
- Training source commit: see `git-provenance.json`. Resolved config and immutable run manifest are included.
|
| 9 |
+
- Generation evaluation: 640 samples across five seeds, temperature 1 without truncation, supplied true level-zero code, scored with FP64 GPT-2 Large and TF32 disabled.
|
| 10 |
+
- Mean generation PPL 111.484 +/- 5.392 standard error; mean token entropy 4.381 nats. 65.3% of samples contain Unicode replacement characters. This is an intermediate research checkpoint with substantial generation defects.
|
| 11 |
+
- Entropy >= 4.5 retains 182/640 samples; pooled generation PPL 203.126. See `evaluation/step-28000-fp64/` for full results and saved generations.
|
| 12 |
+
|
| 13 |
+
The full checkpoint is `milestone-iter-28000.pt`; component weights can be exported with `utils.iclr_training.export_training_checkpoint` from the matching source checkout. Local absolute paths in saved configuration record the original run environment and are not portable defaults.
|
| 14 |
+
|
| 15 |
+
See `../provenance/` at the run root for the SHA-256 inventory. The backup does not include training data or credentials. Training continues independently of this snapshot.
|
iclr-debug-decoder-135b-20260916/decoder-mstok/config.yaml
ADDED
|
@@ -0,0 +1,236 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
task: mstok-next-concept
|
| 2 |
+
experiment: iclr-debug-decoder-135b-20260916-decoder-mstok
|
| 3 |
+
experiment_dir: /home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok
|
| 4 |
+
dataset:
|
| 5 |
+
train_source: /home/ubuntu/data/full_owt/train_gpt2.bin
|
| 6 |
+
validate_source: /home/ubuntu/data/full_owt/valid_gpt2.bin
|
| 7 |
+
pad_token_id: 50257
|
| 8 |
+
pre_tokenizer:
|
| 9 |
+
name: hf
|
| 10 |
+
tokenizer: hf
|
| 11 |
+
model_id: gpt2
|
| 12 |
+
special_tokens:
|
| 13 |
+
eos_token: <|endoftext|>
|
| 14 |
+
pad_token: <|pad|>
|
| 15 |
+
codec:
|
| 16 |
+
n_layers: 6
|
| 17 |
+
context_length: 256
|
| 18 |
+
embed_dim: 384
|
| 19 |
+
in_vocab_size: 50304
|
| 20 |
+
vqvae_vocab_size:
|
| 21 |
+
- 16384
|
| 22 |
+
- 16384
|
| 23 |
+
- 16384
|
| 24 |
+
- 16384
|
| 25 |
+
- 16384
|
| 26 |
+
- 16384
|
| 27 |
+
- 16384
|
| 28 |
+
- 16384
|
| 29 |
+
- 16384
|
| 30 |
+
- 16384
|
| 31 |
+
- 16384
|
| 32 |
+
- 16384
|
| 33 |
+
- 16384
|
| 34 |
+
- 16384
|
| 35 |
+
- 16384
|
| 36 |
+
- 16384
|
| 37 |
+
compression_factor: 1
|
| 38 |
+
dropout: 0.1
|
| 39 |
+
pre_quant_groupnorm: 4
|
| 40 |
+
pre_quant_dropout: 0.2
|
| 41 |
+
vector_quantizer_config:
|
| 42 |
+
clss: multiscale_residual_vector_quantizer
|
| 43 |
+
decay: 0.99
|
| 44 |
+
epsilon: 1.0e-05
|
| 45 |
+
commitment_cost: 0.25
|
| 46 |
+
learned_l1_sampling: true
|
| 47 |
+
learned_all_sampling: false
|
| 48 |
+
quant_resi:
|
| 49 |
+
enabled: true
|
| 50 |
+
ratio: 0.5
|
| 51 |
+
share_mode: 0
|
| 52 |
+
learnable_ratio: false
|
| 53 |
+
levels:
|
| 54 |
+
use_manual_levels: true
|
| 55 |
+
manual_levels:
|
| 56 |
+
- 1
|
| 57 |
+
- 4
|
| 58 |
+
- 9
|
| 59 |
+
- 16
|
| 60 |
+
- 25
|
| 61 |
+
- 36
|
| 62 |
+
- 49
|
| 63 |
+
- 64
|
| 64 |
+
- 81
|
| 65 |
+
- 100
|
| 66 |
+
- 121
|
| 67 |
+
- 144
|
| 68 |
+
- 169
|
| 69 |
+
- 196
|
| 70 |
+
- 225
|
| 71 |
+
- 256
|
| 72 |
+
aux:
|
| 73 |
+
fine_drop:
|
| 74 |
+
prob: 0.5
|
| 75 |
+
min_keep: 1
|
| 76 |
+
generator:
|
| 77 |
+
n_layer: 12
|
| 78 |
+
n_head: 12
|
| 79 |
+
bias: true
|
| 80 |
+
dropout: 0.1
|
| 81 |
+
n_embd: 768
|
| 82 |
+
context_length: 256
|
| 83 |
+
vocab_size:
|
| 84 |
+
- 16384
|
| 85 |
+
- 16384
|
| 86 |
+
- 16384
|
| 87 |
+
- 16384
|
| 88 |
+
- 16384
|
| 89 |
+
- 16384
|
| 90 |
+
- 16384
|
| 91 |
+
- 16384
|
| 92 |
+
- 16384
|
| 93 |
+
- 16384
|
| 94 |
+
- 16384
|
| 95 |
+
- 16384
|
| 96 |
+
- 16384
|
| 97 |
+
- 16384
|
| 98 |
+
- 16384
|
| 99 |
+
attn_pattern: block_diagonal
|
| 100 |
+
use_positional_encoding: true
|
| 101 |
+
use_level_encoding: true
|
| 102 |
+
prefix_len: 0
|
| 103 |
+
use_rope: true
|
| 104 |
+
rope_base: 10000
|
| 105 |
+
use_qk_norm: true
|
| 106 |
+
post_upsample_conv:
|
| 107 |
+
enabled: true
|
| 108 |
+
kernel_size: 3
|
| 109 |
+
use_relu2: true
|
| 110 |
+
use_flex_attention: true
|
| 111 |
+
shared_output_head: true
|
| 112 |
+
shared_head_per_level_bias: false
|
| 113 |
+
shared_head_per_level_scale: false
|
| 114 |
+
shared_head_adapter_rank: 0
|
| 115 |
+
objectives:
|
| 116 |
+
codec_weight: 1.0
|
| 117 |
+
ncp_weight: 1.0
|
| 118 |
+
soft_assignment_temperature: 1.0
|
| 119 |
+
prediction_temperature: 1.0
|
| 120 |
+
mstok_weight: 0.25
|
| 121 |
+
residual_weight: 1.0
|
| 122 |
+
reconstruction_weight: 1.0
|
| 123 |
+
teacher_ema_decay: 0.999
|
| 124 |
+
teacher_ema_warmup_steps: 1000
|
| 125 |
+
mstok_warmup_steps: 500
|
| 126 |
+
optimization:
|
| 127 |
+
codec_lr: 0.001
|
| 128 |
+
codec_min_lr: 0.0001
|
| 129 |
+
codec_warmup_iters: 0
|
| 130 |
+
codec_lr_decay_iters: 171662
|
| 131 |
+
generator_lr: 0.0005
|
| 132 |
+
generator_min_lr: 1.0e-05
|
| 133 |
+
generator_warmup_iters: 300
|
| 134 |
+
generator_lr_decay_iters: 171662
|
| 135 |
+
beta_1: 0.9
|
| 136 |
+
codec_beta_2: 0.99
|
| 137 |
+
generator_beta_2: 0.99
|
| 138 |
+
weight_decay: 0.1
|
| 139 |
+
max_grad_norm: 1.0
|
| 140 |
+
level_loss_alpha: 1.0
|
| 141 |
+
grad_accumulation_steps: 12
|
| 142 |
+
corruption:
|
| 143 |
+
mode: per_level
|
| 144 |
+
per_level_probs:
|
| 145 |
+
- 0.85
|
| 146 |
+
- 0.8321428571
|
| 147 |
+
- 0.8142857143
|
| 148 |
+
- 0.7964285714
|
| 149 |
+
- 0.7785714286
|
| 150 |
+
- 0.7607142857
|
| 151 |
+
- 0.7428571429
|
| 152 |
+
- 0.725
|
| 153 |
+
- 0.7071428571
|
| 154 |
+
- 0.6892857143
|
| 155 |
+
- 0.6714285714
|
| 156 |
+
- 0.6535714286
|
| 157 |
+
- 0.6357142857
|
| 158 |
+
- 0.6178571429
|
| 159 |
+
- 0.6
|
| 160 |
+
skip_level0: true
|
| 161 |
+
training:
|
| 162 |
+
log_interval: 10
|
| 163 |
+
eval_interval: 1000
|
| 164 |
+
checkpoint_interval: 1000
|
| 165 |
+
eval_batch_size: 4
|
| 166 |
+
val_iters: 25
|
| 167 |
+
keep_last: 3
|
| 168 |
+
seed: 55
|
| 169 |
+
codec_initialization_seed: 42
|
| 170 |
+
generator_initialization_seed: 55
|
| 171 |
+
total_iters: 171662
|
| 172 |
+
batch_size: 32
|
| 173 |
+
expected_world_size: 8
|
| 174 |
+
expected_global_batch_size: 3072
|
| 175 |
+
resume_checkpoint: null
|
| 176 |
+
milestone_steps:
|
| 177 |
+
- 1272
|
| 178 |
+
- 6358
|
| 179 |
+
- 17167
|
| 180 |
+
- 34333
|
| 181 |
+
- 68665
|
| 182 |
+
- 102997
|
| 183 |
+
- 137330
|
| 184 |
+
- 171662
|
| 185 |
+
generation_eval_steps:
|
| 186 |
+
- 17167
|
| 187 |
+
- 34333
|
| 188 |
+
- 68665
|
| 189 |
+
- 102997
|
| 190 |
+
- 137330
|
| 191 |
+
- 171662
|
| 192 |
+
torch_compile:
|
| 193 |
+
enable: true
|
| 194 |
+
scope: loss
|
| 195 |
+
mode: default
|
| 196 |
+
dynamic: false
|
| 197 |
+
fullgraph: true
|
| 198 |
+
backend: inductor
|
| 199 |
+
mixed_precision:
|
| 200 |
+
enable: true
|
| 201 |
+
wandb:
|
| 202 |
+
entity: mstok
|
| 203 |
+
project: iclr-debug
|
| 204 |
+
group: full-owt-ctx256-16sq-135b-v1
|
| 205 |
+
enable: true
|
| 206 |
+
id: iclr-debug-decoder-135b-20260916-decoder-mstok
|
| 207 |
+
resume: allow
|
| 208 |
+
gradients_and_params:
|
| 209 |
+
enable: false
|
| 210 |
+
log: all
|
| 211 |
+
log_freq: 1000
|
| 212 |
+
export:
|
| 213 |
+
vqvae_config_template: /home/ubuntu/mstok-runs/iclr-debug-joint-20260916/config/repro-ctx256/vqvae.yaml
|
| 214 |
+
ncp_config_template: /home/ubuntu/mstok-runs/iclr-debug-joint-20260916/config/repro-ctx256/ncp-sharedhead.yaml
|
| 215 |
+
semantic:
|
| 216 |
+
enabled: true
|
| 217 |
+
mode: eostok
|
| 218 |
+
weight: 0.5
|
| 219 |
+
warmup_steps: 500
|
| 220 |
+
teacher_id: FacebookAI/roberta-base
|
| 221 |
+
teacher_revision: e2da8e2f811d1448a5b465c236feacd80ffbac7b
|
| 222 |
+
tokenizer_revision: 607a30d783dfa663caf39e06633721c8d4cfcd7e
|
| 223 |
+
teacher_dim: 768
|
| 224 |
+
feature_layer: 6
|
| 225 |
+
projector_dim: 2048
|
| 226 |
+
projector_seed: 56
|
| 227 |
+
alignment_site: decoder
|
| 228 |
+
iclr_debug:
|
| 229 |
+
stage: decoder-mstok
|
| 230 |
+
target_positions: 135000000000
|
| 231 |
+
lr_horizon_positions: 135000000000
|
| 232 |
+
stop_step: 171662
|
| 233 |
+
pilot: false
|
| 234 |
+
tokenizer_checkpoint: null
|
| 235 |
+
tokenizer_steps: null
|
| 236 |
+
budget_unit: input positions including padding; log non-padding tokens separately
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/controller.py
ADDED
|
@@ -0,0 +1,54 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import datetime
|
| 2 |
+
import json
|
| 3 |
+
import math
|
| 4 |
+
import os
|
| 5 |
+
from pathlib import Path
|
| 6 |
+
import subprocess
|
| 7 |
+
import sys
|
| 8 |
+
|
| 9 |
+
REPO = Path('/home/ubuntu/mstok-runs/iclr-debug-joint-20260916')
|
| 10 |
+
STAGE = Path('/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok')
|
| 11 |
+
OUT = STAGE / 'evaluation/step-28000-fp64'
|
| 12 |
+
EXPORT = STAGE / 'exports/step-28000'
|
| 13 |
+
ENV = dict(os.environ, CUDA_VISIBLE_DEVICES='0', OMP_NUM_THREADS='1',
|
| 14 |
+
TOKENIZERS_PARALLELISM='false', PYTORCH_ALLOC_CONF='expandable_segments:True')
|
| 15 |
+
|
| 16 |
+
def status(state, **extra):
|
| 17 |
+
payload = dict(state=state, updated_utc=datetime.datetime.now(datetime.timezone.utc).isoformat(),
|
| 18 |
+
controller_pid=os.getpid(), **extra)
|
| 19 |
+
temporary = OUT / 'manual-eval-status.tmp'
|
| 20 |
+
temporary.write_text(json.dumps(payload, indent=2) + '\n')
|
| 21 |
+
temporary.replace(OUT / 'manual-eval-status.json')
|
| 22 |
+
print(json.dumps(payload), flush=True)
|
| 23 |
+
|
| 24 |
+
try:
|
| 25 |
+
for seed in range(5):
|
| 26 |
+
result = OUT / f'random-seed-{seed}.json'
|
| 27 |
+
if not result.exists():
|
| 28 |
+
status('evaluating', seed=seed)
|
| 29 |
+
code = "import runpy,torch; torch.cuda.set_per_process_memory_fraction(0.15); runpy.run_path('/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/evaluator.py',run_name='__main__')"
|
| 30 |
+
command = [sys.executable, '-u', '-c', code, '--ckpt', str(EXPORT/'ncp.pt'),
|
| 31 |
+
'--codec', str(EXPORT/'vqvae.pt'), '--validation-source',
|
| 32 |
+
'/home/ubuntu/data/full_owt/valid_gpt2.bin', '--ref-model', 'gpt2-large',
|
| 33 |
+
'--ref-dtype', 'float64', '--seed', str(seed), '--n-samples', '128',
|
| 34 |
+
'--batch-size', '4', '--temperature', '1', '--top-k', '0', '--top-p', '1',
|
| 35 |
+
'--save-all-generations', '--out', str(result)]
|
| 36 |
+
with result.with_suffix('.log').open('a') as log:
|
| 37 |
+
subprocess.run(command, cwd=REPO, env=ENV, stdout=log, stderr=subprocess.STDOUT,
|
| 38 |
+
check=True, timeout=1800)
|
| 39 |
+
row = json.loads(result.read_text())
|
| 40 |
+
assert row['checkpoint_step'] == 28000 and row['seed'] == seed
|
| 41 |
+
assert row['reference_dtype'] == 'float64' and not row['tf32_allowed']
|
| 42 |
+
assert row['protocol']['n_levels'] == 16 and row['protocol']['n_provided_levels'] == 1
|
| 43 |
+
assert row['requested_samples'] == row['scored_samples'] == 128
|
| 44 |
+
assert row['scored_tokens'] > 0
|
| 45 |
+
assert math.isfinite(row['mean_ppl']) and math.isfinite(row['mean_entropy_nats'])
|
| 46 |
+
status('seed-complete', seed=seed, mean_ppl=row['mean_ppl'], entropy=row['mean_entropy_nats'])
|
| 47 |
+
with (OUT/'summarize.log').open('w') as log:
|
| 48 |
+
subprocess.run([sys.executable, 'evaluation/summarize_gen_ppl.py', '--input-glob',
|
| 49 |
+
str(OUT/'random-seed-[0-4].json'), '--out', str(OUT/'summary.json')],
|
| 50 |
+
cwd=REPO, env=ENV, stdout=log, stderr=subprocess.STDOUT, check=True)
|
| 51 |
+
status('complete', summary=str(OUT/'summary.json'))
|
| 52 |
+
except BaseException as exc:
|
| 53 |
+
status('failed', error=str(exc))
|
| 54 |
+
raise
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/evaluator.py
ADDED
|
@@ -0,0 +1,187 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
"""Evaluate random-sampling generation PPL from a completed NCP checkpoint."""
|
| 2 |
+
|
| 3 |
+
import argparse
|
| 4 |
+
import json
|
| 5 |
+
import math
|
| 6 |
+
import sys
|
| 7 |
+
from collections import Counter
|
| 8 |
+
from pathlib import Path
|
| 9 |
+
|
| 10 |
+
REPO_ROOT = Path("/home/ubuntu/mstok-runs/iclr-debug-joint-20260916")
|
| 11 |
+
if str(REPO_ROOT) not in sys.path:
|
| 12 |
+
sys.path.insert(0, str(REPO_ROOT))
|
| 13 |
+
|
| 14 |
+
import hydra
|
| 15 |
+
import torch
|
| 16 |
+
from omegaconf import OmegaConf
|
| 17 |
+
from transformers import AutoModelForCausalLM, AutoTokenizer
|
| 18 |
+
|
| 19 |
+
from data.serialized_dataset import SerializedDataset
|
| 20 |
+
from evaluation.gen_ppl import generation_perplexity
|
| 21 |
+
from utils.misc import set_manual_seed
|
| 22 |
+
from utils.registry import trainers
|
| 23 |
+
|
| 24 |
+
|
| 25 |
+
def main():
|
| 26 |
+
parser = argparse.ArgumentParser()
|
| 27 |
+
parser.add_argument("--ckpt", required=True)
|
| 28 |
+
parser.add_argument("--codec")
|
| 29 |
+
parser.add_argument("--validation-source")
|
| 30 |
+
parser.add_argument("--ref-model", default="gpt2-large")
|
| 31 |
+
parser.add_argument("--ref-dtype", choices=("bfloat16", "float32", "float64"), default="bfloat16",
|
| 32 |
+
help="Legacy default retained; new ICLR runs explicitly use float32")
|
| 33 |
+
parser.add_argument("--seed", type=int, default=0)
|
| 34 |
+
parser.add_argument("--n-samples", type=int, default=128)
|
| 35 |
+
parser.add_argument("--batch-size", type=int, default=8)
|
| 36 |
+
parser.add_argument("--temperature", type=float, default=1.0)
|
| 37 |
+
parser.add_argument("--top-k", type=int, default=0)
|
| 38 |
+
parser.add_argument("--top-p", type=float, default=1.0)
|
| 39 |
+
parser.add_argument("--out", required=True)
|
| 40 |
+
parser.add_argument("--save-all-generations", action="store_true")
|
| 41 |
+
args = parser.parse_args()
|
| 42 |
+
|
| 43 |
+
device = "cuda"
|
| 44 |
+
checkpoint = torch.load(args.ckpt, map_location="cpu", weights_only=False)
|
| 45 |
+
cfg = checkpoint["args"]
|
| 46 |
+
OmegaConf.set_struct(cfg, False)
|
| 47 |
+
if args.codec:
|
| 48 |
+
cfg.tokenizer.vqvae.checkpoint_path = args.codec
|
| 49 |
+
if args.validation_source:
|
| 50 |
+
cfg.dataset.validate_source = args.validation_source
|
| 51 |
+
cfg.optimization.token_corruption_prob = 0.0
|
| 52 |
+
if getattr(cfg.optimization, "stochastic_encode", None) is not None:
|
| 53 |
+
cfg.optimization.stochastic_encode.enabled = False
|
| 54 |
+
else:
|
| 55 |
+
cfg.optimization.stochastic_encode = {"enabled": False}
|
| 56 |
+
|
| 57 |
+
set_manual_seed(args.seed)
|
| 58 |
+
# Seed setup enables TF32 by default, so enforce evaluation precision after it.
|
| 59 |
+
if args.ref_dtype in ("float32", "float64"):
|
| 60 |
+
torch.backends.cuda.matmul.allow_tf32 = False
|
| 61 |
+
torch.backends.cudnn.allow_tf32 = False
|
| 62 |
+
Trainer = hydra.utils.get_class(trainers[cfg.task])
|
| 63 |
+
trainer = Trainer(cfg, device)
|
| 64 |
+
trainer.model.load_state_dict(checkpoint["model"])
|
| 65 |
+
model = trainer.model.to(device).eval()
|
| 66 |
+
tokenizer = trainer.tokenizer
|
| 67 |
+
levels = list(model.levels)
|
| 68 |
+
prefix_len = int(cfg.model.prefix_len)
|
| 69 |
+
|
| 70 |
+
dataset = SerializedDataset(
|
| 71 |
+
cfg.dataset.validate_source,
|
| 72 |
+
int(cfg.model.context_length),
|
| 73 |
+
cfg.task,
|
| 74 |
+
tokenizer.pad_token_id,
|
| 75 |
+
)
|
| 76 |
+
full_sequences, _ = dataset.get_rand_batch(args.n_samples)
|
| 77 |
+
prefix_sequences = full_sequences[:, :prefix_len]
|
| 78 |
+
target_sequences = full_sequences[:, prefix_len:]
|
| 79 |
+
|
| 80 |
+
prompts = []
|
| 81 |
+
generations = []
|
| 82 |
+
with torch.no_grad():
|
| 83 |
+
for start in range(0, args.n_samples, args.batch_size):
|
| 84 |
+
end = min(start + args.batch_size, args.n_samples)
|
| 85 |
+
prefix = prefix_sequences[start:end].to(device)
|
| 86 |
+
target_idx = tokenizer.encode_pre_tokenized_idx(
|
| 87 |
+
target_sequences[start:end].to(device)
|
| 88 |
+
)
|
| 89 |
+
level0 = target_idx[:, : levels[0]]
|
| 90 |
+
generated_idx = model.generate(
|
| 91 |
+
level0,
|
| 92 |
+
prefix,
|
| 93 |
+
temperature=args.temperature,
|
| 94 |
+
top_k=args.top_k,
|
| 95 |
+
top_p=args.top_p,
|
| 96 |
+
)
|
| 97 |
+
reconstructions = tokenizer.decode_multiscales(
|
| 98 |
+
generated_idx, skip_special_tokens=True
|
| 99 |
+
)
|
| 100 |
+
generations.extend(
|
| 101 |
+
reconstructions[row][-1] for row in range(len(reconstructions))
|
| 102 |
+
)
|
| 103 |
+
prompts.extend(
|
| 104 |
+
tokenizer.pre_tokenizer.decode_batch(
|
| 105 |
+
prefix.cpu().tolist(), skip_special_tokens=True
|
| 106 |
+
)
|
| 107 |
+
)
|
| 108 |
+
|
| 109 |
+
if args.save_all_generations:
|
| 110 |
+
# Preserve every generated sample even if scoring fails or some are empty.
|
| 111 |
+
samples_path = Path(args.out).with_suffix(".samples.json")
|
| 112 |
+
samples_path.parent.mkdir(parents=True, exist_ok=True)
|
| 113 |
+
samples_path.write_text(json.dumps(dict(seed=args.seed, checkpoint=args.ckpt,
|
| 114 |
+
prompts=prompts, generations=generations), indent=2) + "\n")
|
| 115 |
+
kept = [(prompt, text) for prompt, text in zip(prompts, generations) if text.strip()]
|
| 116 |
+
prompts = [prompt for prompt, _ in kept]
|
| 117 |
+
generations = [text for _, text in kept]
|
| 118 |
+
|
| 119 |
+
entropies = []
|
| 120 |
+
for text in generations:
|
| 121 |
+
ids = tokenizer.pre_tokenizer.encode(text)
|
| 122 |
+
if ids:
|
| 123 |
+
counts = Counter(ids)
|
| 124 |
+
entropies.append(
|
| 125 |
+
-sum((count / len(ids)) * math.log(count / len(ids)) for count in counts.values())
|
| 126 |
+
)
|
| 127 |
+
|
| 128 |
+
del model, trainer, tokenizer
|
| 129 |
+
torch.cuda.empty_cache()
|
| 130 |
+
|
| 131 |
+
ref_tokenizer = AutoTokenizer.from_pretrained(args.ref_model)
|
| 132 |
+
ref_model = AutoModelForCausalLM.from_pretrained(
|
| 133 |
+
args.ref_model, torch_dtype=getattr(torch, args.ref_dtype),
|
| 134 |
+
**({"attn_implementation": "eager"} if args.ref_dtype == "float64" else {})
|
| 135 |
+
).to(device).eval()
|
| 136 |
+
assert all(p.dtype == getattr(torch, args.ref_dtype) for p in ref_model.parameters())
|
| 137 |
+
result = generation_perplexity(
|
| 138 |
+
ref_model,
|
| 139 |
+
ref_tokenizer,
|
| 140 |
+
prompts,
|
| 141 |
+
generations,
|
| 142 |
+
score_only_generated=True,
|
| 143 |
+
batch_size=args.batch_size,
|
| 144 |
+
device=device,
|
| 145 |
+
)
|
| 146 |
+
|
| 147 |
+
assert result.nll_sum.dtype == getattr(torch, args.ref_dtype)
|
| 148 |
+
assert bool(torch.isfinite(result.ppl).all()) and bool((result.token_count > 0).all())
|
| 149 |
+
payload = {
|
| 150 |
+
"checkpoint": str(Path(args.ckpt).resolve()),
|
| 151 |
+
"checkpoint_step": int(checkpoint["step"]),
|
| 152 |
+
"codec": str(Path(cfg.tokenizer.vqvae.checkpoint_path).resolve()),
|
| 153 |
+
"reference_model": args.ref_model,
|
| 154 |
+
"reference_dtype": args.ref_dtype,
|
| 155 |
+
"reference_attention_implementation": ref_model.config._attn_implementation,
|
| 156 |
+
"per_sample_nll_sum": result.nll_sum.tolist(),
|
| 157 |
+
"per_sample_token_count": result.token_count.tolist(),
|
| 158 |
+
"per_sample_ppl": result.ppl.tolist(),
|
| 159 |
+
"per_sample_entropy_nats": entropies,
|
| 160 |
+
"tf32_allowed": torch.backends.cuda.matmul.allow_tf32,
|
| 161 |
+
"protocol": {
|
| 162 |
+
"sampling": "random" if args.top_k == 0 and args.top_p == 1.0 else "truncated",
|
| 163 |
+
"temperature": args.temperature,
|
| 164 |
+
"top_k": args.top_k,
|
| 165 |
+
"top_p": args.top_p,
|
| 166 |
+
"n_provided_levels": 1,
|
| 167 |
+
"n_levels": len(levels),
|
| 168 |
+
"document_aware_validation_sampling": dataset.doc_offsets is not None,
|
| 169 |
+
},
|
| 170 |
+
"seed": args.seed,
|
| 171 |
+
"requested_samples": args.n_samples,
|
| 172 |
+
"scored_samples": len(generations),
|
| 173 |
+
"skipped_empty_samples": args.n_samples - len(generations),
|
| 174 |
+
"scored_tokens": int(result.token_count.sum()),
|
| 175 |
+
"mean_ppl": float(result.mean_ppl),
|
| 176 |
+
"median_ppl": float(result.ppl.median()),
|
| 177 |
+
"mean_entropy_nats": sum(entropies) / len(entropies),
|
| 178 |
+
"sample_generations": generations if args.save_all_generations else generations[:8],
|
| 179 |
+
}
|
| 180 |
+
output_path = Path(args.out)
|
| 181 |
+
output_path.parent.mkdir(parents=True, exist_ok=True)
|
| 182 |
+
output_path.write_text(json.dumps(payload, indent=2) + "\n")
|
| 183 |
+
print(json.dumps({key: value for key, value in payload.items() if key != "sample_generations"}, indent=2))
|
| 184 |
+
|
| 185 |
+
|
| 186 |
+
if __name__ == "__main__":
|
| 187 |
+
main()
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/manifest.json
ADDED
|
@@ -0,0 +1,25 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"checkpoint_step": 28000,
|
| 3 |
+
"input_positions": 22020096000,
|
| 4 |
+
"source_checkpoint": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/checkpoint-iter-28000.pt",
|
| 5 |
+
"retained_checkpoint": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/source-checkpoint.pt",
|
| 6 |
+
"checkpoint_sha256": "d9eb43b22050308d6a75176a2978dadc8b0eeba38d905afa73521385e4a373c0",
|
| 7 |
+
"evaluator_sha256": "e2c87f66f1001853f58b9e78a646fc4e089fc894d05fa6c215a8baf493a84055",
|
| 8 |
+
"checkout": "/home/ubuntu/mstok-runs/iclr-debug-joint-20260916",
|
| 9 |
+
"reference_dtype": "float64",
|
| 10 |
+
"reference_attention": "eager",
|
| 11 |
+
"tf32_allowed": false,
|
| 12 |
+
"memory_fraction": 0.15,
|
| 13 |
+
"gpu": 0,
|
| 14 |
+
"batch_size": 4,
|
| 15 |
+
"seeds": [
|
| 16 |
+
0,
|
| 17 |
+
1,
|
| 18 |
+
2,
|
| 19 |
+
3,
|
| 20 |
+
4
|
| 21 |
+
],
|
| 22 |
+
"samples_per_seed": 128,
|
| 23 |
+
"sampling": "temperature-1-untruncated-supplied-true-level0",
|
| 24 |
+
"concurrent_with_training": true
|
| 25 |
+
}
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/quality-summary.json
ADDED
|
@@ -0,0 +1,87 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"checkpoint_step": 28000,
|
| 3 |
+
"reference_dtype": "float64",
|
| 4 |
+
"tf32_allowed": false,
|
| 5 |
+
"unfiltered": {
|
| 6 |
+
"samples": 640,
|
| 7 |
+
"scored_tokens": 160778,
|
| 8 |
+
"pooled_token_weighted_ppl": 110.93052440167119,
|
| 9 |
+
"mean_entropy_nats": 4.381471981070202,
|
| 10 |
+
"fraction_samples_with_replacement_character": 0.653125,
|
| 11 |
+
"adjacent_repeated_word_fraction": 0.03943009422883846
|
| 12 |
+
},
|
| 13 |
+
"entropy_ge_4p5": {
|
| 14 |
+
"samples": 182,
|
| 15 |
+
"scored_tokens": 45589,
|
| 16 |
+
"pooled_token_weighted_ppl": 203.12568343002383,
|
| 17 |
+
"mean_entropy_nats": 4.5968029107321735,
|
| 18 |
+
"fraction_samples_with_replacement_character": 0.6208791208791209,
|
| 19 |
+
"adjacent_repeated_word_fraction": 0.03950939465147932
|
| 20 |
+
},
|
| 21 |
+
"acceptance_rate": 0.284375,
|
| 22 |
+
"comparison_step17167": {
|
| 23 |
+
"checkpoint_step": 17167,
|
| 24 |
+
"reference_dtype": "float64",
|
| 25 |
+
"threshold": 4.5,
|
| 26 |
+
"entropy_definition": "Empirical GPT-2 unigram entropy in nats within each generated sample",
|
| 27 |
+
"selection": "Filter existing 640 generations; no new generation; all accepted samples pooled",
|
| 28 |
+
"accepted": {
|
| 29 |
+
"samples": 115,
|
| 30 |
+
"scored_tokens": 28860,
|
| 31 |
+
"pooled_token_weighted_ppl": 189.63357736496337,
|
| 32 |
+
"mean_entropy_nats": 4.57744085314838,
|
| 33 |
+
"median_sample_ppl": 171.1510152167491,
|
| 34 |
+
"adjacent_repeated_word_fraction": 0.04295705405426942,
|
| 35 |
+
"samples_with_replacement_character": 63,
|
| 36 |
+
"fraction_samples_with_replacement_character": 0.5478260869565217
|
| 37 |
+
},
|
| 38 |
+
"unfiltered": {
|
| 39 |
+
"samples": 640,
|
| 40 |
+
"scored_tokens": 160898,
|
| 41 |
+
"pooled_token_weighted_ppl": 109.72590001293689,
|
| 42 |
+
"mean_entropy_nats": 4.3665410349745,
|
| 43 |
+
"median_sample_ppl": 99.25832710923365,
|
| 44 |
+
"adjacent_repeated_word_fraction": 0.04761292563004243,
|
| 45 |
+
"samples_with_replacement_character": 431,
|
| 46 |
+
"fraction_samples_with_replacement_character": 0.6734375
|
| 47 |
+
},
|
| 48 |
+
"acceptance_rate": 0.1796875,
|
| 49 |
+
"per_seed": [
|
| 50 |
+
{
|
| 51 |
+
"seed": 0,
|
| 52 |
+
"accepted": 26,
|
| 53 |
+
"proposed": 128,
|
| 54 |
+
"ppl": 198.79796628571316
|
| 55 |
+
},
|
| 56 |
+
{
|
| 57 |
+
"seed": 1,
|
| 58 |
+
"accepted": 21,
|
| 59 |
+
"proposed": 128,
|
| 60 |
+
"ppl": 171.26448994329286
|
| 61 |
+
},
|
| 62 |
+
{
|
| 63 |
+
"seed": 2,
|
| 64 |
+
"accepted": 25,
|
| 65 |
+
"proposed": 128,
|
| 66 |
+
"ppl": 201.94449304567118
|
| 67 |
+
},
|
| 68 |
+
{
|
| 69 |
+
"seed": 3,
|
| 70 |
+
"accepted": 21,
|
| 71 |
+
"proposed": 128,
|
| 72 |
+
"ppl": 201.29939871172186
|
| 73 |
+
},
|
| 74 |
+
{
|
| 75 |
+
"seed": 4,
|
| 76 |
+
"accepted": 22,
|
| 77 |
+
"proposed": 128,
|
| 78 |
+
"ppl": 174.01773580284262
|
| 79 |
+
}
|
| 80 |
+
],
|
| 81 |
+
"mean_of_seed_ppls": 189.46481675784833,
|
| 82 |
+
"ppl_definition": "exp(sum saved FP64 reference NLL / sum scored tokens)",
|
| 83 |
+
"conditioning": "supplied true level-zero code",
|
| 84 |
+
"fp64_source": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-17167-fp64",
|
| 85 |
+
"entropy_source": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-17167-diversity-audit/step_17167-samples.json"
|
| 86 |
+
}
|
| 87 |
+
}
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-0.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-0.samples.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-1.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-1.samples.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-2.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-2.samples.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-3.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-3.samples.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-4.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-4.samples.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/report.md
ADDED
|
@@ -0,0 +1,44 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
# Step 28,000 generation evaluation
|
| 2 |
+
|
| 3 |
+
FP64 GPT-2 Large generation PPL: 111.484 +/- 5.392 standard error across five seeds.
|
| 4 |
+
|
| 5 |
+
640 generations, temperature 1 without truncation, supplied true level-zero code. TF32 disabled.
|
| 6 |
+
|
| 7 |
+
## Quality measures
|
| 8 |
+
|
| 9 |
+
| Metric | All samples | Entropy >= 4.5 |
|
| 10 |
+
|---|---:|---:|
|
| 11 |
+
| Samples | 640.000 | 182.000 |
|
| 12 |
+
| Pooled token-weighted PPL | 110.931 | 203.126 |
|
| 13 |
+
| Mean entropy (nats) | 4.381 | 4.597 |
|
| 14 |
+
| Samples with replacement character (%) | 65.312 | 62.088 |
|
| 15 |
+
| Adjacent repeated words (%) | 3.943 | 3.951 |
|
| 16 |
+
|
| 17 |
+
## First three seed-0 samples
|
| 18 |
+
|
| 19 |
+
Samples shown in original order, without selecting for quality.
|
| 20 |
+
|
| 21 |
+
### Seed 0, index 0 — PPL 126.29, entropy 3.747
|
| 22 |
+
|
| 23 |
+
64 - of. 2017 28 - February. - 13 (10 -: 8 2019 - - 5 July 2 1 2015 - - 5 5 - - 15 1015 2009 - 528 - 29 29 2014 - September 11 8. 14 -1141. 10. March 2016 - 9x0 2014 January - 6 2018 10 25 May07 - 24E 2017 2 35 - 61 pm 2019 - 4 -10 - February 7 -. June 2017 - 2017 8 8 412. 2017 - the 5 February, 2013 - 3.8, 5, - 3 2 7 - 2017 16 3.10 5 - April 2 -.. 660 - February 4an 2015. - 7 17 17 -10 - 12, July. - 8 C 8 2016 - 5 18a 9 8 December 1025 - September 1023,th the - November October17 October 2016 4.7 - the September -12 - May 14 --30 miles - 17 March - - March March 7 4 the17 6 416. 12 - December - 3 - 3 3 2015 -. -16 16 4 16 2016 -20 to to - 4 4 2017 - - the - 1201 - June 8 - 2017 6 other06 2018 June 12 AM 19 19 19 2015- 102 - 29
|
| 24 |
+
|
| 25 |
+
### Seed 0, index 1 — PPL 119.24, entropy 4.375
|
| 26 |
+
|
| 27 |
+
K about that.
|
| 28 |
+
|
| 29 |
+
That’s true, they’re not wrong, that they get. But that’s the issue to itself.”
|
| 30 |
+
|
| 31 |
+
We need you to this reason, Dallas quarterback. Ryan Smith second shot, and then that that means and with to hindsight success.
|
| 32 |
+
|
| 33 |
+
HURWRY: All all. There’s no other reason to say his reason reason are probably thinks he should be be on the issue.
|
| 34 |
+
|
| 35 |
+
“If think this the run have guy once, a guy still still throw it stuff out there it’s just just going to be be the challenge I’m running or or I� ave running running on that day — if that Portland�s case, its�s true. That used�s not the reason for it because things I do not think. But it’s not that who everything, I can’t have I not not to stop stop or get better a real challenge.”
|
| 36 |
+
|
| 37 |
+
But then think this is, guess that we a got the offense, as whole game can do.
|
| 38 |
+
|
| 39 |
+
|
| 40 |
+
STHY? Well, it would be mean, if you don� MIDt, be game game. You I
|
| 41 |
+
|
| 42 |
+
### Seed 0, index 2 — PPL 97.97, entropy 4.454
|
| 43 |
+
|
| 44 |
+
to take me very different time; but at the same time else. When I older parents would get to play out, all every time, I do that, something my life is much better than any else, anything around it, gets gets crazy for it, now it's now. There's the moment. If you just come out. It works out. If just did it, and it's always the main thing. want.. ItIt like naming, "I'd always thank for me." "�That's, 'Oh shit. How Can to it, right now is when you came out right when I was, then you really really come superb out. That could not get things great. So I was allowed to get for the work out. — "I've been in a but we was have a a sort of a done". Yeah, I just felt like something that because I was this that I got to get back. You could last was to see what I was getting in the process. So I'm not okay. At the point, want want to know that you want to let people do you, it's not obvious that you find, something else heal something is a negotiation. It's a. cool. Every person will get trust who someone
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/summarize-quality.py
ADDED
|
@@ -0,0 +1,52 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
import json
|
| 2 |
+
import math
|
| 3 |
+
from pathlib import Path
|
| 4 |
+
import re
|
| 5 |
+
import statistics
|
| 6 |
+
|
| 7 |
+
OUT = Path(__file__).resolve().parent
|
| 8 |
+
rows=[]
|
| 9 |
+
for seed in range(5):
|
| 10 |
+
result=json.loads((OUT/f'random-seed-{seed}.json').read_text())
|
| 11 |
+
assert result['reference_dtype']=='float64' and not result['tf32_allowed']
|
| 12 |
+
assert result['checkpoint_step']==28000 and result['scored_samples']==128
|
| 13 |
+
for i,text in enumerate(result['sample_generations']):
|
| 14 |
+
words=re.findall(r'\b\w+\b',text.lower())
|
| 15 |
+
row=dict(seed=seed,index=i,text=text,entropy_nats=result['per_sample_entropy_nats'][i],
|
| 16 |
+
nll_sum=result['per_sample_nll_sum'][i],scored_tokens=result['per_sample_token_count'][i],
|
| 17 |
+
ppl=result['per_sample_ppl'][i],has_replacement_character=int('\ufffd' in text),
|
| 18 |
+
adjacent_repeated_word_fraction=sum(a==b for a,b in zip(words,words[1:]))/max(1,len(words)-1))
|
| 19 |
+
assert math.isfinite(row['ppl']) and row['scored_tokens']>0
|
| 20 |
+
rows.append(row)
|
| 21 |
+
|
| 22 |
+
def metrics(group):
|
| 23 |
+
if not group:
|
| 24 |
+
return dict(samples=0)
|
| 25 |
+
return dict(samples=len(group),scored_tokens=sum(r['scored_tokens'] for r in group),
|
| 26 |
+
pooled_token_weighted_ppl=math.exp(sum(r['nll_sum'] for r in group)/sum(r['scored_tokens'] for r in group)),
|
| 27 |
+
mean_entropy_nats=statistics.mean(r['entropy_nats'] for r in group),
|
| 28 |
+
fraction_samples_with_replacement_character=statistics.mean(r['has_replacement_character'] for r in group),
|
| 29 |
+
adjacent_repeated_word_fraction=statistics.mean(r['adjacent_repeated_word_fraction'] for r in group))
|
| 30 |
+
|
| 31 |
+
accepted=[r for r in rows if r['entropy_nats']>=4.5]
|
| 32 |
+
quality=dict(checkpoint_step=28000,reference_dtype='float64',tf32_allowed=False,
|
| 33 |
+
unfiltered=metrics(rows),entropy_ge_4p5=metrics(accepted),acceptance_rate=len(accepted)/len(rows),
|
| 34 |
+
comparison_step17167=json.loads((OUT.parent/'step-17167-entropy-ge-4p5/summary.json').read_text()))
|
| 35 |
+
(OUT/'quality-summary.json').write_text(json.dumps(quality,indent=2)+'\n')
|
| 36 |
+
(OUT/'accepted-entropy-ge-4p5.json').write_text(json.dumps(accepted,indent=2)+'\n')
|
| 37 |
+
summary=json.loads((OUT/'summary.json').read_text())
|
| 38 |
+
summary.update(reference_dtype='float64',tf32_allowed=False,quality_summary=str(OUT/'quality-summary.json'),
|
| 39 |
+
generation_tf32_allowed=False,total_requested_samples=640,total_skipped_empty_samples=0)
|
| 40 |
+
(OUT/'summary.json').write_text(json.dumps(summary,indent=2)+'\n')
|
| 41 |
+
lines=['# Step 28,000 generation evaluation','',
|
| 42 |
+
f"FP64 GPT-2 Large generation PPL: {summary['mean_gen_ppl']:.3f} +/- {summary['gen_ppl_se']:.3f} standard error across five seeds.",
|
| 43 |
+
'', '640 generations, temperature 1 without truncation, supplied true level-zero code. TF32 disabled.',
|
| 44 |
+
'', '## Quality measures','', '| Metric | All samples | Entropy >= 4.5 |','|---|---:|---:|']
|
| 45 |
+
for label,key,scale in [('Samples','samples',1),('Pooled token-weighted PPL','pooled_token_weighted_ppl',1),('Mean entropy (nats)','mean_entropy_nats',1),('Samples with replacement character (%)','fraction_samples_with_replacement_character',100),('Adjacent repeated words (%)','adjacent_repeated_word_fraction',100)]:
|
| 46 |
+
lines.append(f"| {label} | {quality['unfiltered'][key]*scale:.3f} | {quality['entropy_ge_4p5'].get(key,float('nan'))*scale:.3f} |")
|
| 47 |
+
lines+=['','## First three seed-0 samples','', 'Samples shown in original order, without selecting for quality.']
|
| 48 |
+
for row in rows[:3]:
|
| 49 |
+
lines+=['',f"### Seed {row['seed']}, index {row['index']} — PPL {row['ppl']:.2f}, entropy {row['entropy_nats']:.3f}",'',row['text']]
|
| 50 |
+
(OUT/'report.md').write_text('\n'.join(lines)+'\n')
|
| 51 |
+
print(json.dumps(summary,indent=2))
|
| 52 |
+
print(json.dumps({k:v for k,v in quality.items() if k!='comparison_step17167'},indent=2))
|
iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/summary.json
ADDED
|
@@ -0,0 +1,41 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"protocol": {
|
| 3 |
+
"sampling": "random",
|
| 4 |
+
"temperature": 1.0,
|
| 5 |
+
"top_k": 0,
|
| 6 |
+
"top_p": 1.0,
|
| 7 |
+
"n_provided_levels": 1,
|
| 8 |
+
"n_levels": 16,
|
| 9 |
+
"document_aware_validation_sampling": true
|
| 10 |
+
},
|
| 11 |
+
"reference_model": "gpt2-large",
|
| 12 |
+
"checkpoint": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/exports/step-28000/ncp.pt",
|
| 13 |
+
"checkpoint_step": 28000,
|
| 14 |
+
"seeds": [
|
| 15 |
+
0,
|
| 16 |
+
1,
|
| 17 |
+
2,
|
| 18 |
+
3,
|
| 19 |
+
4
|
| 20 |
+
],
|
| 21 |
+
"samples_per_seed": 128,
|
| 22 |
+
"total_scored_samples": 640,
|
| 23 |
+
"total_scored_tokens": 160778,
|
| 24 |
+
"per_seed_mean_ppl": {
|
| 25 |
+
"0": 124.87831485116675,
|
| 26 |
+
"1": 114.46117735800877,
|
| 27 |
+
"2": 119.77330723462428,
|
| 28 |
+
"3": 102.35317455545147,
|
| 29 |
+
"4": 95.95460863688895
|
| 30 |
+
},
|
| 31 |
+
"mean_gen_ppl": 111.48411652722804,
|
| 32 |
+
"gen_ppl_se": 5.392206594348971,
|
| 33 |
+
"mean_of_seed_median_ppl": 105.71566681551471,
|
| 34 |
+
"mean_entropy_nats": 4.381471981070202,
|
| 35 |
+
"reference_dtype": "float64",
|
| 36 |
+
"tf32_allowed": false,
|
| 37 |
+
"quality_summary": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/quality-summary.json",
|
| 38 |
+
"generation_tf32_allowed": false,
|
| 39 |
+
"total_requested_samples": 640,
|
| 40 |
+
"total_skipped_empty_samples": 0
|
| 41 |
+
}
|
iclr-debug-decoder-135b-20260916/decoder-mstok/milestone-iter-28000.pt
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:d9eb43b22050308d6a75176a2978dadc8b0eeba38d905afa73521385e4a373c0
|
| 3 |
+
size 4879455521
|
iclr-debug-decoder-135b-20260916/decoder-mstok/teacher-provenance.json
ADDED
|
@@ -0,0 +1,6 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"teacher_id": "FacebookAI/roberta-base",
|
| 3 |
+
"teacher_revision": "e2da8e2f811d1448a5b465c236feacd80ffbac7b",
|
| 4 |
+
"tokenizer_revision": "607a30d783dfa663caf39e06633721c8d4cfcd7e",
|
| 5 |
+
"token_map_sha256": "0d5aa4e0a98157722f03cf9a0c7a1f17178c370dd728896090077ea9efe0b593"
|
| 6 |
+
}
|
iclr-debug-decoder-135b-20260916/git-provenance.json
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"commit": "1b5226d4bbdd34653cb5fdd58d436355fb76ac8c"
|
| 3 |
+
}
|
iclr-debug-decoder-135b-20260916/manifest.json
ADDED
|
@@ -0,0 +1,525 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"format": "iclr-debug-production-v1",
|
| 3 |
+
"variant": "decoder-mstok",
|
| 4 |
+
"configs": {
|
| 5 |
+
"decoder-mstok": {
|
| 6 |
+
"task": "mstok-next-concept",
|
| 7 |
+
"experiment": "iclr-debug-decoder-135b-20260916-decoder-mstok",
|
| 8 |
+
"experiment_dir": "/home/ubuntu/mstok-results/iclr-debug-decoder-135b-20260916/decoder-mstok",
|
| 9 |
+
"dataset": {
|
| 10 |
+
"train_source": "/home/ubuntu/data/full_owt/train_gpt2.bin",
|
| 11 |
+
"validate_source": "/home/ubuntu/data/full_owt/valid_gpt2.bin",
|
| 12 |
+
"pad_token_id": 50257
|
| 13 |
+
},
|
| 14 |
+
"pre_tokenizer": {
|
| 15 |
+
"name": "hf",
|
| 16 |
+
"tokenizer": "hf",
|
| 17 |
+
"model_id": "gpt2",
|
| 18 |
+
"special_tokens": {
|
| 19 |
+
"eos_token": "<|endoftext|>",
|
| 20 |
+
"pad_token": "<|pad|>"
|
| 21 |
+
}
|
| 22 |
+
},
|
| 23 |
+
"codec": {
|
| 24 |
+
"n_layers": 6,
|
| 25 |
+
"context_length": 256,
|
| 26 |
+
"embed_dim": 384,
|
| 27 |
+
"in_vocab_size": 50304,
|
| 28 |
+
"vqvae_vocab_size": [
|
| 29 |
+
16384,
|
| 30 |
+
16384,
|
| 31 |
+
16384,
|
| 32 |
+
16384,
|
| 33 |
+
16384,
|
| 34 |
+
16384,
|
| 35 |
+
16384,
|
| 36 |
+
16384,
|
| 37 |
+
16384,
|
| 38 |
+
16384,
|
| 39 |
+
16384,
|
| 40 |
+
16384,
|
| 41 |
+
16384,
|
| 42 |
+
16384,
|
| 43 |
+
16384,
|
| 44 |
+
16384
|
| 45 |
+
],
|
| 46 |
+
"compression_factor": 1,
|
| 47 |
+
"dropout": 0.1,
|
| 48 |
+
"pre_quant_groupnorm": 4,
|
| 49 |
+
"pre_quant_dropout": 0.2,
|
| 50 |
+
"vector_quantizer_config": {
|
| 51 |
+
"clss": "multiscale_residual_vector_quantizer",
|
| 52 |
+
"decay": 0.99,
|
| 53 |
+
"epsilon": 1e-05,
|
| 54 |
+
"commitment_cost": 0.25,
|
| 55 |
+
"learned_l1_sampling": true,
|
| 56 |
+
"learned_all_sampling": false,
|
| 57 |
+
"quant_resi": {
|
| 58 |
+
"enabled": true,
|
| 59 |
+
"ratio": 0.5,
|
| 60 |
+
"share_mode": 0,
|
| 61 |
+
"learnable_ratio": false
|
| 62 |
+
},
|
| 63 |
+
"levels": {
|
| 64 |
+
"use_manual_levels": true,
|
| 65 |
+
"manual_levels": [
|
| 66 |
+
1,
|
| 67 |
+
4,
|
| 68 |
+
9,
|
| 69 |
+
16,
|
| 70 |
+
25,
|
| 71 |
+
36,
|
| 72 |
+
49,
|
| 73 |
+
64,
|
| 74 |
+
81,
|
| 75 |
+
100,
|
| 76 |
+
121,
|
| 77 |
+
144,
|
| 78 |
+
169,
|
| 79 |
+
196,
|
| 80 |
+
225,
|
| 81 |
+
256
|
| 82 |
+
]
|
| 83 |
+
},
|
| 84 |
+
"aux": {
|
| 85 |
+
"fine_drop": {
|
| 86 |
+
"prob": 0.5,
|
| 87 |
+
"min_keep": 1
|
| 88 |
+
}
|
| 89 |
+
}
|
| 90 |
+
}
|
| 91 |
+
},
|
| 92 |
+
"generator": {
|
| 93 |
+
"n_layer": 12,
|
| 94 |
+
"n_head": 12,
|
| 95 |
+
"bias": true,
|
| 96 |
+
"dropout": 0.1,
|
| 97 |
+
"n_embd": 768,
|
| 98 |
+
"context_length": 256,
|
| 99 |
+
"vocab_size": [
|
| 100 |
+
16384,
|
| 101 |
+
16384,
|
| 102 |
+
16384,
|
| 103 |
+
16384,
|
| 104 |
+
16384,
|
| 105 |
+
16384,
|
| 106 |
+
16384,
|
| 107 |
+
16384,
|
| 108 |
+
16384,
|
| 109 |
+
16384,
|
| 110 |
+
16384,
|
| 111 |
+
16384,
|
| 112 |
+
16384,
|
| 113 |
+
16384,
|
| 114 |
+
16384
|
| 115 |
+
],
|
| 116 |
+
"attn_pattern": "block_diagonal",
|
| 117 |
+
"use_positional_encoding": true,
|
| 118 |
+
"use_level_encoding": true,
|
| 119 |
+
"prefix_len": 0,
|
| 120 |
+
"use_rope": true,
|
| 121 |
+
"rope_base": 10000,
|
| 122 |
+
"use_qk_norm": true,
|
| 123 |
+
"post_upsample_conv": {
|
| 124 |
+
"enabled": true,
|
| 125 |
+
"kernel_size": 3
|
| 126 |
+
},
|
| 127 |
+
"use_relu2": true,
|
| 128 |
+
"use_flex_attention": true,
|
| 129 |
+
"shared_output_head": true,
|
| 130 |
+
"shared_head_per_level_bias": false,
|
| 131 |
+
"shared_head_per_level_scale": false,
|
| 132 |
+
"shared_head_adapter_rank": 0
|
| 133 |
+
},
|
| 134 |
+
"objectives": {
|
| 135 |
+
"codec_weight": 1.0,
|
| 136 |
+
"ncp_weight": 1.0,
|
| 137 |
+
"soft_assignment_temperature": 1.0,
|
| 138 |
+
"prediction_temperature": 1.0,
|
| 139 |
+
"mstok_weight": 0.25,
|
| 140 |
+
"residual_weight": 1.0,
|
| 141 |
+
"reconstruction_weight": 1.0,
|
| 142 |
+
"teacher_ema_decay": 0.999,
|
| 143 |
+
"teacher_ema_warmup_steps": 1000,
|
| 144 |
+
"mstok_warmup_steps": 500
|
| 145 |
+
},
|
| 146 |
+
"optimization": {
|
| 147 |
+
"codec_lr": 0.001,
|
| 148 |
+
"codec_min_lr": 0.0001,
|
| 149 |
+
"codec_warmup_iters": 0,
|
| 150 |
+
"codec_lr_decay_iters": 171662,
|
| 151 |
+
"generator_lr": 0.0005,
|
| 152 |
+
"generator_min_lr": 1e-05,
|
| 153 |
+
"generator_warmup_iters": 300,
|
| 154 |
+
"generator_lr_decay_iters": 171662,
|
| 155 |
+
"beta_1": 0.9,
|
| 156 |
+
"codec_beta_2": 0.99,
|
| 157 |
+
"generator_beta_2": 0.99,
|
| 158 |
+
"weight_decay": 0.1,
|
| 159 |
+
"max_grad_norm": 1.0,
|
| 160 |
+
"level_loss_alpha": 1.0,
|
| 161 |
+
"grad_accumulation_steps": 12,
|
| 162 |
+
"corruption": {
|
| 163 |
+
"mode": "per_level",
|
| 164 |
+
"per_level_probs": [
|
| 165 |
+
0.85,
|
| 166 |
+
0.8321428571,
|
| 167 |
+
0.8142857143,
|
| 168 |
+
0.7964285714,
|
| 169 |
+
0.7785714286,
|
| 170 |
+
0.7607142857,
|
| 171 |
+
0.7428571429,
|
| 172 |
+
0.725,
|
| 173 |
+
0.7071428571,
|
| 174 |
+
0.6892857143,
|
| 175 |
+
0.6714285714,
|
| 176 |
+
0.6535714286,
|
| 177 |
+
0.6357142857,
|
| 178 |
+
0.6178571429,
|
| 179 |
+
0.6
|
| 180 |
+
],
|
| 181 |
+
"skip_level0": true
|
| 182 |
+
}
|
| 183 |
+
},
|
| 184 |
+
"training": {
|
| 185 |
+
"log_interval": 10,
|
| 186 |
+
"eval_interval": 1000,
|
| 187 |
+
"checkpoint_interval": 1000,
|
| 188 |
+
"eval_batch_size": 4,
|
| 189 |
+
"val_iters": 25,
|
| 190 |
+
"keep_last": 3,
|
| 191 |
+
"seed": 55,
|
| 192 |
+
"codec_initialization_seed": 42,
|
| 193 |
+
"generator_initialization_seed": 55,
|
| 194 |
+
"total_iters": 171662,
|
| 195 |
+
"batch_size": 32,
|
| 196 |
+
"expected_world_size": 8,
|
| 197 |
+
"expected_global_batch_size": 3072,
|
| 198 |
+
"resume_checkpoint": null,
|
| 199 |
+
"milestone_steps": [
|
| 200 |
+
1272,
|
| 201 |
+
6358,
|
| 202 |
+
17167,
|
| 203 |
+
34333,
|
| 204 |
+
68665,
|
| 205 |
+
102997,
|
| 206 |
+
137330,
|
| 207 |
+
171662
|
| 208 |
+
],
|
| 209 |
+
"generation_eval_steps": [
|
| 210 |
+
17167,
|
| 211 |
+
34333,
|
| 212 |
+
68665,
|
| 213 |
+
102997,
|
| 214 |
+
137330,
|
| 215 |
+
171662
|
| 216 |
+
]
|
| 217 |
+
},
|
| 218 |
+
"torch_compile": {
|
| 219 |
+
"enable": true,
|
| 220 |
+
"scope": "loss",
|
| 221 |
+
"mode": "default",
|
| 222 |
+
"dynamic": false,
|
| 223 |
+
"fullgraph": true,
|
| 224 |
+
"backend": "inductor"
|
| 225 |
+
},
|
| 226 |
+
"mixed_precision": {
|
| 227 |
+
"enable": true
|
| 228 |
+
},
|
| 229 |
+
"wandb": {
|
| 230 |
+
"entity": "mstok",
|
| 231 |
+
"project": "iclr-debug",
|
| 232 |
+
"group": "full-owt-ctx256-16sq-135b-v1",
|
| 233 |
+
"enable": true,
|
| 234 |
+
"id": "iclr-debug-decoder-135b-20260916-decoder-mstok",
|
| 235 |
+
"resume": "allow",
|
| 236 |
+
"gradients_and_params": {
|
| 237 |
+
"enable": false,
|
| 238 |
+
"log": "all",
|
| 239 |
+
"log_freq": 1000
|
| 240 |
+
}
|
| 241 |
+
},
|
| 242 |
+
"export": {
|
| 243 |
+
"vqvae_config_template": "/home/ubuntu/mstok-runs/iclr-debug-joint-20260916/config/repro-ctx256/vqvae.yaml",
|
| 244 |
+
"ncp_config_template": "/home/ubuntu/mstok-runs/iclr-debug-joint-20260916/config/repro-ctx256/ncp-sharedhead.yaml"
|
| 245 |
+
},
|
| 246 |
+
"semantic": {
|
| 247 |
+
"enabled": true,
|
| 248 |
+
"mode": "eostok",
|
| 249 |
+
"weight": 0.5,
|
| 250 |
+
"warmup_steps": 500,
|
| 251 |
+
"teacher_id": "FacebookAI/roberta-base",
|
| 252 |
+
"teacher_revision": "e2da8e2f811d1448a5b465c236feacd80ffbac7b",
|
| 253 |
+
"tokenizer_revision": "607a30d783dfa663caf39e06633721c8d4cfcd7e",
|
| 254 |
+
"teacher_dim": 768,
|
| 255 |
+
"feature_layer": 6,
|
| 256 |
+
"projector_dim": 2048,
|
| 257 |
+
"projector_seed": 56,
|
| 258 |
+
"alignment_site": "decoder"
|
| 259 |
+
},
|
| 260 |
+
"iclr_debug": {
|
| 261 |
+
"stage": "decoder-mstok",
|
| 262 |
+
"target_positions": 135000000000,
|
| 263 |
+
"lr_horizon_positions": 135000000000,
|
| 264 |
+
"stop_step": 171662,
|
| 265 |
+
"pilot": false,
|
| 266 |
+
"tokenizer_checkpoint": null,
|
| 267 |
+
"tokenizer_steps": null,
|
| 268 |
+
"budget_unit": "input positions including padding; log non-padding tokens separately"
|
| 269 |
+
}
|
| 270 |
+
}
|
| 271 |
+
},
|
| 272 |
+
"sources": {
|
| 273 |
+
"config/repro-ctx256/alignment-generator.yaml": "f62f36d078a3a373a19ea716a3ebc7eda5ae2fb7d670c5b935c734826bb76895",
|
| 274 |
+
"config/repro-ctx256/alignment-joint.yaml": "bb511950d20335e4237e406464ab26760a7fd9c9c9472c0507addd88d33d5178",
|
| 275 |
+
"config/repro-ctx256/alignment-tokenizer.yaml": "2d32955296f7d15990f284328c1cf8f52071abdba28033a65e272bb49cf5dd56",
|
| 276 |
+
"config/repro-ctx256/alignment.yaml": "3a486fa2ed28eca4fc406232669399b0625d274376ac43271f4ccf187b1e9077",
|
| 277 |
+
"config/repro-ctx256/joint.yaml": "8b4889111687a52e2d7d71995bda45a7f01f2777674c88191f338dca4c1732d2",
|
| 278 |
+
"config/repro-ctx256/mstok-semantic-eostok-matched10ep.yaml": "ae8a72c0c9f1cb5d2e301d83a1cdc08985a1f895bf627a5f75b1b3ce627f5e58",
|
| 279 |
+
"config/repro-ctx256/mstok-semantic-eostok.yaml": "2c9c48b21325c0e479990701b28116d78f611101724846747d173e52145fe15a",
|
| 280 |
+
"config/repro-ctx256/mstok-semantic-gear-matched10ep.yaml": "164cb8ac39fa407bb337f559ad8a9c62a34708fb900b8829d3ccf182961f2d0c",
|
| 281 |
+
"config/repro-ctx256/mstok-semantic-gear.yaml": "8cb8c998551daef0bf7fd9ec77b13b34c06662028d8820767e1432b5a9e8be54",
|
| 282 |
+
"config/repro-ctx256/mstok-semantic.yaml": "e0b00dbfa3eb838103f657c19bde20e36375c7fbe692623452ea093003f1ac98",
|
| 283 |
+
"config/repro-ctx256/mstok-w1-pilot.yaml": "21ee22eeccaaea2563c92ccc84f6df3871ea5da6a5ce980656088ea577dc1438",
|
| 284 |
+
"config/repro-ctx256/mstok-w1.yaml": "b6bf6079114b3c997fab89e85a53318bb2ec5dae55678513599d84152eab701b",
|
| 285 |
+
"config/repro-ctx256/mstok.yaml": "d9bb216f4eeb58857d72be399d5f8847d197ce70d0e504879f0788045483dc87",
|
| 286 |
+
"config/repro-ctx256/ncp-sharedhead.yaml": "86c9efbcb2b4f74b7253aa3bc4393de4a6a05e20ea30da573d1ca094ec600645",
|
| 287 |
+
"config/repro-ctx256/ncp.yaml": "93663654c4ce5ad0931db052edf507fdcb7c0ff9fc9722f4018820deee5c5acf",
|
| 288 |
+
"config/repro-ctx256/substitution-generator.yaml": "064f55887b9b2b8f3b0b52c094d4a193510dc1148ad7a5c7f53e317971f87c7c",
|
| 289 |
+
"config/repro-ctx256/substitution-joint.yaml": "b1576f4ecfe6fc6029a0cd868271b80a1d988b7f3667c6b1a8d24a6a4015df3b",
|
| 290 |
+
"config/repro-ctx256/substitution-tokenizer.yaml": "e29be5046f0fafbfe4c06973be4cd036a004e845348911cc0e6e05a7fca5d425",
|
| 291 |
+
"config/repro-ctx256/substitution.yaml": "ebba1c7e664de0e84913bdfd38e794837781c42c1b48cde90135ff9417856a95",
|
| 292 |
+
"config/repro-ctx256/vqvae.yaml": "05fde46e57c3e6b411e8a6afcac251f178e4b8a0856a0b99332bb284fb4710ce",
|
| 293 |
+
"data/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 294 |
+
"data/serialized_dataset.py": "5622c2d94236cd1b9b4ca18e8d5c272a86b54d193d99048fb9b7cd2687e8dbc6",
|
| 295 |
+
"data/tinysentences/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 296 |
+
"data/tinysentences/clean.py": "918a9a91aae484a5e06d6d212c2c0974946a4032e061c62a7b65851735bfc8dd",
|
| 297 |
+
"data/tinysentences/clean_parsed.py": "009a087f967e171bc6098484dbfd58635302900f6b07629629898f1d00fa70a3",
|
| 298 |
+
"data/tinysentences/generate_meta.py": "662714f1370e05567d412add168a7c3a05f4b1c9adb047fad374900bc1936a7e",
|
| 299 |
+
"data/tinysentences/generate_metadata_char.py": "afd3336ea03ff0653ccaf41c121d9b93debf7a380db6cd881b62b308ee6ee68c",
|
| 300 |
+
"data/tinysentences/generate_training_data_word.py": "82e2edcb7ea4db0609ccf2b047a0b207e7b8fcff4fc0b4ceee521ff278e8a8e8",
|
| 301 |
+
"data/tinysentences/parse_summaries.py": "2317159364dbbe7206f252d2939cab5d375395e0a536be0a1f60c5df88b31702",
|
| 302 |
+
"data/tinysentences/prepare.py": "35263b287f93337fda4207ed1991e9e8a820ab75139411bb7dd5f0a8a25bcbd9",
|
| 303 |
+
"data/tinysentences/sentences_only.py": "5e4d72bf612df27dad46cf606a40402606d897043ae24274dc2dfa88ec0b3e79",
|
| 304 |
+
"data/tinysentences/stats.py": "2665e9e6fac4a8f6c3e5bd8b07cbfd941ba6359a43981d89bdef58c136fdd3a9",
|
| 305 |
+
"data/tinysentences/summarize/merge_summaries_char.py": "bce94f4d9a824c2d25e45448023d59b7fbc42aba22b33efbc067cdbba93e1934",
|
| 306 |
+
"data/tinysentences/summarize/merge_summaries_word.py": "f6079110332cc38b2dd3571c30bedcca97098c39498f45eec1ef6e60cce7ac81",
|
| 307 |
+
"data/tinysentences/summarize/summarize_gemini.py": "282a438efce0dee351fd9833a59c56e4480bed4314dae40844201794ada3bd96",
|
| 308 |
+
"data/tinysentences/summarize/summarize_gpt.py": "fc08cbc37d6161e5b0057a1ede43272d96f7274524917d6c4d42c22f505c6e09",
|
| 309 |
+
"data/tinysentences/summarize/summarize_llama.py": "7f3f8230652a5b88ea2856b273d497cd5b633580b430c909e4f996cd0232832d",
|
| 310 |
+
"data/tinysentences/summarize/summarize_qwen.py": "b07e3c226ab0fdf5b21b025263adfa7c2625d640f0fac3c787bc4a2c9e32107f",
|
| 311 |
+
"data/tinysentences/summarize/summarize_together.py": "8c26ea052f2a64c2653d39f1ea5c13983dfbc9547f21cd7d7ea9d9b39033bffe",
|
| 312 |
+
"data/tinysentences/test.py": "9749244da77c078f550a78751664b45bb92067682f3aa27c2338419f1186c844",
|
| 313 |
+
"data/tinysentences/test_regex.py": "46652e189a0d5f5c0c7a2c90b857082004bf6727ef1d3ad42b631ac7bb0aa88f",
|
| 314 |
+
"data/tinysentences/tinysentences_char_dataset.py": "152ab9c30aec669a9db14df9e90412597870105eff9cfb737b876aad2323a206",
|
| 315 |
+
"data/tinysentences/tinysentences_dataset.py": "f90251d21871f6a7ad5ebd2ad0a6e59cd0b02ca245f7e46f93e6d970bf657bc7",
|
| 316 |
+
"data/tinysentences/tinysentences_vae.py": "30f7adde5413348faf5066448706747fe26a7fedd510ffea720aa0bf8283f6f8",
|
| 317 |
+
"data/tinysentences/to_csv.py": "c06c74648ed95b4494835c631c94fe1a00a9ff47a272f5e38546a7f94ff9be8b",
|
| 318 |
+
"data/tinysentences/tokenize/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 319 |
+
"data/tinysentences/tokenize/common.py": "e877c74df2a48d15991492b126cd0c5fe26731d61f3aa36d83ebf2cb9bda34a0",
|
| 320 |
+
"data/tinysentences/tokenize/tokenizer.py": "c16abc50b92a163f6978542b36b03c9a14aa00153c73117c4d52084d5de25c2a",
|
| 321 |
+
"data/tinysentences/tokenize/train_bpe.py": "fa8e55978dd5ca90117200648b468fb22822bf6981f15f0394c815e72eae1bdb",
|
| 322 |
+
"data/tinystories/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 323 |
+
"data/tinystories/download.py": "5afce37e912f6965b6722f0bc028076a655fea78304f1a213e4862dbe223cdcb",
|
| 324 |
+
"data/tinystories/stats.py": "216650fe19b78e8431c4512ff87a5c84af64d9e56ab8fdaf3952ca035725ea1b",
|
| 325 |
+
"data/wikitext/parse_wikitext.py": "ab583e0328f82afc3ba242410e8076be42ef1a16a51d66f3eb46a713cca591d0",
|
| 326 |
+
"evaluation/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 327 |
+
"evaluation/eval_random_gen_ppl_checkpoint.py": "221ea1917dc2d3bc19be39da0e0b8cca49295d0400a05da276cd50956876da8d",
|
| 328 |
+
"evaluation/gen_ppl.py": "e6853b14d27f93407505eb7607a1bd9fed49a7add1511ceca2b4a3076d533859",
|
| 329 |
+
"evaluation/gen_ppl_ncp.py": "96d68adee372e5186d0470659171eb03829e3d37f611e8f5917b45fb61e3a4ad",
|
| 330 |
+
"evaluation/gen_ppl_ntp.py": "b2638622bc041fa88e9d538be4229b3723ebec19c5b99fd6e5b73db4776dc952",
|
| 331 |
+
"evaluation/guess.py": "c27803225dd21b2deff1e7e494b9e3666180b93776340164ce6e953dacc50656",
|
| 332 |
+
"evaluation/ppl.py": "f3b581cb4849b0b2aa518729b8338838c7055697c3a26028dedb819c335ac838",
|
| 333 |
+
"evaluation/run_genppl_multiseed.py": "f8c9ac982f4caa4335b5f4d1a22d70f56dc11a706106a3dfb31d344eb8c93c2e",
|
| 334 |
+
"evaluation/run_genppl_release_ts.py": "d1bc5c00089118d9cf6c05c77459329ad40c2f8bc347b5aeedda8f5318de8796",
|
| 335 |
+
"evaluation/run_genppl_sweep.py": "ade4f39e373125848fdd32938e8d7bf86a0ec5108cca9bd75f0a966d092a798e",
|
| 336 |
+
"evaluation/show_8k_samples.py": "30ad825d5a24375e178750dfaaac45bc486bc3587f2e7d6d305d53bf37f8d226",
|
| 337 |
+
"evaluation/show_low_sample.py": "d696f1e54ca67e87565dce5864e0e6e2ad4194108cfb1c9c1e06eebd97008da6",
|
| 338 |
+
"evaluation/summarize_gen_ppl.py": "307f3d0d85884879295d5268aaf7f399f654b1b330aaa5728d6824ad24c6d15a",
|
| 339 |
+
"evaluation/sweep_gen_ppl_checkpoints.py": "3a9141e27d84b49c5493194dfc5a70bfa0a915cdf5ceb7a46afcf27cc6f33acb",
|
| 340 |
+
"evaluation/sweep_gen_ppl_ncp.py": "d91c2959cfed51d224ccb1c0d0f2c3b9a35c4027cc9ab8c65aa7603d587de761",
|
| 341 |
+
"evaluation/sweep_gen_ppl_ncp_scaling.py": "05249c1e00c5eb05042b6ea5b1ca008ecb8f0e4a77537ee755a27c153c25a17c",
|
| 342 |
+
"evaluation/sweep_gen_ppl_ncp_yolo.py": "4b26ef25b27af2f7fcbcb5ff3c312d47d2a4746ba759155d28802e978a681755",
|
| 343 |
+
"evaluation/sweep_gen_ppl_ntp.py": "14a04074741ee44580ec6557870d42534dd172d26c5be3e943ed5baf1b727303",
|
| 344 |
+
"evaluation/sweep_random_llama_yolo.py": "4062b54289fcd9407d15feb1ffd241b8722fd717c700350c22472fecec5ae46c",
|
| 345 |
+
"evaluation/val_oracle_ppl.py": "1a23f3c2897b6a6a3d239ed533979de35352d303bae094f4006bef4156f935c0",
|
| 346 |
+
"evaluation/val_vqvae_ppl.py": "4c293c502d3faf982d4ad66566b7a037faa7d28191d929de1621cb67f5ae8b5f",
|
| 347 |
+
"models/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 348 |
+
"models/common.py": "99aa0da356963a282936413d738ee1da72a3e3f277fecc28eddb553051582a4b",
|
| 349 |
+
"models/masked_lm.py": "80279614d74efc7b4c25306789d74e2f71045c54e4e4f485ee365f664705d7f6",
|
| 350 |
+
"models/next_concept.py": "743667aac3bdceb7bb7476a580700185b2638ddd9575e6357db57acc7d699339",
|
| 351 |
+
"models/next_token.py": "d0d292b9f5c032a351cbca1a8f0a6d41c3062890cb3e3f5b67d4e23a7cffeb77",
|
| 352 |
+
"models/quant.py": "b7b73abd4e55e10454cb680dea82124c38c222d37236c4c76b5c7d21123ac1b1",
|
| 353 |
+
"models/semantic_input.py": "b06658d4bc641397cbb973255eadfd2d6087fde244d801d3c70d29e3ce841bf2",
|
| 354 |
+
"models/vae.py": "02e2213a8878920374baf27dd10dfc03847dce866557c97e816c427d230c9ded",
|
| 355 |
+
"models/vqvae.py": "8eb3503eb1d990121b7d6138a4f8216aecfc72246d4a92bbecf98025c8a3d504",
|
| 356 |
+
"scripts/benchmark_iclr_debug.py": "dba86e8d578eec2188c633a479ccf673d58f3e5580b9074869cae1a352b8a043",
|
| 357 |
+
"scripts/benchmark_mstok_compile.py": "aad909da6f05ba25ac3d8f6b3cf450103bed55d0b93b1490764c0ef54f35b606",
|
| 358 |
+
"scripts/eval_owt_ckpts_gpt2_large.py": "4977a2d614698f60032fb6fd7fef614eb503d5b3883f499146a4fc048214a509",
|
| 359 |
+
"scripts/evaluate_mstok_sidecar.py": "9ba33e865fa0fad34bd2ab375d7fe38f37c441e2d96348f0e389e29dd9643455",
|
| 360 |
+
"scripts/evaluations/sweep_gen_ppl_shared_vocab_ncp.py": "2311ac71f9beba282426faa21390a5898c7d183a95f529755d6e1d6e8fc24050",
|
| 361 |
+
"scripts/neurips/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 362 |
+
"scripts/neurips/eval_gen_ppl_ncp_owt.py": "b92eff1ce198eb4ac8c88da7113b41d095c12acad92db541685175993db3535c",
|
| 363 |
+
"scripts/neurips/sweep_gen_ppl_owt_gpt2.py": "02adb1629cfa40ff51a75152e4da9c9ee00ead42b7ecda9e78cf7101e6826a67",
|
| 364 |
+
"scripts/preflight_owtsmall_mstok_ctx256.py": "a4f80921435b6cd136d089ec39bfa5442359d0dffbeefaaa1019178b8364657a",
|
| 365 |
+
"scripts/prepare_full_owt.py": "b1cb6b74397625c2b0643c4ee744ca3c89dfa6bbb7eeb7e9681e3a36f91a2c24",
|
| 366 |
+
"scripts/r1_lab/diagnose_at_40k.py": "b1582dcc29ccebc9255eaa77a7d56a7dbdef589ec5d5a9b0af150dede98874ab",
|
| 367 |
+
"scripts/r1_lab/inspect_summaries.py": "4154e8a6c1b3f68df5c623a57a1a7632b594c46c5f6d1d85cb08c476d6aa9cea",
|
| 368 |
+
"scripts/r1_lab/overnight_pipeline.py": "53f4e03ce1a5fb43fbca3ae455288f45ae8350b0a16f3d5318763011de6a8f96",
|
| 369 |
+
"scripts/r1_lab/prime_rope_qknorm_score_cache.py": "a791cdec03b415bfccdfd8d89c7e659aff8a2a5696d16ab7b4350ff608d0391b",
|
| 370 |
+
"scripts/r1_lab/sweep_gen_ppl_r1lab.py": "c7c0acf9aa6f07bb0a985cf0d5e5692c696e658ebb0470b7c92d9332565f9c06",
|
| 371 |
+
"scripts/r1_lab/sweep_gen_ppl_r1lab_v15.py": "524aa8680aa1b9fea2689509848c19e5a74a3ad72e78b0388bbca7acd43be22e",
|
| 372 |
+
"scripts/r1_lab/sweep_gen_ppl_rope_qknorm.py": "7822df057355b0a3a71ea5145742107c6e30f906859096237ff61d765fe36d50",
|
| 373 |
+
"scripts/r1_lab/train_variant.py": "9d0a2091784609ac9728226afd82f56fd4c1ac2dc08e89e1b17b60d34bb3d55e",
|
| 374 |
+
"scripts/run_alignment.py": "6312cd9d0976899763589fdcbd258c7681f6e9e96d7396cf35ff022d4695b862",
|
| 375 |
+
"scripts/run_iclr_debug.py": "a48a14067af036a0eba46d8de69922c615c05f6bbbab11587dca292ea57682ce",
|
| 376 |
+
"scripts/run_mstok_semantic.py": "8a7a38a50dd6694950bfa0e7b79718ec653c7e70be99fe13cbf2ee05c28f46a0",
|
| 377 |
+
"scripts/run_mstok_w1_pilot.py": "35224e462d0d11eb3e4326b0daa8bad7c99e1e30cdc6383310f8e2c8b5bb708f",
|
| 378 |
+
"scripts/run_substitution.py": "2cd22eaa58dab8b5231984eee340a3a2bdd0b88cc83333f7133447471e7f73f1",
|
| 379 |
+
"scripts/smoke_test_masked_l0.py": "d24f1270b3219a31212b68bcbb786f649c7c72f96e53f602d87ad6c17f3f73e5",
|
| 380 |
+
"scripts/test0_diagnostics.py": "ed389c666beef015ccac004f314a5a1428def303b46ac0e442f0236337344549",
|
| 381 |
+
"scripts/test0_r1_similarity.py": "4a486878bebd418a3b5454ddbf02585726ccbacc2700fdf591aa524766f9c9c9",
|
| 382 |
+
"scripts/test0_r1_uniqueness.py": "3ee59d4397d5b268151dbb8ca779c7f9bf8461d129d97fcd1a517c48ec2a69a1",
|
| 383 |
+
"scripts/test0_residual_norms.py": "3c4f68e5960e84796dc0bca4fd4583bf3b598552654f9b565fd5fcabd420f9e1",
|
| 384 |
+
"scripts/test_gen_equiv.py": "ae0e9c513422ee58bdc050ed6218ee54fc555dfb15ea8968c9ffc62a7a7ea99a",
|
| 385 |
+
"scripts/training_budget.py": "bb94c19c6dfbdd542946b10476b8d6958b8a189abd141cf62288fdb5fa9a2d38",
|
| 386 |
+
"scripts/vqvae_utilization.py": "5f90327f1507a5042b2f77b119fe4fc0df4e2131e5bbc21069dffd1fadc059c5",
|
| 387 |
+
"scripts/xsum/codec_ceiling.py": "ddb751e9152d5c88966e588d89cdceee4c5d91c4139aaedf369f65a45ecf2276",
|
| 388 |
+
"scripts/xsum/eval_merge.py": "6e313d6ed25227d0418b7422d58e09ceda4fc6dfae7d0fd663417b306241135c",
|
| 389 |
+
"scripts/xsum/eval_q0.py": "82fe04dccce28d7fc9c6ec8b4c9c91404a49527f9437913f44f51a6853ba9abc",
|
| 390 |
+
"scripts/xsum/eval_shard.py": "07192f6ba6219da1e8123ceeb99748c86da85322859ed5fa25cebcc3816f7ea8",
|
| 391 |
+
"scripts/xsum/eval_shard_cfg.py": "95ac4a8e2e3a73a21fa4303c62df80c086f23bcfe5ab34eb07ae90b585052ecb",
|
| 392 |
+
"scripts/xsum/eval_test.py": "7f92d578c31893e788e5fa7d9eef3098ba5602b3642873ce2ce031342c8be868",
|
| 393 |
+
"scripts/xsum/eval_uncond_rep.py": "ba4740b29bc091d866e2dfd91881e9a7d5c46a1bfa993d13eb63c0467397ec28",
|
| 394 |
+
"scripts/xsum/make_warmstart.py": "78793d6ef11fbf658c2837ef12cb843fc4817add8ba73c822b5e408901de3318",
|
| 395 |
+
"scripts/xsum/mbr_fast.py": "3db6e21d2dcce2e3bcf6f22474535d7eba0232b071b4271d2b172f787aed5c85",
|
| 396 |
+
"scripts/xsum/mbr_select.py": "107a061014c62395b7f5fd986fba9e55dd738054e2474022005cb7e9e0c9b8fb",
|
| 397 |
+
"scripts/xsum/mbr_shard_cfg.py": "3e7162375028c391276149eff288e9de40c874a3108ac53de0248198f4d08791",
|
| 398 |
+
"scripts/xsum/pool_append.py": "e50252f48c4d8d905be0de3eb56ca1d720f8f8545c6fdcb58ed27f0360c4e2ad",
|
| 399 |
+
"scripts/xsum/prep_xsum_le48.py": "3688ed1abc49d18f75d34f8382ff68a5afbd753b36cefdc0652d0bffb6a5f347",
|
| 400 |
+
"scripts/xsum/prep_xsum_mix30.py": "4bd47e52637c4b6bca5fd4956c1787b506a7f51b8a5bb0b68f19f728c4d48afb",
|
| 401 |
+
"scripts/xsum/prep_xsum_mix_split.py": "ec2fe74f08068ae8fda4ed59a86537a1d3f48276d205099bbfe16f8428ca36d5",
|
| 402 |
+
"scripts/xsum/prep_xsum_pass_uncond.py": "82ea50248d811e5182e63d0b4dfd6f995dd2f47eb963b5bb1c5d09bc533db7c1",
|
| 403 |
+
"scripts/xsum/prep_xsum_test_all.py": "7389fbdf5612ba524801c5b04109f810b139152ca50101b60acbb9b1e767b256",
|
| 404 |
+
"scripts/xsum/score_checkpoints.py": "41b412c06cd92cbd4caa6b586da7ddc7e913607a47c1a3aac0ee0a36a20a36be",
|
| 405 |
+
"tokenization/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 406 |
+
"tokenization/bpe_tokenizer.py": "0c8af84868a6aa40f981768f8ab49d28e501d55cce444edb1ad315de53c3445a",
|
| 407 |
+
"tokenization/build_word_vocab.py": "778f9666f4e7c642dc471773ee81c92c26975db37de1ebfffb605cae211351b5",
|
| 408 |
+
"tokenization/byte_tokenizer.py": "eb567a1772ae69738884ed9b94330aa39387c49e54d76e4b9752e67674fad810",
|
| 409 |
+
"tokenization/compression_ratio.py": "ed07ab1cc38d6176362631ce655bc974dd3623b87693f51279d5047e6bc6890f",
|
| 410 |
+
"tokenization/hf_tokenizer.py": "2ca8fae097f9c640bcb52b4788c37de8e0d9bb2fa971e30a5abb39f90ff97d4c",
|
| 411 |
+
"tokenization/tokenize.py": "f8ab54338613b6bd84513fdda85be173fd32274c35d97ef39a5c17d584f8108f",
|
| 412 |
+
"tokenization/tokenize_doc_level.py": "a06c81cf088fb81d2317ed1c3a1417a0e48ae0ea04962824f13f5d3beddc8902",
|
| 413 |
+
"tokenization/tokenize_vqvae.py": "14730eec3da87f6528e5fbd1166d5da79e10ac1700f7cd0be73bc40716bdebfc",
|
| 414 |
+
"tokenization/tokenizer.py": "e527d47e4ed8634c673d48a872a07c9cc8828d30988331de2d4bd3b7c21ddcb7",
|
| 415 |
+
"tokenization/train_bpe.py": "bf4e2a29eeba874024ed95f32fe720ae2320f72cb978361692db73e71be98683",
|
| 416 |
+
"tokenization/vqvae_tokenizer.py": "d7fa9d595865aa3e0b0d185bfd04721eb297fe5f657c4fca4ce69fc96aca308f",
|
| 417 |
+
"tokenization/word_tokenizer.py": "a405621bf35ed6cbb29abadd9d1d651e0783e6801fc7645b35e799d5de381ac2",
|
| 418 |
+
"train_mstok.py": "077b7c99c12daad2bbd986918d0f487411779de05a472af71124098da399b79d",
|
| 419 |
+
"trainer/alignment_trainer.py": "54f8ee7cca4d407223f2ce5146e6fc561ad748e20f2c702beb12ae040cd290c4",
|
| 420 |
+
"trainer/corruption.py": "f0f24c94652e6e9cf8e4ef30ae1a302e3232d64e765c931696d454db96e8b57d",
|
| 421 |
+
"trainer/joint_trainer.py": "8c8ebb55d21b6e223be1f7595c51f1b9257ccac40f8cff28260c1088656e2e49",
|
| 422 |
+
"trainer/mstok_trainer.py": "90712e396472d5dac9679b62af4f0b62afa6390a839d4d9bd9dae57406fe76a4",
|
| 423 |
+
"trainer/ncp_trainer.py": "73fc23f78f403dc39800b924dfb18ad42ffeab847fed30e931183c3d3acaee68",
|
| 424 |
+
"trainer/ntp_trainer.py": "5076d7aa29fac70869e1e309a15af6c4d3a91d6de8e00ac113a586841a9866e1",
|
| 425 |
+
"trainer/onpolicy.py": "6ff1fe4103d71001afe2d73e574d20c7242381e420929db14020141464ec6821",
|
| 426 |
+
"trainer/plain_two_stage.py": "4c5973ed87100b1b478aa1df0bb16cd3d1eb5844dffe5894003dd67c463f6c9a",
|
| 427 |
+
"trainer/scheduled_sampling.py": "f53515e23c2cec1231265f21a341ff77d11fa8a6129e0e382d540cc85571f5ed",
|
| 428 |
+
"trainer/semantic_mstok_trainer.py": "ebc02e80d279a445c99d8f09505cdfaf7632f1a661d6880f09cc1bd0fa90ef30",
|
| 429 |
+
"trainer/semantic_teacher.py": "7983b5da86d25954ed8267eee03d144c421305967705e0a242f4f025794ceeae",
|
| 430 |
+
"trainer/substitution_trainer.py": "33c7833b2e8dfabbb9ee56b258708055d3194b8d7453a1b8d93f59b44ff91465",
|
| 431 |
+
"utils/__init__.py": "e3b0c44298fc1c149afbf4c8996fb92427ae41e4649b934ca495991b7852b855",
|
| 432 |
+
"utils/arg_utils.py": "e9eb7971f6273e9630b56fd87a03725b29183a8c80abb014c4c84f8b1bdb6445",
|
| 433 |
+
"utils/benchmark.py": "ccf436ef72a4e626c7f671d8e63a683057bf502a3852e5a7142410d1dc3e5501",
|
| 434 |
+
"utils/data.py": "78cce11e777813755cd171ccaa3247f89ab99ac4079e0157e55c72ce8a5a7fe5",
|
| 435 |
+
"utils/dist.py": "b2849d3ce03a8b1737e3dd61216e76e4a9e82d63ad1ea74cddf72e2285eaca64",
|
| 436 |
+
"utils/experiment_config.py": "616a70fefaac99ab0d856a1b6c36787a28fb9af8a4c55fc6d4765012fbaaeb38",
|
| 437 |
+
"utils/iclr_training.py": "10c6857fa2319952fe0df18067730b42e3ff755244e2b059f6c95882f8d27c8e",
|
| 438 |
+
"utils/logging.py": "901e67515cc65a35f575f05d16a3b9ef4a5b2cb0ea252cd7abbc80ecce9e88e8",
|
| 439 |
+
"utils/lr.py": "4d1ea2c25340afc4bdde8b582dab36f174e0b2965fa7bafa2a7539019dbe2b7e",
|
| 440 |
+
"utils/misc.py": "be4a323116a0956a77960d1baf613f268057d6a36023663d2af9f8e4728cbe77",
|
| 441 |
+
"utils/plot_levels.py": "916b39885ced191960c26c334a3b16f9c99751f6c29c46bd69d14fee414965f3",
|
| 442 |
+
"utils/registry.py": "f7a44c6d18433100873d5df462b9f08a8bdeb11cf3d92226334c7e5ece318318",
|
| 443 |
+
"utils/train.py": "43eef988000f3a5e20856653d725337a0fe0dbe276f0cf80e1b200292abe812d",
|
| 444 |
+
"train_iclr_debug.py": "e0a0adf4a30ef55925f9017de24d434299749332e58a507a6424265f7a5f796e",
|
| 445 |
+
"config/iclr-debug/README.md": "1c757d2ff37026d08c22324ec44cfa760b0599af62c8ad9b728ada471a8f2905",
|
| 446 |
+
"config/iclr-debug/benchmark-16sq.yaml": "18c9cef7032f9d3806d8b0c8b20fe0fc7d1ea98ad1d2440e24a67ee873c97f9b",
|
| 447 |
+
"config/iclr-debug/full-owt-audit.json": "fefa394834f6e9f65944d3b514633c29061071ebe6ca60fce21a6edef9748c1e",
|
| 448 |
+
"config/iclr-debug/training-defaults.yaml": "3b144b8d53f96fe035e36edca96aac456e6cf21202373a467a3cd875fe86b474",
|
| 449 |
+
"config/iclr-debug/NCM-BUDGET-AUDIT.md": "e734caeab7edb612de7ff508a8c0ec52445530ab7a35895b5b793d6a7b752e4b",
|
| 450 |
+
"config/iclr-debug/TRAINING.md": "91555457c5cf6af078075673ce96cb440bc8e22aeae5211c252d94f92ea87ac7",
|
| 451 |
+
"config/iclr-debug/BENCHMARK.md": "2781ea506212e12a0640a4b0d2a2c881355ec85756ca4014bccb1b1aedccdf59",
|
| 452 |
+
"config/iclr-debug/HF-MODEL-CARD.md": "de1fe176ebde262eacdf32af7d3f8c6497a764dd98e5a9094dad26de4aade38e"
|
| 453 |
+
},
|
| 454 |
+
"dataset": {
|
| 455 |
+
"dataset": "hazyresearch/ncm-tokenized-datasets",
|
| 456 |
+
"revision": "951e04517fd53d57521d8e09abd6511a662787d0",
|
| 457 |
+
"dtype": "little-endian uint16",
|
| 458 |
+
"eos_id": 50256,
|
| 459 |
+
"offset_units": "token positions; [0, each EOS position + 1] with final sentinel",
|
| 460 |
+
"sampling": "length-weighted within-segment windows; pad short segments with 50257",
|
| 461 |
+
"caveat": "EOS segments are not independently verified source-document identities",
|
| 462 |
+
"splits": {
|
| 463 |
+
"train": {
|
| 464 |
+
"bytes": 18070847356,
|
| 465 |
+
"tokens": 9035423678,
|
| 466 |
+
"sha256": "b24b78a07f56974fd0c9855dc16f25b836d2cb172be23eb5d3c650cd6d73f56a",
|
| 467 |
+
"min_token_id": 0,
|
| 468 |
+
"max_token_id": 50256,
|
| 469 |
+
"segments": 8009770,
|
| 470 |
+
"trailing_tokens": 0,
|
| 471 |
+
"eos_only_segments": 0,
|
| 472 |
+
"length_percentiles": {
|
| 473 |
+
"min": 132.0,
|
| 474 |
+
"p25": 416.0,
|
| 475 |
+
"p50": 718.0,
|
| 476 |
+
"p75": 1249.0,
|
| 477 |
+
"p99": 7610.0,
|
| 478 |
+
"max": 131288.0
|
| 479 |
+
},
|
| 480 |
+
"segments_shorter_than_256": 673011,
|
| 481 |
+
"expected_padding_fraction_at_256": 1.5955357567300495e-05,
|
| 482 |
+
"offsets_status": "created",
|
| 483 |
+
"offsets_sha256": "7e84564946b8cf8190dbdc1e86bdf2fe8cfd9235a517328c80b4df038ea0028b",
|
| 484 |
+
"offset_entries": 8009771
|
| 485 |
+
},
|
| 486 |
+
"valid": {
|
| 487 |
+
"bytes": 9186834,
|
| 488 |
+
"tokens": 4593417,
|
| 489 |
+
"sha256": "1f2ee17e08327a39d4e2417948d02c7b998d626d2dcb2661197a6e208f58266f",
|
| 490 |
+
"min_token_id": 0,
|
| 491 |
+
"max_token_id": 50256,
|
| 492 |
+
"segments": 3999,
|
| 493 |
+
"trailing_tokens": 0,
|
| 494 |
+
"eos_only_segments": 0,
|
| 495 |
+
"length_percentiles": {
|
| 496 |
+
"min": 143.0,
|
| 497 |
+
"p25": 433.0,
|
| 498 |
+
"p50": 734.0,
|
| 499 |
+
"p75": 1292.5,
|
| 500 |
+
"p99": 7122.719999999997,
|
| 501 |
+
"max": 24868.0
|
| 502 |
+
},
|
| 503 |
+
"segments_shorter_than_256": 363,
|
| 504 |
+
"expected_padding_fraction_at_256": 1.6064002171981108e-05,
|
| 505 |
+
"offsets_status": "created",
|
| 506 |
+
"offsets_sha256": "09a748ce1dcec7de36c77c48e39ccbf5c0293cc073689fa7c5f48bdfb9bb9445",
|
| 507 |
+
"offset_entries": 4000
|
| 508 |
+
}
|
| 509 |
+
},
|
| 510 |
+
"overlap": {
|
| 511 |
+
"exact_train_segment_matches": 0,
|
| 512 |
+
"unique_validation_segments_in_train": 0,
|
| 513 |
+
"unique_validation_segments": 3999,
|
| 514 |
+
"near_duplicates_checked": false
|
| 515 |
+
}
|
| 516 |
+
},
|
| 517 |
+
"hf_repo": null,
|
| 518 |
+
"generation_evaluation": true,
|
| 519 |
+
"versions": {
|
| 520 |
+
"torch": "2.13.0+cu126",
|
| 521 |
+
"transformers": "4.44.2",
|
| 522 |
+
"numpy": "1.26.4",
|
| 523 |
+
"hydra-core": "1.3.6"
|
| 524 |
+
}
|
| 525 |
+
}
|
iclr-debug-decoder-135b-20260916/provenance/hf-backup-step-28000-manifest.json
ADDED
|
@@ -0,0 +1,116 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"repository": "iskhare/iclr-debug",
|
| 3 |
+
"checkpoint_step": 28000,
|
| 4 |
+
"files": [
|
| 5 |
+
{
|
| 6 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/milestone-iter-28000.pt",
|
| 7 |
+
"bytes": 4879455521,
|
| 8 |
+
"sha256": "d9eb43b22050308d6a75176a2978dadc8b0eeba38d905afa73521385e4a373c0"
|
| 9 |
+
},
|
| 10 |
+
{
|
| 11 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/manifest.json",
|
| 12 |
+
"bytes": 28971,
|
| 13 |
+
"sha256": "53dc5b666685e03efc016e4311b89ab94bb6babe710405c28ae58356406582d0"
|
| 14 |
+
},
|
| 15 |
+
{
|
| 16 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/git-provenance.json",
|
| 17 |
+
"bytes": 59,
|
| 18 |
+
"sha256": "96f601db3a051fd136086ca5f54c87ce2a72449e528d826a3b882541953e3fa2"
|
| 19 |
+
},
|
| 20 |
+
{
|
| 21 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/config.yaml",
|
| 22 |
+
"bytes": 4827,
|
| 23 |
+
"sha256": "073bd5a69f1d0a39019efffd79f4b7bd645e0c1792edead7e31b524515ef134f"
|
| 24 |
+
},
|
| 25 |
+
{
|
| 26 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/teacher-provenance.json",
|
| 27 |
+
"bytes": 270,
|
| 28 |
+
"sha256": "96e0063afe4211b1109509690a697fd8a7366941988c894f1027e1ead45938cf"
|
| 29 |
+
},
|
| 30 |
+
{
|
| 31 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/summary.json",
|
| 32 |
+
"bytes": 1174,
|
| 33 |
+
"sha256": "2bee9a7e46cdf0402e56360d2dadf306d3347f835432b6685241bcab04f712ba"
|
| 34 |
+
},
|
| 35 |
+
{
|
| 36 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/quality-summary.json",
|
| 37 |
+
"bytes": 2945,
|
| 38 |
+
"sha256": "3441e3ff7d5cce050e74ff6daf146514e0fc26621e1675d005e8a177a75e1d1c"
|
| 39 |
+
},
|
| 40 |
+
{
|
| 41 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/report.md",
|
| 42 |
+
"bytes": 3444,
|
| 43 |
+
"sha256": "4f15ced1ff8b7aeb09895f1e01ab27b27865d4bbd692d15874ccada09991e652"
|
| 44 |
+
},
|
| 45 |
+
{
|
| 46 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/manifest.json",
|
| 47 |
+
"bytes": 916,
|
| 48 |
+
"sha256": "993caaf29b702767a24bbddf23e27fb961b652fdcf5ce5771de39cd6f0dd1568"
|
| 49 |
+
},
|
| 50 |
+
{
|
| 51 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/evaluator.py",
|
| 52 |
+
"bytes": 7709,
|
| 53 |
+
"sha256": "e2c87f66f1001853f58b9e78a646fc4e089fc894d05fa6c215a8baf493a84055"
|
| 54 |
+
},
|
| 55 |
+
{
|
| 56 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/controller.py",
|
| 57 |
+
"bytes": 3094,
|
| 58 |
+
"sha256": "1de01fb6ca96049e786f86c61d07dae9a73897824e73c5a8be1aa47f2e41edd7"
|
| 59 |
+
},
|
| 60 |
+
{
|
| 61 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/summarize-quality.py",
|
| 62 |
+
"bytes": 3757,
|
| 63 |
+
"sha256": "60ba459f27cde464bd4ed8bd2f5518e5b97d120d315042788d9ad37fbfd059d1"
|
| 64 |
+
},
|
| 65 |
+
{
|
| 66 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-0.json",
|
| 67 |
+
"bytes": 139308,
|
| 68 |
+
"sha256": "60889f234d7ce2324f2f97782dc936971ad62ba8764fe0cc2902bc3ffc15aab6"
|
| 69 |
+
},
|
| 70 |
+
{
|
| 71 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-0.samples.json",
|
| 72 |
+
"bytes": 129390,
|
| 73 |
+
"sha256": "ccb0e33aba18b71c5c1d77c8c957f146604fc9b68f244e9db03508f60c188861"
|
| 74 |
+
},
|
| 75 |
+
{
|
| 76 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-1.json",
|
| 77 |
+
"bytes": 141566,
|
| 78 |
+
"sha256": "40b7bd89e42f6bc5b7cb12022373ffcdf743e21d3308a56ccddbb58d3d48fc4f"
|
| 79 |
+
},
|
| 80 |
+
{
|
| 81 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-1.samples.json",
|
| 82 |
+
"bytes": 131645,
|
| 83 |
+
"sha256": "e5f7af200bdbf470e43cdca1a689c245c2927846e51bb8ddd45fe0e7c35dafcd"
|
| 84 |
+
},
|
| 85 |
+
{
|
| 86 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-2.json",
|
| 87 |
+
"bytes": 143228,
|
| 88 |
+
"sha256": "77fac889717c95593b412a885ce8fde84fa05a56445e94cf951346f3e6fcb2f3"
|
| 89 |
+
},
|
| 90 |
+
{
|
| 91 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-2.samples.json",
|
| 92 |
+
"bytes": 133321,
|
| 93 |
+
"sha256": "8da15a810f2766becb09dc5ed56de844c61c0470eeb13ac376f362bd420a168a"
|
| 94 |
+
},
|
| 95 |
+
{
|
| 96 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-3.json",
|
| 97 |
+
"bytes": 141349,
|
| 98 |
+
"sha256": "f63cf95471df8e843388ff0c53335db775e17874ee3dab00ec80485c969fc2f6"
|
| 99 |
+
},
|
| 100 |
+
{
|
| 101 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-3.samples.json",
|
| 102 |
+
"bytes": 131426,
|
| 103 |
+
"sha256": "02fb5a585fde4f9188e8d9be168ed17b72f51d4586eb5d9e909dc2286f2295f7"
|
| 104 |
+
},
|
| 105 |
+
{
|
| 106 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-4.json",
|
| 107 |
+
"bytes": 140771,
|
| 108 |
+
"sha256": "f413441516002fd050c00d2283d3d9ac6a86cc90414d49fe514da110d8969d5f"
|
| 109 |
+
},
|
| 110 |
+
{
|
| 111 |
+
"path_in_repo": "iclr-debug-decoder-135b-20260916/decoder-mstok/evaluation/step-28000-fp64/random-seed-4.samples.json",
|
| 112 |
+
"bytes": 130875,
|
| 113 |
+
"sha256": "0a4d08f610465b45b68e31cdc448f092bb18fe447f5556ec5a1cea114f2338ab"
|
| 114 |
+
}
|
| 115 |
+
]
|
| 116 |
+
}
|