keepsloading commited on
Commit
1f623ee
·
verified ·
1 Parent(s): 6a26953

Upload folder using huggingface_hub

Browse files
Files changed (24) hide show
  1. outputs/eval/logs/token_t0525/babilong.log +286 -286
  2. outputs/eval/logs/token_t0525/ruler_a.log +253 -251
  3. outputs/eval/logs/token_t0525/ruler_b.log +205 -211
  4. outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/results_2026-07-18T16-10-15.187546.json +505 -0
  5. outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa1_2026-07-18T16-10-15.187546.jsonl +3 -0
  6. outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa2_2026-07-18T16-10-15.187546.jsonl +3 -0
  7. outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa3_2026-07-18T16-10-15.187546.jsonl +3 -0
  8. outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa4_2026-07-18T16-10-15.187546.jsonl +3 -0
  9. outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa5_2026-07-18T16-10-15.187546.jsonl +3 -0
  10. outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/results_2026-07-18T15-51-37.349704.json +821 -0
  11. outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multikey_2_2026-07-18T15-51-37.349704.jsonl +3 -0
  12. outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multiquery_2026-07-18T15-51-37.349704.jsonl +3 -0
  13. outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_single_1_2026-07-18T15-51-37.349704.jsonl +3 -0
  14. outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_single_3_2026-07-18T15-51-37.349704.jsonl +3 -0
  15. outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_fwe_2026-07-18T15-51-37.349704.jsonl +3 -0
  16. outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_qa_hotpot_2026-07-18T15-51-37.349704.jsonl +3 -0
  17. outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_vt_2026-07-18T15-51-37.349704.jsonl +3 -0
  18. outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/results_2026-07-18T16-04-28.086981.json +714 -0
  19. outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multikey_1_2026-07-18T16-04-28.086981.jsonl +3 -0
  20. outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multikey_3_2026-07-18T16-04-28.086981.jsonl +3 -0
  21. outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multivalue_2026-07-18T16-04-28.086981.jsonl +3 -0
  22. outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_single_2_2026-07-18T16-04-28.086981.jsonl +3 -0
  23. outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_cwe_2026-07-18T16-04-28.086981.jsonl +3 -0
  24. outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_qa_squad_2026-07-18T16-04-28.086981.jsonl +3 -0
outputs/eval/logs/token_t0525/babilong.log CHANGED
@@ -1,301 +1,301 @@
1
- 2026-07-18:13:09:04 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
- 2026-07-18:13:09:08 INFO [_cli.run:376] Selected Tasks: ['babilong_longctx']
3
- 2026-07-18:13:09:09 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
- 2026-07-18:13:09:09 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
- 2026-07-18:13:09:12 INFO [models.huggingface:161] Using device 'cuda:0'
6
- 2026-07-18:13:09:12 INFO [models.huggingface:548] Model type cannot be determined. Using default model type 'causal'
7
- 2026-07-18:13:09:13 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
8
- 2026-07-18:13:09:14 WARNING [api.task:856] babilong_qa5: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
9
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
10
- 2026-07-18:13:09:15 WARNING [api.task:856] babilong_qa4: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
11
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
12
- 2026-07-18:13:09:15 WARNING [api.task:856] babilong_qa3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
13
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
14
- 2026-07-18:13:09:15 WARNING [api.task:856] babilong_qa2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
15
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
16
- 2026-07-18:13:09:15 WARNING [api.task:856] babilong_qa1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
17
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
18
- 2026-07-18:13:09:16 INFO [tasks:700] Selected tasks:
19
- 2026-07-18:13:09:16 INFO [tasks:703] Group: babilong_longctx
20
- 2026-07-18:13:09:16 INFO [tasks:726] ConfigurableGroup(group=babilong_longctx,group_alias=None): {'babilong_qa1': ConfigurableTask(task_name=babilong_qa1,output_type=generate_until,num_fewshot=2,num_samples=1000), 'babilong_qa2': ConfigurableTask(task_name=babilong_qa2,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa3': ConfigurableTask(task_name=babilong_qa3,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa4': ConfigurableTask(task_name=babilong_qa4,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa5': ConfigurableTask(task_name=babilong_qa5,output_type=generate_until,num_fewshot=2,num_samples=999)}
21
- 2026-07-18:13:09:16 INFO [evaluator:314] babilong_qa1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
22
- 2026-07-18:13:09:16 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa1 from 2 to 2
23
- 2026-07-18:13:09:16 INFO [evaluator:314] babilong_qa2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
24
- 2026-07-18:13:09:16 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa2 from 2 to 2
25
- 2026-07-18:13:09:16 INFO [evaluator:314] babilong_qa3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
26
- 2026-07-18:13:09:16 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa3 from 2 to 2
27
- 2026-07-18:13:09:16 INFO [evaluator:314] babilong_qa4: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
28
- 2026-07-18:13:09:16 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa4 from 2 to 2
29
- 2026-07-18:13:09:16 INFO [evaluator:314] babilong_qa5: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
30
- 2026-07-18:13:09:16 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa5 from 2 to 2
31
- 2026-07-18:13:09:16 INFO [api.task:311] Building contexts for babilong_qa1 on rank 0...
32
 
33
-
34
  0%| | 0/50 [00:00<?, ?it/s]
35
- 2026-07-18:13:09:16 INFO [api.task:311] Building contexts for babilong_qa2 on rank 0...
36
  0%| | 0/50 [00:00<?, ?it/s]
 
37
 
38
-
39
  0%| | 0/50 [00:00<?, ?it/s]
40
- 2026-07-18:13:09:16 INFO [api.task:311] Building contexts for babilong_qa3 on rank 0...
41
  0%| | 0/50 [00:00<?, ?it/s]
 
42
 
43
-
44
  0%| | 0/50 [00:00<?, ?it/s]
45
- 2026-07-18:13:09:16 INFO [api.task:311] Building contexts for babilong_qa4 on rank 0...
46
  0%| | 0/50 [00:00<?, ?it/s]
 
47
 
48
-
49
  0%| | 0/50 [00:00<?, ?it/s]
50
- 2026-07-18:13:09:16 INFO [api.task:311] Building contexts for babilong_qa5 on rank 0...
51
  0%| | 0/50 [00:00<?, ?it/s]
 
52
 
53
-
54
  0%| | 0/50 [00:00<?, ?it/s]
55
- 2026-07-18:13:09:16 INFO [evaluator:584] Running generate_until requests
56
  0%| | 0/50 [00:00<?, ?it/s]
 
57
 
58
 
59
-
60
-
61
-
62
-
63
-
64
-
65
-
66
-
67
-
68
-
69
-
70
-
71
-
72
-
73
-
74
-
75
-
76
-
77
-
78
-
79
-
80
-
81
-
82
-
83
-
84
-
85
-
86
-
87
-
88
-
89
-
90
-
91
-
92
-
93
-
94
-
95
-
96
-
97
-
98
-
99
-
100
-
101
-
102
-
103
-
104
-
105
-
106
-
107
-
108
-
109
-
110
-
111
-
112
-
113
-
114
-
115
-
116
-
117
-
118
-
119
-
120
-
121
-
122
-
123
-
124
-
125
-
126
-
127
-
128
-
129
-
130
-
131
-
132
-
133
-
134
-
135
-
136
-
137
-
138
-
139
-
140
-
141
-
142
-
143
-
144
-
145
-
146
-
147
-
148
-
149
-
150
-
151
-
152
-
153
-
154
-
155
-
156
-
157
-
158
-
159
-
160
-
161
-
162
-
163
-
164
-
165
-
166
-
167
-
168
-
169
-
170
-
171
-
172
-
173
-
174
-
175
-
176
-
177
-
178
-
179
-
180
-
181
-
182
-
183
-
184
-
185
-
186
-
187
-
188
-
189
-
190
-
191
-
192
-
193
-
194
-
195
-
196
-
197
-
198
-
199
-
200
-
201
-
202
-
203
-
204
-
205
-
206
-
207
-
208
-
209
-
210
-
211
-
212
-
213
-
214
-
215
-
216
-
217
-
218
-
219
-
220
-
221
-
222
-
223
-
224
-
225
-
226
-
227
-
228
-
229
-
230
-
231
-
232
-
233
-
234
-
235
-
236
-
237
-
238
-
239
-
240
-
241
-
242
-
243
-
244
-
245
-
246
-
247
-
248
-
249
-
250
-
251
-
252
-
253
-
254
-
255
-
256
-
257
-
258
-
259
-
260
-
261
-
262
-
263
-
264
-
265
-
266
-
267
-
268
-
269
-
270
-
271
-
272
-
273
-
274
-
275
-
276
-
277
-
278
-
279
-
280
-
281
-
282
-
283
-
284
-
285
-
286
-
287
-
288
-
289
-
290
-
291
-
292
-
293
-
294
-
295
-
296
-
297
-
298
-
299
-
300
-
301
-
302
-
303
-
304
-
305
-
306
-
307
- 2026-07-18:13:14:07 INFO [loggers.evaluation_tracker:247] Saving results aggregated
308
- 2026-07-18:13:14:07 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/*.jsonl
309
  hf ({'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 50.0, num_fewshot: 2, batch_size: 1
310
  | Tasks |Version|Filter|n-shot|Metric| |Value| |Stderr|
311
  |----------------|------:|------|-----:|------|---|----:|---|-----:|
 
1
+ 2026-07-18:16:04:30 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
+ 2026-07-18:16:04:34 INFO [_cli.run:376] Selected Tasks: ['babilong_longctx']
3
+ 2026-07-18:16:04:35 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
+ 2026-07-18:16:04:35 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
+ 2026-07-18:16:04:38 INFO [models.huggingface:161] Using device 'cuda:0'
6
+ 2026-07-18:16:04:38 INFO [models.huggingface:548] Model type cannot be determined. Using default model type 'causal'
7
+ 2026-07-18:16:04:39 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
8
+ 2026-07-18:16:04:40 WARNING [api.task:856] babilong_qa5: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
9
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
10
+ 2026-07-18:16:04:41 WARNING [api.task:856] babilong_qa4: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
11
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
12
+ 2026-07-18:16:04:41 WARNING [api.task:856] babilong_qa3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
13
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
14
+ 2026-07-18:16:04:41 WARNING [api.task:856] babilong_qa2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
15
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
16
+ 2026-07-18:16:04:41 WARNING [api.task:856] babilong_qa1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
17
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
18
+ 2026-07-18:16:04:41 INFO [tasks:700] Selected tasks:
19
+ 2026-07-18:16:04:41 INFO [tasks:703] Group: babilong_longctx
20
+ 2026-07-18:16:04:41 INFO [tasks:726] ConfigurableGroup(group=babilong_longctx,group_alias=None): {'babilong_qa1': ConfigurableTask(task_name=babilong_qa1,output_type=generate_until,num_fewshot=2,num_samples=1000), 'babilong_qa2': ConfigurableTask(task_name=babilong_qa2,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa3': ConfigurableTask(task_name=babilong_qa3,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa4': ConfigurableTask(task_name=babilong_qa4,output_type=generate_until,num_fewshot=2,num_samples=999), 'babilong_qa5': ConfigurableTask(task_name=babilong_qa5,output_type=generate_until,num_fewshot=2,num_samples=999)}
21
+ 2026-07-18:16:04:41 INFO [evaluator:314] babilong_qa1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
22
+ 2026-07-18:16:04:41 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa1 from 2 to 2
23
+ 2026-07-18:16:04:41 INFO [evaluator:314] babilong_qa2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
24
+ 2026-07-18:16:04:41 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa2 from 2 to 2
25
+ 2026-07-18:16:04:41 INFO [evaluator:314] babilong_qa3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
26
+ 2026-07-18:16:04:41 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa3 from 2 to 2
27
+ 2026-07-18:16:04:41 INFO [evaluator:314] babilong_qa4: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
28
+ 2026-07-18:16:04:41 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa4 from 2 to 2
29
+ 2026-07-18:16:04:41 INFO [evaluator:314] babilong_qa5: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 16, 'until': []}
30
+ 2026-07-18:16:04:41 WARNING [evaluator:333] Overwriting default num_fewshot of babilong_qa5 from 2 to 2
31
+ 2026-07-18:16:04:41 INFO [api.task:311] Building contexts for babilong_qa1 on rank 0...
32
 
 
33
  0%| | 0/50 [00:00<?, ?it/s]
34
+
35
  0%| | 0/50 [00:00<?, ?it/s]
36
+ 2026-07-18:16:04:41 INFO [api.task:311] Building contexts for babilong_qa2 on rank 0...
37
 
 
38
  0%| | 0/50 [00:00<?, ?it/s]
39
+
40
  0%| | 0/50 [00:00<?, ?it/s]
41
+ 2026-07-18:16:04:42 INFO [api.task:311] Building contexts for babilong_qa3 on rank 0...
42
 
 
43
  0%| | 0/50 [00:00<?, ?it/s]
44
+
45
  0%| | 0/50 [00:00<?, ?it/s]
46
+ 2026-07-18:16:04:42 INFO [api.task:311] Building contexts for babilong_qa4 on rank 0...
47
 
 
48
  0%| | 0/50 [00:00<?, ?it/s]
49
+
50
  0%| | 0/50 [00:00<?, ?it/s]
51
+ 2026-07-18:16:04:42 INFO [api.task:311] Building contexts for babilong_qa5 on rank 0...
52
 
 
53
  0%| | 0/50 [00:00<?, ?it/s]
54
+
55
  0%| | 0/50 [00:00<?, ?it/s]
56
+ 2026-07-18:16:04:42 INFO [evaluator:584] Running generate_until requests
57
 
58
 
59
+
60
+
61
+
62
+
63
+
64
+
65
+
66
+
67
+
68
+
69
+
70
+
71
+
72
+
73
+
74
+
75
+
76
+
77
+
78
+
79
+
80
+
81
+
82
+
83
+
84
+
85
+
86
+
87
+
88
+
89
+
90
+
91
+
92
+
93
+
94
+
95
+
96
+
97
+
98
+
99
+
100
+
101
+
102
+
103
+
104
+
105
+
106
+
107
+
108
+
109
+
110
+
111
+
112
+
113
+
114
+
115
+
116
+
117
+
118
+
119
+
120
+
121
+
122
+
123
+
124
+
125
+
126
+
127
+
128
+
129
+
130
+
131
+
132
+
133
+
134
+
135
+
136
+
137
+
138
+
139
+
140
+
141
+
142
+
143
+
144
+
145
+
146
+
147
+
148
+
149
+
150
+
151
+
152
+
153
+
154
+
155
+
156
+
157
+
158
+
159
+
160
+
161
+
162
+
163
+
164
+
165
+
166
+
167
+
168
+
169
+
170
+
171
+
172
+
173
+
174
+
175
+
176
+
177
+
178
+
179
+
180
+
181
+
182
+
183
+
184
+
185
+
186
+
187
+
188
+
189
+
190
+
191
+
192
+
193
+
194
+
195
+
196
+
197
+
198
+
199
+
200
+
201
+
202
+
203
+
204
+
205
+
206
+
207
+
208
+
209
+
210
+
211
+
212
+
213
+
214
+
215
+
216
+
217
+
218
+
219
+
220
+
221
+
222
+
223
+
224
+
225
+
226
+
227
+
228
+
229
+
230
+
231
+
232
+
233
+
234
+
235
+
236
+
237
+
238
+
239
+
240
+
241
+
242
+
243
+
244
+
245
+
246
+
247
+
248
+
249
+
250
+
251
+
252
+
253
+
254
+
255
+
256
+
257
+
258
+
259
+
260
+
261
+
262
+
263
+
264
+
265
+
266
+
267
+
268
+
269
+
270
+
271
+
272
+
273
+
274
+
275
+
276
+
277
+
278
+
279
+
280
+
281
+
282
+
283
+
284
+
285
+
286
+
287
+
288
+
289
+
290
+
291
+
292
+
293
+
294
+
295
+
296
+
297
+
298
+
299
+
300
+
301
+
302
+
303
+
304
+
305
+
306
+
307
+ 2026-07-18:16:10:15 INFO [loggers.evaluation_tracker:247] Saving results aggregated
308
+ 2026-07-18:16:10:15 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/*.jsonl
309
  hf ({'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 50.0, num_fewshot: 2, batch_size: 1
310
  | Tasks |Version|Filter|n-shot|Metric| |Value| |Stderr|
311
  |----------------|------:|------|-----:|------|---|----:|---|-----:|
outputs/eval/logs/token_t0525/ruler_a.log CHANGED
@@ -1,65 +1,66 @@
1
- 2026-07-18:12:44:28 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
- 2026-07-18:12:44:32 INFO [_cli.run:376] Selected Tasks: ['niah_single_1', 'niah_single_3', 'niah_multikey_2', 'niah_multiquery', 'ruler_vt', 'ruler_fwe', 'ruler_qa_hotpot']
3
- 2026-07-18:12:44:34 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
- 2026-07-18:12:44:34 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
- 2026-07-18:12:44:37 INFO [models.huggingface:161] Using device 'cuda:0'
6
- 2026-07-18:12:44:37 INFO [models.huggingface:548] Model type cannot be determined. Using default model type 'causal'
7
- 2026-07-18:12:44:37 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
8
- 2026-07-18:12:44:46 WARNING [api.task:856] niah_single_1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
9
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
10
- 2026-07-18:12:44:46 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
11
 
12
 
13
-
14
-
15
-
16
-
17
-
18
-
19
-
20
- 2026-07-18:12:44:55 WARNING [api.task:856] niah_single_3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
21
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
22
 
23
 
24
-
25
-
26
-
27
-
28
-
29
-
30
-
31
-
32
- 2026-07-18:12:45:04 WARNING [api.task:856] niah_multikey_2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
33
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
34
 
35
 
36
-
37
-
38
-
39
-
40
-
41
-
42
- 2026-07-18:12:45:11 WARNING [api.task:856] niah_multiquery: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
43
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
44
 
45
 
46
-
47
-
48
-
49
-
50
-
51
-
52
-
53
-
54
- 2026-07-18:12:45:20 WARNING [api.task:856] ruler_vt: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
55
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
56
- 2026-07-18:12:45:20 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
57
  Max length 500 | Current length 307 | Noises: 5
58
  Max length 500 | Current length 410 | Noises: 10
59
  Max length 500 | Current length 531 | Noises: 15
60
  Num noises: 10
61
 
62
-
63
  0%| | 0/1 [00:00<?, ?it/s]
 
64
  0%| | 0/1 [00:00<?, ?it/s]
65
  Max length 8192 | Current length 791 | Noises: 10
66
  Max length 8192 | Current length 1025 | Noises: 20
67
  Max length 8192 | Current length 1273 | Noises: 30
@@ -95,231 +96,232 @@ Max length 8192 | Current length 8230 | Noises: 320
95
  Num noises: 310
96
 
97
 
98
  0%| | 0/500 [00:00<?, ?it/s]
99
-
100
  13%|█▎ | 66/500 [00:01<00:06, 65.54it/s]
101
-
102
  26%|██▋ | 132/500 [00:02<00:05, 65.39it/s]
103
-
104
  40%|███▉ | 198/500 [00:03<00:04, 64.83it/s]
105
-
106
  53%|█████▎ | 265/500 [00:04<00:03, 65.67it/s]
107
-
108
  66%|██████▋ | 332/500 [00:05<00:02, 65.84it/s]
109
-
110
  80%|████████ | 400/500 [00:06<00:01, 66.25it/s]
111
-
112
  94%|█████████▎| 468/500 [00:07<00:00, 66.59it/s]
113
- 2026-07-18:12:45:28 WARNING [api.task:856] ruler_fwe: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
114
  14%|█▎ | 68/500 [00:01<00:06, 67.09it/s]
 
115
  27%|██▋ | 136/500 [00:02<00:05, 67.09it/s]
 
116
  41%|████ | 204/500 [00:03<00:04, 66.64it/s]
 
117
  54%|█████▍ | 271/500 [00:04<00:03, 66.60it/s]
 
118
  68%|██████▊ | 338/500 [00:05<00:02, 66.54it/s]
 
119
  81%|████████ | 406/500 [00:06<00:01, 66.70it/s]
 
120
  95%|█████████▍| 473/500 [00:07<00:00, 66.69it/s]
 
121
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
122
 
123
 
124
-
125
-
126
-
127
-
128
-
129
-
130
-
131
-
132
-
133
-
134
-
135
- 2026-07-18:12:45:41 WARNING [api.task:856] ruler_qa_hotpot: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
136
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
137
 
138
 
139
-
140
-
141
-
142
-
143
-
144
-
145
-
146
-
147
-
148
-
149
-
150
-
151
-
152
-
153
-
154
-
155
-
156
-
157
- 2026-07-18:12:46:03 INFO [tasks:700] Selected tasks:
158
- 2026-07-18:12:46:03 INFO [tasks:691] Task: ruler_qa_hotpot (ruler/qa_hotpot.yaml)
159
- 2026-07-18:12:46:03 INFO [tasks:691] Task: ruler_fwe (ruler/fwe.yaml)
160
- 2026-07-18:12:46:03 INFO [tasks:691] Task: ruler_vt (ruler/vt.yaml)
161
- 2026-07-18:12:46:03 INFO [tasks:691] Task: niah_multiquery (ruler/niah_multiquery.yaml)
162
- 2026-07-18:12:46:03 INFO [tasks:691] Task: niah_multikey_2 (ruler/niah_multikey_2.yaml)
163
- 2026-07-18:12:46:03 INFO [tasks:691] Task: niah_single_3 (ruler/niah_single_3.yaml)
164
- 2026-07-18:12:46:03 INFO [tasks:691] Task: niah_single_1 (ruler/niah_single_1.yaml)
165
- 2026-07-18:12:46:03 INFO [evaluator:314] ruler_qa_hotpot: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 32, 'until': []}
166
- 2026-07-18:12:46:03 INFO [evaluator:314] ruler_fwe: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 50, 'until': []}
167
- 2026-07-18:12:46:03 INFO [evaluator:314] ruler_vt: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 30, 'until': []}
168
- 2026-07-18:12:46:03 INFO [evaluator:314] niah_multiquery: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
169
- 2026-07-18:12:46:03 INFO [evaluator:314] niah_multikey_2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
170
- 2026-07-18:12:46:03 INFO [evaluator:314] niah_single_3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
171
- 2026-07-18:12:46:03 INFO [evaluator:314] niah_single_1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
172
- 2026-07-18:12:46:03 INFO [api.task:311] Building contexts for ruler_qa_hotpot on rank 0...
173
 
174
-
175
  0%| | 0/20 [00:00<?, ?it/s]
176
- 2026-07-18:12:46:03 INFO [api.task:311] Building contexts for ruler_fwe on rank 0...
177
  0%| | 0/20 [00:00<?, ?it/s]
 
178
 
179
-
180
  0%| | 0/20 [00:00<?, ?it/s]
181
- 2026-07-18:12:46:03 INFO [api.task:311] Building contexts for ruler_vt on rank 0...
182
  0%| | 0/20 [00:00<?, ?it/s]
 
183
 
184
-
185
  0%| | 0/20 [00:00<?, ?it/s]
186
- 2026-07-18:12:46:03 INFO [api.task:311] Building contexts for niah_multiquery on rank 0...
187
  0%| | 0/20 [00:00<?, ?it/s]
 
188
 
189
-
190
  0%| | 0/20 [00:00<?, ?it/s]
191
- 2026-07-18:12:46:03 INFO [api.task:311] Building contexts for niah_multikey_2 on rank 0...
192
  0%| | 0/20 [00:00<?, ?it/s]
 
193
 
194
-
195
  0%| | 0/20 [00:00<?, ?it/s]
196
- 2026-07-18:12:46:03 INFO [api.task:311] Building contexts for niah_single_3 on rank 0...
197
  0%| | 0/20 [00:00<?, ?it/s]
 
198
 
199
-
200
  0%| | 0/20 [00:00<?, ?it/s]
201
- 2026-07-18:12:46:03 INFO [api.task:311] Building contexts for niah_single_1 on rank 0...
202
  0%| | 0/20 [00:00<?, ?it/s]
 
203
 
204
-
205
  0%| | 0/20 [00:00<?, ?it/s]
206
- 2026-07-18:12:46:03 INFO [evaluator:584] Running generate_until requests
207
  0%| | 0/20 [00:00<?, ?it/s]
 
208
 
209
 
210
-
211
-
212
-
213
-
214
-
215
-
216
-
217
-
218
-
219
-
220
-
221
-
222
-
223
-
224
-
225
-
226
-
227
-
228
-
229
-
230
-
231
-
232
 
233
-
234
-
235
-
236
-
237
-
238
-
239
-
240
-
241
-
242
-
243
-
244
-
245
-
246
-
247
-
248
-
249
-
250
-
251
-
252
-
253
-
254
-
255
-
256
-
257
 
258
 
259
-
260
-
261
-
262
-
263
-
264
-
265
-
266
-
267
-
268
-
269
-
270
-
271
-
272
-
273
-
274
-
275
-
276
-
277
-
278
-
279
-
280
-
281
-
282
-
283
-
284
-
285
-
286
-
287
-
288
-
289
-
290
-
291
-
292
-
293
-
294
-
295
-
296
-
297
-
298
-
299
-
300
-
301
-
302
-
303
-
304
-
305
-
306
-
307
-
308
-
309
-
310
-
311
-
312
-
313
-
314
-
315
-
316
-
317
-
318
-
319
-
320
-
321
-
322
-
323
-
324
-
325
-
326
-
327
-
328
-
329
-
330
-
331
-
332
-
333
-
334
-
335
-
336
-
337
-
338
-
339
-
340
-
341
-
342
-
343
-
344
-
345
-
346
-
347
-
348
-
349
-
350
- 2026-07-18:12:56:47 INFO [loggers.evaluation_tracker:247] Saving results aggregated
351
- 2026-07-18:12:56:47 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/*.jsonl
352
  hf ({'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 20.0, num_fewshot: None, batch_size: 1
353
  | Tasks |Version|Filter|n-shot|Metric| | Value | |Stderr|
354
  |---------------|------:|------|-----:|-----:|---|------:|---|------|
 
1
+ 2026-07-18:15:39:08 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
+ 2026-07-18:15:39:12 INFO [_cli.run:376] Selected Tasks: ['niah_single_1', 'niah_single_3', 'niah_multikey_2', 'niah_multiquery', 'ruler_vt', 'ruler_fwe', 'ruler_qa_hotpot']
3
+ 2026-07-18:15:39:13 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
+ 2026-07-18:15:39:13 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
+ 2026-07-18:15:39:16 INFO [models.huggingface:161] Using device 'cuda:0'
6
+ 2026-07-18:15:39:16 INFO [models.huggingface:548] Model type cannot be determined. Using default model type 'causal'
7
+ 2026-07-18:15:39:17 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
8
+ 2026-07-18:15:39:25 WARNING [api.task:856] niah_single_1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
9
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
10
+ 2026-07-18:15:39:25 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
11
 
12
 
13
+
14
+
15
+
16
+
17
+
18
+
19
+
20
+ 2026-07-18:15:39:34 WARNING [api.task:856] niah_single_3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
21
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
22
 
23
 
24
+
25
+
26
+
27
+
28
+
29
+
30
+
31
+
32
+ 2026-07-18:15:39:43 WARNING [api.task:856] niah_multikey_2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
33
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
34
 
35
 
36
+
37
+
38
+
39
+
40
+
41
+
42
+
43
+ 2026-07-18:15:39:51 WARNING [api.task:856] niah_multiquery: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
44
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
45
 
46
 
47
+
48
+
49
+
50
+
51
+
52
+
53
+
54
+
55
+ 2026-07-18:15:39:59 WARNING [api.task:856] ruler_vt: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
56
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
57
+ 2026-07-18:15:39:59 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
58
  Max length 500 | Current length 307 | Noises: 5
59
  Max length 500 | Current length 410 | Noises: 10
60
  Max length 500 | Current length 531 | Noises: 15
61
  Num noises: 10
62
 
 
63
  0%| | 0/1 [00:00<?, ?it/s]
64
+
65
  0%| | 0/1 [00:00<?, ?it/s]
66
  Max length 8192 | Current length 791 | Noises: 10
67
  Max length 8192 | Current length 1025 | Noises: 20
68
  Max length 8192 | Current length 1273 | Noises: 30
 
96
  Num noises: 310
97
 
98
 
99
  0%| | 0/500 [00:00<?, ?it/s]
 
100
  13%|█▎ | 66/500 [00:01<00:06, 65.54it/s]
 
101
  26%|██▋ | 132/500 [00:02<00:05, 65.39it/s]
 
102
  40%|███▉ | 198/500 [00:03<00:04, 64.83it/s]
 
103
  53%|█████▎ | 265/500 [00:04<00:03, 65.67it/s]
 
104
  66%|██████▋ | 332/500 [00:05<00:02, 65.84it/s]
 
105
  80%|████████ | 400/500 [00:06<00:01, 66.25it/s]
 
106
  94%|█████████▎| 468/500 [00:07<00:00, 66.59it/s]
107
+
108
  14%|█▎ | 68/500 [00:01<00:06, 67.09it/s]
109
+
110
  27%|██▋ | 136/500 [00:02<00:05, 67.09it/s]
111
+
112
  41%|████ | 204/500 [00:03<00:04, 66.64it/s]
113
+
114
  54%|█████▍ | 271/500 [00:04<00:03, 66.60it/s]
115
+
116
  68%|██████▊ | 338/500 [00:05<00:02, 66.54it/s]
117
+
118
  81%|████████ | 406/500 [00:06<00:01, 66.70it/s]
119
+
120
  95%|█████████▍| 473/500 [00:07<00:00, 66.69it/s]
121
+ 2026-07-18:15:40:08 WARNING [api.task:856] ruler_fwe: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
122
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
123
 
124
 
125
+
126
+
127
+
128
+
129
+
130
+
131
+
132
+
133
+
134
+
135
+
136
+
137
+ 2026-07-18:15:40:21 WARNING [api.task:856] ruler_qa_hotpot: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
138
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
139
 
140
 
141
+
142
+
143
+
144
+
145
+
146
+
147
+
148
+
149
+
150
+
151
+
152
+
153
+
154
+
155
+
156
+
157
+
158
+
159
+ 2026-07-18:15:40:43 INFO [tasks:700] Selected tasks:
160
+ 2026-07-18:15:40:43 INFO [tasks:691] Task: ruler_qa_hotpot (ruler/qa_hotpot.yaml)
161
+ 2026-07-18:15:40:43 INFO [tasks:691] Task: ruler_fwe (ruler/fwe.yaml)
162
+ 2026-07-18:15:40:43 INFO [tasks:691] Task: ruler_vt (ruler/vt.yaml)
163
+ 2026-07-18:15:40:43 INFO [tasks:691] Task: niah_multiquery (ruler/niah_multiquery.yaml)
164
+ 2026-07-18:15:40:43 INFO [tasks:691] Task: niah_multikey_2 (ruler/niah_multikey_2.yaml)
165
+ 2026-07-18:15:40:43 INFO [tasks:691] Task: niah_single_3 (ruler/niah_single_3.yaml)
166
+ 2026-07-18:15:40:43 INFO [tasks:691] Task: niah_single_1 (ruler/niah_single_1.yaml)
167
+ 2026-07-18:15:40:43 INFO [evaluator:314] ruler_qa_hotpot: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 32, 'until': []}
168
+ 2026-07-18:15:40:43 INFO [evaluator:314] ruler_fwe: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 50, 'until': []}
169
+ 2026-07-18:15:40:43 INFO [evaluator:314] ruler_vt: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 30, 'until': []}
170
+ 2026-07-18:15:40:43 INFO [evaluator:314] niah_multiquery: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
171
+ 2026-07-18:15:40:43 INFO [evaluator:314] niah_multikey_2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
172
+ 2026-07-18:15:40:43 INFO [evaluator:314] niah_single_3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
173
+ 2026-07-18:15:40:43 INFO [evaluator:314] niah_single_1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
174
+ 2026-07-18:15:40:43 INFO [api.task:311] Building contexts for ruler_qa_hotpot on rank 0...
175
 
 
176
  0%| | 0/20 [00:00<?, ?it/s]
177
+
178
  0%| | 0/20 [00:00<?, ?it/s]
179
+ 2026-07-18:15:40:43 INFO [api.task:311] Building contexts for ruler_fwe on rank 0...
180
 
 
181
  0%| | 0/20 [00:00<?, ?it/s]
182
+
183
  0%| | 0/20 [00:00<?, ?it/s]
184
+ 2026-07-18:15:40:43 INFO [api.task:311] Building contexts for ruler_vt on rank 0...
185
 
 
186
  0%| | 0/20 [00:00<?, ?it/s]
187
+
188
  0%| | 0/20 [00:00<?, ?it/s]
189
+ 2026-07-18:15:40:43 INFO [api.task:311] Building contexts for niah_multiquery on rank 0...
190
 
 
191
  0%| | 0/20 [00:00<?, ?it/s]
192
+
193
  0%| | 0/20 [00:00<?, ?it/s]
194
+ 2026-07-18:15:40:43 INFO [api.task:311] Building contexts for niah_multikey_2 on rank 0...
195
 
 
196
  0%| | 0/20 [00:00<?, ?it/s]
197
+
198
  0%| | 0/20 [00:00<?, ?it/s]
199
+ 2026-07-18:15:40:43 INFO [api.task:311] Building contexts for niah_single_3 on rank 0...
200
 
 
201
  0%| | 0/20 [00:00<?, ?it/s]
202
+
203
  0%| | 0/20 [00:00<?, ?it/s]
204
+ 2026-07-18:15:40:43 INFO [api.task:311] Building contexts for niah_single_1 on rank 0...
205
 
 
206
  0%| | 0/20 [00:00<?, ?it/s]
207
+
208
  0%| | 0/20 [00:00<?, ?it/s]
209
+ 2026-07-18:15:40:43 INFO [evaluator:584] Running generate_until requests
210
 
211
 
212
+
213
+
214
+
215
+
216
+
217
+
218
+
219
+
220
+
221
+
222
+
223
+
224
+
225
+
226
+
227
+
228
+
229
+
230
+
231
+
232
+
233
+
234
 
235
+
236
+
237
+
238
+
239
+
240
+
241
+
242
+
243
+
244
+
245
+
246
+
247
+
248
+
249
+
250
+
251
+
252
+
253
+
254
+
255
+
256
+
257
+
258
+
259
 
260
 
261
+
262
+
263
+
264
+
265
+
266
+
267
+
268
+
269
+
270
+
271
+
272
+
273
+
274
+
275
+
276
+
277
+
278
+
279
+
280
+
281
+
282
+
283
+
284
+
285
+
286
+
287
+
288
+
289
+
290
+
291
+
292
+
293
+
294
+
295
+
296
+
297
+
298
+
299
+
300
+
301
+
302
+
303
+
304
+
305
+
306
+
307
+
308
+
309
+
310
+
311
+
312
+
313
+
314
+
315
+
316
+
317
+
318
+
319
+
320
+
321
+
322
+
323
+
324
+
325
+
326
+
327
+
328
+
329
+
330
+
331
+
332
+
333
+
334
+
335
+
336
+
337
+
338
+
339
+
340
+
341
+
342
+
343
+
344
+
345
+
346
+
347
+
348
+
349
+
350
+
351
+
352
+ 2026-07-18:15:51:37 INFO [loggers.evaluation_tracker:247] Saving results aggregated
353
+ 2026-07-18:15:51:37 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/*.jsonl
354
  hf ({'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 20.0, num_fewshot: None, batch_size: 1
355
  | Tasks |Version|Filter|n-shot|Metric| | Value | |Stderr|
356
  |---------------|------:|------|-----:|-----:|---|------:|---|------|
outputs/eval/logs/token_t0525/ruler_b.log CHANGED
@@ -1,240 +1,234 @@
1
- 2026-07-18:12:56:49 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
- 2026-07-18:12:56:53 INFO [_cli.run:376] Selected Tasks: ['niah_single_2', 'niah_multikey_1', 'niah_multikey_3', 'niah_multivalue', 'ruler_cwe', 'ruler_qa_squad']
3
- 2026-07-18:12:56:54 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
- 2026-07-18:12:56:54 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
- 2026-07-18:12:56:58 INFO [models.huggingface:161] Using device 'cuda:0'
6
- 2026-07-18:12:56:58 INFO [models.huggingface:548] Model type cannot be determined. Using default model type 'causal'
7
- 2026-07-18:12:56:58 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
8
- 2026-07-18:12:57:07 WARNING [api.task:856] niah_single_2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
9
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
10
- 2026-07-18:12:57:08 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
11
 
12
 
13
-
14
-
15
-
16
-
17
-
18
-
19
-
20
-
21
-
22
- 2026-07-18:12:57:18 WARNING [api.task:856] niah_multikey_1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
23
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
24
 
25
 
26
-
27
-
28
-
29
-
30
-
31
-
32
-
33
-
34
-
35
- 2026-07-18:12:57:28 WARNING [api.task:856] niah_multikey_3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
36
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
37
 
38
 
39
-
40
-
41
-
42
-
43
-
44
-
45
- 2026-07-18:12:57:34 WARNING [api.task:856] niah_multivalue: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
46
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
47
 
48
 
49
-
50
-
51
-
52
-
53
-
54
-
55
-
56
-
57
-
58
- 2026-07-18:12:57:44 WARNING [api.task:856] ruler_cwe: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
59
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
60
- 2026-07-18:12:57:44 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
61
 
62
 
63
-
64
-
65
-
66
-
67
-
68
-
69
- 2026-07-18:12:57:52 WARNING [api.task:856] ruler_qa_squad: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
70
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
71
 
72
 
73
-
74
-
75
-
76
-
77
-
78
-
79
-
80
-
81
-
82
- 2026-07-18:12:58:03 INFO [tasks:700] Selected tasks:
83
- 2026-07-18:12:58:03 INFO [tasks:691] Task: ruler_qa_squad (ruler/qa_squad.yaml)
84
- 2026-07-18:12:58:03 INFO [tasks:691] Task: ruler_cwe (ruler/cwe.yaml)
85
- 2026-07-18:12:58:03 INFO [tasks:691] Task: niah_multivalue (ruler/niah_multivalue.yaml)
86
- 2026-07-18:12:58:03 INFO [tasks:691] Task: niah_multikey_3 (ruler/niah_multikey_3.yaml)
87
- 2026-07-18:12:58:03 INFO [tasks:691] Task: niah_multikey_1 (ruler/niah_multikey_1.yaml)
88
- 2026-07-18:12:58:03 INFO [tasks:691] Task: niah_single_2 (ruler/niah_single_2.yaml)
89
- 2026-07-18:12:58:03 INFO [evaluator:314] ruler_qa_squad: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 32, 'until': []}
90
- 2026-07-18:12:58:03 INFO [evaluator:314] ruler_cwe: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 120, 'until': []}
91
- 2026-07-18:12:58:03 INFO [evaluator:314] niah_multivalue: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
92
- 2026-07-18:12:58:03 INFO [evaluator:314] niah_multikey_3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
93
- 2026-07-18:12:58:03 INFO [evaluator:314] niah_multikey_1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
94
- 2026-07-18:12:58:03 INFO [evaluator:314] niah_single_2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
95
- 2026-07-18:12:58:03 INFO [api.task:311] Building contexts for ruler_qa_squad on rank 0...
96
 
97
-
98
  0%| | 0/20 [00:00<?, ?it/s]
99
- 2026-07-18:12:58:03 INFO [api.task:311] Building contexts for ruler_cwe on rank 0...
100
  0%| | 0/20 [00:00<?, ?it/s]
 
101
 
102
-
103
  0%| | 0/20 [00:00<?, ?it/s]
104
- 2026-07-18:12:58:03 INFO [api.task:311] Building contexts for niah_multivalue on rank 0...
105
  0%| | 0/20 [00:00<?, ?it/s]
 
106
 
107
-
108
  0%| | 0/20 [00:00<?, ?it/s]
109
- 2026-07-18:12:58:03 INFO [api.task:311] Building contexts for niah_multikey_3 on rank 0...
110
  0%| | 0/20 [00:00<?, ?it/s]
 
111
 
112
-
113
  0%| | 0/20 [00:00<?, ?it/s]
114
- 2026-07-18:12:58:03 INFO [api.task:311] Building contexts for niah_multikey_1 on rank 0...
115
  0%| | 0/20 [00:00<?, ?it/s]
 
116
 
117
-
118
  0%| | 0/20 [00:00<?, ?it/s]
119
- 2026-07-18:12:58:03 INFO [api.task:311] Building contexts for niah_single_2 on rank 0...
120
  0%| | 0/20 [00:00<?, ?it/s]
 
121
 
122
-
123
  0%| | 0/20 [00:00<?, ?it/s]
124
- 2026-07-18:12:58:03 INFO [evaluator:584] Running generate_until requests
125
  0%| | 0/20 [00:00<?, ?it/s]
 
126
 
127
 
128
-
129
-
130
-
131
-
132
-
133
-
134
-
135
-
136
-
137
-
138
-
139
-
140
-
141
-
142
-
143
-
144
-
145
-
146
-
147
-
148
-
149
-
150
-
151
-
152
-
153
-
154
-
155
-
156
-
157
-
158
-
159
-
160
-
161
-
162
-
163
-
164
-
165
-
166
-
167
-
168
-
169
-
170
-
171
-
172
-
173
-
174
-
175
-
176
-
177
-
178
-
179
-
180
-
181
-
182
-
183
-
184
-
185
-
186
-
187
-
188
-
189
-
190
-
191
-
192
-
193
-
194
-
195
-
196
-
197
-
198
-
199
-
200
-
201
-
202
-
203
-
204
-
205
-
206
-
207
-
208
-
209
-
210
-
211
-
212
-
213
-
214
-
215
-
216
-
217
-
218
-
219
-
220
-
221
-
222
-
223
-
224
-
225
-
226
-
227
-
228
-
229
-
230
-
231
-
232
-
233
-
234
-
235
-
236
-
237
-
238
-
239
-
240
-
241
-
242
-
243
-
244
-
245
-
246
-
247
-
248
- 2026-07-18:13:09:02 INFO [loggers.evaluation_tracker:247] Saving results aggregated
249
- 2026-07-18:13:09:02 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/*.jsonl
250
  hf ({'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 20.0, num_fewshot: None, batch_size: 1
251
  | Tasks |Version|Filter|n-shot|Metric| | Value | |Stderr|
252
  |---------------|------:|------|-----:|-----:|---|------:|---|------|
 
1
+ 2026-07-18:15:51:39 WARNING [config.evaluate_config:281] --limit SHOULD ONLY BE USED FOR TESTING. REAL METRICS SHOULD NOT BE COMPUTED USING LIMIT.
2
+ 2026-07-18:15:51:43 INFO [_cli.run:376] Selected Tasks: ['niah_single_2', 'niah_multikey_1', 'niah_multikey_3', 'niah_multivalue', 'ruler_cwe', 'ruler_qa_squad']
3
+ 2026-07-18:15:51:45 INFO [evaluator:211] Setting random seed to 0 | Setting numpy seed to 1234 | Setting torch manual seed to 1234 | Setting fewshot manual seed to 1234
4
+ 2026-07-18:15:51:45 INFO [evaluator:236] Initializing hf model, with arguments: {'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'trust_remote_code': True, 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}
5
+ 2026-07-18:15:51:47 INFO [models.huggingface:161] Using device 'cuda:0'
6
+ 2026-07-18:15:51:48 INFO [models.huggingface:548] Model type cannot be determined. Using default model type 'causal'
7
+ 2026-07-18:15:51:48 INFO [models.huggingface:423] Model parallel was set to False, max memory was not set, and device map was set to {'': 'cuda:0'}
8
+ 2026-07-18:15:51:57 WARNING [api.task:856] niah_single_2: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
9
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
10
+ 2026-07-18:15:51:57 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
11
 
12
 
13
+
14
+
15
+
16
+
17
+
18
+
19
+
20
+
21
+ 2026-07-18:15:52:06 WARNING [api.task:856] niah_multikey_1: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
22
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
23
 
24
 
25
+
26
+
27
+
28
+
29
+
30
+
31
+
32
+
33
+ 2026-07-18:15:52:15 WARNING [api.task:856] niah_multikey_3: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
34
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
35
 
36
 
37
+
38
+
39
+
40
+
41
+
42
+ 2026-07-18:15:52:21 WARNING [api.task:856] niah_multivalue: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
43
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
44
 
45
 
46
+
47
+
48
+
49
+
50
+
51
+
52
+
53
+
54
+ 2026-07-18:15:52:30 WARNING [api.task:856] ruler_cwe: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
55
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
56
+ 2026-07-18:15:52:30 INFO [tasks.ruler.common_utils:26] Using tokenizer /workspace/outputs/l2a_style/stage2/checkpoint-25 for synthetic tasks.
57
 
58
 
59
+
60
+
61
+
62
+
63
+
64
+ 2026-07-18:15:52:37 WARNING [api.task:856] ruler_qa_squad: Custom kwargs can be passed to `--metadata` in console (as json string) or to the TaskManager.
 
65
  For example --metadata='{"max_seq_lengths":[4096, 8192]}'. For details see task Readme.
66
 
67
 
68
+
69
+
70
+
71
+
72
+
73
+
74
+
75
+
76
+ 2026-07-18:15:52:46 INFO [tasks:700] Selected tasks:
77
+ 2026-07-18:15:52:46 INFO [tasks:691] Task: ruler_qa_squad (ruler/qa_squad.yaml)
78
+ 2026-07-18:15:52:46 INFO [tasks:691] Task: ruler_cwe (ruler/cwe.yaml)
79
+ 2026-07-18:15:52:46 INFO [tasks:691] Task: niah_multivalue (ruler/niah_multivalue.yaml)
80
+ 2026-07-18:15:52:46 INFO [tasks:691] Task: niah_multikey_3 (ruler/niah_multikey_3.yaml)
81
+ 2026-07-18:15:52:46 INFO [tasks:691] Task: niah_multikey_1 (ruler/niah_multikey_1.yaml)
82
+ 2026-07-18:15:52:46 INFO [tasks:691] Task: niah_single_2 (ruler/niah_single_2.yaml)
83
+ 2026-07-18:15:52:46 INFO [evaluator:314] ruler_qa_squad: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 32, 'until': []}
84
+ 2026-07-18:15:52:46 INFO [evaluator:314] ruler_cwe: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 120, 'until': []}
85
+ 2026-07-18:15:52:46 INFO [evaluator:314] niah_multivalue: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
86
+ 2026-07-18:15:52:46 INFO [evaluator:314] niah_multikey_3: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
87
+ 2026-07-18:15:52:46 INFO [evaluator:314] niah_multikey_1: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
88
+ 2026-07-18:15:52:46 INFO [evaluator:314] niah_single_2: Using gen_kwargs: {'do_sample': False, 'temperature': 0.0, 'max_gen_toks': 128, 'until': []}
89
+ 2026-07-18:15:52:47 INFO [api.task:311] Building contexts for ruler_qa_squad on rank 0...
 
90
 
 
91
  0%| | 0/20 [00:00<?, ?it/s]
92
+
93
  0%| | 0/20 [00:00<?, ?it/s]
94
+ 2026-07-18:15:52:47 INFO [api.task:311] Building contexts for ruler_cwe on rank 0...
95
 
 
96
  0%| | 0/20 [00:00<?, ?it/s]
97
+
98
  0%| | 0/20 [00:00<?, ?it/s]
99
+ 2026-07-18:15:52:47 INFO [api.task:311] Building contexts for niah_multivalue on rank 0...
100
 
 
101
  0%| | 0/20 [00:00<?, ?it/s]
102
+
103
  0%| | 0/20 [00:00<?, ?it/s]
104
+ 2026-07-18:15:52:47 INFO [api.task:311] Building contexts for niah_multikey_3 on rank 0...
105
 
 
106
  0%| | 0/20 [00:00<?, ?it/s]
107
+
108
  0%| | 0/20 [00:00<?, ?it/s]
109
+ 2026-07-18:15:52:47 INFO [api.task:311] Building contexts for niah_multikey_1 on rank 0...
110
 
 
111
  0%| | 0/20 [00:00<?, ?it/s]
112
+
113
  0%| | 0/20 [00:00<?, ?it/s]
114
+ 2026-07-18:15:52:47 INFO [api.task:311] Building contexts for niah_single_2 on rank 0...
115
 
 
116
  0%| | 0/20 [00:00<?, ?it/s]
117
+
118
  0%| | 0/20 [00:00<?, ?it/s]
119
+ 2026-07-18:15:52:47 INFO [evaluator:584] Running generate_until requests
120
 
121
 
122
+
123
+
124
+
125
+
126
+
127
+
128
+
129
+
130
+
131
+
132
+
133
+
134
+
135
+
136
+
137
+
138
+
139
+
140
+
141
+
142
+
143
+
144
+
145
+
146
+
147
+
148
+
149
+
150
+
151
+
152
+
153
+
154
+
155
+
156
+
157
+
158
+
159
+
160
+
161
+
162
+
163
+
164
+
165
+
166
+
167
+
168
+
169
+
170
+
171
+
172
+
173
+
174
+
175
+
176
+
177
+
178
+
179
+
180
+
181
+
182
+
183
+
184
+
185
+
186
+
187
+
188
+
189
+
190
+
191
+
192
+
193
+
194
+
195
+
196
+
197
+
198
+
199
+
200
+
201
+
202
+
203
+
204
+
205
+
206
+
207
+
208
+
209
+
210
+
211
+
212
+
213
+
214
+
215
+
216
+
217
+
218
+
219
+
220
+
221
+
222
+
223
+
224
+
225
+
226
+
227
+
228
+
229
+
230
+
231
+
232
+
233
+
234
+
235
+
236
+
237
+
238
+
239
+
240
+
241
+
242
+ 2026-07-18:16:04:28 INFO [loggers.evaluation_tracker:247] Saving results aggregated
243
+ 2026-07-18:16:04:28 INFO [loggers.evaluation_tracker:119] Saving per-task samples to /workspace/outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/*.jsonl
244
  hf ({'pretrained': '/workspace/outputs/l2a_style/stage2/checkpoint-25', 'dtype': 'bfloat16', 'max_length': 16384, 'attn_implementation': 'sdpa'}), gen_kwargs: ({}), limit: 20.0, num_fewshot: None, batch_size: 1
245
  | Tasks |Version|Filter|n-shot|Metric| | Value | |Stderr|
246
  |---------------|------:|------|-----:|-----:|---|------:|---|------|
outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/results_2026-07-18T16-10-15.187546.json ADDED
@@ -0,0 +1,505 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "results": {
3
+ "babilong_longctx": {
4
+ "acc,none": 0.608,
5
+ "acc_stderr,none": 0.029548990797366614,
6
+ "alias": "babilong_longctx"
7
+ },
8
+ "babilong_qa1": {
9
+ "alias": " - babilong_qa1",
10
+ "acc,none": 0.74,
11
+ "acc_stderr,none": 0.06266203485560373
12
+ },
13
+ "babilong_qa2": {
14
+ "alias": " - babilong_qa2",
15
+ "acc,none": 0.52,
16
+ "acc_stderr,none": 0.0713714056959817
17
+ },
18
+ "babilong_qa3": {
19
+ "alias": " - babilong_qa3",
20
+ "acc,none": 0.34,
21
+ "acc_stderr,none": 0.06767268161329719
22
+ },
23
+ "babilong_qa4": {
24
+ "alias": " - babilong_qa4",
25
+ "acc,none": 0.72,
26
+ "acc_stderr,none": 0.06414269805898185
27
+ },
28
+ "babilong_qa5": {
29
+ "alias": " - babilong_qa5",
30
+ "acc,none": 0.72,
31
+ "acc_stderr,none": 0.06414269805898185
32
+ }
33
+ },
34
+ "groups": {
35
+ "babilong_longctx": {
36
+ "acc,none": 0.608,
37
+ "acc_stderr,none": 0.029548990797366614,
38
+ "alias": "babilong_longctx"
39
+ }
40
+ },
41
+ "group_subtasks": {
42
+ "babilong_longctx": [
43
+ "babilong_qa1",
44
+ "babilong_qa2",
45
+ "babilong_qa3",
46
+ "babilong_qa4",
47
+ "babilong_qa5"
48
+ ]
49
+ },
50
+ "configs": {
51
+ "babilong_qa1": {
52
+ "task": "babilong_qa1",
53
+ "custom_dataset": "def load_dataset(**kwargs):\n config_name = kwargs.get(\"max_seq_lengths\", \"0k\")\n\n # Get specific qa split\n qa_split = kwargs.get(\"qa_split\")\n\n eval_logger.info(\n f\"Loading babilong dataset: max_seq_lengths={config_name}, split={qa_split}\"\n )\n dataset = datasets.load_dataset(\n \"RMT-team/babilong-1k-samples\", name=config_name, split=qa_split\n )\n return {qa_split: dataset}\n",
54
+ "dataset_path": "RMT-team/babilong-1k-samples",
55
+ "dataset_kwargs": {
56
+ "qa_split": "qa1"
57
+ },
58
+ "test_split": "qa1",
59
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
60
+ "doc_to_target": "{{target}}",
61
+ "unsafe_code": false,
62
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n pred = postprocess_pred(results)\n target = doc.get(\"target\", \"\").strip()\n\n # String match\n score = 1.0 if target.lower() in pred[0].lower() else 0.0\n\n return {\"acc\": score}\n",
63
+ "description": "I will give you context with the facts about positions of different persons hidden in some random text and a question. You need to answer the question based only on the information from the facts. If a person was in different locations, use the latest location to answer the question.\nAlways return your answer in the following format:\nThe most recent location of 'person' is 'location'. Do not write anything else after that.\n\n",
64
+ "target_delimiter": " ",
65
+ "fewshot_delimiter": "\n\n",
66
+ "fewshot_config": {
67
+ "sampler": "first_n",
68
+ "split": null,
69
+ "process_docs": null,
70
+ "fewshot_indices": null,
71
+ "samples": [
72
+ {
73
+ "input": "Charlie went to the hallway. Judith come back to the kitchen. Charlie travelled to balcony.",
74
+ "question": "Where is Charlie?",
75
+ "target": "The most recent location of Charlie is balcony."
76
+ },
77
+ {
78
+ "input": "Alan moved to the garage. Charlie went to the beach. Alan went to the shop. Rouse travelled to balcony.",
79
+ "question": "Where is Alan?",
80
+ "target": "The most recent location of Alan is shop."
81
+ }
82
+ ],
83
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
84
+ "doc_to_choice": null,
85
+ "doc_to_target": "{{target}}",
86
+ "gen_prefix": null,
87
+ "fewshot_delimiter": "\n\n",
88
+ "target_delimiter": " "
89
+ },
90
+ "num_fewshot": 2,
91
+ "metric_list": [
92
+ {
93
+ "metric": "acc",
94
+ "aggregation": "mean",
95
+ "higher_is_better": true
96
+ }
97
+ ],
98
+ "output_type": "generate_until",
99
+ "generation_kwargs": {
100
+ "do_sample": false,
101
+ "temperature": 0.0,
102
+ "max_gen_toks": 16,
103
+ "until": []
104
+ },
105
+ "repeats": 1,
106
+ "should_decontaminate": false,
107
+ "metadata": {
108
+ "version": 0.0,
109
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
110
+ "trust_remote_code": true,
111
+ "dtype": "bfloat16",
112
+ "max_length": 16384,
113
+ "attn_implementation": "sdpa",
114
+ "max_seq_lengths": "8k"
115
+ }
116
+ },
117
+ "babilong_qa2": {
118
+ "task": "babilong_qa2",
119
+ "custom_dataset": "def load_dataset(**kwargs):\n config_name = kwargs.get(\"max_seq_lengths\", \"0k\")\n\n # Get specific qa split\n qa_split = kwargs.get(\"qa_split\")\n\n eval_logger.info(\n f\"Loading babilong dataset: max_seq_lengths={config_name}, split={qa_split}\"\n )\n dataset = datasets.load_dataset(\n \"RMT-team/babilong-1k-samples\", name=config_name, split=qa_split\n )\n return {qa_split: dataset}\n",
120
+ "dataset_path": "RMT-team/babilong-1k-samples",
121
+ "dataset_kwargs": {
122
+ "qa_split": "qa2"
123
+ },
124
+ "test_split": "qa2",
125
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
126
+ "doc_to_target": "{{target}}",
127
+ "unsafe_code": false,
128
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n pred = postprocess_pred(results)\n target = doc.get(\"target\", \"\").strip()\n\n # String match\n score = 1.0 if target.lower() in pred[0].lower() else 0.0\n\n return {\"acc\": score}\n",
129
+ "description": "I will give you context with the facts about locations and actions of different persons hidden in some random text and a question. You need to answer the question based only on the information from the facts. If a person got an item in the first location and travelled to the second location the item is also in the second location. If a person dropped an item in the first location and moved to the second location the item remains in the first location.\nAlways return your answer in the following format:\nThe 'item' is in 'location'. Do not write anything else after that.\n\n",
130
+ "target_delimiter": " ",
131
+ "fewshot_delimiter": "\n\n",
132
+ "fewshot_config": {
133
+ "sampler": "first_n",
134
+ "split": null,
135
+ "process_docs": null,
136
+ "fewshot_indices": null,
137
+ "samples": [
138
+ {
139
+ "input": "Charlie went to the kitchen. Charlie got a bottle. Charlie moved to the balcony.",
140
+ "question": "Where is the bottle?",
141
+ "target": "The bottle is in the balcony."
142
+ },
143
+ {
144
+ "input": "Alan moved to the garage. Alan got a screw driver. Alan moved to the kitchen.",
145
+ "question": "Where is the screw driver?",
146
+ "target": "The screw driver is in the kitchen."
147
+ }
148
+ ],
149
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
150
+ "doc_to_choice": null,
151
+ "doc_to_target": "{{target}}",
152
+ "gen_prefix": null,
153
+ "fewshot_delimiter": "\n\n",
154
+ "target_delimiter": " "
155
+ },
156
+ "num_fewshot": 2,
157
+ "metric_list": [
158
+ {
159
+ "metric": "acc",
160
+ "aggregation": "mean",
161
+ "higher_is_better": true
162
+ }
163
+ ],
164
+ "output_type": "generate_until",
165
+ "generation_kwargs": {
166
+ "do_sample": false,
167
+ "temperature": 0.0,
168
+ "max_gen_toks": 16,
169
+ "until": []
170
+ },
171
+ "repeats": 1,
172
+ "should_decontaminate": false,
173
+ "metadata": {
174
+ "version": 0.0,
175
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
176
+ "trust_remote_code": true,
177
+ "dtype": "bfloat16",
178
+ "max_length": 16384,
179
+ "attn_implementation": "sdpa",
180
+ "max_seq_lengths": "8k"
181
+ }
182
+ },
183
+ "babilong_qa3": {
184
+ "task": "babilong_qa3",
185
+ "custom_dataset": "def load_dataset(**kwargs):\n config_name = kwargs.get(\"max_seq_lengths\", \"0k\")\n\n # Get specific qa split\n qa_split = kwargs.get(\"qa_split\")\n\n eval_logger.info(\n f\"Loading babilong dataset: max_seq_lengths={config_name}, split={qa_split}\"\n )\n dataset = datasets.load_dataset(\n \"RMT-team/babilong-1k-samples\", name=config_name, split=qa_split\n )\n return {qa_split: dataset}\n",
186
+ "dataset_path": "RMT-team/babilong-1k-samples",
187
+ "dataset_kwargs": {
188
+ "qa_split": "qa3"
189
+ },
190
+ "test_split": "qa3",
191
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
192
+ "doc_to_target": "{{target}}",
193
+ "unsafe_code": false,
194
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n pred = postprocess_pred(results)\n target = doc.get(\"target\", \"\").strip()\n\n # String match\n score = 1.0 if target.lower() in pred[0].lower() else 0.0\n\n return {\"acc\": score}\n",
195
+ "description": "I give you context with the facts about locations and actions of different persons hidden in some random text and a question. You need to answer the question based only on the information from the facts. If a person got an item in the first location and travelled to the second location the item is also in the second location. If a person dropped an item in the first location and moved to the second location the item remains in the first location.\nAlways return your answer in the following format:\nBefore the $location_1$ the $item$ was in the $location_2$. Do not write anything else after that.\n\n",
196
+ "target_delimiter": " ",
197
+ "fewshot_delimiter": "\n\n",
198
+ "fewshot_config": {
199
+ "sampler": "first_n",
200
+ "split": null,
201
+ "process_docs": null,
202
+ "fewshot_indices": null,
203
+ "samples": [
204
+ {
205
+ "input": "John journeyed to the bedroom. Mary grabbed the apple. Mary went back to the bathroom. Daniel journeyed to the bedroom. Daniel moved to the garden. Mary travelled to the kitchen.",
206
+ "question": "Where was the apple before the kitchen?",
207
+ "target": "Before the kitchen the apple was in the bathroom."
208
+ },
209
+ {
210
+ "input": "John went back to the bedroom. John went back to the garden. John went back to the kitchen. Sandra took the football. Sandra travelled to the garden. Sandra journeyed to the bedroom.",
211
+ "question": "Where was the football before the bedroom?",
212
+ "target": "Before the bedroom the football was in the garden."
213
+ }
214
+ ],
215
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
216
+ "doc_to_choice": null,
217
+ "doc_to_target": "{{target}}",
218
+ "gen_prefix": null,
219
+ "fewshot_delimiter": "\n\n",
220
+ "target_delimiter": " "
221
+ },
222
+ "num_fewshot": 2,
223
+ "metric_list": [
224
+ {
225
+ "metric": "acc",
226
+ "aggregation": "mean",
227
+ "higher_is_better": true
228
+ }
229
+ ],
230
+ "output_type": "generate_until",
231
+ "generation_kwargs": {
232
+ "do_sample": false,
233
+ "temperature": 0.0,
234
+ "max_gen_toks": 16,
235
+ "until": []
236
+ },
237
+ "repeats": 1,
238
+ "should_decontaminate": false,
239
+ "metadata": {
240
+ "version": 0.0,
241
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
242
+ "trust_remote_code": true,
243
+ "dtype": "bfloat16",
244
+ "max_length": 16384,
245
+ "attn_implementation": "sdpa",
246
+ "max_seq_lengths": "8k"
247
+ }
248
+ },
249
+ "babilong_qa4": {
250
+ "task": "babilong_qa4",
251
+ "custom_dataset": "def load_dataset(**kwargs):\n config_name = kwargs.get(\"max_seq_lengths\", \"0k\")\n\n # Get specific qa split\n qa_split = kwargs.get(\"qa_split\")\n\n eval_logger.info(\n f\"Loading babilong dataset: max_seq_lengths={config_name}, split={qa_split}\"\n )\n dataset = datasets.load_dataset(\n \"RMT-team/babilong-1k-samples\", name=config_name, split=qa_split\n )\n return {qa_split: dataset}\n",
252
+ "dataset_path": "RMT-team/babilong-1k-samples",
253
+ "dataset_kwargs": {
254
+ "qa_split": "qa4"
255
+ },
256
+ "test_split": "qa4",
257
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
258
+ "doc_to_target": "{{target}}",
259
+ "unsafe_code": false,
260
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n pred = postprocess_pred(results)\n target = doc.get(\"target\", \"\").strip()\n\n # String match\n score = 1.0 if target.lower() in pred[0].lower() else 0.0\n\n return {\"acc\": score}\n",
261
+ "description": "I will give you context with the facts about different people, their location and actions, hidden in some random text and a question. You need to answer the question based only on the information from the facts.\nYour answer should contain only one word - location. Do not write anything else after that.\n\n",
262
+ "target_delimiter": " ",
263
+ "fewshot_delimiter": "\n\n",
264
+ "fewshot_config": {
265
+ "sampler": "first_n",
266
+ "split": null,
267
+ "process_docs": null,
268
+ "fewshot_indices": null,
269
+ "samples": [
270
+ {
271
+ "input": "The hallway is south of the kitchen. The bedroom is north of the kitchen.",
272
+ "question": "What is the kitchen south of?",
273
+ "target": "bedroom"
274
+ },
275
+ {
276
+ "input": "The garden is west of the bedroom. The bedroom is west of the kitchen.",
277
+ "question": "What is west of the bedroom?",
278
+ "target": "garden"
279
+ }
280
+ ],
281
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
282
+ "doc_to_choice": null,
283
+ "doc_to_target": "{{target}}",
284
+ "gen_prefix": null,
285
+ "fewshot_delimiter": "\n\n",
286
+ "target_delimiter": " "
287
+ },
288
+ "num_fewshot": 2,
289
+ "metric_list": [
290
+ {
291
+ "metric": "acc",
292
+ "aggregation": "mean",
293
+ "higher_is_better": true
294
+ }
295
+ ],
296
+ "output_type": "generate_until",
297
+ "generation_kwargs": {
298
+ "do_sample": false,
299
+ "temperature": 0.0,
300
+ "max_gen_toks": 16,
301
+ "until": []
302
+ },
303
+ "repeats": 1,
304
+ "should_decontaminate": false,
305
+ "metadata": {
306
+ "version": 0.0,
307
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
308
+ "trust_remote_code": true,
309
+ "dtype": "bfloat16",
310
+ "max_length": 16384,
311
+ "attn_implementation": "sdpa",
312
+ "max_seq_lengths": "8k"
313
+ }
314
+ },
315
+ "babilong_qa5": {
316
+ "task": "babilong_qa5",
317
+ "custom_dataset": "def load_dataset(**kwargs):\n config_name = kwargs.get(\"max_seq_lengths\", \"0k\")\n\n # Get specific qa split\n qa_split = kwargs.get(\"qa_split\")\n\n eval_logger.info(\n f\"Loading babilong dataset: max_seq_lengths={config_name}, split={qa_split}\"\n )\n dataset = datasets.load_dataset(\n \"RMT-team/babilong-1k-samples\", name=config_name, split=qa_split\n )\n return {qa_split: dataset}\n",
318
+ "dataset_path": "RMT-team/babilong-1k-samples",
319
+ "dataset_kwargs": {
320
+ "qa_split": "qa5"
321
+ },
322
+ "test_split": "qa5",
323
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
324
+ "doc_to_target": "{{target}}",
325
+ "unsafe_code": false,
326
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n pred = postprocess_pred(results)\n target = doc.get(\"target\", \"\").strip()\n\n # String match\n score = 1.0 if target.lower() in pred[0].lower() else 0.0\n\n return {\"acc\": score}\n",
327
+ "description": "I will give you context with the facts about locations and their relations hidden in some random text and a question. You need to answer the question based only on the information from the facts.\nYour answer should contain only one word. Do not write anything else after that. Do not explain your answer.\n\n",
328
+ "target_delimiter": " ",
329
+ "fewshot_delimiter": "\n\n",
330
+ "fewshot_config": {
331
+ "sampler": "first_n",
332
+ "split": null,
333
+ "process_docs": null,
334
+ "fewshot_indices": null,
335
+ "samples": [
336
+ {
337
+ "input": "Mary picked up the apple there. Mary gave the apple to Fred. Mary moved to the bedroom. Bill took the milk there.",
338
+ "question": "Who did Mary give the apple to?",
339
+ "target": "Fred"
340
+ },
341
+ {
342
+ "input": "Jeff took the football there. Jeff passed the football to Fred. Jeff got the milk there. Bill travelled to the bedroom.",
343
+ "question": "Who gave the football?",
344
+ "target": "Jeff"
345
+ },
346
+ {
347
+ "input": "Fred picked up the apple there. Fred handed the apple to Bill. Bill journeyed to the bedroom. Jeff went back to the garden.",
348
+ "question": "What did Fred give to Bill?",
349
+ "target": "apple"
350
+ }
351
+ ],
352
+ "doc_to_text": "{{input.strip()}}\n{{question.strip()}}",
353
+ "doc_to_choice": null,
354
+ "doc_to_target": "{{target}}",
355
+ "gen_prefix": null,
356
+ "fewshot_delimiter": "\n\n",
357
+ "target_delimiter": " "
358
+ },
359
+ "num_fewshot": 2,
360
+ "metric_list": [
361
+ {
362
+ "metric": "acc",
363
+ "aggregation": "mean",
364
+ "higher_is_better": true
365
+ }
366
+ ],
367
+ "output_type": "generate_until",
368
+ "generation_kwargs": {
369
+ "do_sample": false,
370
+ "temperature": 0.0,
371
+ "max_gen_toks": 16,
372
+ "until": []
373
+ },
374
+ "repeats": 1,
375
+ "should_decontaminate": false,
376
+ "metadata": {
377
+ "version": 0.0,
378
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
379
+ "trust_remote_code": true,
380
+ "dtype": "bfloat16",
381
+ "max_length": 16384,
382
+ "attn_implementation": "sdpa",
383
+ "max_seq_lengths": "8k"
384
+ }
385
+ }
386
+ },
387
+ "versions": {
388
+ "babilong_longctx": 0.0,
389
+ "babilong_qa1": 0.0,
390
+ "babilong_qa2": 0.0,
391
+ "babilong_qa3": 0.0,
392
+ "babilong_qa4": 0.0,
393
+ "babilong_qa5": 0.0
394
+ },
395
+ "n-shot": {
396
+ "babilong_qa1": 2,
397
+ "babilong_qa2": 2,
398
+ "babilong_qa3": 2,
399
+ "babilong_qa4": 2,
400
+ "babilong_qa5": 2
401
+ },
402
+ "higher_is_better": {
403
+ "babilong_longctx": {
404
+ "acc": true
405
+ },
406
+ "babilong_qa1": {
407
+ "acc": true
408
+ },
409
+ "babilong_qa2": {
410
+ "acc": true
411
+ },
412
+ "babilong_qa3": {
413
+ "acc": true
414
+ },
415
+ "babilong_qa4": {
416
+ "acc": true
417
+ },
418
+ "babilong_qa5": {
419
+ "acc": true
420
+ }
421
+ },
422
+ "n-samples": {
423
+ "babilong_qa1": {
424
+ "original": 1000,
425
+ "effective": 50
426
+ },
427
+ "babilong_qa2": {
428
+ "original": 999,
429
+ "effective": 50
430
+ },
431
+ "babilong_qa3": {
432
+ "original": 999,
433
+ "effective": 50
434
+ },
435
+ "babilong_qa4": {
436
+ "original": 999,
437
+ "effective": 50
438
+ },
439
+ "babilong_qa5": {
440
+ "original": 999,
441
+ "effective": 50
442
+ }
443
+ },
444
+ "config": {
445
+ "model": "hf",
446
+ "model_args": {
447
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
448
+ "trust_remote_code": true,
449
+ "dtype": "bfloat16",
450
+ "max_length": 16384,
451
+ "attn_implementation": "sdpa"
452
+ },
453
+ "model_num_parameters": 2031969308,
454
+ "model_dtype": "torch.bfloat16",
455
+ "model_revision": "main",
456
+ "model_sha": "",
457
+ "batch_size": "1",
458
+ "batch_sizes": [],
459
+ "device": "cuda:0",
460
+ "use_cache": null,
461
+ "limit": 50.0,
462
+ "bootstrap_iters": 100000,
463
+ "gen_kwargs": {},
464
+ "random_seed": 0,
465
+ "numpy_seed": 1234,
466
+ "torch_seed": 1234,
467
+ "fewshot_seed": 1234
468
+ },
469
+ "git_hash": null,
470
+ "date": 1784390674.4487982,
471
+ "pretty_env_info": "PyTorch version: 2.9.1+cu128\nIs debug build: False\nCUDA used to build PyTorch: 12.8\nROCM used to build PyTorch: N/A\n\nOS: Ubuntu 22.04.5 LTS (x86_64)\nGCC version: Could not collect\nClang version: Could not collect\nCMake version: version 4.1.2\nLibc version: glibc-2.35\n\nPython version: 3.11.14 | packaged by conda-forge | (main, Oct 22 2025, 22:46:25) [GCC 14.3.0] (64-bit runtime)\nPython platform: Linux-6.12.90-120.164.amzn2023.x86_64-x86_64-with-glibc2.35\nIs CUDA available: True\nCUDA runtime version: Could not collect\nCUDA_MODULE_LOADING set to: \nGPU models and configuration: GPU 0: NVIDIA A100-SXM4-80GB\nNvidia driver version: 580.159.03\ncuDNN version: Could not collect\nIs XPU available: False\nHIP runtime version: N/A\nMIOpen runtime version: N/A\nIs XNNPACK available: True\n\nCPU:\nArchitecture: x86_64\nCPU op-mode(s): 32-bit, 64-bit\nAddress sizes: 46 bits physical, 48 bits virtual\nByte Order: Little Endian\nCPU(s): 96\nOn-line CPU(s) list: 0-95\nVendor ID: GenuineIntel\nModel name: Intel(R) Xeon(R) Platinum 8275CL CPU @ 3.00GHz\nCPU family: 6\nModel: 85\nThread(s) per core: 2\nCore(s) per socket: 24\nSocket(s): 2\nStepping: 7\nBogoMIPS: 5999.99\nFlags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon rep_good nopl xtopology nonstop_tsc cpuid aperfmperf tsc_known_freq pni pclmulqdq monitor ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand hypervisor lahf_lm abm 3dnowprefetch pti fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid mpx avx512f avx512dq rdseed adx smap clflushopt clwb avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves ida arat pku ospke\nHypervisor vendor: KVM\nVirtualization type: full\nL1d cache: 1.5 MiB (48 instances)\nL1i cache: 1.5 MiB (48 instances)\nL2 cache: 48 MiB (48 instances)\nL3 cache: 71.5 MiB (2 instances)\nNUMA node(s): 2\nNUMA node0 CPU(s): 0-23,48-71\nNUMA node1 CPU(s): 24-47,72-95\nVulnerability Gather data sampling: Unknown: Dependent on hypervisor status\nVulnerability Indirect target selection: Mitigation; Aligned branch/return thunks\nVulnerability Itlb multihit: KVM: Mitigation: VMX unsupported\nVulnerability L1tf: Mitigation; PTE Inversion\nVulnerability Mds: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown\nVulnerability Meltdown: Mitigation; PTI\nVulnerability Mmio stale data: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown\nVulnerability Reg file data sampling: Not affected\nVulnerability Retbleed: Vulnerable\nVulnerability Spec rstack overflow: Not affected\nVulnerability Spec store bypass: Vulnerable\nVulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization\nVulnerability Spectre v2: Mitigation; Retpolines; STIBP disabled; RSB filling; PBRSB-eIBRS Not affected; BHI Retpoline\nVulnerability Srbds: Not affected\nVulnerability Tsa: Not affected\nVulnerability Tsx async abort: Not affected\nVulnerability Vmscape: Not affected\n\nVersions of relevant libraries:\n[pip3] numpy==2.3.4\n[pip3] nvidia-cublas-cu12==12.8.4.1\n[pip3] nvidia-cuda-cupti-cu12==12.8.90\n[pip3] nvidia-cuda-nvrtc-cu12==12.8.93\n[pip3] nvidia-cuda-runtime-cu12==12.8.90\n[pip3] nvidia-cudnn-cu12==9.10.2.21\n[pip3] nvidia-cufft-cu12==11.3.3.83\n[pip3] nvidia-curand-cu12==10.3.9.90\n[pip3] nvidia-cusolver-cu12==11.7.3.90\n[pip3] nvidia-cusparse-cu12==12.5.8.93\n[pip3] nvidia-cusparselt-cu12==0.7.1\n[pip3] nvidia-nccl-cu12==2.27.5\n[pip3] nvidia-nvjitlink-cu12==12.8.93\n[pip3] nvidia-nvtx-cu12==12.8.90\n[pip3] optree==0.17.0\n[pip3] torch==2.9.1+cu128\n[pip3] torchaudio==2.9.1+cu128\n[pip3] torchelastic==0.2.2\n[pip3] torchvision==0.24.1+cu128\n[pip3] triton==3.5.1\n[conda] numpy 2.3.4 py311h2e04523_0 conda-forge\n[conda] nvidia-cublas-cu12 12.8.4.1 pypi_0 pypi\n[conda] nvidia-cuda-cupti-cu12 12.8.90 pypi_0 pypi\n[conda] nvidia-cuda-nvrtc-cu12 12.8.93 pypi_0 pypi\n[conda] nvidia-cuda-runtime-cu12 12.8.90 pypi_0 pypi\n[conda] nvidia-cudnn-cu12 9.10.2.21 pypi_0 pypi\n[conda] nvidia-cufft-cu12 11.3.3.83 pypi_0 pypi\n[conda] nvidia-curand-cu12 10.3.9.90 pypi_0 pypi\n[conda] nvidia-cusolver-cu12 11.7.3.90 pypi_0 pypi\n[conda] nvidia-cusparse-cu12 12.5.8.93 pypi_0 pypi\n[conda] nvidia-cusparselt-cu12 0.7.1 pypi_0 pypi\n[conda] nvidia-nccl-cu12 2.27.5 pypi_0 pypi\n[conda] nvidia-nvjitlink-cu12 12.8.93 pypi_0 pypi\n[conda] nvidia-nvtx-cu12 12.8.90 pypi_0 pypi\n[conda] optree 0.17.0 pypi_0 pypi\n[conda] torch 2.9.1+cu128 pypi_0 pypi\n[conda] torchaudio 2.9.1+cu128 pypi_0 pypi\n[conda] torchelastic 0.2.2 pypi_0 pypi\n[conda] torchvision 0.24.1+cu128 pypi_0 pypi\n[conda] triton 3.5.1 pypi_0 pypi",
472
+ "transformers_version": "4.54.0",
473
+ "lm_eval_version": "0.4.11",
474
+ "upper_git_hash": null,
475
+ "tokenizer_pad_token": [
476
+ "<|endoftext|>",
477
+ "151643"
478
+ ],
479
+ "tokenizer_eos_token": [
480
+ "<|im_end|>",
481
+ "151645"
482
+ ],
483
+ "tokenizer_bos_token": [
484
+ null,
485
+ "None"
486
+ ],
487
+ "eot_token_id": 151645,
488
+ "max_length": 16384,
489
+ "task_hashes": {
490
+ "babilong_qa1": "bbdcd109bea1bb245e118e814bd0563de6384bc875a21df5218ab4ed75b27380",
491
+ "babilong_qa2": "fa4f8d12bdd481e02a1051414304811f648caf7bce8f89dcf03ecbf5b86f1e92",
492
+ "babilong_qa3": "737b64396c8b2e5f96da9796bcd5e448f08e5a13e32ae7b610d31a60269b6e48",
493
+ "babilong_qa4": "298a25727c69c9babddcd78f438cae1aa21a388e14e7db90b2206eff753c90c0",
494
+ "babilong_qa5": "e60afcdcdd1c4b3e24b002ffbf8f87085f5ec4a02fcd72318f042dcdafe230f4"
495
+ },
496
+ "model_source": "hf",
497
+ "model_name": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
498
+ "model_name_sanitized": "__workspace__outputs__l2a_style__stage2__checkpoint-25",
499
+ "system_instruction": null,
500
+ "system_instruction_sha": null,
501
+ "fewshot_as_multiturn": null,
502
+ "chat_template": null,
503
+ "chat_template_sha": null,
504
+ "total_evaluation_time_seconds": "344.00240219797706"
505
+ }
outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa1_2026-07-18T16-10-15.187546.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:894b7b485f12742da28f607918847453d7bbffaf5c672ac7535694d830422e83
3
+ size 3164092
outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa2_2026-07-18T16-10-15.187546.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:954df79f19d6f4037ef1ba1e46baf6dc9c2d30a727f2cf4085997f2ff006dd4e
3
+ size 3147516
outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa3_2026-07-18T16-10-15.187546.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:3ef1a161d57d61ed368024971f3a0e608ce70d56782ec8c97b20cc3f3ecfccea
3
+ size 3126962
outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa4_2026-07-18T16-10-15.187546.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:0c40e39f34fff163a0fe5c131806e61937bb4e5d6e7b343c832691b828c3e54f
3
+ size 3137620
outputs/eval/token_t0525/babilong8k_qa1_qa5_n50/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_babilong_qa5_2026-07-18T16-10-15.187546.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5dddf97f2f532bec678e21782ffb214a1d795c62adb06c291d394b8f43b287c6
3
+ size 3135600
outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/results_2026-07-18T15-51-37.349704.json ADDED
@@ -0,0 +1,821 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "results": {
3
+ "niah_multikey_2": {
4
+ "alias": "niah_multikey_2",
5
+ "4096,none": -1,
6
+ "4096_stderr,none": "N/A",
7
+ "8192,none": 1.0,
8
+ "8192_stderr,none": "N/A"
9
+ },
10
+ "niah_multiquery": {
11
+ "alias": "niah_multiquery",
12
+ "4096,none": -1,
13
+ "4096_stderr,none": "N/A",
14
+ "8192,none": 1.0,
15
+ "8192_stderr,none": "N/A"
16
+ },
17
+ "niah_single_1": {
18
+ "alias": "niah_single_1",
19
+ "4096,none": -1,
20
+ "4096_stderr,none": "N/A",
21
+ "8192,none": 1.0,
22
+ "8192_stderr,none": "N/A"
23
+ },
24
+ "niah_single_3": {
25
+ "alias": "niah_single_3",
26
+ "4096,none": -1,
27
+ "4096_stderr,none": "N/A",
28
+ "8192,none": 1.0,
29
+ "8192_stderr,none": "N/A"
30
+ },
31
+ "ruler_fwe": {
32
+ "alias": "ruler_fwe",
33
+ "4096,none": -1,
34
+ "4096_stderr,none": "N/A",
35
+ "8192,none": 0.7666666666666666,
36
+ "8192_stderr,none": "N/A"
37
+ },
38
+ "ruler_qa_hotpot": {
39
+ "alias": "ruler_qa_hotpot",
40
+ "4096,none": -1,
41
+ "4096_stderr,none": "N/A",
42
+ "8192,none": 0.4,
43
+ "8192_stderr,none": "N/A"
44
+ },
45
+ "ruler_vt": {
46
+ "alias": "ruler_vt",
47
+ "4096,none": -1,
48
+ "4096_stderr,none": "N/A",
49
+ "8192,none": 0.8800000000000002,
50
+ "8192_stderr,none": "N/A"
51
+ }
52
+ },
53
+ "group_subtasks": {
54
+ "niah_single_1": [],
55
+ "niah_single_3": [],
56
+ "niah_multikey_2": [],
57
+ "niah_multiquery": [],
58
+ "ruler_vt": [],
59
+ "ruler_fwe": [],
60
+ "ruler_qa_hotpot": []
61
+ },
62
+ "configs": {
63
+ "niah_multikey_2": {
64
+ "task": "niah_multikey_2",
65
+ "tag": [
66
+ "longcxt"
67
+ ],
68
+ "custom_dataset": "def niah_multikey_2(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"needle\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"needle\",\n type_needle_k=\"words\",\n type_needle_v=\"numbers\",\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
69
+ "dataset_path": "",
70
+ "dataset_name": "",
71
+ "test_split": "test",
72
+ "doc_to_text": "{{input}}",
73
+ "doc_to_target": "{{outputs}}",
74
+ "unsafe_code": false,
75
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
76
+ "description": "",
77
+ "target_delimiter": " ",
78
+ "fewshot_delimiter": "\n\n",
79
+ "fewshot_config": {
80
+ "sampler": "default",
81
+ "split": null,
82
+ "process_docs": null,
83
+ "fewshot_indices": null,
84
+ "samples": null,
85
+ "doc_to_text": "{{input}}",
86
+ "doc_to_choice": null,
87
+ "doc_to_target": "{{outputs}}",
88
+ "gen_prefix": "{{gen_prefix}}",
89
+ "fewshot_delimiter": "\n\n",
90
+ "target_delimiter": " "
91
+ },
92
+ "num_fewshot": 0,
93
+ "metric_list": [
94
+ {
95
+ "metric": "4096",
96
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
97
+ "higher_is_better": true
98
+ },
99
+ {
100
+ "metric": "8192",
101
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
102
+ "higher_is_better": true
103
+ },
104
+ {
105
+ "metric": "16384",
106
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
107
+ "higher_is_better": true
108
+ },
109
+ {
110
+ "metric": "32768",
111
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
112
+ "higher_is_better": true
113
+ },
114
+ {
115
+ "metric": "65536",
116
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
117
+ "higher_is_better": true
118
+ },
119
+ {
120
+ "metric": "131072",
121
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
122
+ "higher_is_better": true
123
+ }
124
+ ],
125
+ "output_type": "generate_until",
126
+ "generation_kwargs": {
127
+ "do_sample": false,
128
+ "temperature": 0.0,
129
+ "max_gen_toks": 128,
130
+ "until": []
131
+ },
132
+ "repeats": 1,
133
+ "should_decontaminate": false,
134
+ "gen_prefix": "{{gen_prefix}}",
135
+ "metadata": {
136
+ "version": 1.0,
137
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
138
+ "trust_remote_code": true,
139
+ "dtype": "bfloat16",
140
+ "max_length": 16384,
141
+ "attn_implementation": "sdpa",
142
+ "max_seq_lengths": [
143
+ 8192
144
+ ]
145
+ }
146
+ },
147
+ "niah_multiquery": {
148
+ "task": "niah_multiquery",
149
+ "tag": [
150
+ "longcxt"
151
+ ],
152
+ "custom_dataset": "def niah_multiquery(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"essay\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"essay\",\n type_needle_k=\"words\",\n type_needle_v=\"numbers\",\n num_needle_q=4,\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
153
+ "dataset_path": "",
154
+ "dataset_name": "",
155
+ "test_split": "test",
156
+ "doc_to_text": "{{input}}",
157
+ "doc_to_target": "{{outputs}}",
158
+ "unsafe_code": false,
159
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
160
+ "description": "",
161
+ "target_delimiter": " ",
162
+ "fewshot_delimiter": "\n\n",
163
+ "fewshot_config": {
164
+ "sampler": "default",
165
+ "split": null,
166
+ "process_docs": null,
167
+ "fewshot_indices": null,
168
+ "samples": null,
169
+ "doc_to_text": "{{input}}",
170
+ "doc_to_choice": null,
171
+ "doc_to_target": "{{outputs}}",
172
+ "gen_prefix": "{{gen_prefix}}",
173
+ "fewshot_delimiter": "\n\n",
174
+ "target_delimiter": " "
175
+ },
176
+ "num_fewshot": 0,
177
+ "metric_list": [
178
+ {
179
+ "metric": "4096",
180
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
181
+ "higher_is_better": true
182
+ },
183
+ {
184
+ "metric": "8192",
185
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
186
+ "higher_is_better": true
187
+ },
188
+ {
189
+ "metric": "16384",
190
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
191
+ "higher_is_better": true
192
+ },
193
+ {
194
+ "metric": "32768",
195
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
196
+ "higher_is_better": true
197
+ },
198
+ {
199
+ "metric": "65536",
200
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
201
+ "higher_is_better": true
202
+ },
203
+ {
204
+ "metric": "131072",
205
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
206
+ "higher_is_better": true
207
+ }
208
+ ],
209
+ "output_type": "generate_until",
210
+ "generation_kwargs": {
211
+ "do_sample": false,
212
+ "temperature": 0.0,
213
+ "max_gen_toks": 128,
214
+ "until": []
215
+ },
216
+ "repeats": 1,
217
+ "should_decontaminate": false,
218
+ "gen_prefix": "{{gen_prefix}}",
219
+ "metadata": {
220
+ "version": 1.0,
221
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
222
+ "trust_remote_code": true,
223
+ "dtype": "bfloat16",
224
+ "max_length": 16384,
225
+ "attn_implementation": "sdpa",
226
+ "max_seq_lengths": [
227
+ 8192
228
+ ]
229
+ }
230
+ },
231
+ "niah_single_1": {
232
+ "task": "niah_single_1",
233
+ "tag": [
234
+ "longcxt"
235
+ ],
236
+ "custom_dataset": "def niah_single_1(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"repeat\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"repeat\",\n type_needle_k=\"words\",\n type_needle_v=\"numbers\",\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
237
+ "dataset_path": "",
238
+ "dataset_name": "",
239
+ "test_split": "test",
240
+ "doc_to_text": "{{input}}",
241
+ "doc_to_target": "{{outputs}}",
242
+ "unsafe_code": false,
243
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
244
+ "description": "",
245
+ "target_delimiter": " ",
246
+ "fewshot_delimiter": "\n\n",
247
+ "fewshot_config": {
248
+ "sampler": "default",
249
+ "split": null,
250
+ "process_docs": null,
251
+ "fewshot_indices": null,
252
+ "samples": null,
253
+ "doc_to_text": "{{input}}",
254
+ "doc_to_choice": null,
255
+ "doc_to_target": "{{outputs}}",
256
+ "gen_prefix": "{{gen_prefix}}",
257
+ "fewshot_delimiter": "\n\n",
258
+ "target_delimiter": " "
259
+ },
260
+ "num_fewshot": 0,
261
+ "metric_list": [
262
+ {
263
+ "metric": "4096",
264
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
265
+ "higher_is_better": true
266
+ },
267
+ {
268
+ "metric": "8192",
269
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
270
+ "higher_is_better": true
271
+ },
272
+ {
273
+ "metric": "16384",
274
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
275
+ "higher_is_better": true
276
+ },
277
+ {
278
+ "metric": "32768",
279
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
280
+ "higher_is_better": true
281
+ },
282
+ {
283
+ "metric": "65536",
284
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
285
+ "higher_is_better": true
286
+ },
287
+ {
288
+ "metric": "131072",
289
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
290
+ "higher_is_better": true
291
+ }
292
+ ],
293
+ "output_type": "generate_until",
294
+ "generation_kwargs": {
295
+ "do_sample": false,
296
+ "temperature": 0.0,
297
+ "max_gen_toks": 128,
298
+ "until": []
299
+ },
300
+ "repeats": 1,
301
+ "should_decontaminate": false,
302
+ "gen_prefix": "{{gen_prefix}}",
303
+ "metadata": {
304
+ "version": 1.0,
305
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
306
+ "trust_remote_code": true,
307
+ "dtype": "bfloat16",
308
+ "max_length": 16384,
309
+ "attn_implementation": "sdpa",
310
+ "max_seq_lengths": [
311
+ 8192
312
+ ]
313
+ }
314
+ },
315
+ "niah_single_3": {
316
+ "task": "niah_single_3",
317
+ "tag": [
318
+ "longcxt"
319
+ ],
320
+ "custom_dataset": "def niah_single_3(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"essay\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"essay\",\n type_needle_k=\"words\",\n type_needle_v=\"uuids\",\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
321
+ "dataset_path": "",
322
+ "dataset_name": "",
323
+ "test_split": "test",
324
+ "doc_to_text": "{{input}}",
325
+ "doc_to_target": "{{outputs}}",
326
+ "unsafe_code": false,
327
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
328
+ "description": "",
329
+ "target_delimiter": " ",
330
+ "fewshot_delimiter": "\n\n",
331
+ "fewshot_config": {
332
+ "sampler": "default",
333
+ "split": null,
334
+ "process_docs": null,
335
+ "fewshot_indices": null,
336
+ "samples": null,
337
+ "doc_to_text": "{{input}}",
338
+ "doc_to_choice": null,
339
+ "doc_to_target": "{{outputs}}",
340
+ "gen_prefix": "{{gen_prefix}}",
341
+ "fewshot_delimiter": "\n\n",
342
+ "target_delimiter": " "
343
+ },
344
+ "num_fewshot": 0,
345
+ "metric_list": [
346
+ {
347
+ "metric": "4096",
348
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
349
+ "higher_is_better": true
350
+ },
351
+ {
352
+ "metric": "8192",
353
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
354
+ "higher_is_better": true
355
+ },
356
+ {
357
+ "metric": "16384",
358
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
359
+ "higher_is_better": true
360
+ },
361
+ {
362
+ "metric": "32768",
363
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
364
+ "higher_is_better": true
365
+ },
366
+ {
367
+ "metric": "65536",
368
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
369
+ "higher_is_better": true
370
+ },
371
+ {
372
+ "metric": "131072",
373
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
374
+ "higher_is_better": true
375
+ }
376
+ ],
377
+ "output_type": "generate_until",
378
+ "generation_kwargs": {
379
+ "do_sample": false,
380
+ "temperature": 0.0,
381
+ "max_gen_toks": 128,
382
+ "until": []
383
+ },
384
+ "repeats": 1,
385
+ "should_decontaminate": false,
386
+ "gen_prefix": "{{gen_prefix}}",
387
+ "metadata": {
388
+ "version": 1.0,
389
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
390
+ "trust_remote_code": true,
391
+ "dtype": "bfloat16",
392
+ "max_length": 16384,
393
+ "attn_implementation": "sdpa",
394
+ "max_seq_lengths": [
395
+ 8192
396
+ ]
397
+ }
398
+ },
399
+ "ruler_fwe": {
400
+ "task": "ruler_fwe",
401
+ "tag": [
402
+ "longcxt"
403
+ ],
404
+ "custom_dataset": "def fwe_download(**kwargs):\n pretrained = kwargs.get(\"tokenizer\", kwargs.get(\"pretrained\", {}))\n df = (\n get_dataset(pretrained, max_seq_length=seq)\n for seq in kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n )\n\n return {\n \"test\": datasets.Dataset.from_list(\n list(itertools.chain.from_iterable(df)), split=datasets.Split.TEST\n )\n }\n",
405
+ "dataset_path": "",
406
+ "dataset_name": "",
407
+ "test_split": "test",
408
+ "doc_to_text": "{{input}}",
409
+ "doc_to_target": "{{outputs}}",
410
+ "unsafe_code": false,
411
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
412
+ "description": "",
413
+ "target_delimiter": " ",
414
+ "fewshot_delimiter": "\n\n",
415
+ "fewshot_config": {
416
+ "sampler": "default",
417
+ "split": null,
418
+ "process_docs": null,
419
+ "fewshot_indices": null,
420
+ "samples": null,
421
+ "doc_to_text": "{{input}}",
422
+ "doc_to_choice": null,
423
+ "doc_to_target": "{{outputs}}",
424
+ "gen_prefix": "{{gen_prefix}}",
425
+ "fewshot_delimiter": "\n\n",
426
+ "target_delimiter": " "
427
+ },
428
+ "num_fewshot": 0,
429
+ "metric_list": [
430
+ {
431
+ "metric": "4096",
432
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
433
+ "higher_is_better": true
434
+ },
435
+ {
436
+ "metric": "8192",
437
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
438
+ "higher_is_better": true
439
+ },
440
+ {
441
+ "metric": "16384",
442
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
443
+ "higher_is_better": true
444
+ },
445
+ {
446
+ "metric": "32768",
447
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
448
+ "higher_is_better": true
449
+ },
450
+ {
451
+ "metric": "65536",
452
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
453
+ "higher_is_better": true
454
+ },
455
+ {
456
+ "metric": "131072",
457
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
458
+ "higher_is_better": true
459
+ }
460
+ ],
461
+ "output_type": "generate_until",
462
+ "generation_kwargs": {
463
+ "do_sample": false,
464
+ "temperature": 0.0,
465
+ "max_gen_toks": 50,
466
+ "until": []
467
+ },
468
+ "repeats": 1,
469
+ "should_decontaminate": false,
470
+ "gen_prefix": "{{gen_prefix}}",
471
+ "metadata": {
472
+ "version": 1.0,
473
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
474
+ "trust_remote_code": true,
475
+ "dtype": "bfloat16",
476
+ "max_length": 16384,
477
+ "attn_implementation": "sdpa",
478
+ "max_seq_lengths": [
479
+ 8192
480
+ ]
481
+ }
482
+ },
483
+ "ruler_qa_hotpot": {
484
+ "task": "ruler_qa_hotpot",
485
+ "tag": [
486
+ "longcxt"
487
+ ],
488
+ "custom_dataset": "def get_hotpotqa(**kwargs):\n return get_qa_dataset(\"hotpotqa\", **kwargs)\n",
489
+ "dataset_path": "",
490
+ "dataset_name": "",
491
+ "test_split": "test",
492
+ "doc_to_text": "{{input}}",
493
+ "doc_to_target": "{{outputs}}",
494
+ "unsafe_code": false,
495
+ "process_results": "def process_results_part(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_part(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
496
+ "description": "",
497
+ "target_delimiter": " ",
498
+ "fewshot_delimiter": "\n\n",
499
+ "fewshot_config": {
500
+ "sampler": "default",
501
+ "split": null,
502
+ "process_docs": null,
503
+ "fewshot_indices": null,
504
+ "samples": null,
505
+ "doc_to_text": "{{input}}",
506
+ "doc_to_choice": null,
507
+ "doc_to_target": "{{outputs}}",
508
+ "gen_prefix": "{{gen_prefix}}",
509
+ "fewshot_delimiter": "\n\n",
510
+ "target_delimiter": " "
511
+ },
512
+ "num_fewshot": 0,
513
+ "metric_list": [
514
+ {
515
+ "metric": "4096",
516
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
517
+ "higher_is_better": true
518
+ },
519
+ {
520
+ "metric": "8192",
521
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
522
+ "higher_is_better": true
523
+ },
524
+ {
525
+ "metric": "16384",
526
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
527
+ "higher_is_better": true
528
+ },
529
+ {
530
+ "metric": "32768",
531
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
532
+ "higher_is_better": true
533
+ },
534
+ {
535
+ "metric": "65536",
536
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
537
+ "higher_is_better": true
538
+ },
539
+ {
540
+ "metric": "131072",
541
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
542
+ "higher_is_better": true
543
+ }
544
+ ],
545
+ "output_type": "generate_until",
546
+ "generation_kwargs": {
547
+ "do_sample": false,
548
+ "temperature": 0.0,
549
+ "max_gen_toks": 32,
550
+ "until": []
551
+ },
552
+ "repeats": 1,
553
+ "should_decontaminate": false,
554
+ "gen_prefix": "{{gen_prefix}}",
555
+ "metadata": {
556
+ "version": 1.0,
557
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
558
+ "trust_remote_code": true,
559
+ "dtype": "bfloat16",
560
+ "max_length": 16384,
561
+ "attn_implementation": "sdpa",
562
+ "max_seq_lengths": [
563
+ 8192
564
+ ]
565
+ }
566
+ },
567
+ "ruler_vt": {
568
+ "task": "ruler_vt",
569
+ "tag": [
570
+ "longcxt"
571
+ ],
572
+ "custom_dataset": "def get_vt_dataset(**kwargs) -> dict[str, datasets.Dataset]:\n pretrained = kwargs.get(\"tokenizer\", kwargs.get(\"pretrained\", \"\"))\n df = (\n get_dataset(tokenizer=get_tokenizer(pretrained), seq=seq)\n for seq in kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n )\n\n return {\n \"test\": datasets.Dataset.from_list(\n list(itertools.chain.from_iterable(df)), split=datasets.Split.TEST\n )\n }\n",
573
+ "dataset_path": "",
574
+ "dataset_name": "",
575
+ "test_split": "test",
576
+ "doc_to_text": "{{input}}",
577
+ "doc_to_target": "{{outputs}}",
578
+ "unsafe_code": false,
579
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
580
+ "description": "",
581
+ "target_delimiter": " ",
582
+ "fewshot_delimiter": "\n\n",
583
+ "fewshot_config": {
584
+ "sampler": "default",
585
+ "split": null,
586
+ "process_docs": null,
587
+ "fewshot_indices": null,
588
+ "samples": null,
589
+ "doc_to_text": "{{input}}",
590
+ "doc_to_choice": null,
591
+ "doc_to_target": "{{outputs}}",
592
+ "gen_prefix": "{{gen_prefix}}",
593
+ "fewshot_delimiter": "\n\n",
594
+ "target_delimiter": " "
595
+ },
596
+ "num_fewshot": 0,
597
+ "metric_list": [
598
+ {
599
+ "metric": "4096",
600
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
601
+ "higher_is_better": true
602
+ },
603
+ {
604
+ "metric": "8192",
605
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
606
+ "higher_is_better": true
607
+ },
608
+ {
609
+ "metric": "16384",
610
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
611
+ "higher_is_better": true
612
+ },
613
+ {
614
+ "metric": "32768",
615
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
616
+ "higher_is_better": true
617
+ },
618
+ {
619
+ "metric": "65536",
620
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
621
+ "higher_is_better": true
622
+ },
623
+ {
624
+ "metric": "131072",
625
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
626
+ "higher_is_better": true
627
+ }
628
+ ],
629
+ "output_type": "generate_until",
630
+ "generation_kwargs": {
631
+ "do_sample": false,
632
+ "temperature": 0.0,
633
+ "max_gen_toks": 30,
634
+ "until": []
635
+ },
636
+ "repeats": 1,
637
+ "should_decontaminate": false,
638
+ "gen_prefix": "{{gen_prefix}}",
639
+ "metadata": {
640
+ "version": 1.0,
641
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
642
+ "trust_remote_code": true,
643
+ "dtype": "bfloat16",
644
+ "max_length": 16384,
645
+ "attn_implementation": "sdpa",
646
+ "max_seq_lengths": [
647
+ 8192
648
+ ]
649
+ }
650
+ }
651
+ },
652
+ "versions": {
653
+ "niah_multikey_2": 1.0,
654
+ "niah_multiquery": 1.0,
655
+ "niah_single_1": 1.0,
656
+ "niah_single_3": 1.0,
657
+ "ruler_fwe": 1.0,
658
+ "ruler_qa_hotpot": 1.0,
659
+ "ruler_vt": 1.0
660
+ },
661
+ "n-shot": {
662
+ "niah_multikey_2": 0,
663
+ "niah_multiquery": 0,
664
+ "niah_single_1": 0,
665
+ "niah_single_3": 0,
666
+ "ruler_fwe": 0,
667
+ "ruler_qa_hotpot": 0,
668
+ "ruler_vt": 0
669
+ },
670
+ "higher_is_better": {
671
+ "niah_multikey_2": {
672
+ "4096": true,
673
+ "8192": true,
674
+ "16384": true,
675
+ "32768": true,
676
+ "65536": true,
677
+ "131072": true
678
+ },
679
+ "niah_multiquery": {
680
+ "4096": true,
681
+ "8192": true,
682
+ "16384": true,
683
+ "32768": true,
684
+ "65536": true,
685
+ "131072": true
686
+ },
687
+ "niah_single_1": {
688
+ "4096": true,
689
+ "8192": true,
690
+ "16384": true,
691
+ "32768": true,
692
+ "65536": true,
693
+ "131072": true
694
+ },
695
+ "niah_single_3": {
696
+ "4096": true,
697
+ "8192": true,
698
+ "16384": true,
699
+ "32768": true,
700
+ "65536": true,
701
+ "131072": true
702
+ },
703
+ "ruler_fwe": {
704
+ "4096": true,
705
+ "8192": true,
706
+ "16384": true,
707
+ "32768": true,
708
+ "65536": true,
709
+ "131072": true
710
+ },
711
+ "ruler_qa_hotpot": {
712
+ "4096": true,
713
+ "8192": true,
714
+ "16384": true,
715
+ "32768": true,
716
+ "65536": true,
717
+ "131072": true
718
+ },
719
+ "ruler_vt": {
720
+ "4096": true,
721
+ "8192": true,
722
+ "16384": true,
723
+ "32768": true,
724
+ "65536": true,
725
+ "131072": true
726
+ }
727
+ },
728
+ "n-samples": {
729
+ "ruler_qa_hotpot": {
730
+ "original": 500,
731
+ "effective": 20
732
+ },
733
+ "ruler_fwe": {
734
+ "original": 500,
735
+ "effective": 20
736
+ },
737
+ "ruler_vt": {
738
+ "original": 500,
739
+ "effective": 20
740
+ },
741
+ "niah_multiquery": {
742
+ "original": 500,
743
+ "effective": 20
744
+ },
745
+ "niah_multikey_2": {
746
+ "original": 500,
747
+ "effective": 20
748
+ },
749
+ "niah_single_3": {
750
+ "original": 500,
751
+ "effective": 20
752
+ },
753
+ "niah_single_1": {
754
+ "original": 500,
755
+ "effective": 20
756
+ }
757
+ },
758
+ "config": {
759
+ "model": "hf",
760
+ "model_args": {
761
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
762
+ "trust_remote_code": true,
763
+ "dtype": "bfloat16",
764
+ "max_length": 16384,
765
+ "attn_implementation": "sdpa"
766
+ },
767
+ "model_num_parameters": 2031969308,
768
+ "model_dtype": "torch.bfloat16",
769
+ "model_revision": "main",
770
+ "model_sha": "",
771
+ "batch_size": "1",
772
+ "batch_sizes": [],
773
+ "device": "cuda:0",
774
+ "use_cache": null,
775
+ "limit": 20.0,
776
+ "bootstrap_iters": 100000,
777
+ "gen_kwargs": {},
778
+ "random_seed": 0,
779
+ "numpy_seed": 1234,
780
+ "torch_seed": 1234,
781
+ "fewshot_seed": 1234
782
+ },
783
+ "git_hash": null,
784
+ "date": 1784389152.295637,
785
+ "pretty_env_info": "PyTorch version: 2.9.1+cu128\nIs debug build: False\nCUDA used to build PyTorch: 12.8\nROCM used to build PyTorch: N/A\n\nOS: Ubuntu 22.04.5 LTS (x86_64)\nGCC version: Could not collect\nClang version: Could not collect\nCMake version: version 4.1.2\nLibc version: glibc-2.35\n\nPython version: 3.11.14 | packaged by conda-forge | (main, Oct 22 2025, 22:46:25) [GCC 14.3.0] (64-bit runtime)\nPython platform: Linux-6.12.90-120.164.amzn2023.x86_64-x86_64-with-glibc2.35\nIs CUDA available: True\nCUDA runtime version: Could not collect\nCUDA_MODULE_LOADING set to: \nGPU models and configuration: GPU 0: NVIDIA A100-SXM4-80GB\nNvidia driver version: 580.159.03\ncuDNN version: Could not collect\nIs XPU available: False\nHIP runtime version: N/A\nMIOpen runtime version: N/A\nIs XNNPACK available: True\n\nCPU:\nArchitecture: x86_64\nCPU op-mode(s): 32-bit, 64-bit\nAddress sizes: 46 bits physical, 48 bits virtual\nByte Order: Little Endian\nCPU(s): 96\nOn-line CPU(s) list: 0-95\nVendor ID: GenuineIntel\nModel name: Intel(R) Xeon(R) Platinum 8275CL CPU @ 3.00GHz\nCPU family: 6\nModel: 85\nThread(s) per core: 2\nCore(s) per socket: 24\nSocket(s): 2\nStepping: 7\nBogoMIPS: 5999.99\nFlags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon rep_good nopl xtopology nonstop_tsc cpuid aperfmperf tsc_known_freq pni pclmulqdq monitor ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand hypervisor lahf_lm abm 3dnowprefetch pti fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid mpx avx512f avx512dq rdseed adx smap clflushopt clwb avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves ida arat pku ospke\nHypervisor vendor: KVM\nVirtualization type: full\nL1d cache: 1.5 MiB (48 instances)\nL1i cache: 1.5 MiB (48 instances)\nL2 cache: 48 MiB (48 instances)\nL3 cache: 71.5 MiB (2 instances)\nNUMA node(s): 2\nNUMA node0 CPU(s): 0-23,48-71\nNUMA node1 CPU(s): 24-47,72-95\nVulnerability Gather data sampling: Unknown: Dependent on hypervisor status\nVulnerability Indirect target selection: Mitigation; Aligned branch/return thunks\nVulnerability Itlb multihit: KVM: Mitigation: VMX unsupported\nVulnerability L1tf: Mitigation; PTE Inversion\nVulnerability Mds: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown\nVulnerability Meltdown: Mitigation; PTI\nVulnerability Mmio stale data: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown\nVulnerability Reg file data sampling: Not affected\nVulnerability Retbleed: Vulnerable\nVulnerability Spec rstack overflow: Not affected\nVulnerability Spec store bypass: Vulnerable\nVulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization\nVulnerability Spectre v2: Mitigation; Retpolines; STIBP disabled; RSB filling; PBRSB-eIBRS Not affected; BHI Retpoline\nVulnerability Srbds: Not affected\nVulnerability Tsa: Not affected\nVulnerability Tsx async abort: Not affected\nVulnerability Vmscape: Not affected\n\nVersions of relevant libraries:\n[pip3] numpy==2.3.4\n[pip3] nvidia-cublas-cu12==12.8.4.1\n[pip3] nvidia-cuda-cupti-cu12==12.8.90\n[pip3] nvidia-cuda-nvrtc-cu12==12.8.93\n[pip3] nvidia-cuda-runtime-cu12==12.8.90\n[pip3] nvidia-cudnn-cu12==9.10.2.21\n[pip3] nvidia-cufft-cu12==11.3.3.83\n[pip3] nvidia-curand-cu12==10.3.9.90\n[pip3] nvidia-cusolver-cu12==11.7.3.90\n[pip3] nvidia-cusparse-cu12==12.5.8.93\n[pip3] nvidia-cusparselt-cu12==0.7.1\n[pip3] nvidia-nccl-cu12==2.27.5\n[pip3] nvidia-nvjitlink-cu12==12.8.93\n[pip3] nvidia-nvtx-cu12==12.8.90\n[pip3] optree==0.17.0\n[pip3] torch==2.9.1+cu128\n[pip3] torchaudio==2.9.1+cu128\n[pip3] torchelastic==0.2.2\n[pip3] torchvision==0.24.1+cu128\n[pip3] triton==3.5.1\n[conda] numpy 2.3.4 py311h2e04523_0 conda-forge\n[conda] nvidia-cublas-cu12 12.8.4.1 pypi_0 pypi\n[conda] nvidia-cuda-cupti-cu12 12.8.90 pypi_0 pypi\n[conda] nvidia-cuda-nvrtc-cu12 12.8.93 pypi_0 pypi\n[conda] nvidia-cuda-runtime-cu12 12.8.90 pypi_0 pypi\n[conda] nvidia-cudnn-cu12 9.10.2.21 pypi_0 pypi\n[conda] nvidia-cufft-cu12 11.3.3.83 pypi_0 pypi\n[conda] nvidia-curand-cu12 10.3.9.90 pypi_0 pypi\n[conda] nvidia-cusolver-cu12 11.7.3.90 pypi_0 pypi\n[conda] nvidia-cusparse-cu12 12.5.8.93 pypi_0 pypi\n[conda] nvidia-cusparselt-cu12 0.7.1 pypi_0 pypi\n[conda] nvidia-nccl-cu12 2.27.5 pypi_0 pypi\n[conda] nvidia-nvjitlink-cu12 12.8.93 pypi_0 pypi\n[conda] nvidia-nvtx-cu12 12.8.90 pypi_0 pypi\n[conda] optree 0.17.0 pypi_0 pypi\n[conda] torch 2.9.1+cu128 pypi_0 pypi\n[conda] torchaudio 2.9.1+cu128 pypi_0 pypi\n[conda] torchelastic 0.2.2 pypi_0 pypi\n[conda] torchvision 0.24.1+cu128 pypi_0 pypi\n[conda] triton 3.5.1 pypi_0 pypi",
786
+ "transformers_version": "4.54.0",
787
+ "lm_eval_version": "0.4.11",
788
+ "upper_git_hash": null,
789
+ "tokenizer_pad_token": [
790
+ "<|endoftext|>",
791
+ "151643"
792
+ ],
793
+ "tokenizer_eos_token": [
794
+ "<|im_end|>",
795
+ "151645"
796
+ ],
797
+ "tokenizer_bos_token": [
798
+ null,
799
+ "None"
800
+ ],
801
+ "eot_token_id": 151645,
802
+ "max_length": 16384,
803
+ "task_hashes": {
804
+ "ruler_qa_hotpot": "fd48cb605efc8c5adac7e3771533db4d4c354b71bf91ce00841c814604ba6404",
805
+ "ruler_fwe": "c4633bcd0acf3ba4b96f905b5ba269ba94cf45a4c8ca0c13aae1468227cc3a10",
806
+ "ruler_vt": "af011da859c8b86f63f57f6256cf59077d3d3d8536084ab67ccf39e18e9a615a",
807
+ "niah_multiquery": "f9a36d2854dfd2bb52ffc7a6bf3f53490fb5303ca9da0046256bc28b55c2adb2",
808
+ "niah_multikey_2": "86e49b8f875843454b2923143b9124ba82e8a2e23b145f792726cd82f4e49c24",
809
+ "niah_single_3": "905e40146a287e5f0c8e2a81f64f082108c45a23862949e6a99bfcbbf561cad2",
810
+ "niah_single_1": "426e8138e992480aa416d01e6b9cef5865a68c00f532400e448a0c20ee026cff"
811
+ },
812
+ "model_source": "hf",
813
+ "model_name": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
814
+ "model_name_sanitized": "__workspace__outputs__l2a_style__stage2__checkpoint-25",
815
+ "system_instruction": null,
816
+ "system_instruction_sha": null,
817
+ "fewshot_as_multiturn": null,
818
+ "chat_template": null,
819
+ "chat_template_sha": null,
820
+ "total_evaluation_time_seconds": "748.3684644349851"
821
+ }
outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multikey_2_2026-07-18T15-51-37.349704.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:a1198521d64dbb83c3bcac7e4b000a43afe8c5dc3c6a6261875cba22529f2a61
3
+ size 987952
outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multiquery_2026-07-18T15-51-37.349704.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:fda48f7c886c33b97a6e90f648b89cd35e4b95b24d3bd08c1cff5bfd40efb2bc
3
+ size 1442196
outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_single_1_2026-07-18T15-51-37.349704.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2b805cb744f2ca51f70ef08f871e84fd7a635d76bd0632761601cf38420eb43b
3
+ size 1231408
outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_single_3_2026-07-18T15-51-37.349704.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:abf7b2d058ae6464aa9140d934b01d08b30c31804b918a3fb4d99484b0c4a910
3
+ size 1435434
outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_fwe_2026-07-18T15-51-37.349704.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:b88939d119cb9876450e15d30af248d129e4eaac0f6d6f5daf63508495d5662c
3
+ size 826859
outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_qa_hotpot_2026-07-18T15-51-37.349704.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1c489a379e7ad7898390535ead6d1a6d23be6f3e40691c74e3546a15ca8b916d
3
+ size 1172240
outputs/eval/token_t0525/ruler8k_splits/a/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_vt_2026-07-18T15-51-37.349704.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:1c13551162068852b669c560c3e9bbc82f84845accd93ed37be9d4123cf1cc1e
3
+ size 1228980
outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/results_2026-07-18T16-04-28.086981.json ADDED
@@ -0,0 +1,714 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "results": {
3
+ "niah_multikey_1": {
4
+ "alias": "niah_multikey_1",
5
+ "4096,none": -1,
6
+ "4096_stderr,none": "N/A",
7
+ "8192,none": 1.0,
8
+ "8192_stderr,none": "N/A"
9
+ },
10
+ "niah_multikey_3": {
11
+ "alias": "niah_multikey_3",
12
+ "4096,none": -1,
13
+ "4096_stderr,none": "N/A",
14
+ "8192,none": 0.75,
15
+ "8192_stderr,none": "N/A"
16
+ },
17
+ "niah_multivalue": {
18
+ "alias": "niah_multivalue",
19
+ "4096,none": -1,
20
+ "4096_stderr,none": "N/A",
21
+ "8192,none": 1.0,
22
+ "8192_stderr,none": "N/A"
23
+ },
24
+ "niah_single_2": {
25
+ "alias": "niah_single_2",
26
+ "4096,none": -1,
27
+ "4096_stderr,none": "N/A",
28
+ "8192,none": 1.0,
29
+ "8192_stderr,none": "N/A"
30
+ },
31
+ "ruler_cwe": {
32
+ "alias": "ruler_cwe",
33
+ "4096,none": -1,
34
+ "4096_stderr,none": "N/A",
35
+ "8192,none": 0.40499999999999997,
36
+ "8192_stderr,none": "N/A"
37
+ },
38
+ "ruler_qa_squad": {
39
+ "alias": "ruler_qa_squad",
40
+ "4096,none": -1,
41
+ "4096_stderr,none": "N/A",
42
+ "8192,none": 0.4208333333333334,
43
+ "8192_stderr,none": "N/A"
44
+ }
45
+ },
46
+ "group_subtasks": {
47
+ "niah_single_2": [],
48
+ "niah_multikey_1": [],
49
+ "niah_multikey_3": [],
50
+ "niah_multivalue": [],
51
+ "ruler_cwe": [],
52
+ "ruler_qa_squad": []
53
+ },
54
+ "configs": {
55
+ "niah_multikey_1": {
56
+ "task": "niah_multikey_1",
57
+ "tag": [
58
+ "longcxt"
59
+ ],
60
+ "custom_dataset": "def niah_multikey_1(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"essay\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"essay\",\n type_needle_k=\"words\",\n type_needle_v=\"numbers\",\n num_needle_k=4,\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
61
+ "dataset_path": "",
62
+ "dataset_name": "",
63
+ "test_split": "test",
64
+ "doc_to_text": "{{input}}",
65
+ "doc_to_target": "{{outputs}}",
66
+ "unsafe_code": false,
67
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
68
+ "description": "",
69
+ "target_delimiter": " ",
70
+ "fewshot_delimiter": "\n\n",
71
+ "fewshot_config": {
72
+ "sampler": "default",
73
+ "split": null,
74
+ "process_docs": null,
75
+ "fewshot_indices": null,
76
+ "samples": null,
77
+ "doc_to_text": "{{input}}",
78
+ "doc_to_choice": null,
79
+ "doc_to_target": "{{outputs}}",
80
+ "gen_prefix": "{{gen_prefix}}",
81
+ "fewshot_delimiter": "\n\n",
82
+ "target_delimiter": " "
83
+ },
84
+ "num_fewshot": 0,
85
+ "metric_list": [
86
+ {
87
+ "metric": "4096",
88
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
89
+ "higher_is_better": true
90
+ },
91
+ {
92
+ "metric": "8192",
93
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
94
+ "higher_is_better": true
95
+ },
96
+ {
97
+ "metric": "16384",
98
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
99
+ "higher_is_better": true
100
+ },
101
+ {
102
+ "metric": "32768",
103
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
104
+ "higher_is_better": true
105
+ },
106
+ {
107
+ "metric": "65536",
108
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
109
+ "higher_is_better": true
110
+ },
111
+ {
112
+ "metric": "131072",
113
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
114
+ "higher_is_better": true
115
+ }
116
+ ],
117
+ "output_type": "generate_until",
118
+ "generation_kwargs": {
119
+ "do_sample": false,
120
+ "temperature": 0.0,
121
+ "max_gen_toks": 128,
122
+ "until": []
123
+ },
124
+ "repeats": 1,
125
+ "should_decontaminate": false,
126
+ "gen_prefix": "{{gen_prefix}}",
127
+ "metadata": {
128
+ "version": 1.0,
129
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
130
+ "trust_remote_code": true,
131
+ "dtype": "bfloat16",
132
+ "max_length": 16384,
133
+ "attn_implementation": "sdpa",
134
+ "max_seq_lengths": [
135
+ 8192
136
+ ]
137
+ }
138
+ },
139
+ "niah_multikey_3": {
140
+ "task": "niah_multikey_3",
141
+ "tag": [
142
+ "longcxt"
143
+ ],
144
+ "custom_dataset": "def niah_multikey_3(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"needle\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"needle\",\n type_needle_k=\"uuids\",\n type_needle_v=\"uuids\",\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
145
+ "dataset_path": "",
146
+ "dataset_name": "",
147
+ "test_split": "test",
148
+ "doc_to_text": "{{input}}",
149
+ "doc_to_target": "{{outputs}}",
150
+ "unsafe_code": false,
151
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
152
+ "description": "",
153
+ "target_delimiter": " ",
154
+ "fewshot_delimiter": "\n\n",
155
+ "fewshot_config": {
156
+ "sampler": "default",
157
+ "split": null,
158
+ "process_docs": null,
159
+ "fewshot_indices": null,
160
+ "samples": null,
161
+ "doc_to_text": "{{input}}",
162
+ "doc_to_choice": null,
163
+ "doc_to_target": "{{outputs}}",
164
+ "gen_prefix": "{{gen_prefix}}",
165
+ "fewshot_delimiter": "\n\n",
166
+ "target_delimiter": " "
167
+ },
168
+ "num_fewshot": 0,
169
+ "metric_list": [
170
+ {
171
+ "metric": "4096",
172
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
173
+ "higher_is_better": true
174
+ },
175
+ {
176
+ "metric": "8192",
177
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
178
+ "higher_is_better": true
179
+ },
180
+ {
181
+ "metric": "16384",
182
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
183
+ "higher_is_better": true
184
+ },
185
+ {
186
+ "metric": "32768",
187
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
188
+ "higher_is_better": true
189
+ },
190
+ {
191
+ "metric": "65536",
192
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
193
+ "higher_is_better": true
194
+ },
195
+ {
196
+ "metric": "131072",
197
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
198
+ "higher_is_better": true
199
+ }
200
+ ],
201
+ "output_type": "generate_until",
202
+ "generation_kwargs": {
203
+ "do_sample": false,
204
+ "temperature": 0.0,
205
+ "max_gen_toks": 128,
206
+ "until": []
207
+ },
208
+ "repeats": 1,
209
+ "should_decontaminate": false,
210
+ "gen_prefix": "{{gen_prefix}}",
211
+ "metadata": {
212
+ "version": 1.0,
213
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
214
+ "trust_remote_code": true,
215
+ "dtype": "bfloat16",
216
+ "max_length": 16384,
217
+ "attn_implementation": "sdpa",
218
+ "max_seq_lengths": [
219
+ 8192
220
+ ]
221
+ }
222
+ },
223
+ "niah_multivalue": {
224
+ "task": "niah_multivalue",
225
+ "tag": [
226
+ "longcxt"
227
+ ],
228
+ "custom_dataset": "def niah_multivalue(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"essay\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"essay\",\n type_needle_k=\"words\",\n type_needle_v=\"numbers\",\n num_needle_v=4,\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
229
+ "dataset_path": "",
230
+ "dataset_name": "",
231
+ "test_split": "test",
232
+ "doc_to_text": "{{input}}",
233
+ "doc_to_target": "{{outputs}}",
234
+ "unsafe_code": false,
235
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
236
+ "description": "",
237
+ "target_delimiter": " ",
238
+ "fewshot_delimiter": "\n\n",
239
+ "fewshot_config": {
240
+ "sampler": "default",
241
+ "split": null,
242
+ "process_docs": null,
243
+ "fewshot_indices": null,
244
+ "samples": null,
245
+ "doc_to_text": "{{input}}",
246
+ "doc_to_choice": null,
247
+ "doc_to_target": "{{outputs}}",
248
+ "gen_prefix": "{{gen_prefix}}",
249
+ "fewshot_delimiter": "\n\n",
250
+ "target_delimiter": " "
251
+ },
252
+ "num_fewshot": 0,
253
+ "metric_list": [
254
+ {
255
+ "metric": "4096",
256
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
257
+ "higher_is_better": true
258
+ },
259
+ {
260
+ "metric": "8192",
261
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
262
+ "higher_is_better": true
263
+ },
264
+ {
265
+ "metric": "16384",
266
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
267
+ "higher_is_better": true
268
+ },
269
+ {
270
+ "metric": "32768",
271
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
272
+ "higher_is_better": true
273
+ },
274
+ {
275
+ "metric": "65536",
276
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
277
+ "higher_is_better": true
278
+ },
279
+ {
280
+ "metric": "131072",
281
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
282
+ "higher_is_better": true
283
+ }
284
+ ],
285
+ "output_type": "generate_until",
286
+ "generation_kwargs": {
287
+ "do_sample": false,
288
+ "temperature": 0.0,
289
+ "max_gen_toks": 128,
290
+ "until": []
291
+ },
292
+ "repeats": 1,
293
+ "should_decontaminate": false,
294
+ "gen_prefix": "{{gen_prefix}}",
295
+ "metadata": {
296
+ "version": 1.0,
297
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
298
+ "trust_remote_code": true,
299
+ "dtype": "bfloat16",
300
+ "max_length": 16384,
301
+ "attn_implementation": "sdpa",
302
+ "max_seq_lengths": [
303
+ 8192
304
+ ]
305
+ }
306
+ },
307
+ "niah_single_2": {
308
+ "task": "niah_single_2",
309
+ "tag": [
310
+ "longcxt"
311
+ ],
312
+ "custom_dataset": "def niah_single_2(**kwargs):\n seq_lengths = kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n return download_dataset(\n generate_samples(\n get_haystack(type_haystack=\"essay\"),\n max_seq_length=seq,\n template=TEMPLATE,\n type_haystack=\"essay\",\n type_needle_k=\"words\",\n type_needle_v=\"numbers\",\n num_samples=500,\n TOKENIZER=get_tokenizer(**kwargs),\n )\n for seq in seq_lengths\n )\n",
313
+ "dataset_path": "",
314
+ "dataset_name": "",
315
+ "test_split": "test",
316
+ "doc_to_text": "{{input}}",
317
+ "doc_to_target": "{{outputs}}",
318
+ "unsafe_code": false,
319
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
320
+ "description": "",
321
+ "target_delimiter": " ",
322
+ "fewshot_delimiter": "\n\n",
323
+ "fewshot_config": {
324
+ "sampler": "default",
325
+ "split": null,
326
+ "process_docs": null,
327
+ "fewshot_indices": null,
328
+ "samples": null,
329
+ "doc_to_text": "{{input}}",
330
+ "doc_to_choice": null,
331
+ "doc_to_target": "{{outputs}}",
332
+ "gen_prefix": "{{gen_prefix}}",
333
+ "fewshot_delimiter": "\n\n",
334
+ "target_delimiter": " "
335
+ },
336
+ "num_fewshot": 0,
337
+ "metric_list": [
338
+ {
339
+ "metric": "4096",
340
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
341
+ "higher_is_better": true
342
+ },
343
+ {
344
+ "metric": "8192",
345
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
346
+ "higher_is_better": true
347
+ },
348
+ {
349
+ "metric": "16384",
350
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
351
+ "higher_is_better": true
352
+ },
353
+ {
354
+ "metric": "32768",
355
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
356
+ "higher_is_better": true
357
+ },
358
+ {
359
+ "metric": "65536",
360
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
361
+ "higher_is_better": true
362
+ },
363
+ {
364
+ "metric": "131072",
365
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
366
+ "higher_is_better": true
367
+ }
368
+ ],
369
+ "output_type": "generate_until",
370
+ "generation_kwargs": {
371
+ "do_sample": false,
372
+ "temperature": 0.0,
373
+ "max_gen_toks": 128,
374
+ "until": []
375
+ },
376
+ "repeats": 1,
377
+ "should_decontaminate": false,
378
+ "gen_prefix": "{{gen_prefix}}",
379
+ "metadata": {
380
+ "version": 1.0,
381
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
382
+ "trust_remote_code": true,
383
+ "dtype": "bfloat16",
384
+ "max_length": 16384,
385
+ "attn_implementation": "sdpa",
386
+ "max_seq_lengths": [
387
+ 8192
388
+ ]
389
+ }
390
+ },
391
+ "ruler_cwe": {
392
+ "task": "ruler_cwe",
393
+ "tag": [
394
+ "longcxt"
395
+ ],
396
+ "custom_dataset": "def get_cw_dataset(**kwargs):\n pretrained = kwargs.get(\"tokenizer\", kwargs.get(\"pretrained\", {}))\n df = (\n get_dataset(pretrained, seq=seq)\n for seq in kwargs.pop(\"max_seq_lengths\", DEFAULT_SEQ_LENGTHS)\n )\n\n return {\n \"test\": datasets.Dataset.from_list(\n list(itertools.chain.from_iterable(df)), split=datasets.Split.TEST\n )\n }\n",
397
+ "dataset_path": "",
398
+ "dataset_name": "",
399
+ "test_split": "test",
400
+ "doc_to_text": "{{input}}",
401
+ "doc_to_target": "{{outputs}}",
402
+ "unsafe_code": false,
403
+ "process_results": "def process_results(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_all(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
404
+ "description": "",
405
+ "target_delimiter": "\n\n",
406
+ "fewshot_delimiter": "\n\n",
407
+ "fewshot_config": {
408
+ "sampler": "default",
409
+ "split": null,
410
+ "process_docs": null,
411
+ "fewshot_indices": null,
412
+ "samples": null,
413
+ "doc_to_text": "{{input}}",
414
+ "doc_to_choice": null,
415
+ "doc_to_target": "{{outputs}}",
416
+ "gen_prefix": "{{gen_prefix}}",
417
+ "fewshot_delimiter": "\n\n",
418
+ "target_delimiter": "\n\n"
419
+ },
420
+ "num_fewshot": 0,
421
+ "metric_list": [
422
+ {
423
+ "metric": "4096",
424
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
425
+ "higher_is_better": true
426
+ },
427
+ {
428
+ "metric": "8192",
429
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
430
+ "higher_is_better": true
431
+ },
432
+ {
433
+ "metric": "16384",
434
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
435
+ "higher_is_better": true
436
+ },
437
+ {
438
+ "metric": "32768",
439
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
440
+ "higher_is_better": true
441
+ },
442
+ {
443
+ "metric": "65536",
444
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
445
+ "higher_is_better": true
446
+ },
447
+ {
448
+ "metric": "131072",
449
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
450
+ "higher_is_better": true
451
+ }
452
+ ],
453
+ "output_type": "generate_until",
454
+ "generation_kwargs": {
455
+ "do_sample": false,
456
+ "temperature": 0.0,
457
+ "max_gen_toks": 120,
458
+ "until": []
459
+ },
460
+ "repeats": 1,
461
+ "should_decontaminate": false,
462
+ "gen_prefix": "{{gen_prefix}}",
463
+ "metadata": {
464
+ "version": 1.0,
465
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
466
+ "trust_remote_code": true,
467
+ "dtype": "bfloat16",
468
+ "max_length": 16384,
469
+ "attn_implementation": "sdpa",
470
+ "max_seq_lengths": [
471
+ 8192
472
+ ]
473
+ }
474
+ },
475
+ "ruler_qa_squad": {
476
+ "task": "ruler_qa_squad",
477
+ "tag": [
478
+ "longcxt"
479
+ ],
480
+ "custom_dataset": "def get_squad(**kwargs):\n return get_qa_dataset(\"squad\", **kwargs)\n",
481
+ "dataset_path": "",
482
+ "dataset_name": "",
483
+ "test_split": "test",
484
+ "doc_to_text": "{{input}}",
485
+ "doc_to_target": "{{outputs}}",
486
+ "unsafe_code": false,
487
+ "process_results": "def process_results_part(doc: dict, results: list[str]) -> dict[str, float]:\n # hacky: set all other lengths to -1\n metrics = {str(length): -1.0 for length in DEFAULT_SEQ_LENGTHS}\n input_len = doc[\"max_length\"]\n pred = postprocess_pred(results)\n score = string_match_part(pred, [doc[\"outputs\"]])\n metrics[str(input_len)] = score\n return metrics\n",
488
+ "description": "",
489
+ "target_delimiter": " ",
490
+ "fewshot_delimiter": "\n\n",
491
+ "fewshot_config": {
492
+ "sampler": "default",
493
+ "split": null,
494
+ "process_docs": null,
495
+ "fewshot_indices": null,
496
+ "samples": null,
497
+ "doc_to_text": "{{input}}",
498
+ "doc_to_choice": null,
499
+ "doc_to_target": "{{outputs}}",
500
+ "gen_prefix": "{{gen_prefix}}",
501
+ "fewshot_delimiter": "\n\n",
502
+ "target_delimiter": " "
503
+ },
504
+ "num_fewshot": 0,
505
+ "metric_list": [
506
+ {
507
+ "metric": "4096",
508
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
509
+ "higher_is_better": true
510
+ },
511
+ {
512
+ "metric": "8192",
513
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
514
+ "higher_is_better": true
515
+ },
516
+ {
517
+ "metric": "16384",
518
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
519
+ "higher_is_better": true
520
+ },
521
+ {
522
+ "metric": "32768",
523
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
524
+ "higher_is_better": true
525
+ },
526
+ {
527
+ "metric": "65536",
528
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
529
+ "higher_is_better": true
530
+ },
531
+ {
532
+ "metric": "131072",
533
+ "aggregation": "def aggregate_metrics(metrics: list[float]) -> float:\n res = [x for x in metrics if x != -1]\n if not res:\n # we don't have any samples with this length\n return -1\n return sum(res) / len(res)\n",
534
+ "higher_is_better": true
535
+ }
536
+ ],
537
+ "output_type": "generate_until",
538
+ "generation_kwargs": {
539
+ "do_sample": false,
540
+ "temperature": 0.0,
541
+ "max_gen_toks": 32,
542
+ "until": []
543
+ },
544
+ "repeats": 1,
545
+ "should_decontaminate": false,
546
+ "gen_prefix": "{{gen_prefix}}",
547
+ "metadata": {
548
+ "version": 1.0,
549
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
550
+ "trust_remote_code": true,
551
+ "dtype": "bfloat16",
552
+ "max_length": 16384,
553
+ "attn_implementation": "sdpa",
554
+ "max_seq_lengths": [
555
+ 8192
556
+ ]
557
+ }
558
+ }
559
+ },
560
+ "versions": {
561
+ "niah_multikey_1": 1.0,
562
+ "niah_multikey_3": 1.0,
563
+ "niah_multivalue": 1.0,
564
+ "niah_single_2": 1.0,
565
+ "ruler_cwe": 1.0,
566
+ "ruler_qa_squad": 1.0
567
+ },
568
+ "n-shot": {
569
+ "niah_multikey_1": 0,
570
+ "niah_multikey_3": 0,
571
+ "niah_multivalue": 0,
572
+ "niah_single_2": 0,
573
+ "ruler_cwe": 0,
574
+ "ruler_qa_squad": 0
575
+ },
576
+ "higher_is_better": {
577
+ "niah_multikey_1": {
578
+ "4096": true,
579
+ "8192": true,
580
+ "16384": true,
581
+ "32768": true,
582
+ "65536": true,
583
+ "131072": true
584
+ },
585
+ "niah_multikey_3": {
586
+ "4096": true,
587
+ "8192": true,
588
+ "16384": true,
589
+ "32768": true,
590
+ "65536": true,
591
+ "131072": true
592
+ },
593
+ "niah_multivalue": {
594
+ "4096": true,
595
+ "8192": true,
596
+ "16384": true,
597
+ "32768": true,
598
+ "65536": true,
599
+ "131072": true
600
+ },
601
+ "niah_single_2": {
602
+ "4096": true,
603
+ "8192": true,
604
+ "16384": true,
605
+ "32768": true,
606
+ "65536": true,
607
+ "131072": true
608
+ },
609
+ "ruler_cwe": {
610
+ "4096": true,
611
+ "8192": true,
612
+ "16384": true,
613
+ "32768": true,
614
+ "65536": true,
615
+ "131072": true
616
+ },
617
+ "ruler_qa_squad": {
618
+ "4096": true,
619
+ "8192": true,
620
+ "16384": true,
621
+ "32768": true,
622
+ "65536": true,
623
+ "131072": true
624
+ }
625
+ },
626
+ "n-samples": {
627
+ "ruler_qa_squad": {
628
+ "original": 500,
629
+ "effective": 20
630
+ },
631
+ "ruler_cwe": {
632
+ "original": 500,
633
+ "effective": 20
634
+ },
635
+ "niah_multivalue": {
636
+ "original": 500,
637
+ "effective": 20
638
+ },
639
+ "niah_multikey_3": {
640
+ "original": 500,
641
+ "effective": 20
642
+ },
643
+ "niah_multikey_1": {
644
+ "original": 500,
645
+ "effective": 20
646
+ },
647
+ "niah_single_2": {
648
+ "original": 500,
649
+ "effective": 20
650
+ }
651
+ },
652
+ "config": {
653
+ "model": "hf",
654
+ "model_args": {
655
+ "pretrained": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
656
+ "trust_remote_code": true,
657
+ "dtype": "bfloat16",
658
+ "max_length": 16384,
659
+ "attn_implementation": "sdpa"
660
+ },
661
+ "model_num_parameters": 2031969308,
662
+ "model_dtype": "torch.bfloat16",
663
+ "model_revision": "main",
664
+ "model_sha": "",
665
+ "batch_size": "1",
666
+ "batch_sizes": [],
667
+ "device": "cuda:0",
668
+ "use_cache": null,
669
+ "limit": 20.0,
670
+ "bootstrap_iters": 100000,
671
+ "gen_kwargs": {},
672
+ "random_seed": 0,
673
+ "numpy_seed": 1234,
674
+ "torch_seed": 1234,
675
+ "fewshot_seed": 1234
676
+ },
677
+ "git_hash": null,
678
+ "date": 1784389903.7655613,
679
+ "pretty_env_info": "PyTorch version: 2.9.1+cu128\nIs debug build: False\nCUDA used to build PyTorch: 12.8\nROCM used to build PyTorch: N/A\n\nOS: Ubuntu 22.04.5 LTS (x86_64)\nGCC version: Could not collect\nClang version: Could not collect\nCMake version: version 4.1.2\nLibc version: glibc-2.35\n\nPython version: 3.11.14 | packaged by conda-forge | (main, Oct 22 2025, 22:46:25) [GCC 14.3.0] (64-bit runtime)\nPython platform: Linux-6.12.90-120.164.amzn2023.x86_64-x86_64-with-glibc2.35\nIs CUDA available: True\nCUDA runtime version: Could not collect\nCUDA_MODULE_LOADING set to: \nGPU models and configuration: GPU 0: NVIDIA A100-SXM4-80GB\nNvidia driver version: 580.159.03\ncuDNN version: Could not collect\nIs XPU available: False\nHIP runtime version: N/A\nMIOpen runtime version: N/A\nIs XNNPACK available: True\n\nCPU:\nArchitecture: x86_64\nCPU op-mode(s): 32-bit, 64-bit\nAddress sizes: 46 bits physical, 48 bits virtual\nByte Order: Little Endian\nCPU(s): 96\nOn-line CPU(s) list: 0-95\nVendor ID: GenuineIntel\nModel name: Intel(R) Xeon(R) Platinum 8275CL CPU @ 3.00GHz\nCPU family: 6\nModel: 85\nThread(s) per core: 2\nCore(s) per socket: 24\nSocket(s): 2\nStepping: 7\nBogoMIPS: 5999.99\nFlags: fpu vme de pse tsc msr pae mce cx8 apic sep mtrr pge mca cmov pat pse36 clflush mmx fxsr sse sse2 ss ht syscall nx pdpe1gb rdtscp lm constant_tsc arch_perfmon rep_good nopl xtopology nonstop_tsc cpuid aperfmperf tsc_known_freq pni pclmulqdq monitor ssse3 fma cx16 pcid sse4_1 sse4_2 x2apic movbe popcnt tsc_deadline_timer aes xsave avx f16c rdrand hypervisor lahf_lm abm 3dnowprefetch pti fsgsbase tsc_adjust bmi1 avx2 smep bmi2 erms invpcid mpx avx512f avx512dq rdseed adx smap clflushopt clwb avx512cd avx512bw avx512vl xsaveopt xsavec xgetbv1 xsaves ida arat pku ospke\nHypervisor vendor: KVM\nVirtualization type: full\nL1d cache: 1.5 MiB (48 instances)\nL1i cache: 1.5 MiB (48 instances)\nL2 cache: 48 MiB (48 instances)\nL3 cache: 71.5 MiB (2 instances)\nNUMA node(s): 2\nNUMA node0 CPU(s): 0-23,48-71\nNUMA node1 CPU(s): 24-47,72-95\nVulnerability Gather data sampling: Unknown: Dependent on hypervisor status\nVulnerability Indirect target selection: Mitigation; Aligned branch/return thunks\nVulnerability Itlb multihit: KVM: Mitigation: VMX unsupported\nVulnerability L1tf: Mitigation; PTE Inversion\nVulnerability Mds: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown\nVulnerability Meltdown: Mitigation; PTI\nVulnerability Mmio stale data: Vulnerable: Clear CPU buffers attempted, no microcode; SMT Host state unknown\nVulnerability Reg file data sampling: Not affected\nVulnerability Retbleed: Vulnerable\nVulnerability Spec rstack overflow: Not affected\nVulnerability Spec store bypass: Vulnerable\nVulnerability Spectre v1: Mitigation; usercopy/swapgs barriers and __user pointer sanitization\nVulnerability Spectre v2: Mitigation; Retpolines; STIBP disabled; RSB filling; PBRSB-eIBRS Not affected; BHI Retpoline\nVulnerability Srbds: Not affected\nVulnerability Tsa: Not affected\nVulnerability Tsx async abort: Not affected\nVulnerability Vmscape: Not affected\n\nVersions of relevant libraries:\n[pip3] numpy==2.3.4\n[pip3] nvidia-cublas-cu12==12.8.4.1\n[pip3] nvidia-cuda-cupti-cu12==12.8.90\n[pip3] nvidia-cuda-nvrtc-cu12==12.8.93\n[pip3] nvidia-cuda-runtime-cu12==12.8.90\n[pip3] nvidia-cudnn-cu12==9.10.2.21\n[pip3] nvidia-cufft-cu12==11.3.3.83\n[pip3] nvidia-curand-cu12==10.3.9.90\n[pip3] nvidia-cusolver-cu12==11.7.3.90\n[pip3] nvidia-cusparse-cu12==12.5.8.93\n[pip3] nvidia-cusparselt-cu12==0.7.1\n[pip3] nvidia-nccl-cu12==2.27.5\n[pip3] nvidia-nvjitlink-cu12==12.8.93\n[pip3] nvidia-nvtx-cu12==12.8.90\n[pip3] optree==0.17.0\n[pip3] torch==2.9.1+cu128\n[pip3] torchaudio==2.9.1+cu128\n[pip3] torchelastic==0.2.2\n[pip3] torchvision==0.24.1+cu128\n[pip3] triton==3.5.1\n[conda] numpy 2.3.4 py311h2e04523_0 conda-forge\n[conda] nvidia-cublas-cu12 12.8.4.1 pypi_0 pypi\n[conda] nvidia-cuda-cupti-cu12 12.8.90 pypi_0 pypi\n[conda] nvidia-cuda-nvrtc-cu12 12.8.93 pypi_0 pypi\n[conda] nvidia-cuda-runtime-cu12 12.8.90 pypi_0 pypi\n[conda] nvidia-cudnn-cu12 9.10.2.21 pypi_0 pypi\n[conda] nvidia-cufft-cu12 11.3.3.83 pypi_0 pypi\n[conda] nvidia-curand-cu12 10.3.9.90 pypi_0 pypi\n[conda] nvidia-cusolver-cu12 11.7.3.90 pypi_0 pypi\n[conda] nvidia-cusparse-cu12 12.5.8.93 pypi_0 pypi\n[conda] nvidia-cusparselt-cu12 0.7.1 pypi_0 pypi\n[conda] nvidia-nccl-cu12 2.27.5 pypi_0 pypi\n[conda] nvidia-nvjitlink-cu12 12.8.93 pypi_0 pypi\n[conda] nvidia-nvtx-cu12 12.8.90 pypi_0 pypi\n[conda] optree 0.17.0 pypi_0 pypi\n[conda] torch 2.9.1+cu128 pypi_0 pypi\n[conda] torchaudio 2.9.1+cu128 pypi_0 pypi\n[conda] torchelastic 0.2.2 pypi_0 pypi\n[conda] torchvision 0.24.1+cu128 pypi_0 pypi\n[conda] triton 3.5.1 pypi_0 pypi",
680
+ "transformers_version": "4.54.0",
681
+ "lm_eval_version": "0.4.11",
682
+ "upper_git_hash": null,
683
+ "tokenizer_pad_token": [
684
+ "<|endoftext|>",
685
+ "151643"
686
+ ],
687
+ "tokenizer_eos_token": [
688
+ "<|im_end|>",
689
+ "151645"
690
+ ],
691
+ "tokenizer_bos_token": [
692
+ null,
693
+ "None"
694
+ ],
695
+ "eot_token_id": 151645,
696
+ "max_length": 16384,
697
+ "task_hashes": {
698
+ "ruler_qa_squad": "8c658a728ccd3d67ddd3f45e8e76527cf15a1ae4119033d8246f293ecc68cf11",
699
+ "ruler_cwe": "f68206129fc7d709951d550ec9269fed05e73bea2fc86ce9f2c758974db5db71",
700
+ "niah_multivalue": "29646655c230b4f155a8de1fbc7ddf22c49e43045bd8fdab5770e7cb45026c7c",
701
+ "niah_multikey_3": "de2cec52885e38d56e0bda0e344e989909969c71ab5f8048b9368a011dd3616b",
702
+ "niah_multikey_1": "338f81e9efa98dbe6979d4f178f7f06cdbd6d367072318667a932b43c5032ab8",
703
+ "niah_single_2": "88ba10e924c2fc6c47543eb646b9e6f7487c7ca026c9391d6cd7054fde59eaad"
704
+ },
705
+ "model_source": "hf",
706
+ "model_name": "/workspace/outputs/l2a_style/stage2/checkpoint-25",
707
+ "model_name_sanitized": "__workspace__outputs__l2a_style__stage2__checkpoint-25",
708
+ "system_instruction": null,
709
+ "system_instruction_sha": null,
710
+ "fewshot_as_multiturn": null,
711
+ "chat_template": null,
712
+ "chat_template_sha": null,
713
+ "total_evaluation_time_seconds": "767.5954433590232"
714
+ }
outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multikey_1_2026-07-18T16-04-28.086981.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:853ac1c1aecd7005e4d2c3a2442e2f95ab2b6195a181e18dd890522231aa15fd
3
+ size 1446074
outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multikey_3_2026-07-18T16-04-28.086981.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:63b35d8f844bf3281a7c84f9bf92e04c7a9f88dfaadc0b9db2c98c140972b0b7
3
+ size 503288
outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_multivalue_2026-07-18T16-04-28.086981.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c47b0de2f9608db51ea260b6b5da89687cf93e04b3104517ee958e5312c24d63
3
+ size 1434824
outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_niah_single_2_2026-07-18T16-04-28.086981.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:83d577e7c704bafad43cc64c986087331cc9b1da48459a533c14cedda1bdcd63
3
+ size 1437596
outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_cwe_2026-07-18T16-04-28.086981.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:c9b295568237b75ee448d27bcfd3773b39c837e618a03e9e2b1c74cb81cb0fb2
3
+ size 635633
outputs/eval/token_t0525/ruler8k_splits/b/lm_eval/__workspace__outputs__l2a_style__stage2__checkpoint-25/samples_ruler_qa_squad_2026-07-18T16-04-28.086981.jsonl ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:2c66ae17c903fd34f65a7b5acce334ff5cde436941a2afd501e48c14ca9b2b99
3
+ size 1387354