Regular-context pretraining releases: documents plus aligned Dolma-2 tokens and masks, one repo per source mix. Token counts on each card.
AI & ML interests
None defined yet.
Recent Activity
View all activity
Long-context (16K packing) pretraining releases, same two-parquet contract as the regular mixes. Token counts on each card.
-
placeholderlabs/pretrain-repository-v2-mix-long-context
Viewer • Updated • 5.16M • 92 -
placeholderlabs/pretrain-academic-mix-long-context
Viewer • Updated • 2M • 166 -
placeholderlabs/pretrain-commits-v2-mix-long-context
Viewer • Updated • 1.98M • 353 -
placeholderlabs/pretrain-nemotron-math-mix-long-context
Viewer • Updated • 84.8k • 68
Regular-context pretraining releases: documents plus aligned Dolma-2 tokens and masks, one repo per source mix. Token counts on each card.
Long-context (16K packing) pretraining releases, same two-parquet contract as the regular mixes. Token counts on each card.
-
placeholderlabs/pretrain-repository-v2-mix-long-context
Viewer • Updated • 5.16M • 92 -
placeholderlabs/pretrain-academic-mix-long-context
Viewer • Updated • 2M • 166 -
placeholderlabs/pretrain-commits-v2-mix-long-context
Viewer • Updated • 1.98M • 353 -
placeholderlabs/pretrain-nemotron-math-mix-long-context
Viewer • Updated • 84.8k • 68
models 0
None public yet
datasets 50
placeholderlabs/pretrain-reasoning-traces-mix
Viewer • Updated • 3.18M • 63
placeholderlabs/pretrain-reasoning-traces-mix-long-context
Viewer • Updated • 1.78k • 30
placeholderlabs/Qwen3.6-27B-Reasoning-Regen-Sharded
Updated • 51
placeholderlabs/GLM-5.2-PerfectBlend-Regen-Sharded
Updated • 65
placeholderlabs/Nemotron-SFT-Math-v4-SE-Sharded
Updated • 60
placeholderlabs/DeepSeek-V4-Pro-Distilled-Math-Sharded
Updated • 27
placeholderlabs/Kimi-K2.7-CodingTraces-Sharded
Updated • 21
placeholderlabs/Kimi-K2.5-Reasoning-Extra-Sharded
Updated • 40
placeholderlabs/GLM-5.2-Conversation-Sharded
Updated • 20
placeholderlabs/Tachibana4-DeepSeek-V4-Pro-Sharded
Updated • 32