|
Download README.md from fromziro/DataBench: direct link, hf CLI and curl.
- Browser
- Download file 2.35 kB
-
https://huggingface.co/fromziro/DataBench/resolve/main/README.md
- Command line
-
hf download hf://fromziro/DataBench/README.md
-
curl -L -o README.md https://huggingface.co/fromziro/DataBench/resolve/main/README.md
2.35 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| tags: | |
| - ablations | |
| - datasets | |
| - 1m | |
| - llama | |
| - custom_tokenizer | |
| # DataBench | |
| Rather than training one model, we trained multiple models, each on a different dataset. Everything else was kept the same, so that the only changing variable would be the dataset itself. | |
| ## Model Architecture | |
| - Base Architecture: `LlamaForCasualLM` | |
| - Tokenizer: [`Harley-ml/Dillionv2-1.3M`](https://huggingface.co/Harley-ml/Dillionv2-1.3M) | |
| - Transformers Version: `5.13.1` | |
| - Hidden Size: `128` | |
| - Vocab Size: 2564 | |
| - Number of Layers: `8` | |
| - Number of Heads: `4` | |
| - Number of KV Heads: `2` | |
| - Intermediate Size: `344` | |
| - Head Dim: `32` | |
| - Max Position Embeddings: `256` | |
| - RoPe Theta: `2500.0` | |
| - Tie Word Embeddings: `true` | |
| - Hidden Activation: `silu` | |
| - MLP Bias: `false` | |
| - Initializer Range: `0.2` | |
| - RMS Norm Eps: `1e-06` | |
| - Pretraining Tp: `1` | |
| - Use Cache: `false` | |
| - **Total Parameters**: `1,780,352` | |
| ## Training Setup | |
| - Epochs: `1` | |
| - Max Steps: `-1.0` | |
| - Batch Size: `400` | |
| - Sequence Length: `256` | |
| - Gradient Accumulation: `2` | |
| - Gradient Clipping: `1.0` | |
| - Gradient Checkpointing: `true` | |
| - Learning Rate: `2.5e-3` | |
| - Eval Split: `0.00165` | |
| - Weight Decay: `0.01` | |
| - Optimizer: `AdamW` | |
| - AdamW Betas: `(0.9, 0.95)` | |
| - AdamW Eps: `1e-8` | |
| - Scheduler: `WSD` | |
| - WSD Warmup Ratio: `0.015` | |
| - WSD Stable Ratio: `0.78` | |
| - WSD Decay Ratio: `0.20` | |
| - WSD Minium LR Ratio: `0.0` | |
| - WSD Number of Cycles: `0.5` | |
| - DType: `float16` | |
| - Torch.Compile: `true` | |
| - DataLoader Workers: `2` | |
| - Seed: `311` | |
| # Results | |
| Accuracy is normalized by length and shown as a percentage. | |
| | Dataset | ARC-Easy | HellaSwag | PIQA | **Avg ↑** | | |
| | :---------------- | ---------: | ---------: | ---------: | ---------: | | |
| | FineWeb | 27.82% | 26.89% | 52.83% | **35.85%** | | |
| | DCLM-1.0-Baseline | 29.08% | 27.06% | 52.23% | **36.12%** | | |
| | DOAB | 31.14% | 28.00% | 52.18% | **37.11%** | | |
| | Project Gutenberg | 25.84% | 24.63% | 50.21% | **33.56%** | | |
| | EOT-2004-Raw | 29.08% | 27.54% | 51.31% | **35.98%** | | |
| | Wikipedia | 28.28% | 27.61% | 51.14% | **35.68%** | | |
| (Will add ArithMark-3.0 later) | |
| ## Notice | |
| This is a work in progress and is currently not completed. By the end of this project, we aim to have tested over 50 datasets. | |
| ## License | |
| Apache 2.0. |