|
Download README.md from GODELEV/Rose-Pro: direct link, hf CLI and curl.
- Browser
- Download file 8.26 kB
-
https://huggingface.co/GODELEV/Rose-Pro/resolve/main/README.md
- Command line
-
hf download hf://GODELEV/Rose-Pro/README.md
-
curl -L -o README.md https://huggingface.co/GODELEV/Rose-Pro/resolve/main/README.md
8.26 kB
| license: apache-2.0 | |
| language: | |
| - en | |
| pipeline_tag: text-generation | |
| tags: | |
| - rose | |
| - custom_arch | |
| - refresh_gates | |
| - original_token_embeddings_reinjection | |
| - 150M_param | |
| - 80B_token_pretrained | |
| - Pretrained_from_scratch | |
| <img src="Banner.png" width="1920"/> | |
| # Rose Pro | |
| ## Architecture graph | |
| <a href="https://hfviewer.com/GODELEV/Rose-Pro?utm_source=huggingface&utm_medium=embedded_model_card&utm_campaign=GODELEV_Rose-Pro_card" target="_blank" rel="noopener"> | |
| <img | |
| src="https://hfviewer.com/api/card.svg?source=GODELEV%2FRose-Pro&granularity=0" | |
| alt="Architecture graph for GODELEV/Rose-Pro. Open in hfviewer" | |
| width="100%" | |
| /> | |
| </a> | |
| ## A Rose That Finally Found the 1100 Club | |
| > I trained a model, watched the numbers climb, | |
| > And told myself, "This one will be prime." | |
| > I widened the model, increased the scale, | |
| > Then watched ArithMark politely fail. | |
| > | |
| > Eighty billion tokens went into the machine, | |
| > Surely this time my model would be clean. | |
| > The Benchmarks smiled, the scores came through, | |
| > And said, "Nice little model. Here's some work for you." | |
| > | |
| > Every time I think I've finally learned, | |
| > Another benchmark shows where I was wrong. | |
| > Truly, I am blessed beyond compare; | |
| > **I keep discovering mistakes I can afford to train.** | |
| --- | |
| ## Introduction | |
| **Rose Pro** is the latest model in the Rose X1 line, with approximately **151.3 million parameters**. | |
| It keeps the same general Rose X1 architecture while increasing the model's capacity, using **24 layers** and a **640-dimensional hidden representation**. The model was trained on approximately **80 billion tokens**. | |
| And yes, I had expectations. | |
| Quite reasonable ones, obviously. | |
| I had a bigger model, more capacity, a mountain of tokens, and enough confidence to become suspicious. | |
| Then the Benchmarks arrived. | |
| It did not exactly hand me a trophy. | |
| It did, however, hand me something much more useful: | |
| **another collection of things to learn.** | |
| --- | |
| ## Model Overview | |
| | Property | Value | | |
| |---|---| | |
| | **Model** | Rose Pro | | |
| | **Architecture** | Rose X1 | | |
| | **Parameters** | 151.274M | | |
| | **Layers** | 24 | | |
| | **Hidden Size** | 640 | | |
| | **Attention Heads** | 10 | | |
| | **KV Heads** | 10 | | |
| | **QK Norm** | Yes | | |
| | **Refresh Gates** | [9, 18] | | |
| | **Native Context Window** | 2048 | | |
| | **Training Tokens** | 79,967,283,200 | | |
| | **Validation Perplexity** | 5.08 | | |
| | **Optimizer** | muon_adamw | | |
| The architecture itself is intentionally not explained in excessive detail here. Rose Pro remains part of the **Rose X1** family, and the interesting part of this release is less about repeating the architecture diagram and more about seeing what the model actually does with its increased capacity. | |
| # The Chemistry of Benchmarks | |
| This is probably my favorite part of the whole release. | |
| The Benchmarks are strange little ecosystem. | |
| You put a model into it with millions of parameters, billions of tokens, a carefully designed architecture, and a considerable amount of optimism. | |
| Then the Benchmarks removes the optimism and gives you percentages. | |
| And suddenly you discover that one model is great at something, another model is better somewhere else, and a benchmark you thought would behave nicely has decided to personally insult you. | |
| That is the **chemistry of Becnhmarks**. | |
| Rose Pro does not dominate every benchmark, and I don't think it needs to. | |
| What matters more is the overall shape of the results and what those results teach me about the model. | |
| # Benchmark Results | |
| The following results are from **0-shot evaluation**. | |
| | Benchmark | Rose-Mini | Rose-Medium | Rose Pro | | |
| |---|---:|---:|---:| | |
| | **HellaSwag** | 28.95% | 35.29% | **38.24%** | | |
| | **PIQA** | 58.54% | 62.95% | **65.18%** | | |
| | **ARC-Easy** | 36.62% | 44.19% | **48.65%** | | |
| | **ARC-Challenge** | 24.32% | 26.19% | **26.96%** | | |
| | **WinoGrande** | 51.54% | 49.80% | **50.43%** | | |
| | **BoolQ** | 61.90% | 55.66% | **60.40%** | | |
| | **OpenBookQA** | 29.40% | 32.60% | **33.60%** | | |
| | **CommonsenseQA** | 19.74% | 21.21% | **19.49%** | | |
| | **LAMBADA** | 22.03% | 31.71% | **32.99%** | | |
| | **BLiMP** | 76.72% | 79.30% | **79.26%** | | |
| | **MMLU** | 23.24% | 23.98% | **26.13%** | | |
| | **SciQ** | 55.90% | 67.10% | **68.80%** | | |
| | **COPA** | 59.00% | 69.00% | **69.00%** | | |
| | **RACE** | 27.18% | 29.09% | **30.24%** | | |
| | **SWAG** | 44.32% | 52.79% | **55.39%** | | |
| | **TruthfulQA MC2** | 44.28% | 41.83% | **39.82%** | | |
| | **WikiText-2 Word PPL** | 46.05 | 27.67 | **25.05** | | |
| | **WikiText-2 Byte PPL** | 2.047 | 1.861 | **1.826** | | |
| The results show a mixed but useful picture. Rose Pro moves strongly on several tasks while some benchmarks barely move and a few go backwards. | |
| Which is precisely why I don't want to turn this README into a victory speech. | |
| The model is good at some things. | |
| It is still learning how to behave itself at others. | |
| # ArithMark 3 & BananaMind | |
| These two deserve their own little section. | |
| | Model | ArithMark 3 | BananaMind Bench 1.1 | | |
| |---|---:|---:| | |
| | **Rose-Mini** | 36.50% | 1037.37 Elo | | |
| | **Rose-Medium** | 38.30% | 1098.7 Elo | | |
| | **Rose Pro** | **38.20%** | **1105 Elo** | | |
| ## ArithMark 3 | |
| Rose Pro scores **38.20%** on ArithMark 3. | |
| Rose-Medium scored **38.30%**. | |
| So I increased the model. | |
| I increased the width. | |
| And ArithMark looked at all of that and said: | |
| **38.20%.** | |
| Thank you. | |
| Very informative. | |
| But that is also exactly why I find these evaluations useful. Parameter count does not magically translate into proportional improvement on every task. | |
| Apparently I have to actually understand what is happening. | |
| How inconvenient. | |
| ## BananaMind Bench 1.1 | |
| And then there is the number I had been waiting for. | |
| **Rose-Medium: 1098.7 Elo** | |
| **Rose Pro: 1105 Elo** | |
| I finally crossed **1100 Elo**. | |
| Is 1100 a scientifically meaningful boundary? | |
| No. | |
| Did 1098.7 bother me enough that I wanted to cross it? | |
| Absolutely. | |
| So I am going to enjoy this completely arbitrary victory. | |
| **1105.** | |
| Beautiful. | |
| # What I Learned This Time | |
| One thing has remained remarkably consistent throughout my model-building journey: | |
| **Every time I train a model, I learn something.** | |
| Sometimes it is about architecture. | |
| Sometimes it is about optimization. | |
| Sometimes it is about evaluation. | |
| And sometimes it is about a mistake that I would very much like to pretend I never made. | |
| But every model leaves me with something new to investigate. | |
| That is probably the most valuable part of building these models. | |
| I start with an idea. | |
| I train it. | |
| I evaluate it. | |
| Then the results politely inform me that my original understanding was only partially correct. | |
| And honestly? | |
| I think I am blessed. | |
| Not because every model performs exactly how I imagined. | |
| Quite the opposite. | |
| I am blessed because every training run gives me another opportunity to discover something I didn't know before. | |
| Apparently this is my preferred method of education. | |
| It is just considerably more expensive than reading a book. | |
| # Training Recipe: One More Thing to Change | |
| The Rose models so far have been trained using **cosine learning-rate scheduling**. | |
| For the next generation, I want to try something different. | |
| The upcoming models will explore **WSD learning-rate scheduling**. | |
| Not because cosine suddenly became illegal. | |
| I simply have another question I want to answer. | |
| And apparently the only way I know how to answer questions is to build another model. | |
| # What's Next? | |
| The Rose X1 journey is not finished. | |
| The next step will be the **Rose 1.5 series**, which will explore further changes to the training recipe and continue the work started with the current Rose models. | |
| Hopefully, the next leaderboard visit will contain fewer surprises. | |
| Knowing me, that statement probably guarantees the opposite. | |
| # Final Thoughts | |
| Rose Pro is not exactly the model I imagined when I started building it. | |
| Some results are encouraging. | |
| Some are competitive. | |
| Some are stubborn. | |
| And some are sitting there quietly reminding me that adding more parameters is not the same thing as solving the problem. | |
| But that is fine. | |
| The point of every model is not to prove that I was right. | |
| Sometimes the point is to find out where I was wrong. | |
| Rose Pro gave me another set of answers. | |
| It also gave me several new questions. | |
| And, most importantly: | |
| **1105 Elo.** | |
| I'll take it. | |
| For now. | |
| --- | |