|
Download README.md from abdi38394/Pyce-MoE-50M-Code: direct link, hf CLI and curl.
- Browser
- Download file 790 Bytes
-
https://huggingface.co/abdi38394/Pyce-MoE-50M-Code/resolve/main/README.md
- Command line
-
hf download hf://abdi38394/Pyce-MoE-50M-Code/README.md
-
curl -L -o README.md https://huggingface.co/abdi38394/Pyce-MoE-50M-Code/resolve/main/README.md
790 Bytes
| license: mit | |
| datasets: | |
| - bigcode/the-stack-smol | |
| - bigcode/the-stack-v2 | |
| - codeparrot/codeparrot-clean | |
| language: | |
| - en | |
| tags: | |
| - moe | |
| - text-generation-inference | |
| Ce modèle est un mini-Transformer de **50 Millions de paramètres** utilisant une architecture **Mixture of Experts (MoE)** avec un routage **Top-1**. Chaque jeton n'active qu'un seul expert parmi 4 à chaque couche, permettant d'avoir la rapidité d'exécution d'un modèle ultra-léger tout en conservant une grande capacité d'apprentissage. | |
| ## Performances | |
| - **Loss Moyenne Globale** : 2.6492 | |
| - **Perplexité Finale (PPL)** : 14.14 | |
| - **Contexte** : 256 tokens | |
| ## Architecture | |
| - Couches : 5 | |
| - Têtes d'attention : 8 | |
| - Experts par couche : 4 | |
| - Dimension du modèle : 512 | |
| - Taille du vocabulaire : 16 000 (BPE Tokenizer) |