|
Download README.md from abdi38394/Pyce-MoE-50M-Code: direct link, hf CLI and curl.
- Browser
- Download file 790 Bytes
-
https://huggingface.co/abdi38394/Pyce-MoE-50M-Code/resolve/main/README.md
- Command line
-
hf download hf://abdi38394/Pyce-MoE-50M-Code/README.md
-
curl -L -o README.md https://huggingface.co/abdi38394/Pyce-MoE-50M-Code/resolve/main/README.md
790 Bytes
metadata
license: mit
datasets:
- bigcode/the-stack-smol
- bigcode/the-stack-v2
- codeparrot/codeparrot-clean
language:
- en
tags:
- moe
- text-generation-inference
Ce modèle est un mini-Transformer de 50 Millions de paramètres utilisant une architecture Mixture of Experts (MoE) avec un routage Top-1. Chaque jeton n'active qu'un seul expert parmi 4 à chaque couche, permettant d'avoir la rapidité d'exécution d'un modèle ultra-léger tout en conservant une grande capacité d'apprentissage.
Performances
- Loss Moyenne Globale : 2.6492
- Perplexité Finale (PPL) : 14.14
- Contexte : 256 tokens
Architecture
- Couches : 5
- Têtes d'attention : 8
- Experts par couche : 4
- Dimension du modèle : 512
- Taille du vocabulaire : 16 000 (BPE Tokenizer)