|
Download README.md from dmsmirnov/diffusion-qa: direct link, hf CLI and curl.
- Browser
- Download file 711 Bytes
-
https://huggingface.co/dmsmirnov/diffusion-qa/resolve/main/README.md
- Command line
-
hf download hf://dmsmirnov/diffusion-qa/README.md
-
curl -L -o README.md https://huggingface.co/dmsmirnov/diffusion-qa/resolve/main/README.md
711 Bytes
| license: mit | |
| tags: | |
| - constant-warmup | |
| - giant | |
| - kaiming | |
| - lamb | |
| - layernorm | |
| - linear | |
| - mixer | |
| - multitask | |
| - swish | |
| - tucker | |
| # inference.py | |
| ## Model Overview | |
| A **giant**-scale implementation of the **mixer** architecture, built for **multitask** tasks. | |
| ## Architecture | |
| - **Architecture**: mixer | |
| - **Scale**: giant | |
| - **Attention**: linear | |
| - **Fusion strategy**: tucker | |
| - **Task head**: multitask | |
| - **Activation**: swish | |
| - **Normalization**: layernorm | |
| - **Initialization**: kaiming | |
| ## Training | |
| - **Optimizer**: lamb | |
| - **LR scheduler**: constant warmup | |
| ## Files | |
| - `inference.py` — main artifact of this repository | |
| ## License | |
| See the license field above. | |