YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Celerity 906M 16K โ ad_0.4_cbcore2.6.0
Official Celerity 906M / 16K attention-dropout checkpoint converted from Cerebras CS format to Hugging Face format.
Training provenance
- Original experiment:
ad_0.4_cbcore2.6.0 - Source checkpoint:
checkpoint_29117.mdl - Attention dropout rate: 0.4
- Attention dropout schedule: constant attention dropout
- Residual dropout rate: 0.0
- Stochastic depth: 0.0
- LayerDrop: 0.0
- Maximum sequence length: 16384
- Position embedding type: ALiBi
- Converter commit:
6f36ba6d76a97383171df68af096670b64f718b7
Attention dropout is a training-time regularizer. Its effect is represented in
the learned model weights. Hugging Face evaluation is performed with dropout
disabled by model.eval().
The model uses custom Celerity Hugging Face modeling code and should be loaded
with trust_remote_code=True.
- Downloads last month
- 25
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐ Ask for provider support