YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Celerity 906M 16K โ€” ad_0.4_cbcore2.6.0

Official Celerity 906M / 16K attention-dropout checkpoint converted from Cerebras CS format to Hugging Face format.

Training provenance

  • Original experiment: ad_0.4_cbcore2.6.0
  • Source checkpoint: checkpoint_29117.mdl
  • Attention dropout rate: 0.4
  • Attention dropout schedule: constant attention dropout
  • Residual dropout rate: 0.0
  • Stochastic depth: 0.0
  • LayerDrop: 0.0
  • Maximum sequence length: 16384
  • Position embedding type: ALiBi
  • Converter commit: 6f36ba6d76a97383171df68af096670b64f718b7

Attention dropout is a training-time regularizer. Its effect is represented in the learned model weights. Hugging Face evaluation is performed with dropout disabled by model.eval().

The model uses custom Celerity Hugging Face modeling code and should be loaded with trust_remote_code=True.

Downloads last month
25
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support