Kite

🎉 You are looking at Kite 8, which is more efficient!

Kite is a small, trained, 7 million parameter language model.

Training

It was trained on the first shard of semran1/cosmopedia-v2-subset, using 1 epoch, 12 batch size, 1e-3 learning rate, and the pika 5 tokenizer.

Limitations

Due to its size, the model is not suitable for production workloads.

Downloads last month
125
Safetensors
Model size
6.79M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train qikp/kite-8-7m-base