v8: continued training to 2.6B tokens (from 1.51B). Improved coherence on narrative prompts. Same architecture (10,052,864 params, 5L d320 5-head SwiGLU RoPE).

#14
Files changed (1) hide show
  1. model.safetensors +2 -2
model.safetensors CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:e5b543e3b0b13dd9444c47d0f264c7b6f97c5f3bea9e7747565eae774188cd87
3
- size 41142264
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:5c3ce3c6c4ac2561c5be3c27f3a12a7a8272338cfee74e29d02bb8826bb6f798
3
+ size 41142232