Are you sure there's enough credits?

#1
by vovaRL - opened
ML intern explorers org

20T is a lot of tokens

ML intern explorers org

There isn't even enough credits currently to do it

ML intern explorers org
•
This comment has been hidden

ha ha no its not 20T

i made a naming mistake sorry

ML intern explorers org

How much then?

its just gonna be 1-2T 20T is too much.

ML intern explorers org

which context

it's in pretokenized chunks

2048 tokens.

its a bit short but i train in phases usally

ML intern explorers org

bro you think you can train the next GPT 4 on 1.4T tokens?

This is for 1.4T tokens
Screenshot_20260815_111743

ML intern explorers org

in 2 days you could only:

8.6B tokens (8.8B with causal-aware attention).

Budget: 2 days × 3.2e15 FLOP/s ≈ 5.5e20 FLOPs, at ~6.5e10 FLOPs/token.

That's ~0.86 tokens/param — far below Chinchilla (20×). For a 2-day run on that node, a ~400M model at 8B tokens is much better matched.

no no! im not planning to use all. i only use a subset im well aware of the time constraints

hmm true

ML intern explorers org

how many then?

i usually get data first so i dont have to worry about it later

ML intern explorers org

If you have $316,800 then go and train on 1.4T tokens for 1 year

i plan the model after the dataa im planning a 3-10B ish parameter model ill test my pipeline first, then train in monthly shards

well i dont lol

your models are quite capable for their size i checked some out previously

ML intern explorers org

I'd really reccomend like a smaller model on more tokens, then you could actually compete with other models. Thats what we do.

your profile says you have over 1900 TFLOPS of compute. Do you own 8 x H200s or is it cloud?

ok thanks for advice

ML intern explorers org

@Bc-AI oops that was a error i once wanted to see how much with 8xH200 and forgot to remove, ty

@Banaxi-Tech can i join your beta testers organisation?

ML intern explorers org

Sure!

@Banaxi-Tech @vovaRL You have inspired me to also try making a novel SLM. As my Larger one takes a looonng time to train i have decided to create a novel SLM combining elements from many sources. I will open source it too. Also @Banaxi-Tech I think you bananamind-2.1 unified concept is quite cool so i reutilised some concepts from it too :)

ML intern explorers org

@Banaxi-Tech @vovaRL You have inspired me to also try making a novel SLM. As my Larger one takes a looonng time to train i have decided to create a novel SLM combining elements from many sources. I will open source it too. Also @Banaxi-Tech I think you bananamind-2.1 unified concept is quite cool so i reutilised some concepts from it too :)

🤗

Sign up or log in to comment