hmm

#1
by Bc-AI - opened
Novi AI org

@GGUFGuy what gpu did you train this on? i think if your gpu supports you can try Bfloat16 in future its what i usually use.

@GGUFGuy what gpu did you train this on? i think if your gpu supports you can try Bfloat16 in future its what i usually use.

molab rtx pro 6000 blackwell, even with that stuff it still took 2 hours

Novi AI org

@GGUFGuy You should use fp32 master weights CUDA autocast to bfloat16 and possibly write your own kernels. i can hit 125K toks per sec on the molab rtx 6000 with that strat

@GGUFGuy You should use fp32 master weights CUDA autocast to bfloat16 and possibly write your own kernels. i can hit 125K toks per sec on the molab rtx 6000 with that strat

During pretraining on Novi-Micro i got ~140K tokens per second, with FP32 btw

Novi AI org

Oh that’s because it was 5M Params. Mine was 134M params

Sign up or log in to comment