Suitability for 8 GB graphics cards

#18
by thermi6 - opened

Thank you so much for your excellent model! It is very useful and usable for small contexts on 8 GB graphics cards. However, as stated in a now closed discussion, the KV cache increases in size very quickly. Then it is not possible anymore to keep the data in VRAM.

Would it be possible to aim at severly limiting the size increase of the KV cache in the near future?

Right now this works on 8 GB cards: 4 bit quant + 32768 tokens context length. Any higher and it is too large for VRAM.

I think you would agree that just 32768 tokens is a bit small. What is your opinion on the matter? Do you have any pointers for improving the situation with the current model?

Sign up or log in to comment