We are building a local inference / fine-tuning solution / app to be used as a GPT alternate. I was wondering if can I take gemma-e2b model and quantize it to 1-bit and then inference it instead of using llama.cpp, using this framework directly
Mehfuz Hossain
mehfuzh
·
AI & ML interests
Co-Founder @smartloop.ai
Recent Activity
commentedon an article about 1 month ago
Falcon-Edge: A series of powerful, universal, fine-tunable 1.58bit language models. updated a dataset almost 2 years ago
smartloop-ai/lexic-ai-tutorial-dataset updated a model almost 2 years ago
mehfuzh/lexic