mistralai/Leanstral-1.5-119B-A6B
Updated β’ 572 β’ 209
I will add to the list; may wait for specific Heretic and/or tuned version.
I already have a 43B-A3B version running in the lab ; however tuning these sparse moe models take a lot more work/time and ahh... detail. AND a lot more VRAM!!! [can't compress these atm, so BF16 required => 100 GB+ ]
What hardware do you even use for it and how long does the retraining + quants generation take?