poolside/Laguna-S-2.1
Text Generation β’ 118B β’ Updated β’ 82.9k β’ β’ 919
I will add to the list; may wait for specific Heretic and/or tuned version.
I already have a 43B-A3B version running in the lab ; however tuning these sparse moe models take a lot more work/time and ahh... detail. AND a lot more VRAM!!! [can't compress these atm, so BF16 required => 100 GB+ ]
What hardware do you even use for it and how long does the retraining + quants generation take?