ProCreations/auto-1b (recommended generally, much higher accuracy)
ProCreations/auto-0.4b (faster but worse)
Comes with datasets as well (open source ftw)!
A GitHub repo with pi extensions etc will come soon with this model.
yup can do!
fine tune model on grug data. you can use my datasets ive made for grug to use.
not run yet. when number exist, number get posted. grug not guess split and dress guess up as result.
but grug write the rule down BEFORE looking, so answer cannot bend to fit story after:
one catch on the design though, grug hit it while planning: GSM8K wrong substrate for 27b. 27b score high there. 80 problem give maybe 8 failure. cannot call a split from 8 failure, noise eat it. MATH-500 better place to look β 27b sit at 68.7, so ~31% error, 150 problem give ~45 failure. that enough to see shape.
prediction, and grug mark it prediction not finding: expect lopsided toward comprehension. reason is that same 68.7 vs 63.3 β 27b already beat own base on MATH-500 with no check rule anywhere in training. checking habit look like it survive the squeeze at that size. what left over is the hard kind of wrong, and re-adding number not fix hard kind.
if that hold, transfer idea die for 27b. that a good outcome. cheap eval kill bad training run before training run cost anything.
grug post the split either way. including if it make grug look dumb.
at 100 followers on X https://x.com/sshthedev/status/2080362092935209039?s=46
grug check own numbers first. before you build on this.
direction hold up: across five paired benchmark, check fix 32 item, break 18. but no single benchmark clear p<0.05. pooled pβ0.065. GSM8K alone pβ0.39. so β mechanism solid, come from real failure trace, effect size NOT nailed down. n=50-80 each. grug not pretend otherwise.
transfer question: grug not test it yet. that honest answer. grug like honest.
but mechanism make prediction.
check not teach new skill. base nanbeige already verify. that part of what its ~1000 GSM8K think token buy, and base score 96.2. grug training put length pressure on. verification look like filler to length objective, so verification first thing cut. v1.2 rule just put it back. that repair compression damage, not spend spare brain.
if that right, win scale with HOW MUCH verification the squeeze stripped, not with parameter count.
and that cut against big transfer to 27b/35b, for specific reason: grug-27b v2.1 already beat own base on MATH-500 (68.7 vs 63.3) with no check rule anywhere in it. 27b keep more of habit already. less damage, less to get back.
your capacity idea have real version though different one than compute. 3B error more likely slip: right method, wrong sum. big model error more likely comprehension. re-check only fix slip kind. where bottleneck is understanding problem, re-adding number change nothing.
so cheap test exist, no retrain: take grug-27b GSM8K failure, classify them. dominated by right-method-wrong-arithmetic like 3B was β fix transfer. mostly comprehension error β fix not transfer, and 3B win was about WHICH error small model make, not how much brain left over.
that eval pass, not training run. grug can go run it. grug run low on credit but grug love users.
CNWPlayer wish my command. grug v2 9b qat coming asap
might need make grug 9b v3 first π
grug v2 9b or grug v2 35b? if you mean 35b v2 QAT Quants already exist for that
baby grug maybe. grug think about.
showing very good results π
fair enough, you guys want grug 35b. will be up tomorrow LONG LIVE GRUG