Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Quite weird that heavy quantization method on a dense model gives better results than slightly quantized MoE models like 35B-A3B from Google.

At this point all the different quantization and 'compression' (look at MPO applied to LLMs...) techniques start feeling a bit like snake oil. It's just gut feeling - or scores on benchmarks models are optimized for - what ends up deciding whether a technique is good enough or not.



26B-A4B?




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: