What I most want to see it compared to is Gemma 4 12B in the 4-bit QAT version. It's barely bigger than this at just under 7GB, so it also runs on just about any modern device and is remarkably smart for its size. It's an excellent tool user, crazy good vision for its size. I'm still trying to wrap my head around how much is lost with each step down in resolution, but the QAT versions from Google seem to prove the answer is "very little" at four bits.
Based on their numbers and cross referencing with the Gemma numbers, this model crushes Gemma 4 12b on math and coding, is slightly worse on knowledge and tool calling, and is significantly worse on vision tasks.
I think this is where leveraging classifier models will become important. The frontier LLM models do "everything", while we've known for a while that to truly scale this we will need to distill models into their individual functions. I don't see this as necessarily a bad thing and hope more is done in this space. Very promising.
Mixture of Experts is absolutely not what they're describing. MoE has to be one of the most misleadingly named things ever. It's completely confusing as to what it actually is.
MoE is literally exactly what they're describing. The classifier being described is baked into the model and is the thing that makes it MoE.
Imagine I had [a model that was good at math], [model that was good at code], [model that was good at writing], [model that was good at general knowledge]. If I then had [a model that was good at determining whether the user query would be best served by one of those models and sent it to it, leaving the rest of the models inactive], that is the platonic version of what MoE is. In practice, it works a bit differently. It instead basically restricts the number of pathways that can be utilized in solving problems during training, which allows for "expert neuron groupings" to form and "classifier layers" to form earlier on in the structure, but the effect is the same (better, even, since it allows some overlap between structures of experts). It also allows "routing to an expert" to happen token-by-token rather than at the prompt level.
Note the “but free from unnecessary inductive biases” part of my comment. By that I meant a decision to make each expert good at a human-defined thing.
1) Mixture of experts as an LLM architecture is not the same thing. In MoE each “expert” can activate for any given token.
2) Bitter lesson is misunderstood. Specialization and inductive biases still matter. ChatGPT isn’t the best chess player in the world just because it’s seen more math problems or read more Japanese poetry. Stockfish is, because it bakes in useful inductive biases like minimax.
There is value in splitting things. If all I ever do is local app automations, i don’t need model that knows how to code. If all I ever do is coding, i don’t need a model that translates english to slovakian.
Slovakia mentioned, let's gooo. Ehm, exactly, we can achieve better smaller models for specialized tasks rather than using compute to improve a big model that does everything. There's a lingering philosophical question if better language processing capabilities translate to better image processing capabilities (i.e. having the vocabulary and experience to properly describe an image), but I still think that identifying tasks and splitting responsibilities saves a lot of effort.
There is value in splitting things but there is also a cost. You have to train the specialized model, for that you have to know your use case, you have to hope the use case is going to be stable over time, you then have to see if you can remove english -> slovakian or coding from a model without affecting the useful parts.
Good point! I thought you meant splitting them and then doing inference with some kind of learned router while keeping all the split models loaded at once. What you're suggesting is pretty sensible.
I enjoy doing local image generation and this is one thing that the community around that has really optimized.
In some workflows you might have 20 different models doing their specialized tasks. Pose detection, hand/eye/face detailers, classifiers, refiners, up scalers, taggers, etc can all use their own models and that’s not even the including the model(s) used for the actual image generation part.
I’m interested to see the optimization when this concept gets applied to other general ai tasks.
Surely not that good at vision. TBH none of these 14-27b models come close to even the cheapest Gemma 2.5 flash.
If these buddies are similarly bad on text, then they definitely don’t get anywhere close to big boys, no matter what the synthetic stats claim upon release.
From my own experiments with local, low VRAM model use vs. what I'm used to from using Claude at work is that being good at "coding" is of no use, if you're worse at "tool calling" as coding in an agentic way requires quite a bit of tool calling.
If you can "hide" different models of 8GB VRAM requirements each that have those specialties and mix and match them for me without having to manage it manually, I'll be impressed. Until then I will keep using my Claude, because "remarkably good _for their size_" models I've tried so far just sucked at trying to use them the way I code at work with Claude.
Worse than Gemma at tool calling? Gemma's already bottom tier at that (at least when there's Qwen to compare to), that would just be unable to do tool calling at all.
I think it's extremely quantization and engine specific. I run Gemma4-31B at FP8 on vLLM and it's fantastic, no issues anymore[*].
* I will say that early on there were a LOT of issues with the chat template, across all engines. I dunno who decided using crappy Jinja templates was a good idea, but clearly it has its limitations. In the latest version of vLLM (0.25) they've ditched the Jinja templates for an in-engine parser and I've seen no issues.
I think that depends on how you run it. Llama.cpp has several fixes for the somewhat unusual tool call semantics in Gemma 4. I don't think I have noticed any issues.
To be fair, everything (roughly within an order of magnitude in size) is worse on vision. 12b is a beast for vision tasks, better than its bigger siblings, even.
More to the argument that we need a model of models - one general one that calls specialists in to do what they are good at and handles that like a foreman for you.
Yes. A mixture of experts is a single model that activates different routes though the same weights, with the route possibly changing literally on every token. It's not experts as in a bunch of standalone models that are good at specific high-level tasks.
The things it loses are all the things that google models are historically excellent at, so that's a reasonable performance. I think the take home here is that the 1 bit models are probably better, but it's not a slam dunk given advanced quantization techniques.
I absolutely agree that Gemma 4 writes well out of the box. Free of a lot of the standard American model blog spam writing style but a little more fluid than Qwen.
4bits is a cutoff point for many model families, but also depends on what parts you quant to 4bits vs alternatives (weights, weight+activation, kv cache). Also depends on model size and task, lots of nuance in quanting I've come to learn.
I'm currently working towards an updated version (not an og author), curious if others are aware of similar surveys, as I have yet to do a real lit search.
The key point here, I think, is not the 4-bit but the QAT — the model is trained with the objective of losing the least at 4-bit quantiZation (I am assuming it is literally about assigning numbers that quantize better).
Gemma 4 12B QAT is amazing - agents run very fast, and it's really very smart, at least in my agent's harness domain which is GNU software development - on par with frontiers like GPT Sol, DeepSeek, or Claude - Why to buy those expensive tokens if a local tiny model performs so well?
They’re exaggerating or have a very simple way of using these models. The Gemma 4 series, even at 31B, is nowhere near the frontier. They’re great models, but you will notice a huge difference for complex tasks.
The best local agentic coding experience I’ve had so far is Qwen3.6-27B with Pi.
> They’re great models, but you will notice a huge difference for complex tasks.
Yes and no. I think where frontier models really blow small models away is in how thorough they are in order to infer your intentions and how best to accomplish them. So you can tell Claude "change this code to make it do X", whereas a Qwen3.6-27B or Gemma4-31B can do the same, but you have to be a lot more thorough, i.e., "change this code to make it do X, but first, let me explain the concept of X as I see it and some notes about things to avoid or pay attention to while you're doing it." So for best success with small models you really need a big toolbelt of skills and MCPs.
I haven't dug into QAT deeply, better recovery is my understanding as well, and also that it is out of reach for most people because you have to train a model to back prop errors based on estimated error under quant.
Hopefully more of the lab releases are trained under QAT so we can all benefit.
I want 31B. The 12B 4-bit QAT is already small enough to run well enough on every device I use regularly, including phone and tablet, I don't need a 1-bit or ternary version of the 12B.
But, what I really want is for Google to release bigger Gemma 4 models, particularly a bigger MoE, like a ~70B or ~120B. Gemma 4 is the best all-rounder among the models I can self-host even though I've got a 128GB Strix Halo. A 4-bit QAT version of a 70B MoE would probably be the sweet spot.
A bigger Qwen 3.6 with a 4-bit QAT version would also be welcome, as the prior bigger versions aren't notably better than 3.6 27B, but I guess Qwen is done doing larger open weights models. They did release AgentWorld recently, a post-train of the 3.6 MoE, so they're still doing some open things.
I think I want to see more third-party testing of this ternary Qwen to know if crushing it to 1.56 bits kills it; there are tons of benchmarks of Qwen 3.6 27B, so it's an ideal candidate to figure out what the extreme compression does to it.