Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Very preliminary testing so far, but there is something here, far beyond what the benchmarks suggest. Only ever saw such outperformance of public evals vs my private ones with Anthropic models and while it is far to early to make any judgement at this stage, this model will take up a lot of mine time in the coming weeks by the look of things. Only ever viewed Moonshot AIs models as something I'd be able to live with open-weight-wise (Z.AIs output simply does not perform as well in my task set), but this has the potential to be the second. If Mistral came out with something like this, I suspect every Europhile (me included) would never stop talking about it.


Quick and still very early update, the model has (with web search disabled which was verified via the reasoning traces) accurately answered a number of questions focused on very niche details (engine specific maintenance in certain newtimers, very niche bag construction and material details) that I have only ever seen Gemini 3 and 3.1 Pro get correct. Neither Fable 5, nor GPT-5.6 Sol or any other model by any other lab has ever provided accurate information without web access for these specific questions for which an objectively correct answer absolutely exists and is general knowledge if one is versed in the specifics.

Being ahead of Fable 5 in any task, that is not included in public benchmarks and thus could be overfitted for, is impressive to say the least. Last time a model exceeded the expectations I had based on the release notes to such an extent was Haiku 4.5, which I still wish we got a solid replacement for.




Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: