Hacker Newsnew | past | comments | ask | show | jobs | submit | ImageXav's commentslogin

This is fantastic. As the little joke I hope it is. Everyone gets their own small disfunctional group, and gets to figure out the challenges of management. You, the manager, are Michael. You know you have to produce something, and you do, but you have no real idea of how. Your diligent agents are Dwight. Overly literal sycophants that are ready to leap to action at your slightest command without any question.

I do think a lot of folk would benefit from the introspection this offers. We've all been given the opportunity to become middle (and middling) managers, and a lot of the challenges we face are those of people who direct. Setting direction is tough. But LLMs are awesome tools.


Insightful. I think I’d do better at this orchestration at this particular stage of technology development if I named all my agents Dwight, maybe with an occasional Creed.


"opportunity to become middle (and middling) managers" or assistant middle managers or assistant to the middle managers


Not feeling this. Is it necessary to call "agents" by human names? Wouldn't objectives be a safer easier to remember approach to naming? Like, say 'Clips' and 'Seeks'.


Safer? From a mental health perspective? Tbh I like human names


um no. from a technical design perspective, obviously.


Gemini tops their vision evals [0] by a mile, with 4/5 top spots going to variants of it. Qwen is the only other contender, likely due to how good it is for object detection, where it crushes the competition [1].

[0] https://playground.roboflow.com/evals

[1] https://playground.roboflow.com/evals/object-detection


Me too. This is an interesting comparison but in my experience Qwen and Gemini have typically been the top contenders for image related tasks. For that reason it would be great to have the comparison here, as I'm not surprised by Gemini's dominance over the other models.


This feels extremely close in nature to being a generalisation of discrete distribution networks.

Paper: https://arxiv.org/abs/2401.00036 Project page: https://discrete-distribution-networks.github.io/

Given that these were published at ICML 2025, at which the author was a top reviewer as per their own website https://alexiglad.github.io/, I wonder how influenced they were to pursue this avenue of the back of it. They do have a related works section in appendix E, but it somehow seems to miss this. Which is odd, as the paper made a splash at the time at the conference and even made it close to the top of hacker news due to the novelty.

Of course, DDNs are a fundamentally new architecture, whereas this is more a generaliseable training strategy, but there is a very similar core.


I've been avidly using Fable since it was re-released and while it has been excellent at building the apps I want, the reasoning has been completely opaque.

Kim, however, has exposed the whole reasoning trace, or enough of it to matter. I'd almost forgotten how nice it is to see this. I've been able to see all of the weird twist and turns it takes and it is joyful. But also, far, far more informative and means I can debug ideas far more thoroughly. Also, at a first glance it seems to have gotten quite far on a niche hobby horse of mine that no LLM has been able to crack. I'll be testing this more for sure.


I have severe complaints about Anthropic's product managers on this front. Their preference for hiding, obscuring, and trying to wrest control from the user are a bit harrowing. It would be wonderful to go back to Claude Code from before March. It seems like every release destroys value for me!


It's a defensive tactic to reduce the effectiveness of distillation.

Say of that what you will, but it's not because they want to wrest control from users.

It's because they don't want Chinese companies to do exactly what Moonshot (Kimi creators) and others have done.


Anthropic’s position being that it is entitled to train models on the creative works of anyone at any time, but its own slop generators’ outputs are sacred jewels that must be protected from being learned from.


Its not surprising a thief would be the most paranoid in securing their spoils.


According to Pat Toulme, the thinking traces and outputs that the Chinese researchers distilled are useful for getting initial trajectories, to prevent a cold start during RL. Once you get those initial correct trajectories (the model actually solving a task), you can generate you own new traces, so the distillation is already done, there is no need for further reasoning traces. Sure they could probably get better alignment with the frontier by distilling fruther, but in any case the damage is already done, at this point hiding reasoning is mostly just hurting users


It's funny because Anthropic is harming my use of their product, to not even stop the supposed theft of their thought chains. Something that any other supposed upstarts could presumably grab from Moonshot or whatever now.

And my complaint also extends to all their tool use explanations, or rather the lack thereof. I get prompted continually for tool use that I can't examine, that's a poorly formatted 1kb bash script, etc. The PM desire to hide valuable information while requiring extensive interaction has really driven the product into a very unusable place compared to where it was a few months ago. (Or perhaps that's just Claude responding to its memory of my use, and I have somehow driven it to be excessively verbose and difficult to use, which would be unfortunate... Perhaps there's a way to reset the memory.)


The reasoning is key as most of the time the summary provided by fable is not enough to understand the choice and correct the logic. You have to either fully trust it or go to an exhaustive code review. This with the fact that you can only use 4.8 to security review the code produce by fable are the reasons I will not renew my anthropic subscription, the current experience is way to degraded.


What will you be replacing it with, if anything?


And recently, since GPT 5.6, OpenAI basically doesn't show anything but a single line, 5 word titles of reasoning traces - titles of summaries of reasoning i presume.

It's effectively just a completely hidden thing now.


Ok, this is really cool. The fact that the robot can use pointing to decide where to go is a great design decision, and robotics really is the next frontier. Definitely cheering on Mistral here!


I think it depends on your workflow. I've had a great experience with the trial. I work in research, and have set up something similar to Kaparthy's auto research. I, with Fable, have managed to get an image generation model down from 80M parameters to 10M and keep the quality of the generated images on par (similar FID). And, importantly, every change was modular, explained by Fable, reviewed by myself, and understood and documented. If not understood, I read relevant docs until I did it didn't accept it as part of the plan. So it ended up being a simple composition of existing ideas which I had previously encountered, but stacked much more rapidly than I could have.

The structure of the code is easily readable as I enforce concenventions followed by good libraries. And I can easily plug in new datasets. It's pretty good frankly.


I read the Economist for over a decade growing up. It was a great way to learn about the world, who was in power where, and the challenges facing economies at the time. I found their exposition to be pretty good given the fact they were restricted to a few pages for important events. However, their proposed solutions were always the same. More market freedom, etc.

I did feel with the change in editorial direction a while back that they lost some of their edge. I've since mostly just stuck to the Financial Times. It feels less worldly, but the content of the articles feels better.


"you know what would fix this problem? put the Torys in charge and let them privatize everything" -- essentially their refrain

that said, they gave generally good briefs about topics but I found whenever it was anything I knew about they tended to be fairly one-sided. Often they'd acknowledge the other view, in half a paragraph, and then immediately ignore the entire argument and keep on harping about whatever their particular axe to grind was.

having it around did help literacy in high school I guess.


Less worldly, how so? Been considering an FT sub.


Agreed. Which is also odd, if you think about it. Surely with the amount of compute Anthropic and others have available, they could test each of the solutions in the SO data they surely have and rank them based on efficiency/elegance/other criteria and remove poor solutions from their training data.


It may feel that way due to the iterative nature of medical improvements, but over the past few decades there has been a consistent reduction in cancer mortality rates across most types of cancer [0]. Treatments really are getting better and more targeted. Immunotherapy has made huge breakthroughs. Combination treatments allow for significantly improved lifespans and better quality of life during treatments. There are a few cancers that remain hard to treat, but I have a lot of confidence that in the coming decades we will make strides in attacking them. That being said, I'm very sorry to hear about the pain you and your family must be going through. I've had a few close loved ones undergo cancer treatment and it was tough.

[0] https://acsjournals.onlinelibrary.wiley.com/doi/10.3322/caac...


Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: