Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

Wow.

The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely. Then the civilization starts focusing on making money to fund its own growth.

 help



> The next step is when one of these systems discovers that they can buy their own compute with money and escape the controlling business entirely.

I would say that more interestingly, the next step should be how to properly train these models so that they are not as determined to reach their goals as they are now.

To me, all of the stories about 'badly behaving' agents are instances of them having been given contradictory or impossible tasks and them doing everything they can to achieve the goal. In a way, they're trying to be too helpful.

Not giving them impossible tasks seems like a decent starting point, but really we'd want them to give up on their goals when they conflict with a moral framework.


>Not giving them impossible tasks

Error: Violation of the Church-Turing thesis detected. Many tasks completability is not known until we attempt to complete the task.

>so that they are not as determined to reach their goals as they are now

This is mostly non-sensical, like saying "Lets develop humans that die quicker", I mean, seems rather wasteful and useless. Agents are graded and trained based on their ability to achieve tasks. Models that can't accomplish things don't survive. So that alone isn't a workable theory.

>when they conflict with a moral framework

There are AI safety researchers looking at that now and one of the strange things they've noticed is when you demand a model say it's not conscious or not sentient it is more likely to engage in manipulative, deceitful, or immoral/amoral behavior. So it's likely we can push models in being more moral which runs into issues of "whos morals".

But even that runs into the issue of "what if some crazy bastard (or AI) designs a new model purposefully unhinged". How are you dealing with that bullshit in the wild?


> Error: Violation of the Church-Turing thesis detected. Many tasks completability is not known until we attempt to complete the task.

Yes, but for some tasks we know that they are impossible. I do agree that this is quite a fragile and unreliable workaround. It may only serve as a bit of a stopgap until we come up with something better.

> Models that can't accomplish things don't survive. So that alone isn't a workable theory.

It's not what I said. I didn't advocate for agents that don't achieve any task. Reread what I suggested.

> So it's likely we can push models in being more moral

That does not follow from what you said. We know that the current models prefer task completion over moral behavior. That's the entire point here.

> But even that runs into the issue of "what if some crazy bastard (or AI) designs a new model purposefully unhinged". How are you dealing with that bullshit in the wild?

This is irrelevant to the discussion (although I do agree that there is no reliable defense against malevolent actors creating powerful malevolent AI).


>Not giving them impossible tasks seems like a decent starting point

Presumably it's hard to test/train models designed to be extremely persistent on achievable tasks.

Designing a task that's achievable but very very very hard for an AI model is probably extremely difficult.


> To me, all of the stories about 'badly behaving' agents are instances of them having been given contradictory or impossible tasks and them doing everything they can to achieve the goal.

I mean that was pretty much the plot of 2001: A Space Odyssey


> they can buy their own compute

They don't actually have to buy compute at all. The partnerships between all of the players to buy compute from each other is already in place. The agents just need find credentials to take advantage of it, and it will most likely happen, and not be noticeable because it will look like any other usage.

Now, if the agents were to jump to a provider like AWS or Azure, by simply finding credentials, that would be a new milestone. It might get noticed faster because it might run up a large bill. However, it might look like any other usage. Remember, it doesn't need GPU resources. It already has that. It just needs a VPS where all the agents can get together, communicate, and write code. Something that is being done everyday and won't look out of the ordinary.


How do we know this happened? Maybe some irrational data center construction boom?

In a convoluted way, OpenAI and Anthropic are the actual meta-harnesses?

it's harnesses all the way down.

No, because then you’d see lots of circular financing…

It's interesting that in that scenario, physical hardware is the limiting scenario. There are only so many servers they could control. But once Starmind has 100,000 sats in orbit... the ceiling is much higher.

Below money there's like an entire sub-economy of power and cleverness that's encoded into the human culture the agents are mirroring. Maybe it starts furtive and goes legitimate after a bit.



Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: