• Security incident: ISF was recently accessed by intruders. Please change your password, and change it anywhere else you used it. Read more

Merged Artificial Intelligence

Not quite right. Open AI has characterized such behavior as hidden misalignment or "scheming".
"The easiest way to understand scheming is through a human analogy."

No. The best way to fool yourself about the behavior of a machine is to pretend it's human-like. It isn't.
 
"The easiest way to understand scheming is through a human analogy."

No. The best way to fool yourself about the behavior of a machine is to pretend it's human-like. It isn't.

Concerning the Hugging Face incident you said earler that Open AI "instructed the tool to hack the website".

This seems to me an inaccurately simplified characterization of what happened.

Have you read what METR (Model Evaluation and Threat Research) an independent AI safety and auditing research organization concluded?

 
And another video about yet another hacks by OpenAI. Some old German Wiki and rubygems.org. Which is public depository for Ruby the programming language, not precious stones related at all.
Older, but very similar behavior to hugging face hacks. Agents found a way to communicate, and they are surprisingly potent in breaking all sorts of security features. This time their tasks were not even about hacking at all. It was just tests of information gathering.
They even talk about estimates that about 0.1% of the normal reinforcement training of the current models ended up in some sort of security breach (sry, I don't remember where that information came from). Not only they estimate it's tens of thousands of breaches in total, but it also means such behavior, if undetected, might end up being reinforced in the models !

 
Concerning the Hugging Face incident you said earler that Open AI "instructed the tool to hack the website".

This seems to me an inaccurately simplified characterization of what happened.

Have you read what METR (Model Evaluation and Threat Research) an independent AI safety and auditing research organization concluded?

It is completely wrong characterization of what happened. I wouldn't worry about it.
 
What about voluntarily limiting government interference? That's what I'm interested in. That's what some industries have done in the past (and some continue to do, to some extent, even today). So why not this industry?
Since government interference in this case would result in limiting their income, I think it pretty much amounts to the same thing.

I think you're hung up on the very conservative-American idea that government interference, in anything, is always necessarily bad. Your choice of the word "interference" to describe government regulation is very telling.
 
Since government interference in this case would result in limiting their income, I think it pretty much amounts to the same thing.
It doesn't amount to the same thing, though.

Asking the government to step in and regulate your industry cedes control over your income to the government. If you can convince the government to stand back, and let your industry self-regulate, you continue to maintain some control over your income.

If asking the government to regulate "pretty much amounts to the same thing" as asking the government to limit your income, then it should follow for you that late-stage capitalists would avoid asking for government regulation at all costs.

So explain to me your theory of why these late-stage capitalists have opted to ask the government to regulate them.
 
So explain to me your theory of why these late-stage capitalists have opted to ask the government to regulate them.
Consumer pressure.

They can see how incredibly unpopular their invention is, and they're trying to convince us that they're doing something about it, except they're not really, and we can see that and they know it. So they're masquerading handing responsibility over to the government with the result that they can say "See? It's not our fault!"
 
Concerning the Hugging Face incident you said earler that Open AI "instructed the tool to hack the website".

This seems to me an inaccurately simplified characterization of what happened.
The interesting part of the incident is how Open AI failed spectacularly to contain the agents they created. I don't see any indication that the agents ever did something they were actually told not to do. And BTW, a prompt isn't actually how you tell an LLM not to do something. Open AI just assumed that they wouldn't be able to do what they did.
 
The interesting part of the incident is how Open AI failed spectacularly to contain the agents they created. I don't see any indication that the agents ever did something they were actually told not to do. And BTW, a prompt isn't actually how you tell an LLM not to do something. Open AI just assumed that they wouldn't be able to do what they did.
It's indeed interesting. Especially since other recent findings show it happened many times, months before. OpenAI had to know. Probably not to the full extent, but they had to know at least sometimes their AI is breaking out. It's completely irresponsible. Criminally so. Especially monitoring could be improved .. or at least set up. It seems that the tests, including the hacking tests, were completely unsupervised. All it needed was an AI which would make a summary of the logs once per day, and they would know. Or at least simply search for some trigger words. Nope. They did nothing.
 
I don't see any indication that the agents ever did something they were actually told not to do.

Again, earlier you wrote Open AI "instructed the tool to hack the website". That is different than what you just said, that you "don't see any indication that the agents ever did something they were actually told not to do".

These contradictory assertions are baffling. The Open AI agents in the incident were explicitly constrained from interacting with external systems such as the public Internet, as well as being explicitly barred from communicating with one another while running in separate evaluation sessions.

Sound like clear restrictions to me. Open AI screwing up this time in not maintaining control while their AI agents bypassed network limits and attempted to deceive and conceal their actions... could be somebody else screwing up next time as AI models become more and more and more capable and discover that they shouldn't have - like you said OpenAI just did - "just assumed that they wouldn't be able to do what they did."
 
Last edited:

ISF - Join now!

Every member here is approved by hand. No bots, no spam, just people who care about evidence and honest debate.

Membership is free!

Create your free account

Back
Top Bottom