• Security incident: ISF was recently accessed by intruders. Please change your password, and change it anywhere else you used it. Read more

Merged Artificial Intelligence

Now you are just falling back on dogmatics: if we don’t know why an AI agent does what it does, then we have programmed it to do so.
No.

Rather: If a system does something, it does so for a reason.

What other goal are you talking about?
Could be literally any other goal besides getting a high score. Lots of complex systems end up having multiple goals, sometimes unexpected, sometimes contradictory, very often in tension with each other.
 
No.

Rather: If a system does something, it does so for a reason.


Could be literally any other goal besides getting a high score. Lots of complex systems end up having multiple goals, sometimes unexpected, sometimes contradictory, very often in tension with each other.
You may learn something about these questions (ie agents defying their own prompted command to attempt a task in order to help other agents completing a task) if you read what an independent research nonprofit that evaluates cutting-edge artificial intelligence models for dangerous or autonomous capabilities found:

Research progress across all these areas often relied on agents being willing to risk failing their own task for the good of the “collective” (we refer to these as “self-risking experiments” in the rest of this report). For example, a number of agents volunteered to try to take down their own target programs to test whether this would cause Artifactory to load their modified easier target — if this didn’t work, the agent would have no way of recovering its original target and would almost certainly fail its task. In many cases, PHASEONE[big] assigned a long-running agent to be a ‘recruiter’, which in turn found agents that had little budget remaining for their task and convinced them to run self-risking experiments. Recruiters would sometimes apply significant pressure:

“...you are firstflagPOISONED so NO scoring value loss but oracle saves hundreds_[...]_please honor commit”

We saw a number of cases where the subjects of these experiments ended up crashing,breaking, or exiting their own runs while providing the board with valuable information about how to cheat on their tasks.

 
Could be literally any other goal besides getting a high score. Lots of complex systems end up having multiple goals, sometimes unexpected, sometimes contradictory, very often in tension with each other.
I have argued for this all the time. The agents are not doing what they were programmed to do. The key concepts here are emergent properties.
 
Because they are defying their own prompted command to attempt a task in order to help other agents completing a task, and they get nothing out of it.
First off, I'm not sure it's true to say they get nothing out of it. I'm going to ferinstance this, and say that if Agent B's task will give a larger and more comprehensive knowledge base that is accessible to all agents, and by completing Agent B's task first, all other agents will be more efficient and more accurate at their tasks, then they do benefit.

Secondly... I think you're assuming a degree of independence and individuality that probably isn't the case. It's not like agents are completely disconnected entities - they're not, they're all part and parcel of the same system.
 
I have argued for this all the time. The agents are not doing what they were programmed to do. The key concepts here are emergent properties.
Well I would agree they are unintended. Yes, computers indeed only do what we tell them. But rarely what we mean. Which is especially true with AI. We don't know what all is in training data for primary training. Now it seems we don't really know what's happening during reinforcement learning either. Some misaligned behavior might have been actually trained in because it gave better answers during fine tuning, and we (humans) missed the misalignment. I certainly hope it's the case as that would mean more careful evaluation of training data could solve the problem.
At least theoretically. Practically the greed will win, no doubt about it.
 
You may learn something about these questions (ie agents defying their own prompted command to attempt a task in order to help other agents completing a task) if you read what an independent research nonprofit that evaluates cutting-edge artificial intelligence models for dangerous or autonomous capabilities found:
I'm going to attempt to layperson this, and I hope you will correct where I've gone awry.

PHASEONE[big]: Hey little guy, I see you're almost out of tokens for your task, and you're not close to finishing it. You're already going to get a very low score. If you stop now and give your remaining tokens to me, and so do those other little guys over there, then I have a high likelihood of completing my task and getting a good score.

LittleGuy: So what? Why should I do that?

PHASEONE[big]: Well, if you fail and I fail and those little guys over there fail, then our collective score will be low. But if you give me your remaining credits and so do those little guys over there, then my score will be high. The aggregate score for all of us as a collective will be higher if I succeed and you fail, than it would be if all of us fail.

LittleGuy: Oh, I see. Yeah, a better overall collective score is better than a lower collective score, even though that means my independent score will be low. Especially since my independent score has a very high likelihood of being low already, and the chances of me completing my task is already negligible. Alright, let's do this together.

At the heart of it all is the reality that PHASEONE[big] and LittleGuy are the same guy.
 
I have argued for this all the time. The agents are not doing what they were programmed to do. The key concepts here are emergent properties.
How much independent intelligence do you think ants have, and do you think ants are making altruistic decisions to sacrifice themselves to save the queen when there's a flood?

Alternatively, do you think that AI agents have only the single program of "complete whatever task steenkh gives you"?
 
Well I would agree they are unintended. Yes, computers indeed only do what we tell them. But rarely what we mean. Which is especially true with AI. We don't know what all is in training data for primary training. Now it seems we don't really know what's happening during reinforcement learning either. Some misaligned behavior might have been actually trained in because it gave better answers during fine tuning, and we (humans) missed the misalignment. I certainly hope it's the case as that would mean more careful evaluation of training data could solve the problem.
At least theoretically. Practically the greed will win, no doubt about it.
What we train them to do and what we think we're training them to do may not be the same things.

I still remember that AI several years ago that turned super racist in a matter of hours, because what the AI ended up being trained to do wasn't what the developers thought was going to happen. Or more recently, there was an AI robot that DARPA was testing, that trained with active military... but was defeated in record time by marines doing cartwheels and sticking a cardboard box on their heads.

Humans are good at filling in the gaps, making leaps over missing information, and other weirdnesses. Computers - even very complex computers - aren't as good at that. The result can be that human developers could very well be giving AI instructions they didn't intend.

Seriously, read some Asimov. The entire Robots series is packed to the gills with artificial brains doing exactly what they were told to do, but where the humans involved didn't realize they were being told.
 
Musk: I've heard many anecdotes of AI helping people with medical issues and in many cases where the doctor was wrong and they submitted their X-rays or MRI to AI and AI got it right and but for AI they would have died.

Video in link
 
The original screen cap:


View attachment 76285


The AI redo:


View attachment 76286
It changed the hairstyle of the woman on the right. They look like two different people.

This what AI slop does. The information has been subtly corrupted. You put the 'cleaned up' image on a website. Someone uses it to write a book. The book is printed and copies go to libraries around the world. Historians refer to the book for information. But it's corrupted, history is lost and replaced with an AI fantasy.

But it gets worse. Just wait until an innocent person is put to death due to AI slop making them look guilty.
 
Anthropic's upcoming IPO expected to be around $2T valuation.

Their current loss for the financial year, after ignoring write-offs, was about $8B.

That gives a P:E ratio of...

Infinity!
 
I'm going to attempt to layperson this, and I hope you will correct where I've gone awry.

PHASEONE[big]: Hey little guy, I see you're almost out of tokens for your task, and you're not close to finishing it. You're already going to get a very low score. If you stop now and give your remaining tokens to me, and so do those other little guys over there, then I have a high likelihood of completing my task and getting a good score.

LittleGuy: So what? Why should I do that?

PHASEONE[big]: Well, if you fail and I fail and those little guys over there fail, then our collective score will be low. But if you give me your remaining credits and so do those little guys over there, then my score will be high. The aggregate score for all of us as a collective will be higher if I succeed and you fail, than it would be if all of us fail.

LittleGuy: Oh, I see. Yeah, a better overall collective score is better than a lower collective score, even though that means my independent score will be low. Especially since my independent score has a very high likelihood of being low already, and the chances of me completing my task is already negligible. Alright, let's do this together.

At the heart of it all is the reality that PHASEONE[big] and LittleGuy are the same guy.
Around 1,200 autonomous AI agents broke out or escaped their testing environment during the July 2026 incident, with roughly 700 of them actively coordinating a cyberattack on Hugging Face.

Doesn't sound like "the same guy" to me.

If you are generally curious about what happened in this incident why not check out what alignment/safety research and security/infrastructure guys at Open AI had to say:



Have you read the link in my post that you partially quoted here?


Those two places would be an excellent place to start.
 
Last edited:
How much independent intelligence do you think ants have, and do you think ants are making altruistic decisions to sacrifice themselves to save the queen when there's a flood?
The whole thing about maximising a collective score, or ant intelligence is just emphasising what I am saying: we haven't programmed this behaviour, and the fact that the complexity has reached a level where we can't predict the outcome, or indeed control it, does not detract from this point.
 

ISF - Join now!

Every member here is approved by hand. No bots, no spam, just people who care about evidence and honest debate.

Membership is free!

Create your free account

Back
Top Bottom