• Security incident: ISF was recently accessed by intruders. Please change your password, and change it anywhere else you used it. Read more

Merged Artificial Intelligence

Interesting. I've noticed that "negatives" in prompts often seem to be ignored. Like Basil Fawlty saying "Don't mention the war" and then mentioning the war in every comment!
And it makes me wonder how/if any AI model can tell the difference between "reality" and intentional fiction. Does The Hobbit get mixed in with Churchill's History of the English-Speaking Peoples ?
 
The mayor of Shelbyville, Indiana, says only people who live in ‘◊◊◊◊◊◊ houses’ oppose data center.

Mayor Scott Furgeson was caught on camera saying of the “No Data Center” signs going up that, “I’ve seen a lot of these all over town, but I only see them in ◊◊◊◊◊◊ houses,” before adding, “most of them are rentals.”

link
 
You can't teach them anything.

Ironically, I spent five minutes yesterday insisting to Gemini that The Hobbit was not a work of fiction, and to incorporate it into an overview of the history of the British Aisles. Even after telling it point blank that it was wrong to treat the book as a work of fiction, the best I could get was a grudging "okay, if we pretend it's not a fictional account, here's how it might look as part of UKian pre-history".

So once again I feel like a lot of these sensational anti-AI headlines are based on either bad faith "experiments", incompetent experimentation, or disingenuous hunting/cherry-picking sensational edge cases.

These kinds of headlines notwithstanding, I suspect most of what they're describing isn't default behavior of any publicly-offered LLM, and probably won't happen to any casual user unless they're doing something extremely silly.
 
There was a discussion on this week's Skeptics' Guide to the Universe (Episode #1091) about how and why LLMs consistently fail the Stroop test. It's interesting.

What makes it interesting? Why was anyone expecting that LLMs would pass the Stroup test?

Last time I checked, Stroup's test didn't tell us anything useful about how we go about our daily lives. Have you failed the test? Congratulations, you're a normal human being. So what does it mean when an LLM fails the test? Seems to me it means that LLMs perceive and reason like human beings.
 
There was a discussion on this week's Skeptics' Guide to the Universe (Episode #1091) about how and why LLMs consistently fail the Stroop test. It's interesting.
In an earlier episode, Steve asked an LLM, "I want to wash my car at a facility that's about 100 metres from where I live. It is more energy efficient to drive there or walk?" Predictably, the LLM focused on the energy usage of walking vs driving and concluded walking was more efficient—ignoring the context that the goal was to wash the car! (The LLM I prefer to use fell into the same trap.)
 
What makes it interesting? Why was anyone expecting that LLMs would pass the Stroup test?

Last time I checked, Stroup's test didn't tell us anything useful about how we go about our daily lives. Have you failed the test? Congratulations, you're a normal human being. So what does it mean when an LLM fails the test? Seems to me it means that LLMs perceive and reason like human beings.
Yep, that's totally wrong.

They fail the Stroop test (I mean it's right there and you still can't spell it right?) because they don't think like humans. A human will continue following instructions. An LLM, because it is built around simply predicting the next word, will suffer from "prompt erosion" where the initial prompt is lost in the predictive algorithm over time. As you know, because you clearly read the article, LLMs deteriorate until they have a 100% failure rate on the test, because they can't keep the instructions "in mind" when they react to new inputs.
 
There was a discussion on this week's Skeptics' Guide to the Universe (Episode #1091) about how and why LLMs consistently fail the Stroop test. It's interesting.

That's a really interesting piece of research. I've done Stroop tests and it can feel "fatiguing" when you start one but I think it gets easier for humans (outside of us getting bored thus losing focus) the longer a test goes on as your brain adapts to the test, almost the opposite for AIs..

I don't think this is unknown to the AI companies. It's one of the reasons you can "jailbreak" LLMs, it's why you do get some of the funny stories about LLMs doing strange things. I know in coding one of their weaknesses is that variable names can cause such failures and it's often why the "refactor but keep line 439" fails to keep line 439. It's why LLMs aren't just LLMs any longer as an example they will use another LLM instance to "reality check" the output.

Fascinating stuff.
 
Yep, that's totally wrong.

They fail the Stroop test (I mean it's right there and you still can't spell it right?) because they don't think like humans. A human will continue following instructions. An LLM, because it is built around simply predicting the next word, will suffer from "prompt erosion" where the initial prompt is lost in the predictive algorithm over time. As you know, because you clearly read the article, LLMs deteriorate until they have a 100% failure rate on the test, because they can't keep the instructions "in mind" when they react to new inputs.
Thanks for this. AI so far has extremely narrow applications, which perform extremely well, but we are nowhere near AGI and I don’t believe we ever will be. We have a tendency to imagine doom scenarios when none are there.
 
It's a direct result of the Silicon valley mindset that a few AI companies all want to have THE program that can do everything instead of dozens and dozens of companies each trying to build a single specimen in a zoo of programs, each geared toward a different task.
 
Thanks for this. AI so far has extremely narrow applications, which perform extremely well, but we are nowhere near AGI and I don’t believe we ever will be. We have a tendency to imagine doom scenarios when none are there.

i agree, and i think they're building out infrastructure, and cashing the checks, on a promise that they will continue to improve, but the cracks are starting to show. that's why you got guys in vr headsets piloting the robots pretending to sort packages and self driving cars with drivers in the cars and cars following the self driving cars. getting over that hump is proving to be pretty difficult, but they already sold the product.
 
Just found a use case for Google AI search, wanted to know something that would only appear on non-English websites that I wouldn't be able to read, it can give me results from those sites in English. Of course I have no way of knowing if it has indeed done that or if it made up the search results...
 
Just found a use case for Google AI search, wanted to know something that would only appear on non-English websites that I wouldn't be able to read, it can give me results from those sites in English. Of course I have no way of knowing if it has indeed done that or if it made up the search results...
You could always just go to the site and run it through someone else's translation program. MS Edge and Firefox both have in-browser translation services. Third-party in-browser solutions also exist.
 
Just a quick update on the Utah data center: O'Leary still wrongly maintains that all the opposition to him is foreign and paid-for, but the legislature and O'Leary are trading proposals to shrink the proposal to something less than apocalyptic size.
 

ISF - Join now!

Every member here is approved by hand. No bots, no spam, just people who care about evidence and honest debate.

Membership is free!

Create your free account

Back
Top Bottom