• Security incident: ISF was recently accessed by intruders. Please change your password, and change it anywhere else you used it. Read more

Merged Artificial Intelligence

IMHO it's just hallucinations, yet again. AI's numeric precision has a limit, the training is basically a lossy compression, and in the last step we add temperature and we don't pick the best possible answer. You can create a generally correct text with that. But you can't play chess with that.
 
I don't think you meant to reference the rules. Moves outside the rules are illegal. I suspect you meant to say they were moves outside the conventional gameplay. To some extent, chess is a formulaic game.

Are they winning moves? Or is the phenomenon simply that they don't follow the traditional openings, gambits, etc.? We've been teaching computers specifically to play chess for decades.

The domain of chess moves is both unambiguously defined and strictly bounded. Show me any chess board with any allowable combination of pieces and I can completely enumerate literally all the possible next moves. And from each of those I can enumerate all the other side's possible following moves. Because the solution set is closed and bounded, it's not the same as an open-ended prompt space of training space. So it's comparatively simple to generate chess moves that fall outside convention, but it's not at all like navigating an open-ended tensor space where convention is baked in.
The problem is that LLM chatbots aren't designed to play chess. Real chess engines are purpose-built. You can ask a chatbot to play a game of chess with you, but it's not good at it and will even propose moves that aren't legal moves. It's a bit like saying that a hammer isn't very useful for brushing your teeth, and a toothbrush isn't very useful for driving nails.
 
would you still think they are intelligent when they still make an illegal move after studying millions of master chess games for centuries in human years? Or would you surmise that they just can't understand the rules?
I was not talking about intelligence as in "Einstein" intelligence, but just intelligence as children have it.

I myself could look at master games for years and not understand anything of it, so I don't understand your point. Contrary to the child, AIs can actually play quite well, despite making the occasional illegal move, so there are differences as well as similarities. I wouldn't rule out an intelligence in AIs solely because they don't have flashes of genius, or don't understand chess rules better than children.

We don't have a framework for understanding, or acknowledging intelligence that is not human, and we tend to think that if it is not human intelligence, then it isn't intelligence at all. And in some sense, it is correct.
 
I think it's worth teasing out a theme here. An LLM based AI can perform task X. A chess engine AI can perform task Y. The more specialized the AI is the better it performs. No single AI is going to perform well at all human activities. So when someone says "AI can do <task>/replace employees doing <task>" etc ask whether the specific AI in question can do all the tasks in a persons job description. We'd do it with humans. "People can write effective SQL" "Yeah but can this specific guy applying for this job do it?"

Also it's standard practice not to bet the farm on new employees. Even with my rather nice CV my last job required a 6 month probation period to prove I could apply all the skills and knowledge I displayed in the interview process to this particular area.

PS excuse typos and grammar. Another extended period of insomnia. Advice not sought.
 
LLM constantly come up with new Chess Moves that the Rules never envisioned ....

I was just riffing on Chanakya's doubts about LLMs being able to come up with something new - when fed with the correct and only rules of Chess, a LLM will nevertheless come up with new, illegal moves that it thinks will be in line with the game rules - because it doesn't understand what a rule is.

I don't know, it might perhaps qualify, in a kind-of-sort-of way. After all, that's kind of like an Einstein thinking up his space-time thing, going outside the rules of what is space and time and gravity. At least we might be inclined to view it as such, maybe. ...On the other hand, it is probably more like a child that hasn't understood the rules of the game and is trying get the queen to jump pieces like it's seen knights do, that's just being dumb about the rules.

I mean: suppose you could, with some justification, argue that LLMs playing illegal moves is like the former, like Einstein in a way, which is to say creativity. But, I'm saying, an equally valid view, in fact probably a way more valid more grounded view, might be that it's simply "hallucinating", as in making up stuff, and "making up stuff" not in terms that make for bona fide creativity but simply in terms of talking nonsense, like producing human figures with seven fingers.

At least that's my layman's take on this.



And the same with Go.

Yep, you'd brought up Go before. How are you looking at this, as an example of creativity, or as an example of critical thinking? I suppose it's a bit of both: but only a very small bit! Here's what I mean:

I don't how you play Go, but I suppose it's like chess, in that at any time the number of possible moves is huge but finite. So that, if via its huge computing recources, an AI model picks on a particular chess move that no human has ever happened to have made before: then, well, I suppose that's creativity of a kind, ...but, I don't know, of a very limited kind?

As far as treating that as an example of critical thinking: Well ditto. Like in endgames, or even in midgame when one finds oneself in a position where some particular piece is trapped, then sometimes we do check out every possible move that is possible in that context, like literally try out mentally all the moves (which are few enough to admit of that kind of treatment) and then choose which move to make. And, I don't know, I suppose that is indeed critical thinking of a kind. But again, a ...very limited kind? Given how ...closed, how exactly the opposite of open-ended, is a chess game?

I don't know, really. As you can tell I'm sure, I'm just thinking my way through this! ...Could you spell out your argument/reasoning fully, about what that Go move you'd referenced earlier, or indeed some unique winning chess move that AI might come up with, says about AI creativity and/or AI critical thinking?
 
Is the critique of the video not an original critique? We know that video is not in its training data, nor was there any critique of that video in its training data.

I don't really know, I feel kind of out of my depth talking about this given my lack of technical knowledge of what goes on inside of the innards of AI systems: but, since you ask, my own tentative answer would be: No, this isn't an original critique, not really. Here's why:

In the critique you've presented, AI's doing two things. One, it's summarizing the vid. Now I don't know how well it's done that, maybe it's done it reasonably well, or maybe it's left out crucial bits. Let's assume it's done it as well as most humans would.

Now, as far as the critique part: as far as I can see, it's identifying subjects that it's seen him, the guy in the vid, talk about, and it's presenting a critique of those component parts. That is, while the video itself is a new one, an original one: but the subjects it's critiquing are things it's encountered in its training material.

And that's ...essentially a parody of critical thinking, isn't it. I don't know how well it's done it. Even if it's done well, even so it's no more than a parody executed well, no more than apparent critical thinking.

...Now sure, you might argue that such parody also is part of what critical thinking is made up of. Like, when you, a human, critique something, then you don't keep reinventing the wheel every time, drawing on such critique as you've encountered before is very much part of what you do as a human. To that extent, sure.

But still, when you come across something completely new, that you haven't ever seen before ---- like, I don't know, if this guy in the video had been talking about some completely new idea never ever presented before --- then would AI have been critique that? I really don't see how it might do that, given my (sketchy) idea of how AI works!

In any case, that AI is able to critique component ideas within this video that it likely has indeed encountered in its training stuff, seems to have no bearing on the question we're actually asking here: which is: Will present-day AI be able to coherently and comprehensively critique some idea it comes across for the very first time?


(TLDR: That's neatly done. But that isn't original critique, not really. Is my tentative take on this. With the emphasis on tentative --- again, I know I'm out of my depth speaking about this, and happy to defer to someone that actually knows their AI innards, including to you if you know about how these things actually work beyond just the sketchy ideas that we're all aware of, about statistical selection of words etc.)
 
Last edited:
I think you're too hung up on the idea that Roko's Basilisk was a wild and crazy new thing that required a quantum leap of creativity to come up with, rather than a straightforward thought experiment that someone once posted on social media on a lazy afternoon.

No, emphatically not! Like I've said more than once, including in so many words in this post: This isn’t about Roko per se. Roko per se is irrelevant. What I’m trying to figure out, via your incidental Roko reference that I happened to come across here the other day, is whether AI is already capable of actual critical thinking and of actual creative ideation, that are not directly derivative of the training material it has absorbed.

(I’d been of the view that it isn’t able to do either. Then I saw it suggested here that it is able to do both. When I was starting to think seriously about updating my view on this basis what the better informed people here were saying, is when there came up further posts here arguing the exact opposite, that is to say making a case for where I was at all along myself. ...So, what we have here is mixed views, clearly. And that’s what I’d like to have cleared up. ...I’ll wait a day or two, if we don’t arrive at some kind of consensus on this already, then I’m thinking a new thread focused on this one single focused question may be an idea--- well, two different but related questions, one on creativity and the other on critical thinking.)


And, yeah: that whole question can be TLDR’d using Roko, by asking: In a world without Roko, or anything similar, can AI of today have come up with that basilisk from completely open-ended prompts? And in a world without Roko, or anything similar, can AI of today have been able to come up with a comprehensive critique of the basilisk if prompted with that completely brand new original idea?

(Roko's just a useful shorthand for discussing this, is all. In the paragraph immediately above, just substitute Roko’s Basilisk with Pascal’s Wager if you like. Or think of AI coming up with the ideation as well as critique of warped-space-time gravity in a world where there’s just Newton’s gravity and nothing else, if you'd prefer something weightier. Or whatever else.)
 
Last edited:
But there are configurations in Chess that are explicitly forbidden (like a Pawn being promoted into a Pawn or moves when in check that do not get you out of Check).
Good points. The quickie data structure I came up with wouldn't do special moves properly.

And why would you need random sampling when you have a library of millions of games, and thus could put probabilities on every configuration based on their frequency of occurrence?
You wouldn't. I just thought it was funny how well it worked compared to more studious solutions.

I don't see what kind of dataset GPT-4 and GPT-5 were trained on, but they both consistently make illegal moves.
The problem is that LLM chatbots aren't designed to play chess. Real chess engines are purpose-built.
That's a clearer statement off the point I was trying to make above. "Artificial intelligence" is an umbrella term. Under the hood, a variety of dissimilar techniques are used depending on the problem. We had such things as Deep Blue that used entirely different technology to play chess than AI uses today. No tensors or matrix solvers. The term "expert system" applied there. That's where humans encode their knowledge explicitly in rules.

LLMs are text-based. They work best by encoding human language and then generating other examples of human language. They can do math, for example. But they don't do math by knowing how to compute. They do math by reproducing other people discussing how to do math.
 
What I’m trying to figure out, ... is whether AI is already capable of actual critical thinking and of actual creative ideation, that are not directly derivative of the training material it has absorbed.
If you have a strong stomach, go through this thread.

Ostensibly it's about really bad photographic analysis. But the OP in that thread relied heavily on AI to evaluate arguments and render opinions on his critics. The first thing you see there is a fair amount of sycophancy. This is a problem with commercial AI. It wants to please the user so that the user prefers that AI over a competitor. Part of deploying a commercial AI is giving it a "personality." And some of it is just straight-up confirmation bias. It's not too hard to write prompts for LLMs that home in on a desired concept.

But it's valid to ask how an LLM can evaluate an argument on a specialized subject that just came up and can't possibly be part of its training data. The short answer is, "badly." But the real question is what's going on in the model.

The model contains not only propositional knowledge on various subjects, but knowledge on the art of critical thinking itself irrespective of the propositions. So in one mode you can ask an AI to evaluate a statement like, "Was the U.S. response to 9/11 appropriate?" and you'll get a synthesis of whatever is in the training data: opinions of journalists, world leaders; the results of polls, etc. But if you give it an excerpt from an argument that you're having with someone online and ask it to evaluate the strength of the arguments on both sides, an LLM has to reach into the part of its training data that comes from such things as textbooks on rhetoric. The LLM can embed such propositions as, "These traits make an argument good (or bad)." And the transformation layer can align those embeddings with the encoding of the argument in the prompt. In this way you can expect an LLM to operate on a higher or more meta level. Its answers are not directly derived from training data.

But you can see how bad the critical analysis is. It says stuff like, "Your critics aren't citing their sources," when the critic is in fact a subject-matter expert speaking from his own knowledge and experience. Yes, citing one's sources is a characteristic of a good argument in general. But it's not the right analysis here. Nor would it be when the proposition is a self-evident mathematical expression, as also occurred. Now of course that particular thread suffers from considerable bad faith from the OP, who probably wasn't giving his AI all the information from the thread that it would need. But the AI is still getting it wrong.
 
And most LLMs have now been forked so you'll get a "coding LLM" - its training is different to the "generic" LLM
 
And most LLMs have now been forked so you'll get a "coding LLM" - its training is different to the "generic" LLM
Yes, as it should be IMHO. If the contemplated use cases are specialized, then so should the training data. And so should be the tokenization and embedding algorithms. A coding dataset can scrape the open-source public repositories such as GitHub without having to digest the collected works of Charles Dickens. And the embedding method can be optimized to represent code concepts at higher granularity.

Having some understanding of how these software products work is helpful in using them wisely. I see lots of people using AI chatbots in a way that suggests they treat them as infallible, impartial oracles. That's dangerous. They're not just using them to write routine emails and reports or generate pictures of Stephen Miller in a bikini. They are expressly relying on AI to substitute for human judgment. If any of the philosophical fears in this thread are ever realized, it will probably be because we voluntarily ceded that power.
 
If you have a strong stomach, go through this thread.

Ostensibly it's about really bad photographic analysis. But the OP in that thread relied heavily on AI to evaluate arguments and render opinions on his critics. The first thing you see there is a fair amount of sycophancy. This is a problem with commercial AI. It wants to please the user so that the user prefers that AI over a competitor. Part of deploying a commercial AI is giving it a "personality." And some of it is just straight-up confirmation bias. It's not too hard to write prompts for LLMs that home in on a desired concept.

Haha, yes, that thread. I'd browsed through bits of pieces of it, back when it was live.

Was he using AI to compose his posts? I didn't get that, I mean I only checked out bits of it, is all. But I think I know what you're talking about here, about the citations and authority thing(y)


But it's valid to ask how an LLM can evaluate an argument on a specialized subject that just came up and can't possibly be part of its training data. The short answer is, "badly."

Yes, badly. But, importantly, it's just a parody, is all. Is my take on it. ...Do you agree? (I want to make sure I'm not making an error myself with how I'm viewing this.) ...It isn't just "badly" critiqued, the fact is that the apparent critique is no more than parody. ...See further:


But the real question is what's going on in the model.

The model contains not only propositional knowledge on various subjects, but knowledge on the art of critical thinking itself irrespective of the propositions. So in one mode you can ask an AI to evaluate a statement like, "Was the U.S. response to 9/11 appropriate?" and you'll get a synthesis of whatever is in the training data: opinions of journalists, world leaders; the results of polls, etc. But if you give it an excerpt from an argument that you're having with someone online and ask it to evaluate the strength of the arguments on both sides, an LLM has to reach into the part of its training data that comes from such things as textbooks on rhetoric. The LLM can embed such propositions as, "These traits make an argument good (or bad)." And the transformation layer can align those embeddings with the encoding of the argument in the prompt. In this way you can expect an LLM to operate on a higher or more meta level. Its answers are not directly derived from training data.

But you can see how bad the critical analysis is. It says stuff like, "Your critics aren't citing their sources," when the critic is in fact a subject-matter expert speaking from his own knowledge and experience. Yes, citing one's sources is a characteristic of a good argument in general. But it's not the right analysis here. Nor would it be when the proposition is a self-evident mathematical expression, as also occurred. Now of course that particular thread suffers from considerable bad faith from the OP, who probably wasn't giving his AI all the information from the thread that it would need. But the AI is still getting it wrong.

So like, sure, it's following the logical-engagement, critical-thinking-engagement thing it's seen in its training: but it's simply ...well, going through the motions. That's my take. Like I said above, that's actually an important qualification, critical even, to understanding what's going on: and, like I said, I'm putting this down here for feedback to make sure I'm not wrong about this myself.

See, if in general someone trips up on that one thing, trips up by insisting on citation-from-authority even when actually referencing an authority, well then, that's kind of subtle. That's a mess-up, sure: but it's a small enough bug. It's a bug that needs fixing, in one's critical thinking repertoire, sure: but that one nuanced bug isn't enough to paint said critical thinker that makes this error as ...as someone incapable of critical thinking.

But in this case, the AI's following its training material in the meta sense like you say, to critique stuff not in its training material: but all it is doing is going through the motions. Like a stopped clock it might sometimes tick all boxes, or it might foul up, but either way, this emphatically isn't critical thinking, and should not, must not, be mistaken for such. Is my take on it, ...unless, well, it turns out I'm mistaken about this.

----------

But in any case:

That's kind of what's happening in everything, isn't it, not just critical thinking. AI's similarly just "going through the motions" when ...when producing visual images as well, isn't it? And yet, it does such a great job of it, already. Sure, it used to get the fingers all wrong, but it's stopped doing that now. ...I'm wondering, in going though the motions of critical thinking, can AI similarly end up somehow improving, getting over this glitch, maybe even this specific glitch of insisting on citations without undertanding why citations are needed, and therefore not understanding why they may not be needed in some cases?

I'm asking, first, if you agree that AI's just going though the motions of critical thinking, just parodying, not the real thing. ...But I'm also asking, in addition, if "just going through the motions" is sufficient to get it to nevertheless improve on the output to where it is difficult to tell the difference, like in many cases it already is.
 
Last edited:

ISF - Join now!

Every member here is approved by hand. No bots, no spam, just people who care about evidence and honest debate.

Membership is free!

Create your free account

Back
Top Bottom