• Security incident: ISF was recently accessed by intruders. Please change your password, and change it anywhere else you used it. Read more

Merged Artificial Intelligence

any math nerds take a look at that chatgpt ai math problem document dump?
Are you talking about Erdős Problem #1196? (ETA: Proof here.)

Or are you talking about using ChatGPT for homework problems?

Or something altogether different? (ETA: Such as the Navier-Stokes Millennium Prize Problem, which has been proof-checked by the Lean proof assistant; the AI solution to Erdős Problem #1196 has also been checked by Lean. Lean, by the way, uses a hygienic macro system, so I can tell myself I made my own extremely tiny contribution to the AI solutions of those two problems.)


Yes there are use cases for LLMs, but the cases where they approach even moderate human capabilities are so few and far between that they'll only ever be hobbyist curios.
In my opinion, that proof of Erdős Problem #1196 approaches moderate human capabilities. To put it even more strongly: I daresay the number of humans who could have come up with that proof on their own is exceeded by the number of humans who could not have done so.

(The existence of a bound was proved in 1935, and a particular bound was conjectured in 1968. That conjecture was not proved by a human until 2022. Prompted by Liam Price, GPT-5.4 Pro found another proof in April 2026.)
 
Last edited:
Are you talking about Erdős Problem #1196? (ETA: Proof here.)

Or are you talking about using ChatGPT for homework problems?

Or something altogether different? (ETA: Such as the Navier-Stokes Millennium Prize Problem, which has been proof-checked by the Lean proof assistant; the AI solution to Erdős Problem #1196 has also been checked by Lean. Lean, by the way, uses a hygienic macro system, so I can tell myself I made my own extremely tiny contribution to the AI solutions of those two problems.)



In my opinion, that proof of Erdős Problem #1196 approaches moderate human capabilities. To put it even more strongly: I daresay the number of humans who could have come up with that proof on their own is exceeded by the number of humans who could not have done so.

(The existence of a bound was proved in 1935, and a particular bound was conjectured in 1968. That conjecture was not proved by a human until 2022. Prompted by Liam Price, GPT-5.4 Pro found another proof in April 2026.)


i was asking about this story.
 
You are all the ENEMY!

Donald J. Trump
@realDonald Trump

The White House considers anyone that uses the term, "Artificial Intelligence," as opposed to the highly accepted new and more accurate term, "Super Intelligence," THE ENEMY!

President DONALD J. TRUMP
 

i was asking about this story.
There are two things going on here.

The proofs themselves appear to have been checked using the Lean programming language. Before editing my previous post, I confirmed that both of the specific proofs I mentioned have been checked using Lean, and that the Lean source code for those proofs is freely available at github. There does not appear to be a great deal of doubt concerning the correctness of these AI-generated and Lean-checked proofs.

The controversial aspects of this concern (1) attribution of credit, (2) use of mathematicians' work (which is often copyrighted), and (3) the use of open (i.e. unsolved) mathematical problems to test AI models.

Attribution of credit and the use of others' work are issues that have been discussed extensively within this thread, and I don't see why those issues are any more or less important for mathematics than for other applications of AI models.

Using open problems to test AI models is worrisome for several reasons. One is that many mathematicians have spent years making progress toward solving some of those problems, and they understandably resent the possibility that some AI model, building on their hard work, may swoop in at any moment to take the credit for a solution.
 
Last edited:
The proofs themselves appear to have been checked using the Lean programming language. Before editing my previous post, I confirmed that both of the specific proofs I mentioned have been checked using Lean, and that the Lean source code for those proofs is freely available at github. There does not appear to be a great of doubt concerning the correctness of these AI-generated and Lean-checked proofs.

Automatic theorem proving is a classical area of artificial intelligence. In 1975, I graded for a course in automatic theorem proving taught by Woody Bledsoe. Several of Bledsoe's graduate students advanced the state of the art. One of those students, Robert S Boyer, collaborated with J Strother Moore to develop the Boyer-Moore theorem prover, which Moore and his team turned into the ACL2 prover that (in 1995) was used to prove the correctness of floating point division in the AMD K5 microprocessor.

Lean 4 is both a functional programming language and a proof assistant. It was developed by Leonardo de Moura, building on the work of previous researchers and systems, notably Coq (now Rocq). Last year, de Moura and three other developers of Lean received ACM SIGPLAN's Programming Languages Software Award.

Without getting into the P=NP question, let's just stipulate the intuitively obvious fact that proof checking is easier than finding a correct proof from scratch. In 1929, however, Kurt Gödel proved his completeness theorem, which asserts the existence of an algorithm capable of proving any valid statement of first order logic. One way to prove that completeness theorem is to describe a specific algorithm that does so. Here is one such algorithm:
  1. Start a process that enumerates all possible sequences of Unicode characters. (That process, left to itself, will never terminate.)
  2. Interrupt that process after each enumerated sequence of characters, and check to see whether that sequence of characters is a proof of the statement you're trying to prove.
  3. If it is, you've found a proof. Otherwise you keep looking.
As a corollary of Gödel's completeness theorem, it is possible to write a computer program that will find a proof of any correct valid statement of any recursively axiomatizable first order theory.

Which is not to say the computer program will find a proof quickly, even when a proof exists. If the thing you're trying to prove isn't valid, then there is no proof, and the computer program might run forever.

So proof checking is computationally feasible, but complete proof procedures may not be.

There is a middle ground, known as a proof assistant. A proof assistant assists mathematicians by checking parts of a proof, and by proving things that a human would find tedious to prove, but the proof assistant probably can't prove significant results without guidance from a mathematician (or AI model) in the form of a suggested proof outline.

Lean 4 is a functional programming language, a proof checker, and a proof assistant. You can write computer programs in Lean 4, just as you can write programs in languages such as Haskell or the functional subset of Scheme. As a proof checker, you can ask it to check your proofs. As a proof assistant, it can help you to come up with those proofs.

Lean 4 can check any proof expressed using first order logic, and can venture beyond first order logic into dependent type theory and a subset of second order logic. Second order logic is inherently problematic, however; for one thing, there is no analogue of Gödel's completeness theorem for second order logic; even worse, there is no fully general algorithm for checking the correctness of proofs expressed using second order logic.

Anyway, I hope my remarks above will help to explain why mathematicians have a fair degree of confidence in proofs generated by AI and fully checked by Lean 4.
 
Last edited:
my additional question with these math equation solutions is if a company had decided to hire a bunch of math guys and given them an unlimited budget, could the problem have been solved?
 
my additional question with these math equation solutions is if a company had decided to hire a bunch of math guys and given them an unlimited budget, could the problem have been solved?
Yes. There is no magic in mathematics, and there is no magic in AI. Any math that can be done by an AI model can be done by human mathematicians.

I am reminded, however, of a song that goes something like this:
​
Anything you can do,​
I can do slower.​
I can do anything​
Slower than you.​

There are many open problems in mathematics that remain unsolved despite decades or centuries of effort by human mathematicians. AI and other tools (such as Lean 4) have been making it easier for human mathematicians to solve some of these open problems. That is a good thing.

Whether it is a good thing for AI agents to solve open problems in mathematics that humans have not yet gotten around to solving, and might not have gotten around to solving for decades or centuries, if ever, appears to be a source of controversy.
 
There are many open problems in mathematics that remain unsolved despite decades or centuries of effort by human mathematicians. AI and other tools (such as Lean 4) have been making it easier for human mathematicians to solve some of these open problems. That is a good thing.
Tools are generally a good thing, which is why we have them. But what if you can't make the tool or buy it yourself? What if it's owned by a trillion dollar company with an enormous computer that has its own gigawatt power station? What if they scraped the entire internet and millions of books to get the required data? What if everybody abandons other tools because this one is supposed to do everything? What if people think it's 'super-intelligent' and rely on it to tell them what to do? What if they have a 'personal' relationship with it that affects their interactions with other people? What if you are forced to use it too because it's everywhere and in everything?

People were rightly concerned about Microsoft having a monopoly on personal computer software, but that's nothing compared to how dependent AI companies want us to be on their 'tool'.

Whether it is a good thing for AI agents to solve open problems in mathematics that humans have not yet gotten around to solving, and might not have gotten around to solving for decades or centuries, if ever, appears to be a source of controversy.
I think we can agree that solving problems in mathematics is a good thing. No reason we shouldn't use AI tools to help with that, so long as we only use them as tools and not treat them as oracles.
 
my additional question with these math equation solutions is if a company had decided to hire a bunch of math guys and given them an unlimited budget, could the problem have been solved?
They could, but their goal is to get us hooked on their product by convincing us that AI is so much better than humans. They have trillions riding on this, so spending a billion or so on 'advertising' is justified. In another industry this would be predatory behavior.
 

ISF - Join now!

Every member here is approved by hand. No bots, no spam, just people who care about evidence and honest debate.

Membership is free!

Create your free account

Back
Top Bottom