What OpenAI is doing with the maths problems is marketing,
I agree with that.
outside of that marketing spend no one is going to be able afford to use these models
I think that remains to be seen. Substantial progress is being made very quickly.
in the way they are doing to solve maths problems which is a brute force approach.
In my opinion, it's a bit misleading to describe the AI models' approach to solving math problems as brute force.
Consider, for example, the
FrontierMath benchmark suite, first described in 2024. (This is one of the AI benchmarks that has been criticized for its inclusion of unsolved problems.) According to that paper, the benchmark's reviewers took care to exclude problems with "answers that aren't easily verifiable, problems where guessing is easier than proving, and cases where simple brute-force methods circumvent the intended difficulty." The developers of the benchmark "conducted interviews with four prominent mathematicians to gather expert perspectives on FrontierMath’s difficulty, significance, and prospects"; three of those four were Fields Medalists. "All four...characterized the research problems in the Frontier Math benchmark as exceptionally challenging, noting that the most difficult questions require deep domain expertise and significant time investment."
It should be noted that "research problems" constitute a minority of the benchmarks, but an appendix gives specific examples of problems rated "high difficulty", "high-medium difficulty", "medium difficulty", "medium-low difficulty", and "low difficulty".
The sample solution for the "low difficulty" problem (A.5) starts with a brute-force calculation to guess a conjecture, followed by use of
Weil conjectures to construct an equation whose unique solution gives the result; proving the uniqueness uses knowledge of Chebyshev polynomials.
The "medium difficulty" problem (A.3) asks the AI to find the smallest prime p ≡ 4 mod 7 for which the function that maps an arbitrary integer n to the sequence of integers satisfying a certain recurrence parameterized by n can be extended to a continuous function. Its solution starts by applying the
Skolem-Mahler-Lech theorem. It then proceeds by considering the
uniformizer of a certain ring and noting that the constant term of the characteristic polynomial, 370639957, is prime (as can easily be determined using brute force). That means something about all roots of the valuation given by the uniformizer are related to primes less than 370639957. (This part of the solution is beyond me, as are several other parts of the solution.) That means "the prime we are looking for" must divide 370639958, whose prime factors are 37, 673, 811, and 9811. Reducing the characteristic polynomial mod p = 9811 yields a 4th degree polynomial with small coefficients. The solution then uses some facts of real analysis to obtain an extension to a continuous function, and explains why such an extension is not possible for other primes. Then, since "the projection map...induces an isomorphism between the group of (p−1)th roots of unities and F
px," "we can find" a certain root of unity that gets us to within a quarter page of equations that solve the problem.
It is my opinion that most humanoids you might interview on a street corner would not be able to solve that problem.
In 2024, when
that paper was written, several then-current "state-of-the-art AI models" were able to solve less than 2% of the benchmark problems. GPT-4 was able to solve about 5%.
GPT-5.4 Pro solves 50% of the undergrad-to-postdoc tier of problems, and solves 38% on research-grade problems. Within the past year, 15 of the open problems have been solved by humans or AI, with AI getting at least some credit for 11 of those solutions.