×
all 27 comments

[–]RealPropRandy 5 points6 points  (0 children)

[–]TLMTGT 3 points4 points  (1 child)

Contrary to the other poster, these are real results as far as I can tell. Note that OpenAI has real mathematicians collaborating with and/or working for them.

I only recognize a few of the problems, but these are/were substantial problems. As for whether the mathematics used to solve/disprove them is useful will take a while to determine. At a cursory glance, the proofs seem to follow the thread of construction/counterexample proofs that LLMs have been successful at recently rather than develop new conceptual ideas, but I'm not 100% sure since these are outside my subfield.

[–]Smooth-Ad8030[S] 0 points1 point  (0 children)

Interesting, yeah these seemed like real results, I figured they were given the sheer number of mathematicians working with them.

What is the functional difference between construction/counterexamples vs new conceptual ideas?

[–]Significant-Green130 1 point2 points  (1 child)

I honestly don’t think anyone outside of math should care much about these things, for a variety of reasons. I don’t think you should listen much to loud opinions on Reddit, including mine. But if you want to know:

1) The results I had heard of here are pretty impressive — I guess some people are most impressed by the sofic group result, but I personally don’t know much about it so I cannot speak to that. So it’s very strange seeing rabid cheerleading or dismissal by many people on Reddit subs that almost certainly have no clue what a sofic group is…

I’d be pretty impressed if a human came up with the ideas on the results I understand better, as the ones I skimmed seem to reinterpret and then improve classical things in ways I find counterintuitive. They seem to use tools that aren’t super unrelated, but that aren’t obviously useful for the problem either. I’d think a human realizing to use these tools for these applications must have had deep intuition to try this route, but it’s not clear to me that logic applies to LLMs that far exceed us at memorization and brute force. But on the other hand, I feel some sort of weird pride that much of the machinery it uses is in the literature in some form; putting it together and realizing it was useful for these problems is not easy by any means, but there’s something nice about the idea that we built up beautiful tools that seem to have more punch than we realized at the time. I suspect it will still be humans that will digest these new arguments and realize where they may lead to new results for other problems. This has already happened for the unit distance conjecture counterexample.

2) Nobody knows what exactly they do behind the scenes, but they have hired many world-class mathematicians/computer scientists. This is very public information, mostly because part of it was about helping their reputation. They also more or less hired any mathematician that publicly extolled their tools on Twitter last year. My understanding is only some of them are directly involved in these math efforts while others are focused on other things there. Many of them are, though, directly involved in helping generate synthetic data to improve their models at math and code, but I don’t know what this entails exactly beyond ripping ArXiv papers and likely generating synthetic reasoning traces in some way. My vague understanding is also that they run their models all the time on these kinds of problems, and if they seem to produce something promising, some of these mathematicians will take a look to see if it is correct or interesting — that could potentially be used to generate more training data to push it towards more promising directions, but who knows. For this batch of problems, my sense is those mathematicians probably helped rewrite the results into a more readable format and to give more context about the argument as LLMs are still pretty bad at that atm. I would guess the models sketched out the main ideas in some form, and then after checking them, the humans would direct Codex to write better related work and proof ideas and so on like a more human paper. But again, just a guess.

3) The true cost is also unknown but the numbers they give are almost certainly misleading for a variety of reasons. It clearly doesn’t account for the massive compute, synthetic data generation, human input, and so on, that goes into training these models precisely towards improving in these directions. Even for these problems, I suspect whatever number they gave is for the successful runs. But that’s obviously not the same thing as the full cost of them trying their models ad nauseum on all problems and then seeing what worked, which likely took considerable human effort as well. They don’t seem to want to provide any clarity about any aspect of their process, but I do find it hard to believe that they have all these brilliant theoreticians and they work as run-of-the-mill SWEs when they probably hadn’t touched code in years.

4) It’s been clear for over a year now that they have viewed math as a source of relatively cheap PR. Whether or not it’s deserved is up to you depending on how you measure their costs vs. achievements, but they certainly care about the effects on their valuation far more than they care about “advancing science” or whatever. They are shockingly nontransparent about anything, despite only lifting off the ground by promising researchers they would do charitable and open science.

[–]Separate-Ear-7258 1 point2 points  (1 child)

I think if you listen to Cal Newport's take on the Erdos problem it will give you insight. I'd love to hear if my perspective is off about these.

Basically, his idea boils down to these points:

1) We already knew these are one of the two sweet spots for LLMs - coding and mathematics. These make sense both are language-based fields that are easily verifiable.

2) It took Open AI paying mathematicians an insane amount of money, throwing an extreme amount of compute, and

3) It did it by proving via counter example, which put plainly, means to finds/provides examples an instance where something fails.

4) Cal thought that smaller, more modular systems or LLMs could really help Mathematicians be more effective. He said he could be probably 2x more effective with these

5) The fact that they are heavily focusing on the impressive math results is a distraction from the fact that it is not a very useful commercial application. Clammy Sam and Wario would much rather prefer useful commercial applications that displace workers than useful math results. They would light all the math people on fire if it meant they could be profitable and justify their bonkers evaluation. to quote Cal:

  • From a business perspective, I actually think this announcement isn’t necessarily good news for OpenAI. There are few markets smaller and less lucrative than professional academic mathematics. The fact that this is the area where OpenAI is dedicating some of their top technical talent (like Noam Brown) underscores the degree to which, like the drunk searching for their keys under the streetlight, their most impressive results are limited to the smaller number of areas that are well-suited to LLMs (i.e., math + computer coding). If this model was brilliant in some more general way, obviously the better examples would be solving problems or automating processes that directly and obviously generate massive revenue or savings for the specific types of companies they hope to make their customers

All in all, AI and LLMs can help mathematicians. Doesn't change anything about the disastrous economics or unit economics.

[–]Smooth-Ad8030[S] 0 points1 point  (0 children)

Great reply, and cal Newport coming in with the best takes as usual, thank you!

[–]Easy_Tie_9380 4 points5 points  (13 children)

It’s fake. The only thing these people can do is lie.

[–]Cold-Environment-634 3 points4 points  (6 children)

I’d really like to know, how can you be so sure? In what way are the findings fake? You mean it’s much more guided by expert humans than they are saying?

[–]Smooth-Ad8030[S] 2 points3 points  (0 children)

That would be my guess. It costing $2,000 is almost assuredly fake for example.

[–]Easy_Tie_9380 0 points1 point  (3 children)

I’m saying that stochastic parrots can’t do math

[–]Quarksperre 2 points3 points  (0 children)

I am on this sub and generally pretty skeptical about the LLM hype. But I don't think these are fake. This should be acknowledged and not blindsighted. There are enough quasi religions around this whole issue already. 

The results are nice. Honestly however I expected something like this to happen way earlier with lean and formalized math. Its a whole set of technologies coming together and in the end math is a game with very strict rules. Its not fuzzy reality. Thats where neural nets generally fail always at some point 

[–]Cold-Environment-634 2 points3 points  (0 children)

Any evidence of this? I’m in no way a booster and that’s why I’m here in the first place, but completely brushing this stuff off as BS seems like a stretch

[–]falken_1983 0 points1 point  (0 children)

What does that mean though?

[–]Smooth-Ad8030[S] 1 point2 points  (4 children)

Oh I’m not saying it’s real, I’m mainly curious about the results without the circle jerk in the math sub. There’s so much vagueness in the announcement it obviously wasn’t a cheap one shot prompt.

Edit: every math result in there, no matter how fake or real, is super important for the field and groundbreaking in the math subreddit which is obviously bullshit.

[–]koveras_backwards 10 points11 points  (2 children)

I've linked this before here, but I guess I can again.

https://bsky.app/profile/gro-tsen.bsky.social/post/3mr3gj6ry622d

The point made is that the purpose of conjectures is not merely to be solved, but to inspire people to develop and explore new mathematical ideas that lead to solving the conjectures. The new mathematics/mathematicians are the important part, not the yes/no solution to conjectures.

It seems that even many working mathematicians do not really understand this point, and there is of course little understanding among the general public.

[–]Smooth-Ad8030[S] 0 points1 point  (0 children)

That’s a great thread, thank you!

[–]SpookyTanuki1 0 points1 point  (0 children)

Great read but did he really write “cum challenge”. Probably should have rethought that phrasing.

[–]ThanklessWaterHeater 0 points1 point  (4 children)

First of all, I personally dislike LLMs, and don’t use them for anything. I don’t want anyone to think that I’m a booster here. But I am an investor, and I think it’s important to keep up with developments in the field.

My brother-in-law is a tenured theoretical mathematician at one of the University of California campuses, and I was talking with him about this earlier this week.

He said that within the world of mathematicians, some of the recent work has been jaw-dropping. Theories that have gone generations without proof have been solved in a matter of minutes by LLMs. He doesn’t use LLMs himself, but he has read some of the papers and says the proofs appear to hold up.

I believe him. But I also think that this is one of very few areas where LLMs work well: 1) a field with very specific, long-established rules that can only be applied in very specific ways in order to create a valid proof. 2) any new theorem (whether generated by a human or by an LLM) is going to be carefully verified by other mathematicians before anyone even mentions it to the outside world. The proofs you read about in the linked article went through the standard peer-review process. My guess is there have been a number of bad proofs generated that nobody heard about because the person who generated it was capable of checking the logic and found it was incorrect.

In fact, as I understand it, LLM’s skill with mathematical proofs has been known long enough that the labs are looking for ways to use it more generally. They are trying to make LLM’s treat every day tasks with the same logic, making sure that every logical step in an argument can be proven the way a mathematical theory can be proven. That said, I read about that work a couple years ago and God knows based on the AI slop I see online every day it doesn’t seem to be working yet.

Anyway, like I said at the top, I’m not a booster here. I do believe my brother-in-law when he tells me this. But I also think this is a very niche skill and that if LLMs are ever going to be 100% reliable in more general uses, they need to be able to do more than generate a mathematical proof.

[–]The-Menhir 0 points1 point  (0 children)

What good does AI proving/disproving theorems bring when proofs are often inscrutable & fail to bring about novel insights, other than pointless dick measuring contests between AI companies who couldn't care less about the field or the intrinsic humanity of it about who has the better model? I don't believe mathematics is a problem to be solved by AI, but a human endeavour.

[–]ThurInTheTrees -2 points-1 points  (2 children)

Your brother-in-law must be dumb as fuck because any actual mathematicians know that this shit is not important at all to the field of math.