×
all 66 comments

[–]Separate-Ear-7258 29 points30 points  (12 children)

I think if you listen to Cal Newport's take on the Erdos problem it will give you insight. I'd love to hear if my perspective is off about these.

Basically, his idea boils down to these points:

1) We already knew these are one of the two sweet spots for LLMs - coding and mathematics. These make sense both are language-based fields that are easily verifiable.

2) It took Open AI paying mathematicians an insane amount of money, throwing an extreme amount of compute, and

3) It did it by proving via counter example, which put plainly, means to finds/provides examples an instance where something fails.

4) Cal thought that smaller, more modular systems or LLMs could really help Mathematicians be more effective. He said he could be probably 2x more effective with these

5) The fact that they are heavily focusing on the impressive math results is a distraction from the fact that it is not a very useful commercial application. Clammy Sam and Wario would much rather prefer useful commercial applications that displace workers than useful math results. They would light all the math people on fire if it meant they could be profitable and justify their bonkers evaluation. to quote Cal:

  • From a business perspective, I actually think this announcement isn’t necessarily good news for OpenAI. There are few markets smaller and less lucrative than professional academic mathematics. The fact that this is the area where OpenAI is dedicating some of their top technical talent (like Noam Brown) underscores the degree to which, like the drunk searching for their keys under the streetlight, their most impressive results are limited to the smaller number of areas that are well-suited to LLMs (i.e., math + computer coding). If this model was brilliant in some more general way, obviously the better examples would be solving problems or automating processes that directly and obviously generate massive revenue or savings for the specific types of companies they hope to make their customers

All in all, AI and LLMs can help mathematicians. Doesn't change anything about the disastrous economics or unit economics.

[–]Significant-Green130 5 points6 points  (2 children)

I broadly agree with most of these points, but not quite (3) and (4). Not all of these results are counterexamples, and I would say that proving that they are in fact counterexamples is not at all easy for some of these examples (the Jacobian conjecture being a big exception). For instance, my limited understanding of the sofic group thing here is people were pretty convinced non-sofic groups exist with a clear candidate, but nobody knew quite how to prove it is non-sofic. This model found a different example, but I don’t know anything beyond that. On (4), it seems to me a lot of the gains are from packing in more and more training on synthetic reasoning traces, as well as more inference-time scaling, so it’s actually not obvious to me that small models could effectively do it. I do think (5) is the most important point though and that one seems to be the right interpretation. Theoretical math has arguably more cultural significance than practical importance, and self-contained sandboxes where they can throw infinite compute at generating training data and do RL with a verifier scales far better than other applications that need to actually touch the world. It probably scales significantly better than coding, since even reading massive, existing codebases takes up precious inference for these models.

[–]Separate-Ear-7258 3 points4 points  (1 child)

fair enough and good points - I have no clue on 3) and 4) and should not be consulted as an expert. I think five is the most imoportant!

[–]hibikir_40k 0 points1 point  (0 children)

Nah, the real issue is that it's a red queen race, so the economics for the companies moving frontier models forward are a real gamble.

There's use cases where people will be happy to pay inference costs when not discounted. But even in the happiest AI story, the fact that there will be massive overinvestment that will sink many companies is guaranteed. The value of having the third best frontier model is not just a matter of arguing about commercial utility of AI: It's just absolute zero regardless. And when that is true, it means companies will overspend to exhaustion, even at a point where they are still spending when the total cost for the eventual winner goes way past total value.

[–]Smooth-Ad8030[S] 3 points4 points  (1 child)

Great reply, and cal Newport coming in with the best takes as usual, thank you!

[–]Separate-Ear-7258 2 points3 points  (0 children)

Thank you my friend! Agreed Cal is a voice of reason amongst the songs being sung by the crazies

[–]larrytheevilbunnie 0 points1 point  (4 children)

Yeah people don’t realize how useless economically a lot of theoretical math is. Like if you take the space of problems mathematicians work on, the probability any of the stuff they come up with actually makes money is on the scale of playing the lottery. We should be doing more math research because learning is cool actually, but let’s not pretend this is gonna magically unlock explosive economic growth. And advertising this as a new model capability is sort of an acknowledgment of failure.

However, how the fuck did Dr Newport not know about the bitter lesson? Like it’s been proven time and time again that more scale leads to better performance than specialized smaller models, to the point where it’s gotten honestly kind of ridiculous.

Also, let’s not forget that the models went from being barely able to do high school math to finding novel shit in 2 years…

[–]65721[🍰] 1 point2 points  (3 children)

However, how the fuck did Dr Newport not know about the bitter lesson? Like it’s been proven time and time again that more scale leads to better performance than specialized smaller models, to the point where it’s gotten honestly kind of ridiculous.

The "bitter lesson" showed that there was quite a bit of potential left in the old approaches, if you gave them enough data and computing power. Researchers, in their intellectual laziness, somehow took that mean "scale indefinitely because this party will go on forever." Now we've all but tapped them dry.

Also, let’s not forget that the models went from being barely able to do high school math to finding novel shit in 2 years…

They still can't do elementary school math, much less high school. What they can do is generate thousands upon thousands of words of slop that OpenAI can pay mathematicians hansomely to rummage through.

[–]larrytheevilbunnie -3 points-2 points  (2 children)

Wait wut, the bitter lesson is not “scaling will work indefinitely.” It’s that general methods that can exploit more compute tend to outperform specialized approaches, which is not the specialized models Dr. Newport is arguing for...

And have you been living under a rock? They've been able to do AIME and IMO for a year, like yeah hallucinations are there sometimes getting you the car wash shit, but you can't hallucinate your way into 90%+ on those tests.

And have any non-paid mathematicians disproved the validations done by the paid ones?

[–]65721[🍰] 3 points4 points  (1 child)

You can read Sutton's Bitter Lesson essay yourself.

Seeking an improvement that makes a difference in the shorter term, researchers seek to leverage their human knowledge of the domain, but the only thing that matters in the long run is the leveraging of computation. These two need not run counter to each other, but in practice they tend to. Time spent on one is time not spent on the other. [...]

One thing that should be learned from the bitter lesson is the great power of general purpose methods, of methods that continue to scale with increased computation even as the available computation becomes very great. The two methods that seem to scale arbitrarily in this way are search and learning.

This "bitter lesson"—and its predecessor in deep learning, and its successor in the "scaling laws"—have been the biggest setbacks for the field in its quest for AGI. Researchers were frustrated by the field's past failures, and the newfound glut of data and computing power validated their temptation. AI was too hard, but the researchers didn't need to suffer through thinking too hard. Existing approaches were enough, and they could scale them forever. It bred intellectual laziness.

Hallucinations aren't there "sometimes." So-called "hallucinations" are LLMs' default and only output. LLMs have no concept of truth, or of anything at all. We choose to call their outputs "capabilities" if they happen to be correct, and "hallucinations" if they happen to be wrong. This is not fixable, by very nature of their architecture.

And you can definitely hallucinate your way into scoring well on a test if you cram enough test-specific data into these things—"teaching to the test," if you will, applicable to nothing but that particular test. Just look at these LLMs' performance on ARC-AGI. OpenAI, Anthropic, Google, etc. will get their models' scores to 90%+ over the years after the test is released. Then ARC will release a new test that asks the same simple fundamentals but in a different format, and every model's score suddenly plummets to like 0.5%. We are now on ARC-AGI-3, and nothing has changed.

[–]RealPropRandy 17 points18 points  (1 child)

[–]FastHotEmu 2 points3 points  (0 children)

Shiiiiiiiiiiiiiiiiinyyyy

[–]Significant-Green130 10 points11 points  (1 child)

I honestly don’t think anyone outside of math should care much about these things, for a variety of reasons. I don’t think you should listen much to loud opinions on Reddit, including mine. But if you want to know:

1) The results I had heard of here are pretty impressive — I guess some people are most impressed by the sofic group result, but I personally don’t know much about it so I cannot speak to that. So it’s very strange seeing rabid cheerleading or dismissal by many people on Reddit subs that almost certainly have no clue what a sofic group is…

I’d be pretty impressed if a human came up with the ideas on the results I understand better, as the ones I skimmed seem to reinterpret and then improve classical things in ways I find counterintuitive. They seem to use tools that aren’t super unrelated, but that aren’t obviously useful for the problem either. I’d think a human realizing to use these tools for these applications must have had deep intuition to try this route, but it’s not clear to me that logic applies to LLMs that far exceed us at memorization and brute force. But on the other hand, I feel some sort of weird pride that much of the machinery it uses is in the literature in some form; putting it together and realizing it was useful for these problems is not easy by any means, but there’s something nice about the idea that we built up beautiful tools that seem to have more punch than we realized at the time. I suspect it will still be humans that will digest these new arguments and realize where they may lead to new results for other problems. This has already happened for the unit distance conjecture counterexample.

2) Nobody knows what exactly they do behind the scenes, but they have hired many world-class mathematicians/computer scientists. This is very public information, mostly because part of it was about helping their reputation. They also more or less hired any mathematician that publicly extolled their tools on Twitter last year. My understanding is only some of them are directly involved in these math efforts while others are focused on other things there. Many of them are, though, directly involved in helping generate synthetic data to improve their models at math and code, but I don’t know what this entails exactly beyond ripping ArXiv papers and likely generating synthetic reasoning traces in some way. My vague understanding is also that they run their models all the time on these kinds of problems, and if they seem to produce something promising, some of these mathematicians will take a look to see if it is correct or interesting — that could potentially be used to generate more training data to push it towards more promising directions, but who knows. For this batch of problems, my sense is those mathematicians probably helped rewrite the results into a more readable format and to give more context about the argument as LLMs are still pretty bad at that atm. I would guess the models sketched out the main ideas in some form, and then after checking them, the humans would direct Codex to write better related work and proof ideas and so on like a more human paper. But again, just a guess.

3) The true cost is also unknown but the numbers they give are almost certainly misleading for a variety of reasons. It clearly doesn’t account for the massive compute, synthetic data generation, human input, and so on, that goes into training these models precisely towards improving in these directions. Even for these problems, I suspect whatever number they gave is for the successful runs. But that’s obviously not the same thing as the full cost of them trying their models ad nauseum on all problems and then seeing what worked, which likely took considerable human effort as well. They don’t seem to want to provide any clarity about any aspect of their process, but I do find it hard to believe that they have all these brilliant theoreticians and they work as run-of-the-mill SWEs when they probably hadn’t touched code in years.

4) It’s been clear for over a year now that they have viewed math as a source of relatively cheap PR. Whether or not it’s deserved is up to you depending on how you measure their costs vs. achievements, but they certainly care about the effects on their valuation far more than they care about “advancing science” or whatever. They are shockingly nontransparent about anything, despite only lifting off the ground by promising researchers they would do charitable and open science.

[–]ksjdragon 10 points11 points  (9 children)

Can we please not post the same thing 50 times. I've responded to something like this many times already...

Mathematician here. In short. Results are cool. AI is not how math will be done. People like insights. AI doesn't give insights. Computationally assisted proofs existed before this. Current AI solving is heavily subsidized and isn't clear the cost it took to do this. Take the same money and hand it out as grants. You'll get 1000x more productivity. These proofs still need real mathematicians to decipher and translate.

Importance: cool, in math, maybe. Overarching impact: none

[–]Smooth-Ad8030[S] 0 points1 point  (1 child)

I checked to make sure this story wasn’t published, but I’ll search harder next time. Wasn’t sure if this was different to previous examples or the same ol hype.

[–]voronaam 1 point2 points  (0 children)

Here is a post from last week here on the same topic: https://old.reddit.com/r/BetterOffline/comments/1v4e0g6/mathematicians_and_real_software_engineers_help/

I do not see a ksjdragon response there, but there is mine there that I am too lazy to repeat.

[–]Snackatron 0 points1 point  (0 children)

Yeah I’m not a mathematician but it feels like AI is a tool that can take existing pre-invented mathematical concepts and trawl through the possible proof paths.

But the thing is, I see innovations like Laplace transform or Fourier transform or a convolution (I’m an engineer) and think that specific insights like that are a totally different ballgame.

I’d find this more interesting if an AI invented entirely new mathematical structures in order to generate a proof

[–]TLMTGT 9 points10 points  (1 child)

Contrary to the other poster, these are real results as far as I can tell. Note that OpenAI has real mathematicians collaborating with and/or working for them.

I only recognize a few of the problems, but these are/were substantial problems. As for whether the mathematics used to solve/disprove them is useful will take a while to determine. At a cursory glance, the proofs seem to follow the thread of construction/counterexample proofs that LLMs have been successful at recently rather than develop new conceptual ideas, but I'm not 100% sure since these are outside my subfield.

[–]Smooth-Ad8030[S] 0 points1 point  (0 children)

Interesting, yeah these seemed like real results, I figured they were given the sheer number of mathematicians working with them.

What is the functional difference between construction/counterexamples vs new conceptual ideas?

[–]ThanklessWaterHeater 5 points6 points  (15 children)

First of all, I personally dislike LLMs, and don’t use them for anything. I don’t want anyone to think that I’m a booster here. But I am an investor, and I think it’s important to keep up with developments in the field.

My brother-in-law is a tenured theoretical mathematician at one of the University of California campuses, and I was talking with him about this earlier this week.

He said that within the world of mathematicians, some of the recent work has been jaw-dropping. Theories that have gone generations without proof have been solved in a matter of minutes by LLMs. He doesn’t use LLMs himself, but he has read some of the papers and says the proofs appear to hold up.

I believe him. But I also think that this is one of very few areas where LLMs work well: 1) a field with very specific, long-established rules that can only be applied in very specific ways in order to create a valid proof. 2) any new theorem (whether generated by a human or by an LLM) is going to be carefully verified by other mathematicians before anyone even mentions it to the outside world. The proofs you read about in the linked article went through the standard peer-review process. My guess is there have been a number of bad proofs generated that nobody heard about because the person who generated it was capable of checking the logic and found it was incorrect.

In fact, as I understand it, LLM’s skill with mathematical proofs has been known long enough that the labs are looking for ways to use it more generally. They are trying to make LLM’s treat every day tasks with the same logic, making sure that every logical step in an argument can be proven the way a mathematical theory can be proven. That said, I read about that work a couple years ago and God knows based on the AI slop I see online every day it doesn’t seem to be working yet.

Anyway, like I said at the top, I’m not a booster here. I do believe my brother-in-law when he tells me this. But I also think this is a very niche skill and that if LLMs are ever going to be 100% reliable in more general uses, they need to be able to do more than generate a mathematical proof.

[–]Smooth-Ad8030[S] 6 points7 points  (6 children)

Interesting, thank you for the reply! I agree, it Sure seems like math and coding are by far the 2 best fields for LLMs.

[–]PatchyWhiskers 4 points5 points  (1 child)

Computers are good at doing computer stuff

[–]Icy-Recognition-7453 0 points1 point  (0 children)

Exactly right. Maths and coding are governed by really stringent rules - "either it is or it isn't", all the way down. Since LLMs are now really, really big computers, and computers have been able to smash themselves against a maths problem continuously faster than people ever since the pocket calculator was invented, it holds that eventually these problems would start to fall.

This is actually proved in part by the Hugging Face breach, in a way. OpenAI's model focused on the solution and ditched the entire process it was meant to follow, and then basically smashed every code package it could call into the problem. That's workslop right there. The LLM can produce anything you like, it's process and verifiability that make it valuable.

[–]ThanklessWaterHeater 2 points3 points  (3 children)

There are uses. Don’t forget Google’s Deep Mind won the Nobel prize for chemistry a few years ago for work on protein folding. I wish they would just keep AI in labs doing that work.

[–]Smooth-Ad8030[S] 2 points3 points  (0 children)

Oh I’m a big fan of modular AI in its current state, I just hate the idea LLMs will solve everything. That seems far fetched and out of reach.

[–]65721[🍰] 1 point2 points  (1 child)

That was a completely different architecture from that of LLMs. AI companies and their "enthusiasts" love to confuse the two and pretend RL's successes in narrow domains are related to the worthless outputs of LLMs.

Also, DeepMind did not win the Nobel Prize. Organizations cannot be awarded the Nobel Prize, unless it's the Peace Prize. They gave the Prize to DeepMind's CEO and VP instead—an utter disgrace.

[–]Ouaiy 2 points3 points  (0 children)

Another piece of it is that LLM math provers work together with automatic proof verifiers like Lean, which make sure the proofs are at least logically consistent. That is comparable to LLM coding producing code which at least compiles correctly. I don't know if mathematicians run into the same problems as coders do, in other words have proofs which are correct but don't properly answer what was asked of them.

[–]The-Menhir 3 points4 points  (1 child)

What good does AI proving/disproving theorems bring when proofs are often inscrutable & fail to bring about novel insights, other than pointless dick measuring contests between AI companies who couldn't care less about the field or the intrinsic humanity of it about who has the better model? I don't believe mathematics is a problem to be solved by AI, but a human endeavour.

[–]Main-Company-5946 0 points1 point  (0 children)

Maybe deriving novel insights from the proofs will be what becomes of the jobs of mathematicians.

After all the counterexample to the Jacobian conjecture alone doesn’t really tell you much beyond that the conjecture is false. But comparing the counterexample against previously made attempts to prove the conjecture was true can show you where the previous attempts failed and help cover potential blind spots.

[–]larrytheevilbunnie 0 points1 point  (0 children)

Just to be clear, these models literally couldn’t do high school level math 2 years ago

[–]creaturefeature16 1 point2 points  (0 children)

LLMs are pattern interpolators. They see connections that no human ever could and are being funded in a way to enable that to happen at a massive scale. This is the kind of stuff I would expect to see happen. They're don't seem like they're discovering as much as uncovering, and that's an important distinction. They're still, and always will be, just supplementary to human efforts.

[–]the-great-defector 0 points1 point  (0 children)

One thing I wonder with this is whether or not OpenAI can now try and pivot this to sell off models for University research PhD projects through something like grant funding? I think Anthropic have also been releasing models for things like biology, so wonder if it's an area they feel they can get some revenue in.

[–]larrytheevilbunnie 0 points1 point  (0 children)

I’m smelling a lot of cope in this thread, like yeah, these particular results aren’t economically important, but the models literally couldn’t do high school level math 2 years ago.

[–]Pale_Neighborhood363 -2 points-1 points  (0 children)

The results are JUST counting. It is useful but trivial. Mathematics has two things the formal and the speculative.

The models are Very Very good at the formal which allows the examination of the speculative.

The results are on the transition of the quantile to the continuum. These points are JUST counting (very complex counting) The machines count much much faster.

Importance is a WEIRD in that mathematically any result is the MOST important. This is the universality of the mathematical philosophy.

The 1980's had computers that started* this trend. The Mandelbrot set & the theory of Chaos date from that time but the mathematics that base those goes back two or more centuries.

This is the mathematics of the 1930's & the 90's being stress tested. 1930's the measurable limits & 1990's the functional machinery.

The mathematical discoveries are esoteric as they are at the edge of our understanding, they become important when technology/physics needs that understanding to grow.

*arbitrary