×
all 119 comments

[–]TerminalObsessions 32 points33 points  (2 children)

LLMs are good at some tasks, and provable outputs in highly-deterministic systems (like mathematics or coding) may wind up being their best use case. Strict controls help counterbalance the tendency of probabilistic systems to confabulate and bullshit. Code compiles or it doesn't. The proof validates or it doesn't. It's not like writing a novel or creating art or even writing a business email, where success is poorly-defined.

You don't need to believe that LLMs are bad at everything or a complete scam. There *are* some applications of the technology. Those applications are probably 1% of the hype, if that, and they aren't going to stop the colossal bubble from collapsing. Anthropic can't pay its bills on mathematical conjectures. While the AI space rivals the crypto space for self-dealing, fraud, and ludicrous hype, it's not *all* fake. There is a new technology, and it can do some things. It just can't do things at the scope or scale being suggested by the folks trying to get rich off of it.

[–]alochmar 3 points4 points  (1 child)

Exactly. To paraphrase Ed, it’s a couple billion dollar industry cosplaying as a trillion dollar industry.

[–]Mashic 59 points60 points  (7 children)

There are current problems in math, that the current knowledge is sufficient to solve them. However, no human has either connected the dots (by trying the right approach among many), or not enough humans decided that it's worth their time.

LLMs in this, especially if guided by humans, can spend so much more time working at higher speed that the human brain to trying multiple approaches to solve these leftover problems.

[–]SoggyMattress2 24 points25 points  (4 children)

Yup. I'm no mathematician so happy to be corrected here but from what I understand LLMs are actually quite good at solving maths problems where it's nothing novel, they just throw enough shit at the wall at such insane volumes something sticks.

The erdos problem it solved earlier in the year just brute forced like billions of tests and eventually something was interesting.

So it's less "omg an LLM solved a maths problem nobody could solve!" And more "LLMs can brute force maths problems not many people were trying to solve."

[–]Summary_Judgment56 18 points19 points  (0 children)

If you're talking about the Erdos problem that was in the news (re: the planar unit distance problem), the slopbot didn't even solve it, it just disproved Erdos's conjecture about what the answer to the problem might be.

[–]ksjdragon 14 points15 points  (1 child)

The sad part is we can most likely write cheap scripts and optimizations to do this sort of search without LLMs, if that was a useful endeavor. It would cost next to nothing, and be done in the background, and potentially find more counterexamples.

They are quite good at solving math that has been given to it. They're pattern matching. And they apply various levels of patterns to language, and most of those patterns in new areas are actually nonsensical. But, go through enough and you'll find one that isn't.

If we imagine the human set of knowledge, there's a lot of abstract concepts that exist in one field, that could be applied to another but simply hasn't for many reasons. That is, we can fill in 'gaps' with preexisting developed blocks. AI can do a form of this, with sufficient iterations and money, but inefficiently as it ultimately searches averages, and not the whole space.

[–]Easy_Tie_9380 0 points1 point  (0 children)

This was simply not a problem that was brute force searchable. Given that, I find it extremely unlikely that it was done by an llm at all.

[–]spellbound1875 0 points1 point  (0 children)

Less brute force and more combining a variety of human found insights together and then running a bunch of tedious math very quickly. It is more targeted than brute force since the LLM's have consumed basically all pu listed math research.

This does have limitations however in that the LLM still can't think so it often only half finishes the work. The Erdos problem it "solved autonomously" earlier this year (really it just disproved an existing solution) was improved by a human within a couple days.

It's a productivity tool with applications in fields with extremely structured syntax.

[–]kiddodeman 15 points16 points  (1 child)

On top of that a russian mathematician apparently was very close to connecting these dots, if I understood correctly. A prposed counterexample that failed is very similar to the provided one.

[–]spellbound1875 4 points5 points  (0 children)

Even the math problem OP brought up apparently had some existing work for the LLM to build off of. The bot is basically just smashing a variety of insights together in novel but very doable ways. It's just addressing the natural human blindspot of not being able to read 100% of all published literature in a field.

[–]TLMTGT 15 points16 points  (5 children)

It was a major open problem, but it's presently not a hot topic in algebraic geometry. This counterexample is important because it provides an opportunity to refine the conjecture. As for the consequences of the conjecture, I'm not sure what it's worth besides generally deepening our understanding of algebraic geometry/math (which is important, of course). However, this is a significant achievement, make no mistake. It falls under the same category as many major announcements in math recently, which is that it's able to come up with interesting examples that would otherwise require too much effort for a mathematician.

[–]AntiqueFigure6 2 points3 points  (4 children)

"...it's able to come up with interesting examples that would otherwise require too much effort for a mathematician."

Which in some sense means that it alters the economics of solving certain problems. It will be interesting to see where the economics land post-bubble - i.e. will it remain cheap enough to continue to be used for solving these problems that were out of reach due to requiring too much human effort. Possibly a long tail of tractable problems that can be solved before expensive re-training required to solve more.

[–]SilchasRuin 2 points3 points  (1 child)

In my former life as a math grad student, this sort of technology would have been incredible. Producing counterexamples can be incredibly hard. You have to thread a needle in obeying the assumptions of the conjecture while violating the conclusion. It can get very demoralizing to do as a human.

[–]AntiqueFigure6 2 points3 points  (0 children)

I imagine that an important reason to give a grad student that kind of work is their relative low cost, and I’d conjecture that post bubble using an LLM for such a task will be significantly more expensive so I don’t see grad students using LLMs to find counter examples is going to become common.

[–]JD_Waterston 2 points3 points  (1 child)

‘Will leading edge models stay cheap-ish?’ Probably not.

‘Will open weight models be able to provide pretty impressive capabilities for pennies?’ Almost assuredly.

Research labs have long offered compute to attract students - but the greater slope of science has benefited from computers without being bleeding edge. AI tools seem very much like that.

[–]AntiqueFigure6 0 points1 point  (0 children)

So the question is in a sense where is the border lie where leading edge is actually necessary? Was bleeding edge really needed for this problem and/ or would it be needed to solve another problem that is in some sense similar wrt what is needed to solve it (which might have different answers depending on what we learned doing this one or other like it in the near future)?

[–]Significant-Green130 20 points21 points  (12 children)

The “weight of the discovery” to most people on Reddit should be “Anthropic employee with math PhD tries to copy the OAI playbook of using math problems to produce PR for their models.” It’s honestly a bit amusing how they copy each other for everything. Anthropic has overall not seemed to care about how the model does on math research compared to code, and in general does not hire nearly as many mathematicians as OAI. I somewhat suspect this is an “internal model” (=slightly better model + massive compute) of Fable for that reason, but idk I didn’t see a transcript this morning and don’t care to check now. I guess the handful of math people at Anthropic maybe feel left out or something and want to get similar PR.

For the .001% of people who actually care about math:
(1) It’s a somewhat famous conjecture. I don’t know much about it to give an opinion, but given how simple the counterexample is, my suspicion is Fable found it by identifying a much smaller class of polynomials that are somehow reasonable to try (maybe they appear in related literature or whatever) and then ran code to try to find one and indeed found it. There probably is a loose “conceptual” reason to look at this family, but until someone looks at the reasoning traces, it’s hard to say. It’s not clear to me for this answer where it is on the “deep guess” vs. “semi-educated brute force” scale, and I don’t know the area at all to have a sense given how simple the counterexample looks. The OAI unit distance conjecture counterexample is a much more intricate construction, but at least I got some vague sense for what it tried compared to this one at the moment, as I didn’t see a transcript this morning.

(2) One thing apparently nobody seems to understand in any of these discussions is that much of mathematics comes from studying interesting questions, not always the answers: the best ones produce a wealth of interesting new frameworks and structures regardless of what the literal answer ends up being. The AI benchmarking/PR by “running models all the time on many problems in case it does something” zeitgeist is deeply irritating and in my view, actually counterproductive to this goal. Others closer to the area might disagree, but just producing a simple counterexample has mostly just the immediate value of stopping people from looking for a proof. But that in itself is not deeply useful unless it leads to a more refined understanding of why and which directions it suggests one should go. At present, humans have a monumental lead in this regard.

(3) I hate the idea that these results are often framed as “AI solved a problem mathematicians have failed to solve for 50 years!!!” Maybe 10% of mathematicians would have heard of this problem any time recently. Of that group, maybe 20% would even care to solve it. Of that group, maybe 20% would ever make a single long and serious effort at producing a counterexample, and these numbers are probably very generous. Mathematicians are much more keen on producing general frameworks and ideas that lead to a broader research program; if you think the result is true, spending time on a counterexample that may not even exist and that may not shed much insight into more general frameworks is not really a good use of time. Most sane mathematicians do not spend their waking hours thinking about one problem, that may or not be too hard, all the time. Rather, many of them pick up tools from reading and talking to people and they try to figure out what they are good for in their area. Obviously, LLMs do not have any such proclivities.

[–]lazier_garlic 5 points6 points  (0 children)

To add to what you said, there were some modest little problems solved in mathematics in the past by amateurs who had a lot of time on their hands and wanted to keep busy.

[–]AntiqueFigure6 5 points6 points  (0 children)

" The AI benchmarking/PR by “running models all the time on many problems in case it does something” zeitgeist..."

something that is likely to stop dead the moment investors stop giving AI hyperscalers 100s of billions of dollars to incinerate, so all in all unless the problems solved in this way themselves provide new tools for mathematical discovery, the legacy is likely to be a list of mildly interesting results and curios.

[–]Impressive-Gene-421 -1 points0 points  (5 children)

As a quick aside, your post being a somewhat nonsensical rant, ~100% of mathematicians would have heard of this problem. Famously so.

[–]cunningjames 7 points8 points  (0 children)

You’re missing the qualification “any time recently”. It’s a statement that this problem is not on the radar most working mathematicians. And that’s true — most of them were not working on the problem and had probably never worked on the problem.

[–]deividragon 4 points5 points  (0 children)

Nope. I am a mathematician and had not heard of it before. Neither had another mathematician coworker of mine. Conjectures that are not hyper-famous like the Riemann hypothesis are often not known of outside of the respective fields.

Regarding the provenance of the polynomial, it seems to be deeply related to a 2D construction by Vitushkin, particularly Example 2 from his paper: "Evaluation of the Jacobian of a rational transformation of C2 and some applications". I don't fully understand the situation, as it's not my field of expertise and the paper is in Russian, but this is what's coming up in maths circles right now.

Also generally, psychologically, I've never been able to draw myself to work on any famous conjectures in my field. It feels like wasted time. Not that many people have the drive to put much effort on working on the most famous, long standing problems. You need to either be exceptionally good (and aware that you are), overconfident or both to have the will to throw yourself hard against something you know plenty of people before you already tried.

[–]Significant-Green130 7 points8 points  (0 children)

Sorry you have difficulty reading. Best of luck.

[–]AntiqueFigure6 5 points6 points  (0 children)

Use of phrase "100% of mathematicians" seems like sufficient proof that a statement is nonsensical.

[–]TribeWars 1 point2 points  (0 children)

You don't need to overhype it either. Most mathematicians don't work anywhere near algebraic geometry and this conjecture is arguably most notable for how easy it is to understand the statement of the conjecture relative to everything else in that subfield.

[–]k_D-5kt7ZA4QGBP68ZBN 0 points1 point  (3 children)

I think you're missing the fact that this went unsolved for a century. And you're dismissing its importance. If it were true you could find out whether a function is injective by computing its jacobian. There's a reason why this blew up so much. I think you have some kind of prejudice towards AI or something, because this is truly extraordinary, and you should spend some time understanding why rather than writing this nonsensical rant. Most mathematicians don't bother trying to find counterexamples due to its unfeasibility, and AI is a cornerstone in that regard.

[–]Significant-Green130 1 point2 points  (2 children)

Serious question: why are you here, commenting on a 4 day old comment?

[–]k_D-5kt7ZA4QGBP68ZBN 0 points1 point  (1 child)

Because this topic is still trending and I wanted to see what the almighty redditors had to say about it? Serious question: why are you here, commenting on a 2h old comment? We need your opinions on more recent topics, fast!

[–]Significant-Green130 0 points1 point  (0 children)

Yes, I got a notification because you apparently got upset about a 4 day old comment and felt compelled to respond. I can't help you with that, I recommend asking Fable or whoever else you have to pay to care. Take care.

[–]Ouaiy 5 points6 points  (4 children)

As I understand it, as an Anthropic employee he had unlimited access to Fable and could run it for hours working on this problem, whereas normally people who ask Claude to work on famous unsolved problems get booted down to a cheaper model (maybe also with time limits) so as not to futilely waste resources.

[–]AntiqueFigure6 1 point2 points  (2 children)

Sounds like a token spend beyond the price range of most university mathematics departments.

[–]Ouaiy 1 point2 points  (0 children)

That, and possibly not even accessible at all.

[–]SethDusek5 0 points1 point  (0 children)

People have still successfully generated novel proofs on public models. https://chatgpt.com/share/69dd1c83-b164-8385-bf2e-8533e9baba9c

The above thread is a guy just saying "solve this problem and either prove or disprove it" and after 80 minutes of thinking it generated a proof

[–]nothingNowhereForNow 6 points7 points  (4 children)

A lot of people in here are trying to downlplay this, but as far as mathematical discoveries go, this is a big one. It's a disproof of a major conjecture that's going to shift how people approach a fairly big question in mathematics. (Wikipedia has already updated to say that it's disproved for dimensions 3 or greater, trivial to prove for dimension 1, and undecided for dimension 2. Focus is probably going to shift to dimension 2 now.)

It's also noteworthy in being a very, very, elegant counterexample. While it looks trivial, finding the right combination of coefficients is actually incredibly difficult due to the search space blowing up, hence why it took nearly 100 years to discover it.

Of all the AI assisted math advances, this is easily the biggest one.

Now what it means for AI, that's unclear. Alpöge and Mathew are both accomplished mathematicians, so it's not some randoms off the street just chucking millenium prize problems into the void and seeing what sticks, which raises a very real possibility that the model had a lot of help getting to the solution. Until Alpöge publishes the prompts he used, we won't know if he just asked it "solve the jacobian conjecture, be sure to think hard about it," "here's a paper that has a pretty good approach to the jacobian conjecture, see if you can do anything with it," or "hey, I've got a routine to solve the jacobian conjecture, go ahead, code this up, see what comes out".

[–]AntiqueFigure6 4 points5 points  (1 child)

"A lot of people in here are trying to downlplay this, but as far as mathematical discoveries go, this is a big one."

I think in many ways the importance of this specific result is less interesting than whether or not this is going to be a tool that can be used widely by the mathematical community. Maybe a maths department at a well endowed Ivy League university or whatever the equivalent is outside the USA could afford it, but could a moderately well funded maths department at a mid tier university afford to find a counterexample using the method used here at a token price that allows for a profit margin after genuinely allowing for all costs?

[–]Sad-Variation-8006 5 points6 points  (0 children)

Do we trust Alpoge to be fully transparent? He is likely being paid more than almost any mathematician on earth, and there are tens if not hundreds of billions of dollars in market cap at stake here for his employer

[–]FairlySadPanda 5 points6 points  (0 children)

The other aspect of this is survivorship bias. "This model COULDN'T find the solution to maths problem X" is not news.

The con of the bubble is being able to tell a story ("we are using sand to create God") and finding ways to keep the valuations going up.

E.g. "all concept artist would be fired once image generation got just a teeny bit better" - invest in image generation companies! Then it didn't get a teeny bit better.

Big discoveries have been made via very dumb methods for a long time. Pluto was discovered by one man spending every night for a month trying to find a single unfixed star by sifting through tens of thousands of photographs of a part of the sky, and only got it so fast by extreme luck.

It's a good thing to have the discovery, it doesn't stop the method being really dumb!

[–]lcnielsen 13 points14 points  (15 children)

The funny thing about this one to me is that it seems like you could easily write a software to programmatically scan for rational, low-degree counterexamples and it would've found this one quite quickly. Just, nobody seems to have tried it.

[–]oaga_strizzi 18 points19 points  (1 child)

“ Just, nobody seems to have tried it.”

Multiple people smarter than me had dedicated their PhD to solve this.

No, simple brute forcing would not work - degree 7 with 3 variables ends up with ~360 coefficients. And you don't know that degree 7 with 3 variables has the counterexample, it might by any other shape.

[–]lcnielsen 1 point2 points  (0 children)

Yeah, I misread the formatting. But there are far more constrained classes of polynomials where it is reasonable to look for counterexamples and there's apparently a good deal of somewhat older work on this, including similar near-counterexamples.

Not quite brute force but not that far off.

[–]LaurenMP74 8 points9 points  (0 children)

Not only that two mathematicians had already shown that if there is a counterexample it has certain properties, most significantly integer coefficients. Which is a pretty strong constraint for a search for counterexamples.

[–]That-Helicopter-9066 4 points5 points  (0 children)

Even constrained, it would have taken an immense amount of time to brute force this solution. Some of the smartest minds dedicated portions of their lives to this problem. You are trivializing this with “nobody seems to have tried it”.

[–]Most-Bookkeeper-950 5 points6 points  (5 children)

This isn't true... the number of polynomials with imteger coefficients blows up extremely quickly

[–]LaurenMP74 2 points3 points  (4 children)

Funny enough in this case, the counterexample doesn't involve an integer greater than 4. So the counterexample was right under everyone's nose the whole time.

[–]Most-Bookkeeper-950 3 points4 points  (3 children)

I see a 12xy2 in the expansion... but yeah, it is an insanely simple counterexample for a problem thats been attempted so many times

[–]lcnielsen 1 point2 points  (2 children)

It factors out though, but I don't know what combinatorics strategies one would typically employ for this search. Likely more than this counterexample exist anyway.

[–]Most-Bookkeeper-950 2 points3 points  (1 child)

The counterexample, fully expanded, is here.

The bivariate case is still open

[–]lcnielsen 2 points3 points  (0 children)

Thanks, that's more readable.

[–]AntiqueFigure6 2 points3 points  (2 children)

“ Just, nobody seems to have tried it.”

Can’t have been a high priority.

[–]lcnielsen 13 points14 points  (1 child)

I think a lot of mathematicians are also not as interested in finding counterexamples that prove something false without necessarily giving more insight.

[–]ksjdragon 8 points9 points  (0 children)

This is generally true. There's a lot of computational feasible unsolved problems that require some computers and scripts, and if a mathematician has sufficient programming background (relatively rare, in my experience) they can find counterexamples. There are a few of these cases, but usually it's not something one in academia would chase, as it's not academically interesting.

The existence of a counterexample is great, but what is better and valuable is an illustration of a class of counterexamples, for instance, and a reason as to why these examples are not valid. Alternatively, a specific counterexample, attached with further narrow criteria for what would make counterexamples would be a good step.

Otherwise we gain very little information outside of a conjecture is false. If so, mathematicians want: if this conjecture is false, what are the bounds in which it is true?

[–]dumnezero 2 points3 points  (0 children)

https://lean-lang.org/

Lean is a functional programming language and theorem prover built for formalizing math and for formal verification, but is flexible enough for general coding. If you’re a beginner, we recommend the Natural Number Game. If you feel ready to dive deeper, there are great textbooks, tutorials and interactive games to be found on this page.

I think they're using stuff like this for doing "Test Driven" efforts.

This is where Lean, Microsoft’s open-source theorem prover, plays a key role. By treating Lean as a formal API for proof verification, we can integrate symbolic logic into LLM reasoning. This enables a neuro-symbolic workflow where models propose ideas and Lean enforces validity, producing auditable proofs https://www.turing.com/resources/lean-and-symbolic-reasoning-in-llms-for-math-problem-solving

[–]sclv 4 points5 points  (0 children)

Its an important and real result, and we should expect that despite all our frustrations with generative AI there are certain very specific applications and fields where it will actually be effective. It requires a lot of factors to coincide -- problems that are easily checkable, but hard to solve, and where solving can be attempted by combining ease of running a lot of tricky subprograms (like computer algebra tools) with searching over a very wide literature. So some problems in math are like this, some specific types of programming are like this, and these sorts of applications for generative AI will exist and remain.

The important thing in my opinion is that this narrow subset of problems doesn't encompass much of interesting contemporary math or programming, although side-elements of these disciplines can be improved by use of such tools. And outside of these fields (which have a lot of human-elbow-grease put in to create a corpus of knowledge and tools for machine checkability) other fields don't tend to have as many problems with these properties, and the further away from these fields we get, the fewer such problems like this we get.

So the utility is real in narrow cases, and increasingly impressive in them, but the narrowness of these cases hasn't shifted, just the capacity within those cases.

[–]AntiqueFigure6 2 points3 points  (4 children)

I guess there will keep being a trickle of similar announcements…until the bubble pops and the money runs out.

[–]Berzerka 0 points1 point  (3 children)

The internet famously went away after the dot com bubble burst, that's why we are having this conversation by fax!

[–]AntiqueFigure6 0 points1 point  (2 children)

Relevant because pets.com sold access to the internet I guess.

[–]Berzerka 0 points1 point  (1 child)

A lot of AI companies will obviously go the way pets.com did, but it seems highly unlikely that all of them will.

Most of the bubble back then was e.g. Cisco, Ericsson and the likes and they did bounce back later.

[–]AntiqueFigure6 0 points1 point  (0 children)

Routers are a lot more useful than LLMs - probably the closest AI bubble equivalent is GPUs, and sure, NVDIA could follow a similar trajectory. 

[–]Lowetheiy 2 points3 points  (1 child)

No one know really knows because this isn't the right audience. I don't think the majority here has taken college level math courses either. Perhaps it would be better to post this on the math subreddit.

[–]PaymentFamiliar375 1 point2 points  (0 children)

you could post this on the math subreddit, but you’ll get basically the same answers as you would by posting it on the singularity subreddit… 

[–]wowbaggerBR 1 point2 points  (0 children)

I have no problem with LLMs being good at this sort of thing. Makes sense: they are better Googles in a way.

[–]LaurenMP74 5 points6 points  (33 children)

While it's nice it's not really all that. What Claude did was find a counterexample ie a number for which the conjecture is false. This doesn't take AI at all, it's just number crunching you can do on any computer. Also all that happened was Claude found a counterexample for one particular case (the conjecture involves polynomials in different dimensions) in this case the Claude counterexample only applies to dimension 3. Which means only that the conjecture is not true for that, it says nothing about the truth of the conjecture for other dimensions and conditions. So the claim Claude disproved the conjecture is not true. It didn't do that, it just showed it's not true in a particular case.

[–]Easy_Tie_9380 0 points1 point  (0 children)

Im pretty confident counterexample holds for all n > 3 as well. Only n=2 remains open now.

[–]lcnielsen 0 points1 point  (29 children)

There's also lemmas that constrain the counterexample space you need to look at. I'm surprised nobody found the counterexample by brute force search before.

[–]absolute-black 11 points12 points  (26 children)

The estimated remaining search space of functions given narrow integer bounds was on the order of 10475 lmao

[–]LaurenMP74 0 points1 point  (22 children)

And yet the counterexample involves the integers 1, 2, 3 and 4. So wasn't a very deep dive to get it.

[–]absolute-black 2 points3 points  (21 children)

If you bind the integers between -10 and 10 the search space is >10475, and you still think this was basic brute force?

edit: for fun I went back and estimated the search space if you knew ahead of time because God told you that |x|<=4 was a viable bound, but didn't tell you to limit the degree further. 9360 ~=10343, lol. What incredible savings!

[–]lcnielsen 2 points3 points  (20 children)

  1. Your implied statistics assume this is the only example of such a polynomial. Entire classes likely exist.

  2. This wasn't a case of integers |n| =< 10 was it now? Well it depends on if you assume the search would happen over the expanded forms or not.

[–]absolute-black 2 points3 points  (19 children)

Yes they do, which does not make the problem brute forceable. I do not think I am the one making the stronger claim of implied statistics here, compared to the people I responded to.

Yes, 1,2,3, and 4 are all <10.

[–]lcnielsen 0 points1 point  (18 children)

Yes, 1,2,3, and 4 are all <10.

They are also < 5.

[–]absolute-black 5 points6 points  (17 children)

Correct. I'm not sure what you think the point is, though. Say someone somehow knew 5 was a plausible bound for integers in the 3 dimension Jacobian (how? why? when? certainly we know that this didn't happen for this result lol). That's only ~10375 options in search space, still about a bajillion kajillion times what could even remotely plausibly be brute forced within a universe made out of matter.

[–]lcnielsen -1 points0 points  (16 children)

Obviously the search happens in some restricted space and you start with dimension 3 if that's the lowest you can go.

And of course more than one counterexample most likely exists.

Or what's your alternative explanation?

[–]lcnielsen -1 points0 points  (1 child)

So you're saying this counterexample wouldn't have been found reasonably quickly? I read one post saying it would've taken a few days 30 years ago which sounds about right.

[–]absolute-black 5 points6 points  (0 children)

It is utterly impossible to imagine finding it with raw brute force inside of the physical universe.

[–]LaurenMP74 -2 points-1 points  (0 children)

Yep, like that if there is a counterexample it has integer coefficients.

[–]Civil_Blueberry4165 0 points1 point  (0 children)

A mathematical conjecture is meant to encourage mathematicians to come up with new mathematical structures or concepts to tackle it. In other words, mathematics is not about theorem proving. So, if a proof or counterexample doesn’t reveal new mathematical ideas, it offers much less scientific values. A case in point is the recent formalized proof of the sphere packing problem in dimension 8 by Math Inc.

[–]Pseudanonymius 0 points1 point  (0 children)

One of the reasons why they are relatively good at math problems, is because there are a hell of a lot of areas of math. Most humans specially in some of them. But math has the fascinating property that surprisingly often solutions to one field can be found by being able to prove the problem is equivalent to a problem in another field. LLM's having consumed all different fields, are quite good at making those cross-connections, while humans often specialize in some fields, but nobody can specialize in all. 

This is combined with the fact that AI can brute-force it's way through some problems, where humans can't. It's the same old every time, LLM's can speed up work when the output has an external validation (so you don't have to trust the AI, you can check if what it did is correct. Math is perfect for that) and if the data it needs to process is very large. 

This is just a subset of math-problebs which this method would work for though. Still, all things considered, this is one of the more impressive accomplishments. 

[–]voronaam 0 points1 point  (5 children)

It is an interesting case, actually. There is a certain bias in the scientific community. Can't remember its name, but it is a lot easier to publish a paper that proves something, than it is to publish a paper that shows that something is wrong.

Most of the scientists formulate hypothesis, then run experiments to confirm it, and if it gets confirmed - publish. However, if the experiment denies it or is inconclusive, nothing gets published. The fact that someone tried to prove a certain hypothesis is lost to history.

So is the case with the Jacobian Conjecture. There was a lot of effort spent on trying to prove it right. But nobody really tried to prove it wrong, because there is no publication and no fame from doing so in the modern scientific world.

You can actually see it yourself: there is no paper, just a short tweet. If not for LLM being involved nobody would've payed any attention really.

It is even possible that some mathematician previously found the counterexample before, but thrown it in a bin because it is not a result you can publish. And moved on to some other problem.

As for the value of the LLM here... It helped us with some math problem. Do you know how much did it cost to make this step forward? I have not seen that information made public. It might very well be that LLM just spent more money on finding this counterexample than all the humanity combined did before it. Mathematicians are not the most well-paid profession. And most people working on this problem were hobbyists.

Is it really a surprise that spending big on advancing science - advances science?

[–]lcnielsen 5 points6 points  (2 children)

As for the value of the LLM here... It helped us with some math problem. Do you know how much did it cost to make this step forward? I have not seen that information made public.

You'd also need to include all the presumably vast resources spent on failing to produce useful results on various issues by these companies and their armies of mathematicians.

[–]voronaam 1 point2 points  (1 child)

I am not a mathematician, but I dip into Application Security quite a lot. A few months ago Anthropic was pushing hard on the marketing at that sphere and the number that struck me: only their external grants to various collaborators were more than all the annual budgets of ALL the bug bounty programs together. And not by small margin, it was about double the amount. And that's only external grants, they also spent untold amounts internally.

To me it is not a big surprise that lots of funding in an area leads to some results. But I take an issue with marketing it as a unique LLM capability - because nobody really tried to spend this much money on any other tech.

[–]lcnielsen 0 points1 point  (0 children)

Yeah, I also work a fair bit with security as a sysadmin/sysops guy, and that was my reaction too. Literally nobody except intelligence agencies used to put this much effort into finding vulns, and they weren't necessarily disclosing them.

There were also a lot of bad practices around like people sharing 1-click exploit scripts as "proofs of concept".

[–]venusisupsidedown 0 points1 point  (1 child)

Whatever you think about LLMs or the practical implications of this result the idea that this is not publishable is totally wrong.

[–]voronaam 0 points1 point  (0 children)

It did not get published, did it?

[–]PensiveinNJ 0 points1 point  (0 children)

"Computers have aided mathematical research for years, as proof assistants that make sure the logical steps in a proof really work and as brute force tools that can chew through huge amounts of data to search for counterexamples to conjectures."

LLMs in math can be used as sort of directed brute force tools. Humans can help narrow down the possibilities so that every single possible permutation of something doesn't need to be considered because there are often situations where the massive amount of compute available wouldn't be able to get through all those possibilities.

They're also better at disproving things than proving them because they have a final target (an answer) already provided. Similar to chess engines searching for checkmate, it's easier if there's already a target to shoot for.

The significance is there was some progress made on an algebraic conjecture, but computers have been used as aids for mathematicians for decades. LLMs are a little better at it than previous tools and they have a lot of compute available but there's nothing else going on beyond that.

If only a really famous mathematician who is involved in machine learning would elucidate for everyone what is actually happening with LLM collaboration it would probably save a large amount of existential anxiety amongst the aspiring and practicing mathematicians around the world.

[–]PaymentFamiliar375 0 points1 point  (1 child)

i'd like to see more of the math proofs that ai has supposedly written discussed on this subreddit. over on the other mathematics subreddits, the general consensus when seeing these proofs is "Singularity is inevitable, ai will take over research math entirely and we'll all be useless, panic and shit yourself." so it's refreshing to see any viewpoints that aren't that.

[–]cooolchild -1 points0 points  (0 children)

exactly, everyone on the math subreddit is insanely doomerist about any post about ai (which seems to comprise 80% of posts on there nowadays) and unwilling to see any nuance. It’s not possible to trivialise this problem, of course it’s an incredible development but it’s also easy to see that the latest string of ai maths proofs are certainly PR stunts to try and get some money from maths departments. For all these impressive proofs, how many bear no fruit? It’s impossible to say.

[–]shachar1000 0 points1 point  (0 children)

The amount of cope in this thread is ridiculous 

[–]FewPage431 0 points1 point  (0 children)

I am glad that this sub exists. Other sub are too pro ai, while this sub is too anti ai. which helps me form a balance view. Although jacobian conjecture has been attempted by many because it's very easy to understand for graduate maths and famous (one way to know is that checking such problems are on r/badmathematics or not, because math crackheads only attempts famous and easy to understand unsolved problem), it cost too much for llm to solve. Right now and in the foreseeable future (20 to 30 years?), mathematicians will be more cost-effective than llm. With enough data and compute, any problem that is solvable by humans can be solved by computer. As time passes and as data and compute will increase, more fields will be cost effectively solved by ai (especially for fields like mathematics, where not only you can generate data easily but also could rl by lean). Machine learning is not a stagnated field. Transformer architecture is just one of the groundbreaking inventions in machine learning. If you just listen to many people who have been working on ml for a long time (Dwarkesh Podcast is great, just ignore ceos), you will be surprised by how different people think to each other. That implies that many people are working in many different directions, contrary to this sub thinking of everyone just on llm pills. This is a great sign of coming progress, plus ai in financial bubble right now, which attracts many students (because of money) to pursue this field, which will increase chances of breakthrough more.