It's good to see validated numerical proofs seeing a resurgence now that they are substantially easier to achieve.
Others might be able to chime in, but my experience is that AI is effectively taking proofs that were once iterative (pushing current ideas further, tightening arguments) but considered time-consuming, and turning it into very low-hanging fruit. So stuff like this, as well as the new lower bound on the asymptotic ratio of zeros on the critical line, really aren't that impressive anymore. But it is likely necessary to do nonetheless, just as low-hanging fruit has always been.
The real frontier remains (even if many were not operating there) the development of new definitions and ideas that push right past the walls that were previously there. AI still seems to be shockingly poor at doing this, and you can feel that when you use it for proper hard problems.
Anyone interested in the subject area has noticed the recent surge of AI-assisted or generated mathematics. From the outside, it is very hard to see, if AI indeed already speeds up progress in mathematics in general - apart from a number of spectacular results, (https://mathoverflow.net/questions/502120/examples-for-the-u...) - or if the quality of the large majority of these "proofs" is so poor that reviewing them is a waste of time.
I would be grateful, if someone familiar with the situation could say a word about this.
As for the proof in question - I'm not sure the author himself can firmly state he understands every detail of what he presented. In a way it's amusing that people who already enjoy a certain level of recognition outside the field are now using it to find (via shiny websites) reviewers for proofs they have worked out as a hobby using Chat GPT and the like.
What we definitely need is more recognition for those who possess the competence and energy to assess the correctness and relevance of such "results".
> reduced to checking the definition of the constant
But this doesn't mean the remaining task is small or doesn't ask much from the reviewer, does it?! Otherwise we'd see a considerable turn-out of new findings on https://palomar-registry.org/ or https://github.com/Vilin97/lean-pool , or not?! (A considerable number of the presented proofs claim new results, as far as I can tell.)
I'm not able to judge this for content, but it feels frustrating to read something so obviously Claude-authored as being written "by" Jude Gomila. It's interesting to see the (presumed correct) reduction in this bound and this presentation is pretty compelling. I learned a bit about the Polymath 15 project from this post and I enjoyed that a lot.
That said, I have no idea what contribution may have been done by Mr. Gomila or not. I don't know if I can ask him for more details. Even if he wrote the whole thing (and either has deeply internalized Claude's voice or had it edited and rewritten by an agent) it's difficult to believe that he's the sort of expert I'd expect of someone who had actually done all this work.
This is basically the same feeling I have about vibe code contributors.
So, maybe this is not quite a fair analogy but a lot of famous modern artists don't do a lot of the actual execution of their work. If Jeff Koons relies on a team of artisans and technicians to make a balloon dog, and if Dale Chihuly depends on a workshop of glass blowers to make a chandelier, and if Paul Cummins used a team of ceramicists to make poppies, what did the headlining artist contribute? Not everything, certainly, but also not nothing. Ideas, vision, feedback, guidance?
I think we're used to the idea in math of generally every idea coming from a specific contributing human. In code, maybe less so, but still we generally we find it useful to attribute at least _lines_ to who touched them last. But maybe, if Claude were just less annoying in the prose it made, and could better imitate an actual human, we'd be able to get used to the idea of humans claiming credit for something that the human guided the software to create?
He posted it. That’s some of the work you question. Presumably he queried the AI and led it in some way. That’s also work.
As far as the exact breakdown we will never know. But questioning the authorship in a way that leads us to believe that the author did nothing is wrong.
It's not wrong, is it? I didn't want to trust myself, Wikipedia agrees that the Riemann hypothesis is equivalent to Λ = 0 (because the lower bound is now 0). This proof lowers the upper bound, but we still don't know if it's 0 or between 0 and that upper bound (inclusive).
Wow, I have a hard time swallowing these Claude generated blogs. So many off-putting llmisms. I wonder if this is something I just need to accept and get used to, or whether it'll pass as models advance.
It's good to see validated numerical proofs seeing a resurgence now that they are substantially easier to achieve.
Others might be able to chime in, but my experience is that AI is effectively taking proofs that were once iterative (pushing current ideas further, tightening arguments) but considered time-consuming, and turning it into very low-hanging fruit. So stuff like this, as well as the new lower bound on the asymptotic ratio of zeros on the critical line, really aren't that impressive anymore. But it is likely necessary to do nonetheless, just as low-hanging fruit has always been.
The real frontier remains (even if many were not operating there) the development of new definitions and ideas that push right past the walls that were previously there. AI still seems to be shockingly poor at doing this, and you can feel that when you use it for proper hard problems.
Anyone interested in the subject area has noticed the recent surge of AI-assisted or generated mathematics. From the outside, it is very hard to see, if AI indeed already speeds up progress in mathematics in general - apart from a number of spectacular results, (https://mathoverflow.net/questions/502120/examples-for-the-u...) - or if the quality of the large majority of these "proofs" is so poor that reviewing them is a waste of time.
I would be grateful, if someone familiar with the situation could say a word about this.
As for the proof in question - I'm not sure the author himself can firmly state he understands every detail of what he presented. In a way it's amusing that people who already enjoy a certain level of recognition outside the field are now using it to find (via shiny websites) reviewers for proofs they have worked out as a hobby using Chat GPT and the like.
What we definitely need is more recognition for those who possess the competence and energy to assess the correctness and relevance of such "results".
> energy to assess the correctness and relevance of such "results".
The author claims to have a proof in Lean. Checking the correctness of the result is now reduced to checking the definition of the constant.
> reduced to checking the definition of the constant
But this doesn't mean the remaining task is small or doesn't ask much from the reviewer, does it?! Otherwise we'd see a considerable turn-out of new findings on https://palomar-registry.org/ or https://github.com/Vilin97/lean-pool , or not?! (A considerable number of the presented proofs claim new results, as far as I can tell.)
In the new world of AI, capable QA will rule the world.
I'm not able to judge this for content, but it feels frustrating to read something so obviously Claude-authored as being written "by" Jude Gomila. It's interesting to see the (presumed correct) reduction in this bound and this presentation is pretty compelling. I learned a bit about the Polymath 15 project from this post and I enjoyed that a lot.
That said, I have no idea what contribution may have been done by Mr. Gomila or not. I don't know if I can ask him for more details. Even if he wrote the whole thing (and either has deeply internalized Claude's voice or had it edited and rewritten by an agent) it's difficult to believe that he's the sort of expert I'd expect of someone who had actually done all this work.
This is basically the same feeling I have about vibe code contributors.
So, maybe this is not quite a fair analogy but a lot of famous modern artists don't do a lot of the actual execution of their work. If Jeff Koons relies on a team of artisans and technicians to make a balloon dog, and if Dale Chihuly depends on a workshop of glass blowers to make a chandelier, and if Paul Cummins used a team of ceramicists to make poppies, what did the headlining artist contribute? Not everything, certainly, but also not nothing. Ideas, vision, feedback, guidance?
I think we're used to the idea in math of generally every idea coming from a specific contributing human. In code, maybe less so, but still we generally we find it useful to attribute at least _lines_ to who touched them last. But maybe, if Claude were just less annoying in the prose it made, and could better imitate an actual human, we'd be able to get used to the idea of humans claiming credit for something that the human guided the software to create?
Personally I found it pleasant to read. Why does it matter if/how much AI was used to write it if the end result is great?
I personally didn't
it starts with concentrated dose of ai-ness, combined with non-human website design
sure, it may become better the longer it goes, but it brings the worst to the front
He posted it. That’s some of the work you question. Presumably he queried the AI and led it in some way. That’s also work.
As far as the exact breakdown we will never know. But questioning the authorship in a way that leads us to believe that the author did nothing is wrong.
Nice, is this the maths equivalent of AI slop?
It's really fun seeing Manim [1] used to illustrate a proof!
(See the video at the bottom)
[1]: https://github.com/3b1b/manim
Apparently the unfiltered Claude voice for mathematics is pretty similar to programming.
10.6% smaller. Not bad.
https://en.wikipedia.org/wiki/De_Bruijn%E2%80%93Newman_const...
There's a lot of 10% improvements between 0.1 and 0.
Λ ≤ 0 in the intro. Great site!
It's not wrong, is it? I didn't want to trust myself, Wikipedia agrees that the Riemann hypothesis is equivalent to Λ = 0 (because the lower bound is now 0). This proof lowers the upper bound, but we still don't know if it's 0 or between 0 and that upper bound (inclusive).
Wow, I have a hard time swallowing these Claude generated blogs. So many off-putting llmisms. I wonder if this is something I just need to accept and get used to, or whether it'll pass as models advance.
The worst part is how in a decade, newer generations will write like this having been exposed to it their whole life.
Every sentence delivered like this.
With overdramatic pacing.
Forever.