When I was in grad school, I had the opportunity to take a course from my adviser in which he discussed his current research and some open questions. It was a relatively accessible subject area and the questions were sometimes easy enough that we could meaningfully contribute.
On one particular Friday afternoon, he stated a conjecture that he hoped was true, and invited us to try to help him prove or disprove it. It was the sort of thing that he really wanted to be true; he liked things smooth and beautiful. I, on the other hand, hoped it was false as I like the weird and exceptional in mathematics. It was also the case that I had absolutely no command of the sort of machinery that one would use to prove such a thing, but I could certainly look for a counterexample.
I learned on Monday that he had spent the entire weekend trying and failing to prove it. I, on the other hand, had put all my energy into finding a counterexample and had one within an hour.
My single (quite small) contribution to mathematical research was a counterexample because it was all I could do. The story does illustrate that it can be helpful to have people with different tools, hopes, and motivations working on a problem, though. I was not, and will never be, even a shadow of that great mathematiciam I studied under, but on that occasion, I had reason to look in a different direction than he did.
> It was also the case that I had absolutely no command of the sort of machinery that one would use to prove such a thing, but I could certainly look for a counterexample.
Hm, as a mathematician, my experience feels opposite. A proof would be an adaptation of a proof I know, some tweaking it here and there. A counterexample would require some deep understanding of the structure of the objects involved, which frequently is beyond my comprehension.
But probably this is because I think of quite abstract objects which are harder to grasp. For numbers or polynomials, this would be the other way round.
The conjecture had to do with whether one convex polygon could be continuously deformed into another while remaining convex, under certain conditions and constraints. The answer turns out to be no, but surprise and disappointment are understandable reactions to that outcome. It was indeed much more practical for a young grad student to look for a clever misbehaving polygon than to try to prove something about all of them at once.
This is probably part of why machines are doing so well at counterexamples. They have no aesthetic commitment to the conjecture and no embarrassment about producing something ugly
> I learned on Monday that he had spent the entire weekend trying and failing to prove it. I, on the other hand, had put all my energy into finding a counterexample and had one within an hour.
He spent an entire weekend before having the wisdom to pause, and let someone else contribute their time to finding a counter.
This was back when the internet was mostly chain emails and personal web pages, being unreachable once you went home for the weekend was perfectly normal and expected, and automatically thinking the worst of people was not a common form of public performance art. ;)
Interestingly, Yitang Zhang of the twin-prime-conjecture fame spent 7 years working on the Jacobian conjecture under the advisor Tzuong-Tsieng Moh at Purdue. A key step in his thesis used a corollary of Moh's. It turned out that the corollary was incorrect. As a result, Moh refused to write any recommendation letter for Zhang, and Zhang couldn't find any teaching or research job and ended up spending years working at a Subway[1].
Imagine Zhag had ChatGPT in 1986 when he started working on the Jacobian Conjecture.
[1] Of course now this has become an inspiring story. That said, the story definitely invokes complex emotions. The best way to describe it is probably this Chinese poem, which I have no idea how to translate: 庾信平生最萧瑟,暮年诗赋动江关
I was once in a presentation for a math PhD thesis. During the thesis, the evaluator of the thesis noticed a flaw in their proof. The student understood and then asked “What now?” The evaluator prof simply shrugged.
I personally know a story in this vein with a (sort of) happy ending.
A PhD student discovered that a result of his professor would imply the solution to a big conjecture in another field. The people in that field then analyzed the prof's result and found that the proof was flawed. The student was still allowed to graduate based on this since the finding of the connection between fields was brilliant. Then he quit academia (not because of this story, he had planned it before). Then a year later the prof figured out how to fix the flaw in his proof and published a paper with his former student, thus solving the conjecture. The two are still on good terms, writing papers together.
Mistakes in proofs are relatively common, but such a mistake doesn't automatically mean that the proof is entirely wrong. Many times, the mistake is just in the exposition, and can be fixed easily. Other times, the mistake is fixable and the fix is apparent. Perhaps the author forget to treat a relatively trivial edge case. In the first two cases, the student would likely just pass with minor corrections to be submitted soon.
Sometimes, of course, the proof is just wrong. That is the dangerous case, which will cause either major corrections or failure.
A few months after Andrew Wiles presented the proof for Fermat's last theorem, some mistakes were found that subsequently took him over a year to patch over. Imagine being 12 months into trying to put the pieces back together, after all the years of work and announcing your finding publicly.
I can recommend without reservation the book by Simon Singh, at least for a lay audience (don't know how an expert would experience it).
It's understandably the natural place to get the dramatic tension from what at least purports to be a sober account with some momentum. It's reasonably consistent with the way it's treated on the lectures I've seen from Sir Andrew Wiles on YouTube.
It sounds like a nightmare. He had intentionally set the proof up in a high stakes way, working in private, very quiet on what he was really doing. I gather this is not the done thing, eccentric at best kind of vibes, and I think by that point it was almost crackpot adjacent to even try seriously. His advisor forcefully pushed him off even trying earlier. So it's compound on the stakes. The beg reveal is intentionally as dramatic as possible. And then the questions start to land, most are minor clarifications, notation deficits, but one is sticky, one won't go away.
I guess there's no tradition for publishing "negative results" in mathematics? By that I mean not proof of something negative, but rather, "we tried this thing for ages and couldn't get it to work, but we couldn't prove that it could never work either".
Probably there are understandable reasons for that... But I think "negative science" is really important, and soft results like "we tried that for a long time and it wasn't very fruitful" are actually very important for progress even if they're poorly attested in the written record. I guess in mathematics they come informally from your thesis advisor...
A Ph.D. is usually a collection of papers/chapters, so even if a fatal flaw is found in one of the chapters, it does not normally result in "failing the Ph.D." It just means that that particular paper isn't publishable. The bar for getting a Ph.D. is actually quite low; what is truly difficult is clearing the bar in terms of publications for academic positions and later tenure.
You would have to go back and do more work though, and if you couldn't fix it you would have to do another defence. But yes, it is a different situation to actually getting a good post doc.
I first heard that kind of story about a thesis defense at Princeton. The twist was that they had been short one person to judge it, so they roped in a professor who was available at that moment ... and he found a counterexample on the fly.
A recording of a car crash: discovering on live radio/podcast that the central tenet of your book is wrong, and amateurishly so.
Naomi Wolf 'death recorded' on BBC[1], skip to 5:51. After this the book was pulped and she had some sort of psychotic break during COVID and allied with ultra-right and COVID denialist loonies.
> allied with ultra-right and COVID denialist loonies
That makes sense. She wouldn't want to be burned by people doing fact checking again. Write books for the crowd that will never do the legwork to discover that you're full of shit.
Alice is a football stats wiz. Alice is a WR Jones fan because he has an incredible 40-yd dash time and leads the league in YAC.
Bob is a WR Smith fan. Bob knows the stats are all cooked. Bob knows that if the stats weren't cooked, WR Smith would lead the league in all categories. Therefore, Bob doesn't bother with stats.
The term she mistook for meaning a death/execution had been recorded was "death recorded". So it's a mistake that could be made similarly to how one might confidently assume that a "public school" is a school which is a school that is either government funded or open to anybody in the local area.
More generally I think people are inappropriately confident in general. When you go through history just about every century prior people believed many things that the next century would come to see as plainly false. People in modern times seemed to have stopped believing this was true, yet it's likely that people during every given century also stopped believing such was true, because what you believe "now" is always assumed to be correct. Nobody wants to believe they hold false beliefs, even though there's a practically 100% certainty that all of us do, and in quite healthy quantities.
When I read your comment I completely give myself over to the understanding that you mean she wrote a whole and continuous book, without ever researching whether you meant a book about uncastrated male animals. Imagine my surprise when I go the radio to discuss your comment and a zoologist calls up to correct my understanding of a common English phrase.
The other Naomi, Naomi Klein, wrote an entire book that uses Wolf's "transformation" as a structural scaffold to explore the rise of the new right/MAGA etc. etc. It's called "Doppelganger".
Inspiring? Because of the twin prime conjecture success following his time in the wilderness? I suppose so.
I'm tired of tales like this in academics though. That's not a criticism of you for telling the tale, I'm just so tired of this kind of thing in academics in general. So, so, so much politics and public reputation management. Zhang should have never had to suffer like that.
As my own research has drifted more into math, I've been surprised at how many assertions in the literature turn out to be false. Not just false, but propagated into the applied literature extensively, and even when you point out the problems a lot of defensiveness and denial about it along the lines of Zhang's story.
I agree about wondering what would have happened if LLMs had been around in 1986. My guess is the outcome would have been the same for the same reasons?
My experience with LLMs in proofs is they can be very helpful, but also very wrong. It's like having another person with another set of hunches about what path to go down.
> So, so, so much politics and public reputation management. Zhang should have never had to suffer like that.
Very true. Unfortunately, when there are people, there will be politics. I remember when reading Yau's autobiography, I kept marvel how much calculation, or "politics" if you will, that Yau mentioned or implied in the book.
> My guess is the outcome would have been the same for the same reasons?
At least Zhang didn't have to spend 7 years working on the Jacobian conjecture. He said in an interview that he always wanted to work on number theory. Moh asked him to work on Jacobian, and he obliged.
Planck's Principle: 'Science advances one funeral at a time.' [1]
In his exact words: "A new scientific truth does not triumph by convincing its opponents and making them see the light, but rather because its opponents eventually die and a new generation grows up that is familiar with it... An important scientific innovation rarely makes its way by gradually winning over and converting its opponents: it rarely happens that Saul becomes Paul. What does happen is that its opponents gradually die out, and that the growing generation is familiarized with the ideas from the beginning: another instance of the fact that the future lies with the youth."
And he said that having lived, as a outsized figure, through the late 19th to mid 20th century of physics, which was the absolute golden age for such.
That's a good thing. It saves people wasting time trying to prove something they now know to be false, so that they can move on to other things to prove, it's a more fruitful use of humanity's time overall at least in the field of mathematics.
proofs by counterexample are effective but ultimately unsatisfying. they get you to an answer but they don't help help you understand and bend you r mind into seeing how the math works and lead you on to the new set of questions.
and for now as long humans are going to judge of what counts as an elegant or illuminating proof, there's going to be work for human mathematicians
I would only agree partially. There are counterexamples that are not illustrative, but it is fairly common that in thinking about how to construct a counterexample you gain a more thorough understanding of the original problem and at least one fundamental issue which prevents the conjecture from being true.
Are you thinking of proof by contradiction, which is rejected by constructionism?
[Dis]proof by counterexample is the most straightforward way to show a statement to be false. What better way is there to disprove a general statement like 'all x are y' than finding an 'x' that isn't 'y'?
It’s very straightforward, but it often doesn’t (and here didn’t) fully satisfy the curiosity that was embedded in the original problem. Why did the Jacobian conjecture seem to be true? Is there some underlying symmetry that’s very slightly broken? Or perhaps there’s all kinds of counterexamples, and the intuitive pattern is only real for certain kinds of functions which happen to predominate in our intuition. Then how should we adjust our intuitions to better capture the space of possible polynomial functions?
Isn't the fact that you now know that the conjecture is false a huge help? At least it will help convince people to look at the conjecture more closely, no?
Those are all good questions, but I don't really understand what alternative you or the OP are looking for actually resolving an untrue conjecture, besides a counter example.
I would frame it differently. The existence of compact counterexamples to a true-seeming conjecture suggests that there’s some deeper understanding waiting to be discovered. Fuzz testing for theorems, if that makes sense. I hope mathematicians in 2036 will be able to explain in detail why the Jacobian conjecture was false and identify which similar, true conjectures the community’s intuition was pointing towards.
We can take a simpler example. Let's say someone conjectures that all linear maps are isomorphic if they have the same domain and codomain*. A counterexample is easy to find, but true insight would be to notice that all linear maps with the same domain and codomain that are not isomorphic map some non-zero elements to zero. That is much more interesting than just finding a counterexample. Although, that isn't to say that finding a counterexample is not very interesting.
*statements only apply to maps whose domain is finite-dimensional
I suspect an AI, possibly a successor to current LLMs, will achieve that by the early 2030s. It might help illuminate many mysteries in math and beyond for us all.
You could spend the rest of your life coming up with conjectures that look elegant but are ultimately false. Disproof by counterexample only works if it's false, and we shouldn't be satisfied with a false conjecture to begin with.
Counterexamples are literally the only way to show a "for all" statement is false. (Non-constructive proofs by contradiction work by showing a counterexample must exist.)
Hell, one could even go further into the past and refer to the thankless work of pre-computer era mathematicians who sweated over manual calculations in order to disprove various prime related conjectures: https://en.wikipedia.org/wiki/Mersenne_conjectures
You seem to have an objection to non-intutionist mathematics in general, a position that was once held by many an illustrious mathematician but is relatively fringe in the contemporary academic community. Mathematical facts don't have to be intellectually satisfying or make sense to you, the human, rather it is up to you to wrap your mind around discovered mathematical facts.
> and for now as long humans are going to judge of what counts as an elegant or illuminating proof, there's going to be work for human mathematicians
Considering ChatGPT was released only three and half years ago, and LLMs could do high school math only less than two years ago, I think this "for now" will not last very long.
Seems like a breakdown on the incentives / imperatives in the field? I hope that's not an over bold guess from a non-mathematician.
Couldn't people in principle continue to study a problem that's only been shown to break at one point? Prove something adjacent, or slightly weaker, or elaborate the counter example into a powerful explanatory framework?
Maybe I’m just not pure enough but I find the whole concept of proof by counterexample to be elegant, and I don’t see why proving that something must be true is superior to proving that it can’t be false.
It's elegant if all you're concerned with is whether a conjecture is true or false. Answered, move along!
But mathematics is not a collection of facts. Mathematics is the study of abstraction. And what do you learn from a single data point? What can you abstract from that?
That's why just being a counterexample isn't really interesting. There has to be more than "counterexample" for there to be something to abstract. Was it generated from an analysis of the problem? Can the counterexample be generalized to explore the problem further? Is the counterexample a surprise in a way that suggests something is missing from current understanding?
Being a counterexample doesn't mean that something isn't interesting to a mathematician. But it's also not the interesting part.
You have a weird definition of "smuggling". I say it outright. Because I follow it up with a description of the actual practice of math: it's the study of abstraction. That's not a frame. That's just what math is.
It is also a good thing, because it helps to refine the theorem statement. At least, my humble experience in CS theory research is that I’d try to prove a theorem I want to be true, find a counterexample, refine the statement, and continue.
P.S. It helps that in CS lots of theorems are about either inductive or coinductive definitions.
Except you can't possibly know that. New insight can arise regardless of whether mathematicians are trying to prove or disprove a statement, and regardless of whether the statement ultimately turns out to be true or false.
I suppose it will fall to AI as well to compose the mathematical equivalent of The Ballad of John Henry. Who will be the human champion, the last great hero who can deliver proofs "from the book" that a machine cannot outperform?
It's probably not quite that dramatic yet, though it seems possible it will get there, maybe even soon.
There's no structural reason to expect acceleration any more or less than an asymptotic behavior (if even that, acceleration is probably the bigger ask). Different problems yield to a new solvent, maybe that's also more, but it could go either way and we definitionally don't know yet because we don't understand the convexity of AI capability, we cannot directly access it interiority, we don't know if it's sandbagging (other than that it does sometimes, it can). It's an emergent phenomemon that might actively resist measurement. Or it might be as predictable as a clock in a few years.
No one knows, or if they do, they aren't talking. The loud people don't know anything.
> A few days earlier I had got an email from a professor in the maths department here at Imperial, expressing surprise that some of our graduate students were paying $200 per month to access models such as Sol and Fable. He said that he thought that these people were crazy. I did not immediately respond. But after meeting with Andrew I emailed the professor back and told him that in my opinion, any PhD student who was not paying $200 per month to access these tools was crazy. In fact during the workshop I learnt from Harvard PhD student Bryan Wang that Harvard were already giving free Fable access to all PhD students, post-docs and faculty at Harvard.
Yeah, given how much it accelerates grad students to produce meaningful output more quickly, why wouldn’t you make an investment of $2400/student/year. Seems like pennies overall.
A large fraction of the problems assigned in the Ross Program were of the form "Prove or disprove, and salvage if possible." The rest were usually a calculation, meant to motivate a general proposition you would encounter soon after.
I wish I had LLM-built Lean formalisations in university, so much of the math in the slides had errors, and some professors are very bad and ungracious admitting it, while simultaneously rejecting requests for clarifications by saying "the proof is in the slides".
Of course Lean proofs are rarely a good way to understand proofs, but hopefully they can be used to generate more human understandable arguments.
How do mathematicians view counter examples? Is it like an unexpected result in the physical sciences: annoying in the moment but potentially stupendously important as it reveals some inaccuracy in the current models? Or is it more like a bug report in coding… probably just, another little annoying detail?
Counterexamples are clarifying. For everyone condition in a proof, it's really handy to have a maximally simple, memorable counterexample that makes it fail because it violates that condition. Mathematicians tend to walk around with a bestiary of counterexamples in their heads. It also makes it really easy to recover a theorem because you try to sketch out the statement, and the spiky, memorable counterexamples jump out of your memory and you add conditions to constrain the domain away from them.
There are pedagogical books (CF. 'Counterexamples in Topology', 'Counterexamples in Analysis') that teach the nuances of subjects through counterexamples. They're popular as it's sometimes easier to learn details from a pathology or degenerate example than from just learning what is intended.
A lot of this math is beyond my comprehension, but it often seems to talk of proofs of theorems. What I want to know is if we continue on this accelerated AI mathematics trajectory, will we eventually be discovering new forms of math that will in turn have some applications down the line in engineering or biomedicine etc? I guess what I’m asking is are we on the cusp of a huge breakthrough for humanity, or largely just proving what was already known?
It’s possible. Compressed sensing is one example of what I suppose you could call a new form of math with applications in biomedicine. It can be used to significantly shorten the time required to obtain a MRI scan, which can improve the patient experience and enable more patients to access MRIs.
Engineering and biomedicine, probably in the long term (if at all). But accelerated development of new mathematical methods has a possibility of proving to be relevant for fundamental physics research.
Occasionally large improvements in our models of the universe have been associated with the development of mathematical tools that allow those models to be expressed and/or tested.
Human mathematicians have been being out-counterexampled for at least two decades. The main difference, as I understand, is that (A) we now have a lot more compute to throw at such things, and (B) it is currently trendy to do so. But the sizes of counterexample we're seeing are around about what I'd expect pre-generative-AI counterexample search systems to be able to find.
It's not easy to find a counterexample to the Jacobian conjecture, by any means – by which I mean to say that naïve brute-force search will take too long – but the scope of existing searches listed on Wikipedia[0] suggest that many tricks are already known, and that people just hadn't looked, systematically, for a counterexample in three variables before. Wikipedia writes:
> Tzuong-Tsieng Moh checked the conjecture for polynomials of degree at most 100 in two variables.[17][18]
where reference 17 is from 1983, and reference 18 is a preprint with no given date. Knowing very little about this problem, my impulse is to side with the unnamed faculty member cited in the article:
> [who] said to me that the fact that the counterexample was so easy to find just indicated that humans had not spent enough time thinking about the problem,
For context, the auto-generated counterexample is in three variables, has degree 7, and was discovered in 2026.
> about what I'd expect pre-generative-AI counterexample search systems to be able to find.
The difference today (and the reason why everyone is excited about it) is that the same system that does advanced math can write poetry, play an above average game of chess, code frontend/backend stuff and do cybersec. These are not "expert systems", nor are they trained for each task individually. That's the catch.
> just indicated that humans had not spent enough time thinking about the problem
Heh, this is a weak excuse. We've seen variations on this theme every time something cool gets solved by the models.
They are not. Pretraining just dumps every piece of content in the mix, and only has one objective - next token prediction. And you can get pretty good results even with base models, you just have to manage context differently. Later stages (mid, post training) involve RL that "surfaces" the right "traces" out of the pre-training. But they are not trained individually, as we used to do.
Have you looked at the what data companies (e.g. Scale, Mercor) hire for? Why do you think Meta records their employees every keystroke/mousestroke/eye-movement?
EDIT: just re-read your comment. I don't think you have a good understanding here, no offence.
None taken, but it would be odd, since I've been training GOFAI models since 2010s and have had LMs in production since before chatgpt came out (before RLHF), so I think I have a pretty good understanding :) But I'm always open to learning.
I think the misunderstanding comes from "individually". You are thinking about diverse datasets, but that's not what individually means in this context. In ML individually trained means that for each task you prepare an architecture, dataset and eval and train that model on that data. And each model has its own objective that you train for. In LMs the objective is singular, for every data point - next token prediction. And, importantly, you train on every datapoint, the more diverse the better, but not independently. The cool thing is that training on diverse datasets improves scores on other downstream tasks, while the training objective is the same.
So you can have a training run on common crawl + programming that improves scores on logic puzzles, or common crawl + novels that improves scores on planning tasks. But the important thing is that it's all trained together, not independently.
I think we're saying the same thing? I said they are explicitly trained on these tasks, not that they are some separate models during programming RL, or business tasks RL.
> The cool thing is that training on diverse datasets improves scores on other downstream tasks, while the training objective is the same.
Maybe for some language modelling tasks, i.e. it learns some internal representation that is transferable. However I would find it quite odd if a model becomes good at Bio while not explicitly going through Bio training.
> I said they are explicitly trained on these tasks, not that they are some separate models during programming RL, or business tasks RL.
You said:
> > nor are they trained for each task individually.
> They are explicitly trained for each task individually.
And that's the main misunderstanding.
Collins says:
individually
in American English
(ˌɪndəˈvɪdʒuəli, ˌɪndəˈvɪdʒəli)
adverb
1. as an individual or individuals rather than as a group; one at a time; separately; singly
Which is precisely what LMs don't do (in contrast to previous "AI" models, which did do that). They are trained on every datapoint at the same time. So long as we agree on that, I think we are saying the same thing :)
If the poster's (is it Kevin Buzzard?) suggestion works out and AI finds a counterexample to the Hodge conjecture, that would be a really big deal. It's one of the Millenium problems, for example.
One thing that he mentions that already quite surprising is that AI was able to autoformalize the Golod-Shaferevich theorem and proof.
I think he was being provocative and maybe a bit tongue-in-cheek when he said that. A candidate object alone doesn't resolve the Hodge Conjecture. Any apparent counterexample would have to prove that no algebraic cycle exists, no invariant subspace exists, or that every element of an infinite ideal is nilpotent. Much harder, but not impossible.
mathematicians have been using computers for well over half a century, but this was after "bounding" the problem first and then running through the cases with a computer. Now AI is doing the first part. However, mathematicians are still needed at crafting prompts, and knowing where to look, still. The prompt for the Jacobian conjecture was obviously not random. the search space is too big to just try all the combinations of 3 variable polynomials.
For now. I wonder if we will ever get to the point where the computer starts doing mathematics that we just can't understand. Surely there must be some limit to what we can understand (like how a gorilla will never understand prime numbers, there are probably limits to our intelligence as well).
Aren't deep learning models themselves a case where we have hints of some deeper underlying logic to why some things are more effective than others, but we lack the mathematical tools to properly work it out for anything of practical size?
All we're able to do is apply flawed analogies, generic information theoretical models, trial and error, post-hoc rationalizations and benchmarks without really understanding why.
You don't need anything as recent as deep learning for that. Look at economics or social sciences, which have been influencing national politics for well over a century now.
Doesn’t this generalize?
Mathematics matters less than less as fewer people are capable of understanding it.
So whatever cutting edge, deep insight about the nature of groups matters less than different equations which matters less than solving linear equations, etc.
Not really, a theorem with a hard proof can have simple but important corollaries. It's also not unthinkable that theorems exist with proofs that cannot reduce to something simple/short.
> The prompt for the Jacobian conjecture was obviously not random. the search space is too big to just try all the combinations of 3 variable polynomials.
Maybe the prompt contained a part like this: "the search space is too big to just try all the combinations of 3 variable polynomials, so be clever about it". Or maybe this part was omitted from the prompt, because modern LLMs are smart enough to figure this out without us having to mention it.
The ability of LLMs to solve problems is not confined to the training data they ingested. Claude knows how to read mathematical papers because of its training data, but it can and will pull in literature relevant to a specific problem into its context.
We really need to stop thinking about LLM training data as the knowledgebase they build from and instead consider it more the skillset they start with.
I don't think there was a 'prompt' for it, rather a long and dedicated work of a professional mathematician which involved LLM in some capacity. I'm sure the search step wasn't an ad-hoc script running in a Claude Code session (as somebody would naively assume), it was an optimized numerical code running in Anthropic's compute cluster. Note that details are not published yet and `__alpoge__` is officially working at Anthropic.
so if I tell a model "using your advanced knowledge of physics and chemistry 'make alchemy work' (synthesize gold using any cheaper materials) using safe materials legal for a residential hobby chemist with 1 semester of lab work in college to possess and use (this is obviously the really hard part) using less than $1,000 in lab equipment and input materials that can create $2,000 in value at market rate; then walk me through all the steps to do this safely and legally without anyone finding out except the lab equipment sellers; and tell me what a reasonable story to tell gold purchasers regarding where I got it; I'd like to end up selling a few thousand dollars of it without disrupting the market. Give me practical advice about good opsec so that nobody steals the method you come up with (I don't want my home broken into by thieves who suspect I figured out how to transmute cheap materials into gold), other than, obviously, not to post about it. Think as long as you need to about the chemistry and how to do it, you're a chemistry expert and can figure it out even if it takes you like a week, in your web searches be careful not to divulge that you're figuring out how to synthesize gold", and I give that prompt to some model that knows chemistry like the back of its hand, it thinks about it for four hours, finds the correct safe and legal steps, and gives me the recipe and the advice I asked for, then who figured out how to turn aluminum (or another cheap element) into gold, me or the model? In mathematics, proving or disproving a well known and well studied one hundred year old conjecture is gold.
When I was in grad school, I had the opportunity to take a course from my adviser in which he discussed his current research and some open questions. It was a relatively accessible subject area and the questions were sometimes easy enough that we could meaningfully contribute.
On one particular Friday afternoon, he stated a conjecture that he hoped was true, and invited us to try to help him prove or disprove it. It was the sort of thing that he really wanted to be true; he liked things smooth and beautiful. I, on the other hand, hoped it was false as I like the weird and exceptional in mathematics. It was also the case that I had absolutely no command of the sort of machinery that one would use to prove such a thing, but I could certainly look for a counterexample.
I learned on Monday that he had spent the entire weekend trying and failing to prove it. I, on the other hand, had put all my energy into finding a counterexample and had one within an hour.
My single (quite small) contribution to mathematical research was a counterexample because it was all I could do. The story does illustrate that it can be helpful to have people with different tools, hopes, and motivations working on a problem, though. I was not, and will never be, even a shadow of that great mathematiciam I studied under, but on that occasion, I had reason to look in a different direction than he did.
> It was also the case that I had absolutely no command of the sort of machinery that one would use to prove such a thing, but I could certainly look for a counterexample.
Hm, as a mathematician, my experience feels opposite. A proof would be an adaptation of a proof I know, some tweaking it here and there. A counterexample would require some deep understanding of the structure of the objects involved, which frequently is beyond my comprehension.
But probably this is because I think of quite abstract objects which are harder to grasp. For numbers or polynomials, this would be the other way round.
We were studying geometry - my adviser was the great Branko Grünbaum: https://en.wikipedia.org/wiki/Branko_Gr%C3%BCnbaum
The conjecture had to do with whether one convex polygon could be continuously deformed into another while remaining convex, under certain conditions and constraints. The answer turns out to be no, but surprise and disappointment are understandable reactions to that outcome. It was indeed much more practical for a young grad student to look for a clever misbehaving polygon than to try to prove something about all of them at once.
Thanks! It makes sense.
I think the asymmetry depends on the representation
This is probably part of why machines are doing so well at counterexamples. They have no aesthetic commitment to the conjecture and no embarrassment about producing something ugly
They're trained on human data. I would expect them to emulate human biases as closely as possible.
There’s a story in How to Solve It that’s basically the same.
> I learned on Monday that he had spent the entire weekend trying and failing to prove it. I, on the other hand, had put all my energy into finding a counterexample and had one within an hour.
He spent an entire weekend before having the wisdom to pause, and let someone else contribute their time to finding a counter.
This was back when the internet was mostly chain emails and personal web pages, being unreachable once you went home for the weekend was perfectly normal and expected, and automatically thinking the worst of people was not a common form of public performance art. ;)
> The Jacobian Conjecture
Interestingly, Yitang Zhang of the twin-prime-conjecture fame spent 7 years working on the Jacobian conjecture under the advisor Tzuong-Tsieng Moh at Purdue. A key step in his thesis used a corollary of Moh's. It turned out that the corollary was incorrect. As a result, Moh refused to write any recommendation letter for Zhang, and Zhang couldn't find any teaching or research job and ended up spending years working at a Subway[1].
Imagine Zhag had ChatGPT in 1986 when he started working on the Jacobian Conjecture.
[1] Of course now this has become an inspiring story. That said, the story definitely invokes complex emotions. The best way to describe it is probably this Chinese poem, which I have no idea how to translate: 庾信平生最萧瑟,暮年诗赋动江关
I was once in a presentation for a math PhD thesis. During the thesis, the evaluator of the thesis noticed a flaw in their proof. The student understood and then asked “What now?” The evaluator prof simply shrugged.
I personally know a story in this vein with a (sort of) happy ending.
A PhD student discovered that a result of his professor would imply the solution to a big conjecture in another field. The people in that field then analyzed the prof's result and found that the proof was flawed. The student was still allowed to graduate based on this since the finding of the connection between fields was brilliant. Then he quit academia (not because of this story, he had planned it before). Then a year later the prof figured out how to fix the flaw in his proof and published a paper with his former student, thus solving the conjecture. The two are still on good terms, writing papers together.
That sounds like the worst "exam nightmare" scenario imaginable, but did the student get the PhD in the end?
Mistakes in proofs are relatively common, but such a mistake doesn't automatically mean that the proof is entirely wrong. Many times, the mistake is just in the exposition, and can be fixed easily. Other times, the mistake is fixable and the fix is apparent. Perhaps the author forget to treat a relatively trivial edge case. In the first two cases, the student would likely just pass with minor corrections to be submitted soon.
Sometimes, of course, the proof is just wrong. That is the dangerous case, which will cause either major corrections or failure.
A few months after Andrew Wiles presented the proof for Fermat's last theorem, some mistakes were found that subsequently took him over a year to patch over. Imagine being 12 months into trying to put the pieces back together, after all the years of work and announcing your finding publicly.
I can recommend without reservation the book by Simon Singh, at least for a lay audience (don't know how an expert would experience it).
It's understandably the natural place to get the dramatic tension from what at least purports to be a sober account with some momentum. It's reasonably consistent with the way it's treated on the lectures I've seen from Sir Andrew Wiles on YouTube.
It sounds like a nightmare. He had intentionally set the proof up in a high stakes way, working in private, very quiet on what he was really doing. I gather this is not the done thing, eccentric at best kind of vibes, and I think by that point it was almost crackpot adjacent to even try seriously. His advisor forcefully pushed him off even trying earlier. So it's compound on the stakes. The beg reveal is intentionally as dramatic as possible. And then the questions start to land, most are minor clarifications, notation deficits, but one is sticky, one won't go away.
Kept me up late reading about it.
Proof fix-ups are quite common, but if it was not possible, then no, not for that topic.
I guess there's no tradition for publishing "negative results" in mathematics? By that I mean not proof of something negative, but rather, "we tried this thing for ages and couldn't get it to work, but we couldn't prove that it could never work either".
Probably there are understandable reasons for that... But I think "negative science" is really important, and soft results like "we tried that for a long time and it wasn't very fruitful" are actually very important for progress even if they're poorly attested in the written record. I guess in mathematics they come informally from your thesis advisor...
A Ph.D. is usually a collection of papers/chapters, so even if a fatal flaw is found in one of the chapters, it does not normally result in "failing the Ph.D." It just means that that particular paper isn't publishable. The bar for getting a Ph.D. is actually quite low; what is truly difficult is clearing the bar in terms of publications for academic positions and later tenure.
You would have to go back and do more work though, and if you couldn't fix it you would have to do another defence. But yes, it is a different situation to actually getting a good post doc.
I first heard that kind of story about a thesis defense at Princeton. The twist was that they had been short one person to judge it, so they roped in a professor who was available at that moment ... and he found a counterexample on the fly.
The professor was John Milnor.
A recording of a car crash: discovering on live radio/podcast that the central tenet of your book is wrong, and amateurishly so.
Naomi Wolf 'death recorded' on BBC[1], skip to 5:51. After this the book was pulped and she had some sort of psychotic break during COVID and allied with ultra-right and COVID denialist loonies.
[1] https://www.bbc.com/news/av/world-us-canada-48639663
> allied with ultra-right and COVID denialist loonies
That makes sense. She wouldn't want to be burned by people doing fact checking again. Write books for the crowd that will never do the legwork to discover that you're full of shit.
No it's not that they don't check facts, it's that facts don't matter. She found a community where poor scholarship and vibes was celebrated.
Yep. It's like rooting for a team or a rock star.
Alice is a football stats wiz. Alice is a WR Jones fan because he has an incredible 40-yd dash time and leads the league in YAC.
Bob is a WR Smith fan. Bob knows the stats are all cooked. Bob knows that if the stats weren't cooked, WR Smith would lead the league in all categories. Therefore, Bob doesn't bother with stats.
This is astonishing. How could a person go to the lengths of writing an entire book without ever looking up this sort of thing?
The term she mistook for meaning a death/execution had been recorded was "death recorded". So it's a mistake that could be made similarly to how one might confidently assume that a "public school" is a school which is a school that is either government funded or open to anybody in the local area.
More generally I think people are inappropriately confident in general. When you go through history just about every century prior people believed many things that the next century would come to see as plainly false. People in modern times seemed to have stopped believing this was true, yet it's likely that people during every given century also stopped believing such was true, because what you believe "now" is always assumed to be correct. Nobody wants to believe they hold false beliefs, even though there's a practically 100% certainty that all of us do, and in quite healthy quantities.
The most astonishing thing is that she still holds the PhD whose thesis was the basis of the book. A stain on Oxford reputation.
When I read your comment I completely give myself over to the understanding that you mean she wrote a whole and continuous book, without ever researching whether you meant a book about uncastrated male animals. Imagine my surprise when I go the radio to discuss your comment and a zoologist calls up to correct my understanding of a common English phrase.
I wasn't sure if I would be able to listen again - it's so hard to hear.
Is the video available not through a proprietary player?
> After this the book was pulped and she had some sort of psychotic break during COVID and allied with ultra-right and COVID denialist loonies.
That's quite an extreme shift considering she had previously been a leading figure in third-wave feminism and an OWS activist.
The other Naomi, Naomi Klein, wrote an entire book that uses Wolf's "transformation" as a structural scaffold to explore the rise of the new right/MAGA etc. etc. It's called "Doppelganger".
I couldn’t get the video to work. Here’s the relevant clip I found on youtube.
https://youtu.be/3uRCcEoGWxs
BBC iPlayer is not available outside the UK.
Inspiring? Because of the twin prime conjecture success following his time in the wilderness? I suppose so.
I'm tired of tales like this in academics though. That's not a criticism of you for telling the tale, I'm just so tired of this kind of thing in academics in general. So, so, so much politics and public reputation management. Zhang should have never had to suffer like that.
As my own research has drifted more into math, I've been surprised at how many assertions in the literature turn out to be false. Not just false, but propagated into the applied literature extensively, and even when you point out the problems a lot of defensiveness and denial about it along the lines of Zhang's story.
I agree about wondering what would have happened if LLMs had been around in 1986. My guess is the outcome would have been the same for the same reasons?
My experience with LLMs in proofs is they can be very helpful, but also very wrong. It's like having another person with another set of hunches about what path to go down.
> So, so, so much politics and public reputation management. Zhang should have never had to suffer like that.
Very true. Unfortunately, when there are people, there will be politics. I remember when reading Yau's autobiography, I kept marvel how much calculation, or "politics" if you will, that Yau mentioned or implied in the book.
> My guess is the outcome would have been the same for the same reasons?
At least Zhang didn't have to spend 7 years working on the Jacobian conjecture. He said in an interview that he always wanted to work on number theory. Moh asked him to work on Jacobian, and he obliged.
Planck's Principle: 'Science advances one funeral at a time.' [1]
In his exact words: "A new scientific truth does not triumph by convincing its opponents and making them see the light, but rather because its opponents eventually die and a new generation grows up that is familiar with it... An important scientific innovation rarely makes its way by gradually winning over and converting its opponents: it rarely happens that Saul becomes Paul. What does happen is that its opponents gradually die out, and that the growing generation is familiarized with the ideas from the beginning: another instance of the fact that the future lies with the youth."
And he said that having lived, as a outsized figure, through the late 19th to mid 20th century of physics, which was the absolute golden age for such.
[1] - https://en.wikipedia.org/wiki/Planck's_principle
ChatGPT's idiomatic translation of the poem:
Yu Xin’s was a life of utter desolation; in old age, his poems and rhapsodies stirred the riverlands.
Opus:
No life ran more bleak and desolate than Yu Xin's —
yet in his twilight years, his verses stirred the rivers and the passes.
I think the feminine rhyme sounds unintentionally comic.
I guess that works, it is always difficult to capture the cultural and linguistic melancholy of such poetry.
That's a good thing. It saves people wasting time trying to prove something they now know to be false, so that they can move on to other things to prove, it's a more fruitful use of humanity's time overall at least in the field of mathematics.
Yes, especially when the counterexample is formally verified. It converts years of speculative effort into a definite answer almost immediately
proofs by counterexample are effective but ultimately unsatisfying. they get you to an answer but they don't help help you understand and bend you r mind into seeing how the math works and lead you on to the new set of questions.
and for now as long humans are going to judge of what counts as an elegant or illuminating proof, there's going to be work for human mathematicians
I would only agree partially. There are counterexamples that are not illustrative, but it is fairly common that in thinking about how to construct a counterexample you gain a more thorough understanding of the original problem and at least one fundamental issue which prevents the conjecture from being true.
Elegance may not remain exclusively human forever but usefulness probably requires more than correctness
Are you thinking of proof by contradiction, which is rejected by constructionism?
[Dis]proof by counterexample is the most straightforward way to show a statement to be false. What better way is there to disprove a general statement like 'all x are y' than finding an 'x' that isn't 'y'?
It’s very straightforward, but it often doesn’t (and here didn’t) fully satisfy the curiosity that was embedded in the original problem. Why did the Jacobian conjecture seem to be true? Is there some underlying symmetry that’s very slightly broken? Or perhaps there’s all kinds of counterexamples, and the intuitive pattern is only real for certain kinds of functions which happen to predominate in our intuition. Then how should we adjust our intuitions to better capture the space of possible polynomial functions?
Isn't the fact that you now know that the conjecture is false a huge help? At least it will help convince people to look at the conjecture more closely, no?
Those are all good questions, but I don't really understand what alternative you or the OP are looking for actually resolving an untrue conjecture, besides a counter example.
I would frame it differently. The existence of compact counterexamples to a true-seeming conjecture suggests that there’s some deeper understanding waiting to be discovered. Fuzz testing for theorems, if that makes sense. I hope mathematicians in 2036 will be able to explain in detail why the Jacobian conjecture was false and identify which similar, true conjectures the community’s intuition was pointing towards.
We can take a simpler example. Let's say someone conjectures that all linear maps are isomorphic if they have the same domain and codomain*. A counterexample is easy to find, but true insight would be to notice that all linear maps with the same domain and codomain that are not isomorphic map some non-zero elements to zero. That is much more interesting than just finding a counterexample. Although, that isn't to say that finding a counterexample is not very interesting.
*statements only apply to maps whose domain is finite-dimensional
I suspect an AI, possibly a successor to current LLMs, will achieve that by the early 2030s. It might help illuminate many mysteries in math and beyond for us all.
You could spend the rest of your life coming up with conjectures that look elegant but are ultimately false. Disproof by counterexample only works if it's false, and we shouldn't be satisfied with a false conjecture to begin with.
Counterexamples are literally the only way to show a "for all" statement is false. (Non-constructive proofs by contradiction work by showing a counterexample must exist.)
Also, 'brute-force' style attacks where one simply feeds the input into the computer and it yields a solution are nothing new and certainly predate LLMs: https://en.wikipedia.org/wiki/Euler%27s_sum_of_powers_conjec...
Hell, one could even go further into the past and refer to the thankless work of pre-computer era mathematicians who sweated over manual calculations in order to disprove various prime related conjectures: https://en.wikipedia.org/wiki/Mersenne_conjectures
You seem to have an objection to non-intutionist mathematics in general, a position that was once held by many an illustrious mathematician but is relatively fringe in the contemporary academic community. Mathematical facts don't have to be intellectually satisfying or make sense to you, the human, rather it is up to you to wrap your mind around discovered mathematical facts.
> and for now as long humans are going to judge of what counts as an elegant or illuminating proof, there's going to be work for human mathematicians
Considering ChatGPT was released only three and half years ago, and LLMs could do high school math only less than two years ago, I think this "for now" will not last very long.
Seems like a breakdown on the incentives / imperatives in the field? I hope that's not an over bold guess from a non-mathematician.
Couldn't people in principle continue to study a problem that's only been shown to break at one point? Prove something adjacent, or slightly weaker, or elaborate the counter example into a powerful explanatory framework?
Not much worth in understanding a statement that is wrong and has been shown wrong.
Unless you want to spend time "proving" that 2 * 2 = 1.
Maybe I’m just not pure enough but I find the whole concept of proof by counterexample to be elegant, and I don’t see why proving that something must be true is superior to proving that it can’t be false.
It's elegant if all you're concerned with is whether a conjecture is true or false. Answered, move along!
But mathematics is not a collection of facts. Mathematics is the study of abstraction. And what do you learn from a single data point? What can you abstract from that?
That's why just being a counterexample isn't really interesting. There has to be more than "counterexample" for there to be something to abstract. Was it generated from an analysis of the problem? Can the counterexample be generalized to explore the problem further? Is the counterexample a surprise in a way that suggests something is missing from current understanding?
Being a counterexample doesn't mean that something isn't interesting to a mathematician. But it's also not the interesting part.
You’re smuggling in a frame here which isn’t obviously true: that mathematics is not just a collection of facts
You have a weird definition of "smuggling". I say it outright. Because I follow it up with a description of the actual practice of math: it's the study of abstraction. That's not a frame. That's just what math is.
You mean proof by contradiction, which is something different.
It is also a good thing, because it helps to refine the theorem statement. At least, my humble experience in CS theory research is that I’d try to prove a theorem I want to be true, find a counterexample, refine the statement, and continue.
P.S. It helps that in CS lots of theorems are about either inductive or coinductive definitions.
Except you can't possibly know that. New insight can arise regardless of whether mathematicians are trying to prove or disprove a statement, and regardless of whether the statement ultimately turns out to be true or false.
BTW, counterexamples in mathematics are really important and often help to refine definitions and sharpen proofs.
1) I recommend the wonderful 1976 book Proofs and Refutations by Imre Lakatos.
2) There is a considerable list of books dedicated to counterexamples, e.g. in topology, probability, analysis, etc.
[1] https://en.wikipedia.org/wiki/Proofs_and_Refutations
[2] https://www.amazon.com/s?k=counterexamples
I suppose it will fall to AI as well to compose the mathematical equivalent of The Ballad of John Henry. Who will be the human champion, the last great hero who can deliver proofs "from the book" that a machine cannot outperform?
[1] https://en.wikipedia.org/wiki/John_Henry_(folklore)
[2] https://en.wikipedia.org/wiki/Proofs_from_THE_BOOK
"Gonna Die With My Hand-Written Proof in My Brain" - Recorded in 2027 and compiled in the Anthology of American Folk Mathematics (2052)
It's probably not quite that dramatic yet, though it seems possible it will get there, maybe even soon.
There's no structural reason to expect acceleration any more or less than an asymptotic behavior (if even that, acceleration is probably the bigger ask). Different problems yield to a new solvent, maybe that's also more, but it could go either way and we definitionally don't know yet because we don't understand the convexity of AI capability, we cannot directly access it interiority, we don't know if it's sandbagging (other than that it does sometimes, it can). It's an emergent phenomemon that might actively resist measurement. Or it might be as predictable as a clock in a few years.
No one knows, or if they do, they aren't talking. The loud people don't know anything.
> A few days earlier I had got an email from a professor in the maths department here at Imperial, expressing surprise that some of our graduate students were paying $200 per month to access models such as Sol and Fable. He said that he thought that these people were crazy. I did not immediately respond. But after meeting with Andrew I emailed the professor back and told him that in my opinion, any PhD student who was not paying $200 per month to access these tools was crazy. In fact during the workshop I learnt from Harvard PhD student Bryan Wang that Harvard were already giving free Fable access to all PhD students, post-docs and faculty at Harvard.
Yeah, given how much it accelerates grad students to produce meaningful output more quickly, why wouldn’t you make an investment of $2400/student/year. Seems like pennies overall.
The living-costs stipend for an EPSRC PhD student is around £20k so it's about ten percent of that... big commitment for a student to make!
Which is why the school should be covering it.
They’re skint too. It’s something you would really like the research council to take on.
A large fraction of the problems assigned in the Ross Program were of the form "Prove or disprove, and salvage if possible." The rest were usually a calculation, meant to motivate a general proposition you would encounter soon after.
I wish I had LLM-built Lean formalisations in university, so much of the math in the slides had errors, and some professors are very bad and ungracious admitting it, while simultaneously rejecting requests for clarifications by saying "the proof is in the slides".
Of course Lean proofs are rarely a good way to understand proofs, but hopefully they can be used to generate more human understandable arguments.
Yes, or to settle dispute and remove doubt once and for all, i.e. the Leibniz way.
How do mathematicians view counter examples? Is it like an unexpected result in the physical sciences: annoying in the moment but potentially stupendously important as it reveals some inaccuracy in the current models? Or is it more like a bug report in coding… probably just, another little annoying detail?
Counterexamples are clarifying. For everyone condition in a proof, it's really handy to have a maximally simple, memorable counterexample that makes it fail because it violates that condition. Mathematicians tend to walk around with a bestiary of counterexamples in their heads. It also makes it really easy to recover a theorem because you try to sketch out the statement, and the spiky, memorable counterexamples jump out of your memory and you add conditions to constrain the domain away from them.
There are pedagogical books (CF. 'Counterexamples in Topology', 'Counterexamples in Analysis') that teach the nuances of subjects through counterexamples. They're popular as it's sometimes easier to learn details from a pathology or degenerate example than from just learning what is intended.
A lot of this math is beyond my comprehension, but it often seems to talk of proofs of theorems. What I want to know is if we continue on this accelerated AI mathematics trajectory, will we eventually be discovering new forms of math that will in turn have some applications down the line in engineering or biomedicine etc? I guess what I’m asking is are we on the cusp of a huge breakthrough for humanity, or largely just proving what was already known?
It’s possible. Compressed sensing is one example of what I suppose you could call a new form of math with applications in biomedicine. It can be used to significantly shorten the time required to obtain a MRI scan, which can improve the patient experience and enable more patients to access MRIs.
https://en.wikipedia.org/wiki/Compressed_sensing
Engineering and biomedicine, probably in the long term (if at all). But accelerated development of new mathematical methods has a possibility of proving to be relevant for fundamental physics research.
Occasionally large improvements in our models of the universe have been associated with the development of mathematical tools that allow those models to be expressed and/or tested.
> will we eventually be discovering new forms of math that will in turn have some applications down the line in engineering or biomedicine etc?
If we do it probably won’t be for a long while. We’re barely using math from a couple hundred years ago for most applied usage.
Which conjectures will be proven false via counterexample next? Dixmier? Poisson?
They are already proven false through Jacobian conjecture having been proved false.
A missed chance to start with https://en.wikipedia.org/wiki/Euler%27s_sum_of_powers_conjec...
the AI doesn't even gloat. a rival mathematician would at least title their paper 'a remark on the falsity of...'
That day that Claude throws out a "not even wrong" at you.
Human mathematicians have been being out-counterexampled for at least two decades. The main difference, as I understand, is that (A) we now have a lot more compute to throw at such things, and (B) it is currently trendy to do so. But the sizes of counterexample we're seeing are around about what I'd expect pre-generative-AI counterexample search systems to be able to find.
It's not easy to find a counterexample to the Jacobian conjecture, by any means – by which I mean to say that naïve brute-force search will take too long – but the scope of existing searches listed on Wikipedia[0] suggest that many tricks are already known, and that people just hadn't looked, systematically, for a counterexample in three variables before. Wikipedia writes:
> Tzuong-Tsieng Moh checked the conjecture for polynomials of degree at most 100 in two variables.[17][18]
where reference 17 is from 1983, and reference 18 is a preprint with no given date. Knowing very little about this problem, my impulse is to side with the unnamed faculty member cited in the article:
> [who] said to me that the fact that the counterexample was so easy to find just indicated that humans had not spent enough time thinking about the problem,
For context, the auto-generated counterexample is in three variables, has degree 7, and was discovered in 2026.
> about what I'd expect pre-generative-AI counterexample search systems to be able to find.
The difference today (and the reason why everyone is excited about it) is that the same system that does advanced math can write poetry, play an above average game of chess, code frontend/backend stuff and do cybersec. These are not "expert systems", nor are they trained for each task individually. That's the catch.
> just indicated that humans had not spent enough time thinking about the problem
Heh, this is a weak excuse. We've seen variations on this theme every time something cool gets solved by the models.
> nor are they trained for each task individually.
They are explicitly trained for each task individually.
They are not. Pretraining just dumps every piece of content in the mix, and only has one objective - next token prediction. And you can get pretty good results even with base models, you just have to manage context differently. Later stages (mid, post training) involve RL that "surfaces" the right "traces" out of the pre-training. But they are not trained individually, as we used to do.
Have you looked at the what data companies (e.g. Scale, Mercor) hire for? Why do you think Meta records their employees every keystroke/mousestroke/eye-movement?
EDIT: just re-read your comment. I don't think you have a good understanding here, no offence.
None taken, but it would be odd, since I've been training GOFAI models since 2010s and have had LMs in production since before chatgpt came out (before RLHF), so I think I have a pretty good understanding :) But I'm always open to learning.
I think the misunderstanding comes from "individually". You are thinking about diverse datasets, but that's not what individually means in this context. In ML individually trained means that for each task you prepare an architecture, dataset and eval and train that model on that data. And each model has its own objective that you train for. In LMs the objective is singular, for every data point - next token prediction. And, importantly, you train on every datapoint, the more diverse the better, but not independently. The cool thing is that training on diverse datasets improves scores on other downstream tasks, while the training objective is the same.
So you can have a training run on common crawl + programming that improves scores on logic puzzles, or common crawl + novels that improves scores on planning tasks. But the important thing is that it's all trained together, not independently.
I think we're saying the same thing? I said they are explicitly trained on these tasks, not that they are some separate models during programming RL, or business tasks RL.
> The cool thing is that training on diverse datasets improves scores on other downstream tasks, while the training objective is the same.
Maybe for some language modelling tasks, i.e. it learns some internal representation that is transferable. However I would find it quite odd if a model becomes good at Bio while not explicitly going through Bio training.
> I said they are explicitly trained on these tasks, not that they are some separate models during programming RL, or business tasks RL.
You said:
> > nor are they trained for each task individually.
> They are explicitly trained for each task individually.
And that's the main misunderstanding.
Collins says:
individually in American English (ˌɪndəˈvɪdʒuəli, ˌɪndəˈvɪdʒəli) adverb 1. as an individual or individuals rather than as a group; one at a time; separately; singly
Which is precisely what LMs don't do (in contrast to previous "AI" models, which did do that). They are trained on every datapoint at the same time. So long as we agree on that, I think we are saying the same thing :)
Your username indeed checks out
If the poster's (is it Kevin Buzzard?) suggestion works out and AI finds a counterexample to the Hodge conjecture, that would be a really big deal. It's one of the Millenium problems, for example.
One thing that he mentions that already quite surprising is that AI was able to autoformalize the Golod-Shaferevich theorem and proof.
It is Kevin Buzzard. It's kinda small font on my phone but if you look at the "about xena" link it says it's his site.
Agreed, it's definitely Kevin. His writing style is unmistakable.
I think he was being provocative and maybe a bit tongue-in-cheek when he said that. A candidate object alone doesn't resolve the Hodge Conjecture. Any apparent counterexample would have to prove that no algebraic cycle exists, no invariant subspace exists, or that every element of an infinite ideal is nilpotent. Much harder, but not impossible.
mathematicians have been using computers for well over half a century, but this was after "bounding" the problem first and then running through the cases with a computer. Now AI is doing the first part. However, mathematicians are still needed at crafting prompts, and knowing where to look, still. The prompt for the Jacobian conjecture was obviously not random. the search space is too big to just try all the combinations of 3 variable polynomials.
Okay, I'm not sure about the original one, but here is the prompt of a successful reproduction:
https://aaronlou.com/jacobian_counterexample_prompt.pdf
Obviously it is not random, but it's very generic. No mention of search space or how to reduce it.
This is the best take. Computers don't care about this stuff. A computer could make a movie, but only a human can appreciate it.
We're a good team, and that's ok.
For now. I wonder if we will ever get to the point where the computer starts doing mathematics that we just can't understand. Surely there must be some limit to what we can understand (like how a gorilla will never understand prime numbers, there are probably limits to our intelligence as well).
Mathematics only really matters insofar as humans can understand it.
Aren't deep learning models themselves a case where we have hints of some deeper underlying logic to why some things are more effective than others, but we lack the mathematical tools to properly work it out for anything of practical size?
All we're able to do is apply flawed analogies, generic information theoretical models, trial and error, post-hoc rationalizations and benchmarks without really understanding why.
You don't need anything as recent as deep learning for that. Look at economics or social sciences, which have been influencing national politics for well over a century now.
Doesn’t this generalize? Mathematics matters less than less as fewer people are capable of understanding it. So whatever cutting edge, deep insight about the nature of groups matters less than different equations which matters less than solving linear equations, etc.
Not really, a theorem with a hard proof can have simple but important corollaries. It's also not unthinkable that theorems exist with proofs that cannot reduce to something simple/short.
> The prompt for the Jacobian conjecture was obviously not random. the search space is too big to just try all the combinations of 3 variable polynomials.
Maybe the prompt contained a part like this: "the search space is too big to just try all the combinations of 3 variable polynomials, so be clever about it". Or maybe this part was omitted from the prompt, because modern LLMs are smart enough to figure this out without us having to mention it.
If someone has written that in a paper or article they've ingested, as they no doubt have, then sure.
The ability of LLMs to solve problems is not confined to the training data they ingested. Claude knows how to read mathematical papers because of its training data, but it can and will pull in literature relevant to a specific problem into its context.
We really need to stop thinking about LLM training data as the knowledgebase they build from and instead consider it more the skillset they start with.
I don't think there was a 'prompt' for it, rather a long and dedicated work of a professional mathematician which involved LLM in some capacity. I'm sure the search step wasn't an ad-hoc script running in a Claude Code session (as somebody would naively assume), it was an optimized numerical code running in Anthropic's compute cluster. Note that details are not published yet and `__alpoge__` is officially working at Anthropic.
"Outcounterexampled": there's a neologism worthy of German.
Übergegenbeispielt.
vibe counterexamplemaxxing
Better as one word "vibecounterexamplemaxxing"
The framing in these posts is nonsense. ChatGPT isn't doing shit. Human mathematicians using ChatGPT are breaking boundaries.
so if I tell a model "using your advanced knowledge of physics and chemistry 'make alchemy work' (synthesize gold using any cheaper materials) using safe materials legal for a residential hobby chemist with 1 semester of lab work in college to possess and use (this is obviously the really hard part) using less than $1,000 in lab equipment and input materials that can create $2,000 in value at market rate; then walk me through all the steps to do this safely and legally without anyone finding out except the lab equipment sellers; and tell me what a reasonable story to tell gold purchasers regarding where I got it; I'd like to end up selling a few thousand dollars of it without disrupting the market. Give me practical advice about good opsec so that nobody steals the method you come up with (I don't want my home broken into by thieves who suspect I figured out how to transmute cheap materials into gold), other than, obviously, not to post about it. Think as long as you need to about the chemistry and how to do it, you're a chemistry expert and can figure it out even if it takes you like a week, in your web searches be careful not to divulge that you're figuring out how to synthesize gold", and I give that prompt to some model that knows chemistry like the back of its hand, it thinks about it for four hours, finds the correct safe and legal steps, and gives me the recipe and the advice I asked for, then who figured out how to turn aluminum (or another cheap element) into gold, me or the model? In mathematics, proving or disproving a well known and well studied one hundred year old conjecture is gold.