On the second day, there are two gaps: “between nothing and zero” and “between zero and nothing”. Two numbers spawn in those two gaps. Call them –1 and 1.
Got lost here. I think I'm officially too dumb for math.
Plot twist: they were both sold on early investment in companies that survived the .com bust. Now they’re VCs that everybody worships as business geniuses even though they’re just lucky idiots, and the sycophantic chatbots finally let them feel as smart as everyone says they are, and a whole bunch of hype-drunk fans are feeding into it.
I think it's just we don't have a way to say/write these numbers. So you make up a way to write them (-1 and 1) and continue.
The numbers don't matter and you could replace -1 and 1 with anything. It's just easier to begin your new fake number at - 1 and 1. Because position does matter.
There has to be some procedure how to come up with "new" numbers though, if you want to have more in the end than just a fancy binary tree - in particular if you want to map your "fake numbers" to the reals, infinity, etc.
(Apparently I was extremely unclear with this text. For clarity: if you want to actually understand surreal numbers, go and read On Numbers and Games, by Conway, which is a delightful book; or get an LLM to talk you through Wikipedia. Original text follows.)
It’s a terrible explanation. A surreal number is defined as a pair of sets of surreal numbers (where you fiddle around the recursion in that definition by defining them in waves, so strictly speaking you’re defining “the surreal numbers born at time T” for each individual T given access to the surreal numbers born at all earlier times, and then you “take the union across all times”, scare quotes because there are too many times for this to result in a set). Zero is a surreal number but the LLM is using the word “zero” to mean “the set containing just the surreal number 0”; “nothing” here is the LLM’s obtuse word for the empty set. Wikipedia may actually be easier to follow.
I mean, I was intending to supply the words that would link the LLM’s explanation to a more normal one, not to explain it; apparently that was extremely unclear. An actual explanation is much longer, as indeed I attempted to indicate by pointing to Wikipedia and saying that it might be more clear.
“Doing better than a totally useless explanation in fewer characters” is in general impossible, of course, eg if the first explanation has only one character.
I'm hoping someone develops an interactive tutor that can teach any subject to any depth.
The tutor should optimize its pedagogy. It should use online RL to adapt to a learner's ideal learning style, model what the student understands and to what degree, and understand what the gaps and next steps are.
I think he just wrote it in a confusing way. The quote before says:
(crucially, “to the left of all” and “to the right of all” also count as “gaps”)
So there are two "nothings" here, left of "all" - i.e. the zero - and right of it.
Though I'm not quite sure how you'd get infinite or irrational numbers by this procedure. Wouldn't you simply get the rational numbers by this?
(unless the "put a number" step is doing more work here than it seems. He doesn't really say which number to put there. In the examples, he mostly did "new number = (left number + right number) / 2", with special cases if any number is "nothing" - but he never actually wrote what the rules are here.)
Yeah I'm curious about the difference between "nothing" and "0", but I just decided to roll with it. Until the Greek letters made my head start to spin as they usually do.
I don't think this is your fault; the description isn't very explicit. Let me try to do a bit better. (I'll also try to go somewhat further, and you should not be discouraged if at some point it stops making sense.)
You can think of the "surreal numbers" as being built up step by step. We start out with no numbers at all, and then we repeatedly do a construction that makes some new numbers.
A surreal number is made from two sets of (pre-existing) surreal numbers. We typically call them L and R, for "left" and "right", and sometimes write it as L|R or {L|R} or something like that. The "left" numbers have to be smaller than the "right" numbers. The resulting number will turn out to be, in a certain sense, the "simplest" number in between all the left numbers and all the right numbers.
Now, as I said, we start out with no numbers at all. It might seem like that gives us no way to proceed, but it does: even given no numbers at all, we can still make a set of numbers, namely the empty set! So we can use that for both L and R, getting ∅|∅. Empty sets on both sides. We call this 0, and it will turn out to behave in the way you'd expect the number 0 to behave.
Now we suddenly have another set available, namely {0}, the set containing only zero. Which means that instead of being able to make one number, maybe we can make four: ∅|∅, ∅|{0}, {0}|∅, {0}|{0}. The first of these we already knew about. The last isn't actually admissible -- remember that the "left" numbers have to be smaller than the "right" numbers, which is "vacuously" true when one of those sets is empty (it means "if you have a number x in the left set, and a number y in the right set, then x<y", and if there are no numbers in the left set or no numbers in the right set then that's trivially true) but isn't true when both sets contain 0 because 0<0 is false.
So actually we get two new numbers: ∅|{0} and {0}|∅. The first fits into what OP calls the gap "between nothing and zero". The second first into what OP calls "the gap between zero and nothing". In both cases, "zero" means a number and "nothing" means a space where we don't yet have any numbers.
The number ∅|{0} is called -1 (it has to lie to the left of 0, and there's no constraint on its left, and -1 is "the simplest number less than 0") and the number {0}|∅ is called +1 (it has to lie to the right of 0, and there's no constraint on its right, and +1 is "the simplest number greater than 0").
I should explicitly acknowledge that I haven't defined what "less than" and "greater than" actually mean for these numbers, nor anything else about how they relate to one another that could possibly justify giving these things the specific names 0, -1, and +1. But there are definitions for "less than" and "greater than" and "plus" and "minus" and so forth, and the whole thing does turn out to work very nicely.
Anyway, once we've got these numbers we have eight possible sets that can go on the left or on the right. The requirement for left-things to be smaller than right-things reduces the possibilities somewhat, and the actual new numbers we get next time around are: ∅|{-1}, which turns out to be -2; {-1}|{0} which turns out to be -1/2; {0}|{+1} which turns out to be +1/2; {+1}|∅ which turns out to be +2. We also get some already-existing numbers in new ways; for instance, {-1}|{+1} is actually equal to 0 ("0 is the simplest number between -1 and +1"). Again, I should explicitly acknowlege that I haven't said anything about how you determine when two of these things are actually equal; again, it does all turn out to work properly.
If you keep going with this construction, you produce all the integers, two at a time, and also all the "dyadic rationals", meaning fractions where the denominator is a power of 2. And then, once you've got all those, at the next stage of construction you abruptly get all the real numbers -- e.g., the square root of 2 is L|R where L = {dyadic rational numbers that are negative or have a square smaller than 2} and R = {dyadic rational numbers that are positive and have a square larger than 2} -- and you also get {0,1,2,3,4,...}|∅, conventionally written as a lower-case Greek letter omega, which is an infinite number, larger than all the integers. (And its negation.) And {0}|{1,1/2,1/3,1/4,...} which is an infinitesimal number, positive but smaller than any ratio of positive integers. And you can then proceed further and construct a vast infinitude of numbers, including all the real numbers (which we've already made) and all of the so-called infinite ordinals (which you can kinda think of as being a sort of "infinite positive integer", though there's more to them than that) and much more, all in a system that lets you do arithmetic and suchlike. It's very elegant, if your brain has been twisted into the mathematician-y shape that finds such things elegant.
This would make a lot more sense to me if "nothing" and "nothing" were instead "-inf" and "+inf"
Some other comments clarify that "nothing" is more accurately "the empty set". This is helpful because at first I wrongly synonomized "nothing" with zero. But now I get tripped up on the "between" language. Maybe it's a lack of background in sets, but I don't know what "between" implies for an integer (zero) and a set (the empty set).
Wow thanks. I'm a mathematician and also got lost at the step. This illustration makes the construction much more clear. The text isn't really describing this process well
In either case I believe people who can put AI to the most value are the mathematicians themselves
The net output of math will increase, and mathematicians have more work now to unravel all this, and make it useful. What AI does is brute force search in the proof-space guiding by some training. The infinite monkey theorem [1] now needs an LLM corollary - Something like: A finite number of LLM monkeys will almost surely find all theorems given an infinite token budget.
Instead, I have a more complicated view, which I actually expressed in my essay The Two Cultures of Mathematics a quarter of a century ago, and which can be summarized by saying that there is a spectrum of attitudes in mathematics to the relationship between problem-solving and conceptual understanding. At one end of the spectrum you have mathematicians who are primarily motivated by the wish to solve problems, who see conceptual understanding as a very important means to that end. At the other you have mathematicians who are primarily motivated by the wish to attain conceptual understanding, who see problem-solving as a very important means to that end.
Before, understand and problem-solving-ability were so interdependent that distinguishing between the two was practically very difficult and probably wouldn’t have changed anyone’s research agenda. Now, they’re not connected, and this guy just did the ultimate meta-experiment of seriously undertaking a project that is intentionally 100% problem-solving and 0% understanding to prove it (maybe 99% and 1% but pretty close. In his transcripts, he never asks ChatGPT about the math, only about its opinions of the math).
As we (as a society) sit around asking ourselves what mathematicians (and software engineers, and anyone in deep technical fields) should be doing all day, we now have this case study to show us how wide our range of options has become.
So, assuming my proof doesn’t rely on a Lean kernel bug, it’s likely to be legit too.
He lacks the understanding to verify his solution properly, and has to lean on those who do have the understanding to verify it, only being able to say himself that it's "likely" to be correct. (And what do those mathematicians get for laboriously checking the generated proof? 40 grand?)
Seems to me problem solving is as dependent on understanding as ever.
Finally, you might be wondering about the token cost. I wasn’t running this project in a particularly token-efficient way and have repeatedly maxed out my 20x Pro subscriptions for both Claude and ChatGPT every week. I also briefly had access to a prerelease model in the last few days, which did not have a usage cap. I was not tracking my actual token usage consistently. Some AI analysis from the recovered logs roughly estimates that we’re totaling around 40 billion tokens, of which around 210 million were output tokens. Over 95% were cache reads.
Who can pay for a 20x Pro subscriptions and also get prerelease models? something weird is going on here.
Several of my coworkers — it’s not that unusual that if you can max out a couple accounts, the companies will obviously notice you (as a high cost customer), and sometimes offer more
I don't know why the author could claim this is "their" proof, and they kept saying "they" did this, "they" built that. but in reality everything is done by the LLM and the author is merely asking it to do things. i guess they did contribute money at least...
Got lost here. I think I'm officially too dumb for math.
I feel the same way. Want to become a dumb and dumber duo? I believe you could be my friend.
Pitch: a mini-series where a Dumb and Dumber duo get access to unlimited tokens via a roommates's account (who is an intern at a frontier lab).
In each episode they make a major science-fiction style breakthrough and grapple with the consequences without revealing themselves.
Plot twist: they were both sold on early investment in companies that survived the .com bust. Now they’re VCs that everybody worships as business geniuses even though they’re just lucky idiots, and the sycophantic chatbots finally let them feel as smart as everyone says they are, and a whole bunch of hype-drunk fans are feeding into it.
I think it's just we don't have a way to say/write these numbers. So you make up a way to write them (-1 and 1) and continue.
The numbers don't matter and you could replace -1 and 1 with anything. It's just easier to begin your new fake number at - 1 and 1. Because position does matter.
There has to be some procedure how to come up with "new" numbers though, if you want to have more in the end than just a fancy binary tree - in particular if you want to map your "fake numbers" to the reals, infinity, etc.
Only as a mental abstraction that's based on our experience/concept of space+time.
(Apparently I was extremely unclear with this text. For clarity: if you want to actually understand surreal numbers, go and read On Numbers and Games, by Conway, which is a delightful book; or get an LLM to talk you through Wikipedia. Original text follows.)
It’s a terrible explanation. A surreal number is defined as a pair of sets of surreal numbers (where you fiddle around the recursion in that definition by defining them in waves, so strictly speaking you’re defining “the surreal numbers born at time T” for each individual T given access to the surreal numbers born at all earlier times, and then you “take the union across all times”, scare quotes because there are too many times for this to result in a set). Zero is a surreal number but the LLM is using the word “zero” to mean “the set containing just the surreal number 0”; “nothing” here is the LLM’s obtuse word for the empty set. Wikipedia may actually be easier to follow.
If you can’t explain it better in the same amount of characters (or fewer), then I don’t think you’re qualified to “nuh-uh!!!” anyone. Sorry buddy.
I mean, I was intending to supply the words that would link the LLM’s explanation to a more normal one, not to explain it; apparently that was extremely unclear. An actual explanation is much longer, as indeed I attempted to indicate by pointing to Wikipedia and saying that it might be more clear.
“Doing better than a totally useless explanation in fewer characters” is in general impossible, of course, eg if the first explanation has only one character.
... wut? :/
I'm hoping someone develops an interactive tutor that can teach any subject to any depth.
The tutor should optimize its pedagogy. It should use online RL to adapt to a learner's ideal learning style, model what the student understands and to what degree, and understand what the gaps and next steps are.
I'd subscribe in a heartbeat.
There's a skill for that (haven't tried myself but intend to; other of author's skills I've used have been a game changer).
- https://github.com/mattpocock/skills/blob/main/skills/produc...
- https://www.youtube.com/watch?v=s5T5oQJcJ6U
The image helped for me at least.
I think he just wrote it in a confusing way. The quote before says:
So there are two "nothings" here, left of "all" - i.e. the zero - and right of it.
Though I'm not quite sure how you'd get infinite or irrational numbers by this procedure. Wouldn't you simply get the rational numbers by this?
(unless the "put a number" step is doing more work here than it seems. He doesn't really say which number to put there. In the examples, he mostly did "new number = (left number + right number) / 2", with special cases if any number is "nothing" - but he never actually wrote what the rules are here.)
Yeah I'm curious about the difference between "nothing" and "0", but I just decided to roll with it. Until the Greek letters made my head start to spin as they usually do.
I don't think this is your fault; the description isn't very explicit. Let me try to do a bit better. (I'll also try to go somewhat further, and you should not be discouraged if at some point it stops making sense.)
You can think of the "surreal numbers" as being built up step by step. We start out with no numbers at all, and then we repeatedly do a construction that makes some new numbers.
A surreal number is made from two sets of (pre-existing) surreal numbers. We typically call them L and R, for "left" and "right", and sometimes write it as L|R or {L|R} or something like that. The "left" numbers have to be smaller than the "right" numbers. The resulting number will turn out to be, in a certain sense, the "simplest" number in between all the left numbers and all the right numbers.
Now, as I said, we start out with no numbers at all. It might seem like that gives us no way to proceed, but it does: even given no numbers at all, we can still make a set of numbers, namely the empty set! So we can use that for both L and R, getting ∅|∅. Empty sets on both sides. We call this 0, and it will turn out to behave in the way you'd expect the number 0 to behave.
Now we suddenly have another set available, namely {0}, the set containing only zero. Which means that instead of being able to make one number, maybe we can make four: ∅|∅, ∅|{0}, {0}|∅, {0}|{0}. The first of these we already knew about. The last isn't actually admissible -- remember that the "left" numbers have to be smaller than the "right" numbers, which is "vacuously" true when one of those sets is empty (it means "if you have a number x in the left set, and a number y in the right set, then x<y", and if there are no numbers in the left set or no numbers in the right set then that's trivially true) but isn't true when both sets contain 0 because 0<0 is false.
So actually we get two new numbers: ∅|{0} and {0}|∅. The first fits into what OP calls the gap "between nothing and zero". The second first into what OP calls "the gap between zero and nothing". In both cases, "zero" means a number and "nothing" means a space where we don't yet have any numbers.
The number ∅|{0} is called -1 (it has to lie to the left of 0, and there's no constraint on its left, and -1 is "the simplest number less than 0") and the number {0}|∅ is called +1 (it has to lie to the right of 0, and there's no constraint on its right, and +1 is "the simplest number greater than 0").
I should explicitly acknowledge that I haven't defined what "less than" and "greater than" actually mean for these numbers, nor anything else about how they relate to one another that could possibly justify giving these things the specific names 0, -1, and +1. But there are definitions for "less than" and "greater than" and "plus" and "minus" and so forth, and the whole thing does turn out to work very nicely.
Anyway, once we've got these numbers we have eight possible sets that can go on the left or on the right. The requirement for left-things to be smaller than right-things reduces the possibilities somewhat, and the actual new numbers we get next time around are: ∅|{-1}, which turns out to be -2; {-1}|{0} which turns out to be -1/2; {0}|{+1} which turns out to be +1/2; {+1}|∅ which turns out to be +2. We also get some already-existing numbers in new ways; for instance, {-1}|{+1} is actually equal to 0 ("0 is the simplest number between -1 and +1"). Again, I should explicitly acknowlege that I haven't said anything about how you determine when two of these things are actually equal; again, it does all turn out to work properly.
If you keep going with this construction, you produce all the integers, two at a time, and also all the "dyadic rationals", meaning fractions where the denominator is a power of 2. And then, once you've got all those, at the next stage of construction you abruptly get all the real numbers -- e.g., the square root of 2 is L|R where L = {dyadic rational numbers that are negative or have a square smaller than 2} and R = {dyadic rational numbers that are positive and have a square larger than 2} -- and you also get {0,1,2,3,4,...}|∅, conventionally written as a lower-case Greek letter omega, which is an infinite number, larger than all the integers. (And its negation.) And {0}|{1,1/2,1/3,1/4,...} which is an infinitesimal number, positive but smaller than any ratio of positive integers. And you can then proceed further and construct a vast infinitude of numbers, including all the real numbers (which we've already made) and all of the so-called infinite ordinals (which you can kinda think of as being a sort of "infinite positive integer", though there's more to them than that) and much more, all in a system that lets you do arithmetic and suchlike. It's very elegant, if your brain has been twisted into the mathematician-y shape that finds such things elegant.
I made a picture, hope this helps: https://excalidraw.com/#json=zfKWWn1h7GzdFca6RDdXl,plr_WeaCt...
Sorry it was confusing.
This would make a lot more sense to me if "nothing" and "nothing" were instead "-inf" and "+inf"
Some other comments clarify that "nothing" is more accurately "the empty set". This is helpful because at first I wrongly synonomized "nothing" with zero. But now I get tripped up on the "between" language. Maybe it's a lack of background in sets, but I don't know what "between" implies for an integer (zero) and a set (the empty set).
Wow thanks. I'm a mathematician and also got lost at the step. This illustration makes the construction much more clear. The text isn't really describing this process well
Someone please vibe-prove that ZFC is inconsistent.
That's awesome! Congratulations!
I'd imagine that in three months when we all have access to communicating agent swarms this should be easier
The net output of math will increase, and mathematicians have more work now to unravel all this, and make it useful. What AI does is brute force search in the proof-space guiding by some training. The infinite monkey theorem [1] now needs an LLM corollary - Something like: A finite number of LLM monkeys will almost surely find all theorems given an infinite token budget.
[1] https://en.wikipedia.org/wiki/Infinite_monkey_theorem
I find this whole post fascinating in the context of https://news.ycombinator.com/item?id=49738091 and particularly this excerpt from Gowers:
Before, understand and problem-solving-ability were so interdependent that distinguishing between the two was practically very difficult and probably wouldn’t have changed anyone’s research agenda. Now, they’re not connected, and this guy just did the ultimate meta-experiment of seriously undertaking a project that is intentionally 100% problem-solving and 0% understanding to prove it (maybe 99% and 1% but pretty close. In his transcripts, he never asks ChatGPT about the math, only about its opinions of the math).
As we (as a society) sit around asking ourselves what mathematicians (and software engineers, and anyone in deep technical fields) should be doing all day, we now have this case study to show us how wide our range of options has become.
Yes, the timing of this post just after Gowers' post is fascinating. It is almost as if the marketing machine is well oiled.
He lacks the understanding to verify his solution properly, and has to lean on those who do have the understanding to verify it, only being able to say himself that it's "likely" to be correct. (And what do those mathematicians get for laboriously checking the generated proof? 40 grand?)
Seems to me problem solving is as dependent on understanding as ever.
Finally, you might be wondering about the token cost. I wasn’t running this project in a particularly token-efficient way and have repeatedly maxed out my 20x Pro subscriptions for both Claude and ChatGPT every week. I also briefly had access to a prerelease model in the last few days, which did not have a usage cap. I was not tracking my actual token usage consistently. Some AI analysis from the recovered logs roughly estimates that we’re totaling around 40 billion tokens, of which around 210 million were output tokens. Over 95% were cache reads.
Who can pay for a 20x Pro subscriptions and also get prerelease models? something weird is going on here.
Several of my coworkers — it’s not that unusual that if you can max out a couple accounts, the companies will obviously notice you (as a high cost customer), and sometimes offer more
ah, if it's anyone it'd be dan abramov
I don't know why the author could claim this is "their" proof, and they kept saying "they" did this, "they" built that. but in reality everything is done by the LLM and the author is merely asking it to do things. i guess they did contribute money at least...
LOL. mood??