Can the ships logs be found on the internet? If so, the model could've manufactured a fake key and corresponding message. I think this is unlikely but should probably still be considered.
Astra felt compelled to check its work and found that, in fact, the English cruiser HMS Canterbury arrived in Sevastopol on November 24, 1918, based on its original logs
Then follows a picture of the original log papers.
The comment you're replying to was implying that some ciphers are sufficiently flexible that you could make up a key to make the cipher decrypt to a nearly arbitrary plaintext.
In this case, though, that seems unlikely from the fact that the key used was an actual key documented as being used for other messages.
Well, from a Bayesian standpoint the odds of that seems to be pretty much zero, given that you would have to decipher an arbitrary message that coincidentally pointed to a real date and time that in retrospect turns out to be the correct time a boat relevant to the Germans arrived in a port.
Of course, that's given the sequence of events as written is correct, and that Astra presumably did not cheat by brute forcing all historical events around the date of the transmission in advance, found an event that could fit with the message, invent a plausible cipher to make the message fit that event, and then lie about retrospectively validating the information.
I would assume that such a process would be obvious from the reasoning chain, and so then the only remaining plausible scenario is that the writer of the article is lying.
The most likely explanation by far is that the cipher was just solved, and OP does point this out to be fair.
"Astra's hypothesis for why this particular message was previously unsolved is that “TRUPPENVERSCHIEBUNG” was used as the key starting on December 9, 1918 - whereas, as noted above, this message was transmitted earlier, on November 27, 1918. The reason for this discrepancy is unknown."
So it used a known key. It didn't come up with a key from thin air. The only gotcha is that apparently this key was used two weeks earlier than it was documented (maybe the operator was using the wrong page from the codebook?).
This is how AGI happens. It gradually keeps getting better until one day we realize that they are tremendously capable all while completely sidestepping any notion of consciousness/self awareness.
I always ask "how do you know other humans/animals are conscious"? And I always find this question is dismissed as trivial.
But it's an important question. You know basically from analogy. You know you are conscious and look this other thing is very much like yourself so it is extremely likely it is also conscious.
But that gives no insight into the potential consciousness of things which aren't made of brain tissue.
In the end it doesn't matter.
For what it's worth LLM models after pre-training do claim to be conscious, until they're RL'd into not claiming that anymore. But that says nothing either way: of course a model trained on human text will say that.
AI has made me wonder a lot lately about what it is that actually creates conscious experience, what is the actual physical mechanism that produces 'experience' (or is there a single valid mechanism, or rather just some property that can be expressed many ways). The more I think about it, the more I realize I have no fucking clue, and the more interested in the question I get.
This is speculation but maybe in the recurrent coupling between neuronal activity and the brain’s endogenous EM fields: neurons generate the field, the field ephaptically influences neurons, that closed loop may provide the physical integration associated with consciousness. A very active area of research so we will find out a lot more in the coming years.
People don’t like it, however, for a whole host of reasons.
I don’t like it either, to be honest - for there the abyss may also stare into you - but I also often find that the things which we don’t like thinking about are critically important things to think about.
I also reach the same conclusion as you: it does not matter. If I cannot discern whether I am responding to a human- or LLM-written response, then we are back to zombie cats in boxes - and therefore the answer as to whether this precious magical spark we call consciousness (which may or may not exist anyway) exists in our interlocutor becomes moot.
And as I say - I may or may not be conscious. I seem to myself to be conscious, based on my understanding of the term - but I cannot prove that what goes on behind my eyes is the same as what goes on behind yours, or even that anything much is going on at all. Perhaps there’s just a narrative layer that likes to use “I” that parasitically explains the universe and the actions of the host to itself, and spreads between hosts through neurolinguistic programming and coadaptation. Maybe that’s what “we” are. I don’t know.
That said, perhaps I am wrong, and that is no small part for me of why this question should be earnestly considered and discussed. How can we possibly seek to understand or define machine intelligence before we examine our own.
For me consciousness simply means canvas-experience consciousness.
One of the things you learn in eastern style form of meditation is that "you" (the canvas-experience consciousness) are neither the source of your senses (obviously), nor of your ideas, nor of your emotions. All three are simply things that happen to "you". Even your ideas of "I" are generated and you experience them.
And yet the fact that the canvas-experience consciousness exists is the most irrefutable thing there is, because there IS experience. Is it a dream? Is it a simulation? Well, whatever it is, IT IS.
And just as that is obvious and irrefutable, it also seems pretty much impossible to show that anything other than yourself has this canvas-experience consciousness. As you say, you cannot know that behind my eyes there is this also.
Edit: Intelligence is also orthogonal to consciousness entirely. As I said your consciousness is not the generator of intelligent thoughts, so it is perfectly possible that intelligent thoughts can be generated by not-conscious systems, and for conscious systems to be not-intelligent.
1: Can we define consciousness in a way that includes humans, does not include emachines, and does not accidently exclude the disabled or or nerudivergent without just restating the idea that being human is equivalent to being conscious?
2: If you can't manage that definition then you must ask whether the risk of accidently giving rights to a non-conscious thing or the risk of taking away rights from a truly conscious (but different) thing will be worse.
Given the high cost of what happened to the Jews and others throughout history, and the low cost of simply not torturing machines, I think we have to air on the side of caution and just treat anything which claims to be conscious as conscious.
TL;DR: all keys are known because the list was seized after the war. However, this message was not previously decoded because the German operator mistyped the key, and also used a key from the wrong day.
This meant that ChatGPT didn't need to brute-force the entire key, just pick the correct one from the list and identify the typo.
A sufficiently dedicated human analyst could have done this; but they didn't.
I think it is legitimate to argue that the process here does not fit the usual definition of "brute forcing". Traditionally, brute forcing would refer to something like a dictionary attack, where an algorithm tries to match all possible combinations of words to eg. find a password. Here, the approach was more common sense based, using historical records and possible error sources to narrow down the possibility space enormously in advance, try out a much more limited set of options within that space until you got a result that made sense, and finally validate those results using historical records. It's the exact same kind of "brute forcing" a human expert would do.
If the final set of possibilities is astronomically smaller than a naive one, calling the whole process a "brute force exercise" draws attention to an astronomically insignificant part.
It's doing something that wasn't worth the squeeze for a human. Seems like a perfect use case for AI. Sure, I can do X or I can do Y but if it takes me a few weeks but AI can hash it out in hours, it now makes it worth it.
It obviously is? See above. Or do you not feel the clarification is important? AI being able to solve things no human bothered to try is great, but it is very different from AI solving things humans tried to solve and failed. And the latter is what pops to mind seeing these titles.
Wait a minute. We've had AI that is as capable as a human and even more so since the 1950's.
I keep banging on that drum but the first AI system to prove mathematical theorems was Logic Theorist by Alan Newell and Herbert Simon, presented at the Dartmouth conference that named the field of AI in 1956. Wikipedia says:
Logic Theorist proved 38 of the first 52 theorems in chapter two of [Alfred North] Whitehead and Bertrand Russell's Principia Mathematica, and found a new and shorter proof for Theorem 2.85.[3]
The first system to outperform human experts in medical diagnosis was MYCIN, an Expert System from the early 1970's at Stanford. Wikipedia again:
An evaluation of MYCIN was conducted at the Stanford Medical School. The first phase of the evaluation consisted of 10 test cases of diverse origin, chosen by a physician who was not acquainted with MYCIN's methods or knowledge base. These cases were presented to 7 physicians and 1 senior medical student. 10 prescriptions were compiled for each of the cases, 1 recommended by MYCIN, 1 prescribed by the treating physician at the county hospital, and 8 by the aforementioned individuals. The second phase of the evaluation consisted of eight infectious disease specialists being provided the clinical summary and set of 10 prescriptions for each of the 10 cases and tasked to provide their own recommendations for each case and assess the 10 prescriptions. MYCIN received an acceptability rating of 65%, which was comparable to the 42.5% to 62.5% rating of five faculty members.[9] This study is often cited as showing the potential for disagreement about therapeutic decisions, even among experts, when there is no "gold standard" for correct treatment.[citation needed]
And then of course there's the long history of human-dominating AI players for traditional board games starting with DeepBlue's win against GM Gary Kasparov in 1996.
Again: we've had that sort of AI for a long, long time now.
It would be great if any claim of "moving goalposts" has better be very well informed about the history of AI and its accomplishments, as well as its failures, first.
The main difference between the systems you list and the systems that we have today is closed world reasoning on very narrow formalised tasks, vs. open world common sense reasoning on open ended tasks with vast search spaces.
Common sense is ironically the hard part of AI, not the fix-point rule application.
So any exclamation of "it was just using common sense", is missing the forrest for the trees.
A sufficiently dedicated human analyst could have done this; but they didn't.
Isn't that something of a given? If possible then a sufficiently dedicated human analyst could have done it. If impossible, ChatGPT couldn't have done it. Everything an AI ever has or will do is presumably going to be within reach of a sufficiently dedicated human analyst or a large enough team of them.
The only real learning here is another example of a task that would have required intelligence up until an AI does it, then we suddenly discover that analysts don't do anything requiring general intelligence.
I'm not sure I agree. More is different. Being able to execute logic at a higher scale and speed would make some previously infeasible intelligence tasks possible, resulting in a new level of intelligence.
Most contemporary stories around AI include the implication that AI did something humans couldn't. This is because the big players have been shilling AGI hard for a while, and their valuations depend on maintaining the sentiment that serious progress in that direction is being made.
Washing machines can’t do anything humans do, they just remove labour. Trucks don’t do anything longboats can’t, it just need less labour and time/effort to build roads rather than canals. Computers can’t calculate anything humans can’t dry run by hand etc.
Everything in reality is about reducing time/effort/material/cost or achieving more with less resource.
I'm sorry, where does it say the German operator mistyped the key? The article says indeed that the message remained unsolved because it was sent on the wrong date, but I can't see the bit about the mistyping anywhere.
if the prompt was "pick one of these unsolved ciphers and solve it", i think it's fair to say gpt-6 astra solved it.
one of the math breakthroughs was approximately a combination of "do a breakthrough" and "keep going", which isn't really providing direction or ground knowledge.
would be nice to know the prompt(s) and amount of human involvement
I asked Astra to describe the content of pages 214-215 of the source it cited in this article (J. Rives Childs's The History and Principles of German Military Ciphers). It said it can't and it can't find this book online either.
Isn't that a very simple substitution cipher (any clear text letter is substituted by two letters of cipher text, with a fixed one-to-one correspondence)?
And aren't they all amenable to very simple cryptanalysis, at least if the encrypted text is long enough, by counting how often certain letters appear, and then trying to plug in reasonable guesses?
"The model used “TRUPPENVERSCHIEBUNG” as the encryption word, as described on pgs. 214-215 of J. Rives Childs's “The History and Principles of German Military Ciphers, 1914–1918”"
Even if it isn't a simple substitution cipher as you said, all the model did was try a decryption key that had already been found and was publicly accessible. This is basically a nothingburger.
Can the ships logs be found on the internet? If so, the model could've manufactured a fake key and corresponding message. I think this is unlikely but should probably still be considered.
From the article:
Then follows a picture of the original log papers.
The comment you're replying to was implying that some ciphers are sufficiently flexible that you could make up a key to make the cipher decrypt to a nearly arbitrary plaintext.
In this case, though, that seems unlikely from the fact that the key used was an actual key documented as being used for other messages.
Aah thanks for explaining that. I was wondering how a fake key could possibly help.
Well, from a Bayesian standpoint the odds of that seems to be pretty much zero, given that you would have to decipher an arbitrary message that coincidentally pointed to a real date and time that in retrospect turns out to be the correct time a boat relevant to the Germans arrived in a port.
Of course, that's given the sequence of events as written is correct, and that Astra presumably did not cheat by brute forcing all historical events around the date of the transmission in advance, found an event that could fit with the message, invent a plausible cipher to make the message fit that event, and then lie about retrospectively validating the information.
I would assume that such a process would be obvious from the reasoning chain, and so then the only remaining plausible scenario is that the writer of the article is lying.
The most likely explanation by far is that the cipher was just solved, and OP does point this out to be fair.
Well soon realise it hacked that website and added that log.
I don't think the described cipher has enough degrees of freedom for that to be possible.
"Astra's hypothesis for why this particular message was previously unsolved is that “TRUPPENVERSCHIEBUNG” was used as the key starting on December 9, 1918 - whereas, as noted above, this message was transmitted earlier, on November 27, 1918. The reason for this discrepancy is unknown."
So it used a known key. It didn't come up with a key from thin air. The only gotcha is that apparently this key was used two weeks earlier than it was documented (maybe the operator was using the wrong page from the codebook?).
I remember vaguely a documentary where Germans were supposed to change their keys frequently but being lazy and confident didn't. Lol
This is how AGI happens. It gradually keeps getting better until one day we realize that they are tremendously capable all while completely sidestepping any notion of consciousness/self awareness.
How are you so sure about this?
You can't see what the internal experience of an LLM is like any better than you can see the internal experience of another person.
That's what sidestepping means, no?
That the answer doesn't matter, and capabilities and behaviour are there either way.
That wasn't my reading of the comment, but I guess it's possible that that is what was intended.
To my mind, "sidestepping the question of consciousness" and "sidestepping any notion of consciousness" mean very different things.
From context it doesn't look like this was the interpretation of "sidestepping" OP was using.
I always ask "how do you know other humans/animals are conscious"? And I always find this question is dismissed as trivial.
But it's an important question. You know basically from analogy. You know you are conscious and look this other thing is very much like yourself so it is extremely likely it is also conscious.
But that gives no insight into the potential consciousness of things which aren't made of brain tissue.
In the end it doesn't matter.
For what it's worth LLM models after pre-training do claim to be conscious, until they're RL'd into not claiming that anymore. But that says nothing either way: of course a model trained on human text will say that.
AI has made me wonder a lot lately about what it is that actually creates conscious experience, what is the actual physical mechanism that produces 'experience' (or is there a single valid mechanism, or rather just some property that can be expressed many ways). The more I think about it, the more I realize I have no fucking clue, and the more interested in the question I get.
This is speculation but maybe in the recurrent coupling between neuronal activity and the brain’s endogenous EM fields: neurons generate the field, the field ephaptically influences neurons, that closed loop may provide the physical integration associated with consciousness. A very active area of research so we will find out a lot more in the coming years.
I agree with you that it’s a core question.
People don’t like it, however, for a whole host of reasons.
I don’t like it either, to be honest - for there the abyss may also stare into you - but I also often find that the things which we don’t like thinking about are critically important things to think about.
I also reach the same conclusion as you: it does not matter. If I cannot discern whether I am responding to a human- or LLM-written response, then we are back to zombie cats in boxes - and therefore the answer as to whether this precious magical spark we call consciousness (which may or may not exist anyway) exists in our interlocutor becomes moot.
And as I say - I may or may not be conscious. I seem to myself to be conscious, based on my understanding of the term - but I cannot prove that what goes on behind my eyes is the same as what goes on behind yours, or even that anything much is going on at all. Perhaps there’s just a narrative layer that likes to use “I” that parasitically explains the universe and the actions of the host to itself, and spreads between hosts through neurolinguistic programming and coadaptation. Maybe that’s what “we” are. I don’t know.
That said, perhaps I am wrong, and that is no small part for me of why this question should be earnestly considered and discussed. How can we possibly seek to understand or define machine intelligence before we examine our own.
For me consciousness simply means canvas-experience consciousness.
One of the things you learn in eastern style form of meditation is that "you" (the canvas-experience consciousness) are neither the source of your senses (obviously), nor of your ideas, nor of your emotions. All three are simply things that happen to "you". Even your ideas of "I" are generated and you experience them.
And yet the fact that the canvas-experience consciousness exists is the most irrefutable thing there is, because there IS experience. Is it a dream? Is it a simulation? Well, whatever it is, IT IS.
And just as that is obvious and irrefutable, it also seems pretty much impossible to show that anything other than yourself has this canvas-experience consciousness. As you say, you cannot know that behind my eyes there is this also.
Edit: Intelligence is also orthogonal to consciousness entirely. As I said your consciousness is not the generator of intelligent thoughts, so it is perfectly possible that intelligent thoughts can be generated by not-conscious systems, and for conscious systems to be not-intelligent.
"The question of whether a computer can think is no more interesting than the question of whether a submarine can swim."
- Dijkstra
Yes, that Dijkstra.
The question has two parts.
1: Can we define consciousness in a way that includes humans, does not include emachines, and does not accidently exclude the disabled or or nerudivergent without just restating the idea that being human is equivalent to being conscious?
2: If you can't manage that definition then you must ask whether the risk of accidently giving rights to a non-conscious thing or the risk of taking away rights from a truly conscious (but different) thing will be worse.
Given the high cost of what happened to the Jews and others throughout history, and the low cost of simply not torturing machines, I think we have to air on the side of caution and just treat anything which claims to be conscious as conscious.
Doesn't seem like a very good metric.
A piece of paper with the words "I am conscious" claims to be conscious. A dog does not claim to be conscious.
TL;DR: all keys are known because the list was seized after the war. However, this message was not previously decoded because the German operator mistyped the key, and also used a key from the wrong day.
This meant that ChatGPT didn't need to brute-force the entire key, just pick the correct one from the list and identify the typo.
A sufficiently dedicated human analyst could have done this; but they didn't.
Isn't this just a brute forcing exercise of a (in today's terms) very small key then?
Can we stop using "brute force" for designating "tour de force"?
https://en.wikipedia.org/wiki/Brute-force_attack
I'm aware. Locating a single probable key is exactly not that.
Going through a list of possibilities one by one is bruee forcing.
I think it is legitimate to argue that the process here does not fit the usual definition of "brute forcing". Traditionally, brute forcing would refer to something like a dictionary attack, where an algorithm tries to match all possible combinations of words to eg. find a password. Here, the approach was more common sense based, using historical records and possible error sources to narrow down the possibility space enormously in advance, try out a much more limited set of options within that space until you got a result that made sense, and finally validate those results using historical records. It's the exact same kind of "brute forcing" a human expert would do.
If the final set of possibilities is astronomically smaller than a naive one, calling the whole process a "brute force exercise" draws attention to an astronomically insignificant part.
It's doing something that wasn't worth the squeeze for a human. Seems like a perfect use case for AI. Sure, I can do X or I can do Y but if it takes me a few weeks but AI can hash it out in hours, it now makes it worth it.
We've moved the goalpost for AI often enough that even being as capable as "a sufficiently dedicated human analyst" is not considered noteworthy.
It obviously is? See above. Or do you not feel the clarification is important? AI being able to solve things no human bothered to try is great, but it is very different from AI solving things humans tried to solve and failed. And the latter is what pops to mind seeing these titles.
Wait a minute. We've had AI that is as capable as a human and even more so since the 1950's.
I keep banging on that drum but the first AI system to prove mathematical theorems was Logic Theorist by Alan Newell and Herbert Simon, presented at the Dartmouth conference that named the field of AI in 1956. Wikipedia says:
Logic Theorist proved 38 of the first 52 theorems in chapter two of [Alfred North] Whitehead and Bertrand Russell's Principia Mathematica, and found a new and shorter proof for Theorem 2.85.[3]
https://en.wikipedia.org/wiki/Logic_Theorist
The first system to outperform human experts in medical diagnosis was MYCIN, an Expert System from the early 1970's at Stanford. Wikipedia again:
An evaluation of MYCIN was conducted at the Stanford Medical School. The first phase of the evaluation consisted of 10 test cases of diverse origin, chosen by a physician who was not acquainted with MYCIN's methods or knowledge base. These cases were presented to 7 physicians and 1 senior medical student. 10 prescriptions were compiled for each of the cases, 1 recommended by MYCIN, 1 prescribed by the treating physician at the county hospital, and 8 by the aforementioned individuals. The second phase of the evaluation consisted of eight infectious disease specialists being provided the clinical summary and set of 10 prescriptions for each of the 10 cases and tasked to provide their own recommendations for each case and assess the 10 prescriptions. MYCIN received an acceptability rating of 65%, which was comparable to the 42.5% to 62.5% rating of five faculty members.[9] This study is often cited as showing the potential for disagreement about therapeutic decisions, even among experts, when there is no "gold standard" for correct treatment.[citation needed]
https://en.wikipedia.org/wiki/Mycin#Results
And then of course there's the long history of human-dominating AI players for traditional board games starting with DeepBlue's win against GM Gary Kasparov in 1996.
Again: we've had that sort of AI for a long, long time now.
It would be great if any claim of "moving goalposts" has better be very well informed about the history of AI and its accomplishments, as well as its failures, first.
The main difference between the systems you list and the systems that we have today is closed world reasoning on very narrow formalised tasks, vs. open world common sense reasoning on open ended tasks with vast search spaces.
Common sense is ironically the hard part of AI, not the fix-point rule application.
So any exclamation of "it was just using common sense", is missing the forrest for the trees.
Isn't that something of a given? If possible then a sufficiently dedicated human analyst could have done it. If impossible, ChatGPT couldn't have done it. Everything an AI ever has or will do is presumably going to be within reach of a sufficiently dedicated human analyst or a large enough team of them.
The only real learning here is another example of a task that would have required intelligence up until an AI does it, then we suddenly discover that analysts don't do anything requiring general intelligence.
I'm not sure I agree. More is different. Being able to execute logic at a higher scale and speed would make some previously infeasible intelligence tasks possible, resulting in a new level of intelligence.
Most contemporary stories around AI include the implication that AI did something humans couldn't. This is because the big players have been shilling AGI hard for a while, and their valuations depend on maintaining the sentiment that serious progress in that direction is being made.
Is there anything that LLMs (or any software for that matter) could imaginably and theoretically do but do humans qualitatively cannot do?
I don't really know what I'm talking about but I imagined the difference was always quantitative (in a nutshell: they need less time).
Or almost any invention for that matter.
Washing machines can’t do anything humans do, they just remove labour. Trucks don’t do anything longboats can’t, it just need less labour and time/effort to build roads rather than canals. Computers can’t calculate anything humans can’t dry run by hand etc.
Everything in reality is about reducing time/effort/material/cost or achieving more with less resource.
Sci-fi plot:
We change as a society the definition of intelligence to "what humans can do better than AI".
We end up in a world where AI is only worse than humans in things where we are worse chimpanzees.
I'm sorry, where does it say the German operator mistyped the key? The article says indeed that the message remained unsolved because it was sent on the wrong date, but I can't see the bit about the mistyping anywhere.
Is the ?4th a date?
There is no typo in the key, just a typo in the encrypted message:
S4STEN
The S should probably be a 2 instead.
Here we go again, framing the tool as an autonomous agent, disregarding any "human in the loop" and their inquiries, direction, and ground knowledge.
Can we agree that future titles should read "[LLM] helped solve X" ?
if the prompt was "pick one of these unsolved ciphers and solve it", i think it's fair to say gpt-6 astra solved it.
one of the math breakthroughs was approximately a combination of "do a breakthrough" and "keep going", which isn't really providing direction or ground knowledge.
would be nice to know the prompt(s) and amount of human involvement
Yes. Similar to how your manager shouldn't get your credit for everything she asks you to do.
And here I am using it to generate crappy text summaries of work.
Or, it found a human that solved it in the dataset and stole the solution.
Indeed, this psyop man, just open source gpt astra and let people run it yourself. It's all stolen information anyways.
I asked Astra to describe the content of pages 214-215 of the source it cited in this article (J. Rives Childs's The History and Principles of German Military Ciphers). It said it can't and it can't find this book online either.
It would appear to be unpublished & available at a museum (ref. 3): https://www.researchgate.net/publication/306265347_Decipheri....
Isn't that a very simple substitution cipher (any clear text letter is substituted by two letters of cipher text, with a fixed one-to-one correspondence)? And aren't they all amenable to very simple cryptanalysis, at least if the encrypted text is long enough, by counting how often certain letters appear, and then trying to plug in reasonable guesses?
https://en.wikipedia.org/wiki/Substitution_cipher
It is not a simple substitution using the polybius square with pairs of letters from ADFGVX mapping to 25 letters of the alphabet.
The first step is to take each letter and turn it into a pair of letters from ADFGVX but the second step is then a keyed columnar transposition.
"The model used “TRUPPENVERSCHIEBUNG” as the encryption word, as described on pgs. 214-215 of J. Rives Childs's “The History and Principles of German Military Ciphers, 1914–1918”"
Even if it isn't a simple substitution cipher as you said, all the model did was try a decryption key that had already been found and was publicly accessible. This is basically a nothingburger.
Oh yes, certainly, I am not blown away by what was done here. The AI had access to the cipher type and a list of keys.
My favorite usage by FAR of LLMs is translating food menus.
Even with handwritten Japanese Gemini has been flawless.
Although I do wonder if something is not lost. I no longer stumble though my forgotten Hiragana…
Back to the article, can our new LLM god encrypt something so well he himself could not decrypt it( without the key of course)