I saw the talk version of that and it has lodged in my memory. The point he made was that "health.gov" was made by a bunch of teams that made sure that their part worked, but nobody was in charge of making the whole thing work.
This is sadly "normal" for large bureaucracies.
Every time I see a giant, multi-million dollar catastrophe, it's always "proper", "documented", "enterprise", "by the book", and... "a total failure". The reason is always that nobody actually cares about the final outcome, only the paperwork in front of them that they need fill out, the checkbox that needs to be ticked, or the compliance requirement that needs to be met.
This, in my opinion, is some part of how we are schooling people - do as you are told and follow the rules and some part of not having stakes in the outcome.
Definitely not on purpose! I keep forgetting that ChatGPT does that now, which is rather irritating because I don’t use it to write comments but the tracker makes it look like I do.
I do use ChatGPT to resolve vague memories of articles read long ago into concrete URLs I can share, which is one of the best uses of the accursed things right now!
I assure you that I am an organic meat human like yourself and the article dates itself to the “before times” and can also be trusted to be free of contamination.
They're pretty readable, I recommend that second one (though it leaves some questions I have unanswered). It's a good example of how latent issues with a large system can go unrecognised for a long time before a seeming unrelated change can trigger a cascading failure
One thing it doesn't make clear – the GPS card only supported the original L1 signal, which is what causes the GPS week rollover issue. The newer L2C signal has a much longer rollover period (157 years vs 19.6 years), which means no rollover until next century. If the GPS card had supported the newer L2C signal, then likely this would not have happened even if the other misconfigurations had still occurred.
Is it really so that Time Division Multiplex timing in cellular networks is based on NTP time? I stopped reading after the article implied this - it doesn't sound plausible at all.
Human made summary: A telecom company had time synchronisation issues in a location of their infrastructure. They eventually used a single hardware GPS clock that fixed the issues but became their only authoritative clock in this location.
One day, the clock was turn off and on and it went 1024 weeks backwards because the time GPS time protocol sucks and use a week counter with too few bits.
Apparently a GPS clock can keep track of the time if it’s up and running when the week counter overflows, otherwise it has to take a wild guess. Apparently their hardware GPS clock used a hardcoded start time from its firmware instead of trusting a less reliable existing clock.
The article finishes with some AI looking suggestions to prevent such an issue to happen again.
Mine would be to have bought one or two more GPS clocks and not from the same provider.
Shout out to the Jeff Geerling video about this
https://www.youtube.com/watch?v=1T9xQy-dsQo
Author needs to decide which end of the stratum stack they want as the top.
This is the perfect description of Telstra as a company.
This reminds me of this: https://obamawhitehouse.archives.gov/blog/2015/03/26/why-we-...
I saw the talk version of that and it has lodged in my memory. The point he made was that "health.gov" was made by a bunch of teams that made sure that their part worked, but nobody was in charge of making the whole thing work.
This is sadly "normal" for large bureaucracies.
Every time I see a giant, multi-million dollar catastrophe, it's always "proper", "documented", "enterprise", "by the book", and... "a total failure". The reason is always that nobody actually cares about the final outcome, only the paperwork in front of them that they need fill out, the checkbox that needs to be ticked, or the compliance requirement that needs to be met.
This, in my opinion, is some part of how we are schooling people - do as you are told and follow the rules and some part of not having stakes in the outcome.
This was inspiring, thanks for sharing.
Why did you include the source tracker in your link?
Definitely not on purpose! I keep forgetting that ChatGPT does that now, which is rather irritating because I don’t use it to write comments but the tracker makes it look like I do.
I do use ChatGPT to resolve vague memories of articles read long ago into concrete URLs I can share, which is one of the best uses of the accursed things right now!
I assure you that I am an organic meat human like yourself and the article dates itself to the “before times” and can also be trusted to be free of contamination.
Lots of bloviating AI-prose. Here's the original report:
https://www.telstra.com.au/exchange/what-we-ve-learned-from-...
https://www.telstra.com.au/content/dam/tcom/dynamic-media-pr...
They're pretty readable, I recommend that second one (though it leaves some questions I have unanswered). It's a good example of how latent issues with a large system can go unrecognised for a long time before a seeming unrelated change can trigger a cascading failure
One thing it doesn't make clear – the GPS card only supported the original L1 signal, which is what causes the GPS week rollover issue. The newer L2C signal has a much longer rollover period (157 years vs 19.6 years), which means no rollover until next century. If the GPS card had supported the newer L2C signal, then likely this would not have happened even if the other misconfigurations had still occurred.
Direct link to the analysis report: https://www.telstra.com.au/content/dam/tcom/dynamic-media-pr...
All I could think of while reading this was “has no one heard of ntptrace?”
It's funny, but I had the same thing happen to me once. One day, overnight, my computer decided it was 2006.
It happened twenty years ago...
;-)
Is it really so that Time Division Multiplex timing in cellular networks is based on NTP time? I stopped reading after the article implied this - it doesn't sound plausible at all.
No, it absolutely is not.
That timing requires a very tight time, frequency and phase reference. It’s derived from local high stability GNSS-disciplined time references.
ai;dr
Human made summary: A telecom company had time synchronisation issues in a location of their infrastructure. They eventually used a single hardware GPS clock that fixed the issues but became their only authoritative clock in this location.
One day, the clock was turn off and on and it went 1024 weeks backwards because the time GPS time protocol sucks and use a week counter with too few bits.
Apparently a GPS clock can keep track of the time if it’s up and running when the week counter overflows, otherwise it has to take a wild guess. Apparently their hardware GPS clock used a hardcoded start time from its firmware instead of trusting a less reliable existing clock.
The article finishes with some AI looking suggestions to prevent such an issue to happen again.
Mine would be to have bought one or two more GPS clocks and not from the same provider.