I saw the talk version of that and it has lodged in my memory. The point he made was that "health.gov" was made by a bunch of teams that made sure that their part worked, but nobody was in charge of making the whole thing work.
This is sadly "normal" for large bureaucracies.
Every time I see a giant, multi-million dollar catastrophe, it's always "proper", "documented", "enterprise", "by the book", and... "a total failure". The reason is always that nobody actually cares about the final outcome, only the paperwork in front of them that they need fill out, the checkbox that needs to be ticked, or the compliance requirement that needs to be met.
They're pretty readable, I recommend that second one (though it leaves some questions I have unanswered). It's a good example of how latent issues with a large system can go unrecognised for a long time before a seeming unrelated change can trigger a cascading failure
Shout out to the Jeff Geerling video about this
https://www.youtube.com/watch?v=1T9xQy-dsQo
Author needs to decide which end of the stratum stack they want as the top.
This is the perfect description of Telstra as a company.
This reminds me of this: https://obamawhitehouse.archives.gov/blog/2015/03/26/why-we-...
I saw the talk version of that and it has lodged in my memory. The point he made was that "health.gov" was made by a bunch of teams that made sure that their part worked, but nobody was in charge of making the whole thing work.
This is sadly "normal" for large bureaucracies.
Every time I see a giant, multi-million dollar catastrophe, it's always "proper", "documented", "enterprise", "by the book", and... "a total failure". The reason is always that nobody actually cares about the final outcome, only the paperwork in front of them that they need fill out, the checkbox that needs to be ticked, or the compliance requirement that needs to be met.
Lots of bloviating AI-prose. Here's the original report:
https://www.telstra.com.au/exchange/what-we-ve-learned-from-...
https://www.telstra.com.au/content/dam/tcom/dynamic-media-pr...
They're pretty readable, I recommend that second one (though it leaves some questions I have unanswered). It's a good example of how latent issues with a large system can go unrecognised for a long time before a seeming unrelated change can trigger a cascading failure
Direct link to the analysis report: https://www.telstra.com.au/content/dam/tcom/dynamic-media-pr...
All I could think of while reading this was “has no one heard of ntptrace?”
It's funny, but I had the same thing happen to me once. One day, overnight, my computer decided it was 2006.
It happened twenty years ago...
;-)