Hacker News

Top stories

Live mirror
30 storiesupdated just nowView source snapshot
  1. Hister: A private search engine for the pages you visit and the files you keep(github.com/asciimoo ↗)
    47comments
  2. Fujitsu launches made-in-Japan next-generation CPU FUJITSU-MONAKA(global.fujitsu ↗)
    145comments
  3. CrowdSec Source Code Leak(crowdsec.net ↗)
    25comments
  4. Rate limits on GitLab.com are changing(about.gitlab.com ↗)
    83comments
  5. T. Rex Had a Body Temperature of 97 Degrees(nytimes.com ↗)
    32comments
  6. Towards Self-Driving Codebases(detail.dev ↗)
    12comments
  7. Infinite-Parameter LLMs: Generating and Adapting Weights from Live Data(arxiv.org ↗)
    8comments
  8. Why I didn’t sign the Fields medallists’ letter(gowers.wordpress.com ↗)
    163comments
  9. How GLM built its own inference infrastructure(z.ai ↗)
    226comments
  10. Zettascale (YC S24) Is Hiring ASIC/FPGA Engineers to Build Chips for ASI(zscc.ai ↗)
    discuss
  11. One year of sponsored Servo development(servo.org ↗)
    127comments
  12. Launch HN: Skillsync (YC W26) – AI chat sessions made portable across agents
    17comments
  13. Grand MS-DOS Gaming General MIDI Showdown(johnnovak.net ↗)
    3comments
  14. The American Religion of Self-Storage Facilities(newyorker.com ↗)
    167comments
  15. CCC invites all model citizens to 40C3(ccc.de ↗)
    113comments
  16. Show HN: Navier-Stokes Visualized as 1kB i386 demos(juandecos.github.io ↗)
    1comments
  17. Ask HN: How to recover Google auth after phone stolen?
    56comments
  18. Running Ubuntu on the Lenovo IdeaPad Duet(vhaudiquet.fr ↗)
    1comments
  19. Show HN: Share your AI Setup, Learn from others(mysetup.ai ↗)
    69comments
  20. The Return of Sail Power: Cargo Ships Are Turning Back to the Wind(gcaptain.com ↗)
    101comments
  21. TSMC revealing details about next gen A14 node(mapyourshow.com ↗)
    6comments
  22. Whoisinspace.com/(whoisinspace.com ↗)
    51comments
  23. Show HN: Craigslist for agent skills, curated by a human(skillbay.sh ↗)
    2comments
  24. LLM Classification Is Feature Engineering(minimallysufficient.com ↗)
    12comments
  25. My temporary PHP fix from 2014 has nearly 20M installs. Today I'm deprecating it(jakeasmith.com ↗)
    85comments
  26. Artificial intelligence now beats some of the best human forecasters(economist.com ↗)
    74comments
  27. Economic policy for AGI(deepmind.com ↗)
    13comments
  28. Stallman: Thousands Dead, Millions Deprived of Liberties (2001)(slashdot.org ↗)
    15comments
  29. Mastering Layout Engines in Graphviz: Dot vs. Neato vs. Twopi vs. Circo(visual-paradigm.com ↗)
    6comments
  30. Vinix – A modern operating system written in V(vinix-os.org ↗)
    42comments

AWS says it can't restore some data from mideast facilities struck by Iran

507 pointsby 1d agowsj.com
378 comments
1d agoHN ↗

Does this mean that even with 3 availability zones for Amazon S3 storage, that some data is lost?

1d agoHN ↗

Obviously what you understand is different from the reality after you actually follow all the footnotes.

1d agoHN ↗

Did they say its S3 data? Could also be single-az EBS or RDS.

1d agoHN ↗

For me-south-1 (Bahrain), all 3 data centres providing the redundancy were blown up by Iran.[1] The redundancy was localised to small geographic area and a single government--something customers of AWS were hopefully aware of when they entrusted AWS with their data.

It's always buyer beware for any claims of availability. Engineers completing a FMECA[2] will (or should) always state upfront what type of failure modes they've deliberately excluded (such as meteor strike) or else every FMECA would be full of failure modes that have never been measured, and are not worth anyone's time worrying about. These exclusions vary by application--a time capsule, seed vault, etc are intended to outlast wars and collapses of empires. Typically a bunch of data centres aren't designed to withstand such failures.

I do think however it'd be reasonable to include the prospect of war for calculating data centre / cloud service availability. Especially in a place such as Bahrain where the country is obviously concerned enough about the prospect of war to have built very permanent and expensive air/missile defence sites. New Zealand on the other hand--maybe not so important to consider.

[1] https://news.ycombinator.com/item?id=49033240

[2] https://en.wikipedia.org/wiki/Failure_Mode,_Effects,_and_Cri...

21h agoHN ↗

As far as I know, the attacks happened at different times. If Amazon knew that they had lost some data redundancy, shouldn’t they have been quickly mirroring that out of the region?

20h agoHN ↗

That would be a legal nightmare. They don't necessarily know what customers' data residency requirements are.

20h agoHN ↗

This. We have (well, had) customers running in me-south-1 and once the first AZ went down we wanted to proactively move their data to other regions even just as cold backups. But our legal department slapped that down pretty quickly.

20h agoHN ↗

Most likely, their own data residency terms prohibit this. It would be interesting to know if, when 2 out of 3 AZs got destroyed, customers got a heads up to move their data to a different region?

20h agoHN ↗

We received repeated, constant heads up to move our data by the first AZ much less second. The problem is that nobody is storing data in Bahrain unless there are data residency requirements for it.

nobody wakes up one morning and chooses to launch instances, CDN or S3 and would choose Bahrain as that without a requirement to, we were contractually and legally forbidden (in the middle as a vendor) to copy even encrypted data where we don't have the key out for redundancy, so the best we could do was tell our subcustomers to download all of their buckets to their office or some employee laptops at their office

3h agoHN ↗

I believe regional DR is the customer's responsibility per their Shared Responsibility Model.

21h agoHN ↗

AZs weren’t meant to be disaster resistant, eg, an earthquake or hurricane could take out a whole region.

Regions were always the scale of disaster isolation on AWS.

20h agoHN ↗

Regions are really the scale of disaster isolation only in extreme cases - such as global catastrophe (meteor strike taking out a city) or in this case, when actively targeted in war. I don't really see the same thing happening to a US or European region.

21h agoHN ↗

I think so.

Although the more paranoid AWS customers who turned on (and pay for) S3 cross region replication or similar cross region DR for other services would be fine.

21h agoHN ↗

I don't think this has any teeth. They don't compensate in the event of loss afaict.

21h agoHN ↗

Considering all the data they have globally, they might still be compliant.

21h agoHN ↗

Even if they have payable SLA on this, most SLAs have Acts of God and Acts of War exemption.

20h agoHN ↗

But do they have Act of Special Operation exemptions?

11h agoHN ↗

I'm going to reword my Terms of Service this second to add "any military operation" next to "acts of war". But I'm sure we'll then have to demonstrate whether paramilitary are assimilated to the military.

4h agoHN ↗

I was just reading an insurance policy and it said "any war, including undeclared wars"

The insurers always know how to weasel out of it.

14h agoHN ↗

US refuses payout for soldiers who died in Iran 'because it isn't a war'

In response to another question, Mr Vance rejected using the word “war” to characterise US operations in Iran, saying there was “no active shooting”.

https://news.ycombinator.com/item?id=49596955

21h agoHN ↗

They say it's "designed for" 11 9s, not guaranteed.

20h agoHN ↗

But if they want to design for extreme probabilities you need to account for tail risks, so their design should have included a missile defense system. At some point you need to start worrying about asteroid defense too.

11h agoHN ↗

Ah so that's why we need a lunar base. To uphold S3's 11 9s of availability

20h agoHN ↗

Making a probabilistic claim while excluding a factor that dominates those statistics is... is quite creative accounting.

20h agoHN ↗

Are you saying that most data loss happens because your data center gets blown up in a shooting war? Like, AWS is the first digital service provider to lose data in decades?

19h agoHN ↗

I'm saying that if you have eliminated more mundane failures like dying harddrives, cosmic rays and so on from your systems and your calculation ends up with 11 nines then actually those "force majeure" events are probable enough that they dominate whatever other residuals are supposedly hiding in those last 0.0000000001%.

The region has seen a bunch of wars in the last 100 years, so the annual war-rate is > 1%. Even if we generously add the assumption that only 1 in 100 wars affects a datacenter you can see that wars become a major source of correlated hardware failures that they need to solve to actually deliver that kind of reliability.

18h agoHN ↗

Cosmic rays and dying hard drives are not force majeure though.

17h agoHN ↗

You don’t want to blend probabilities like this, because the tactics you use as a consumer vary between the two. If you consider 11 9s like “object AFR”, you might build systems that are resilient to very occasional single object loss. And it’s useful to know at what rate that might occur.

Whereas with these force majeure events you’d want a complete DR setup, and it’s typically an async recovery. Here it is useful to understand the fault domain (single server or single building or multi-building) so you can plan.

Blending the two numbers doesn’t help you build better against the systems. And the force majeure events are rare enough that they won’t happen … until they do. I’m not sure that knowing the precise probability that Iran would attack a gulf nation would change the fact that if they do, you need to have a DR story.

7h agoHN ↗

Seems like begging the question to me. You can't blend the numbers because amazon didn't blend the numbers. If they did and miraculously still arrived at 11 9s then that would also cover things such as wars and natural catastrophes, e.g. because they do offsite backups internally.

5h agoHN ↗

I’m not saying you can’t blend the numbers, I’m saying you shouldn’t blend the numbers. Because one number doesn’t communicate what you actually need to know to build.

You want to know how reliable the service is in steady state. For example it’s useful to know that S3 is effectively lossless in steady state whereas EBS volumes have an AFR of about 0.1%. You build your apps very differently between S3 and EBS knowing this. You can build highly resilient applications on each, but you code them differently, informed by these design goals.

You separately want to understand the failure modes that will require you to fully recover from backup. For example knowing that cloud storage is resilient to everything but region failure would inform you that your backups should be out of the region, not just a bucket in the same region. You don’t get that perspective from just a 9s number.

19h agoHN ↗

Force majeure carveouts are really common in every type of contract.

You should check your home insurance contract, for instance... It likely would not cover an ICBM strike.

19h agoHN ↗

Somehow I feel like the biggest post-apocalyptic problem will be the loss of home equity due to uninsured damage causing a collapse of financial markets.

19h agoHN ↗

I think it's what most people comparing provider SLAs would expect

19h agoHN ↗

Are you suggesting that their technical documents have separate availability numbers to predict geopolitical events and war?

18h agoHN ↗

The sales pitch should change from "probabilistically we will NEVER lose your data" to "you are most likely to lose your data due to wars, terrorists, software bugs, someone losing the master encryption key, the government forcing us to...".

Offsite backups are sadly rare these days, and aws sales is the main reason why.

19h agoHN ↗

Creative accounting works and is good because it works. If your customers give you more money because you lied to them, but it's legal, then it's good.

15h agoHN ↗

Excluding war as force majeure in the Middle East is the same as excluding high tide as force majeure building a sandcastle at low tide.

9h agoHN ↗

This is very on point. While I am not legally trained in US law, and especially not in any Middle Eastern law, I have a enough knowledge on the law in my country here where I live…

Invoking force majeure requires the entity to prove all three following to be true:

A. That the event was unexpected and therefore unavoidable.

B. That the event was outside the control of the entity.

C. That the event made it impossible for the company to resolve the issue.

War in the region is as you say rather common unfortunately. The fact that AWS is used by the IDF (https://www.972mag.com/cloud-israeli-army-gaza-amazon-google...) should be considered a factor whether or not their data centre became a more likely target or not. What remains is the ability or not for AWS to do multi-location reduncancy.

3h agoHN ↗

As someone who ran a tech company in Northern Ireland at a time when violence was much more common than it is today, I can tell you that our force majeure clauses always disclaimed liability for "riot, violent disorder, civil commotion, and acts of unlawful civil unrest and terrorism".

How could it be otherwise - do you expect AWS to have its own army and missile defence system?

2h agoHN ↗

The difference between you and Amazon is hundreds of millions of dollars in political donations and lobbying. Amazon certainly has the political connections to make defense of its datacenters a national security priority of the strongest military on earth if it made it a priority.

21h agoHN ↗

The footnote which says that is the design durability against equipment failure literally begins:

In the unlikely case of the loss or damage to all or part of an AWS Availability Zone, data in a One Zone storage class may be lost. For example, events like fire and water damage could result in data loss

16h agoHN ↗

It seems unreasonable to blame Amazon here. The AZ was destroyed. Are they supposed to have missile/drone defense?

I'm not going to complain to DoorDash if my order is delayed due to a car crash

13h agoHN ↗

If a senior leader at Doordash swore to God that your sandwich would 100% guaranteed make it to you, regardless of whether or not there was a car crash, then yeah, maybe you should complain

13h agoHN ↗

Yeah I agree with you. If they made a specific promise to some unlikely case out of their control, then I would expect compensation.

Did Amazon make such a promise? They didn't as far as I know. My understanding is they provide specific guarantees like given an AZ outage, your data is still safe (provided you architect correctly).

4h agoHN ↗

Correct analogy would be: they promised to make 3 copies of my sandwich, from 3 different restaurants, and deliver them by 3 different cars, taking different routes, in case one of them gets in a crash. But in this case all 3 restaurants were bombed by Iran!

2h agoHN ↗

Wild berries has missle defence now, why not Amazon?

15h agoHN ↗

Hot take: if you have a service that is 11 nines reliable, but there is an underlying component whose reliability is lower, cap the nines to that component.

12h agoHN ↗

If they haven't changed it recently, 11 nines is the durability target by design, but it is not set in any SLA. the S3 SLA is focused on availability.

1d agoHN ↗

I'm sorry but your data is in another castle.

21h agoHN ↗

This should have been in that super Dario game.

21h agoHN ↗

This is the flipside of data residency requirements that countries are now starting to require. If the EU wants to keep data in the EU, then great, but when the war comes and energy and infrastructure are hit, people would have wished for backups in North America, Asia, and the middle east.

21h agoHN ↗

For long-term backups you want offline cold storage in an underground facility in a friendly jurisdiction, not a datacenter.

19h agoHN ↗

We've beer-o-clock wargamed this a bit.

If I had an "important enough" client, I think I'd store all out local (Sydney + Melbourne AWS cross region) data to AWS Singapore (to protect against Australian jurisdictional and political risks) and to a non AWS cloud provider in the EU somewhere. I reckon thatd be close to as resilient a pile of hard drives in an underground bunker, for significantly less setup and ongoing cost, while also being much more available when needed. (Can you imagine the queue at the underground bunker when multiple AWS regions get bombed? Or even imagine getting to the bunker in "a friendly jurisdiction" while a shooting war is taking place?)

We haven't worked out a decent solution to Visa and Mastercard payment network going down for more than a couple of cloud billing cycles though.

19h agoHN ↗

More likely than the network going down is you getting banned from the network because someone thought you were selling porn.

21h agoHN ↗

War did not randomly came. America intentionally caused it.

And has lawless goverment and unaccountable tech industry making it bad place for data.

21h agoHN ↗

I wonder if Amazon can sue the US govt bc they basically caused this material loss to their business. I'm sure they cannot. Maybe a lawyer can explain why?

21h agoHN ↗

In US courts you can sue anyone for anything but you might not win. The US government has sovereign immunity from most civil liability. As for the legal system in Bahrain I have no idea but hypothetically even if Amazon could somehow win a judgment they wouldn't be able to collect.

19h agoHN ↗

Why wouldn't they be able to collect? As long as Amazon does business in a country the legal system in that country can collect.

19h agoHN ↗

How would the Bahrain legal system force the USA to pay reparations for the war it started? Seize all US assets, the way we did with Russia?

10h agoHN ↗

I mean that is how it works in all legal systems. If you refuse to pay your assets will be seized. Which is why they would almost certainly not refuse.

9h agoHN ↗

What assets does the US have in Bahrain? Military bases? How do you seize a military base without getting blown up?

The US would obviously retaliate by seizing Bahrain's assets in the US, which probably includes 100% of Bahrain's money (small countries don't get to have independent financial systems).

13h agoHN ↗

I'm not a lawyer, but this seems like a difficult case.

For one thing, regardless of what the US did, they did not launch the attack that damaged the datacenters. That attack was provoked by US actions, but the attack was an intentional and voluntary act by Iran/IRGC. I don't think any of the usual settings for liability for the acts of others really apply here, but I'm not a lawyer.

You'd also need to find a court of competent jurisdiction. The US is out because of soveriegn immunity. The Military Claims Act prohibits claims that "arise from action by an enemy or result directly or indirectly from an act of the armed forces of the United States in combat" [1]. The Federal Tort Claims Act bars "any claim arising in a foreign country" as well as "any claim arising out of the combatant activities of the military or naval forces, or the Coast Guard, during time of war" [2] and precedent recognize a state of war without a declaration.

The courts in the country where the data centers were attacked might not be interested in addressing conduct of the US that didn't happen in their country. The courts in Iran are likely to consider the attack on the data center a legitimate act of war and not a tort; anyway good luck collecting against the US with a judgement from an Iranian court.

Amazon might have a better time making a claim against Iran, in the court where the attacks took place, but good luck collecting a judgement. Maybe a case in the International Court of Justice, but those have to be submitted by a country and involved countries must consent.

[1] https://uscode.house.gov/view.xhtml?path=/prelim@title10/sub... chapter section 2734 (b) (3)

[2] see page 28 and 29 of https://www.congress.gov/crs_external_products/R/PDF/R45732/...

21h agoHN ↗

If you can not restore data from EU based Amazon Datacenters because it’s destroyed … you will definitely have better things to do like packing your go bag or buying the last groceries for a while.

19h agoHN ↗

Border between France and Germany is a particularly well know peaceful area, especially Alsace-Lorraine.

21h agoHN ↗

I don't think this is a flipside, unless you're thinking from the PoV of data itself rather than its owners.

20h agoHN ↗

data residency requirements in the EU don't categorically exclude data storage in other countries. The EDPB explicitly recognizes encrypted backups, for example, as valid as long as the keys remain in the EU and there's a secure transfer mechanism (p. 30) exactly for reasons such as disaster recovery.

https://www.edpb.europa.eu/system/files/documents/2021-06/ed...

20h agoHN ↗

The EU is big enough to house multiple regions for multiple cloud providers, all in the same jurisdiction (so they can actually be used within data sovereignty requirements). Not true for most of the other places in Asia/UK/South America/etc.

19h agoHN ↗

I'm comfortable enough with the Sydney and Melbourne AWS regions - about 700km (400 miles) apart and with (at least) 3 AZs in each. If something takes out enough AWS datacenters to lose some of work's or client data stored across all that, the uptime and resilience of the CRUD platforms I'm responsible for will not be very high on my personal priority list. (At least on AWS datacenter is within 10km of my home. I'm hoping that well before Australia gets involved in the sort of geopolitical conflict that might mean missile strikes against civilian infrastructure, I'll have headed bush to hang out with my off grid friends)

21h agoHN ↗

No disaster recovery plan? No offsite backups? Someone failed to applied the most basic principles that have existed for decades.

21h agoHN ↗

The more dramatic contingency you have to plan for, the more expensive the plan gets.

Earlier this week I mentioned that if we lose enough data centres to bring our operation down, the first items in the to-do list becomes securing weapons, vehicles and fuel.

21h agoHN ↗

"Daddy, where were you when the flames reached our house?"

"I was in the office, reviewing Terraform plans"

20h agoHN ↗

I had the same discussion with a manager about the backups of financial contracts for cleaning school facilities.

He just couldn't get past the notion that if the six copies in four buildings across two states were all simultaneously physically destroyed, then most likely there are also no more schools left standing, and hence the contracts to clean them are null and void. Also, payment is now in booze and ammunition, not dollars.

20h agoHN ↗

to-do list becomes securing weapons, vehicles and fuel.

I toured a datacenter once back in the early 2000s and they showed me 30 days of generator fuel storage. When i asked them why 30 days and not 35 they replied "we're such a major customer of both electricity and fuel that if we don't get electricity or fuel for 30 days there's way bigger problems than your website not being online" hah.

20h agoHN ↗

That's probably already true for 7 days or less

20h agoHN ↗

7 days of unreliable electricity wouldn't be unheard of for a very large storm

20h agoHN ↗

Large storms regularly result in some customers with lack of utility power for more than 7 days for some customers. When storms take out major transmission lines and roads and bridges, you can end up with some pretty lengthy outages, and fuel deliveries will also be difficult.

Look at data center responses from Hurricanes Katrina and Sandy. This guy [1] was onsite for Katrina. Lost utility power on August 29. They did have some access to fuel on Sep 1, but pretty spotty until maybe the 3rd. Looks like power started coming back in some places Sep 8, and maybe widely restored on Sep 14th.

I would say, by 7 days in you'll probably have a good idea of if 30 days might not be enough.

[1] https://web.archive.org/web/20061206181753/http://interdicto...

19h agoHN ↗

You also need to understand whether the generator backup actually runs everything. Where I work it doesn't. Only "essential" systems get backup power.

And if the data center is more than about 5 years old it almost certainly was not planned with adequate backup power to run racks of GPUs.

10h agoHN ↗

I had a similar discussion in the low 1990s at a large electronic manufacturer that had a pair of Unisys mainframes, one in São Paulo and another in Manaus, in the Amazon region communicating over a satellite link (two large 3m dishes on both places). There was one question of what happens if both fail at the same time, and I pointed that anything that takes out São Paulo and Manaus at the same time is a civilisation ending event, and we shouldn’t worry too much about that.

Anyway, we had a load shedding agreement with a large bank across the street from São Paulo and we could switch over to their Unisys mainframe in a matter of minutes, and vice versa.

21h agoHN ↗

That depends on the data. If this is EBS or single-AZ S3, then from Amazon's perspective this was correct. Backup responsibly (for any data that does need to be backed up) lives with the customer, and Amazon has no way of knowing about that. EBS data data is unrecoverable, and that's what's reported.

Now if this was multi-AZ S3 or whatever then this would be significant.

The article does not tell us what products were impacted.

20h agoHN ↗

I was unaware that Amazon even sold single AZ S3. 20% discount. Doesn't seem worth it. By the time I commit to purchasing S3 space, it has to be important data.

I get that S3 is convenient and reasonably performant, but it is not cheap at all.

20h agoHN ↗

That’s simply not true. I use S3 (well GCS mostly) for data that I wouldn’t be upset if it’s lost. And I pay the zonal discount for it.

19h agoHN ↗

Google internally has lots of possible redundancy levels for data.

They don't sell any of the lower and less reliable levels to the public, I suspect simply because the reputational damage from losing user data is so bad, and the news will take no notice of the fact the user got a discount for less reliable storage.

17h agoHN ↗

Most of Google's customers wouldn't know how to choose anyways, if these were exposed. My memory was quite hazy but I recall having a discussion with my colleague on choosing which Reed–Solomon code for our Colossus files, and apparently the choice was down to RS(8,3) or RS(9,3). I don't think even as Googlers we really had enough information to make an informed choice. Comparatively it was much easier to decide which cells to use for multi-location replication in Placer.

19h agoHN ↗

You call it "discount" but it's a 20% discount on a 10x inflated price, so it's an 8x inflated price

19h agoHN ↗

Single AZ S3 has other benefits. The point isn't the price, it's that it's _highly performant_ since you can keep all of your reads in the same AZ

17h agoHN ↗

Yeah agreed - any ephemeral stuff I need is generally in DynamoDB - S3 (and database) are for permanent storage.

17h agoHN ↗

It's a great service for large caches. For example, we process a lot of imagery that we download from third-party providers. We save a lot of latency by storing the data in a single-AZ S3.

If it dies, we will just have to re-download the data.

18h agoHN ↗

EBS data data (sic) is unrecoverable, and that's what's reported.

I don't see where this is reported? TFA does not mention EBS. In fact, TFA seems to be nigh content-free, beyond "AWS (allegedly, and is uncited) says they cannot restore some data."

The article does not tell us what products were impacted.

… right … which conflicts with EBS being what's reported …

(I would agree with your point that if EBS, or some AZ-level data was lost, then, yeah, that's the contract.)

21h agoHN ↗

But that is something the customer needs to consider. AWS doesnt offer that as standard if your data is in one zone, and during a war even multiple zones in the same region may not be sufficient.

21h agoHN ↗

Nothing is ever real-time. Eventual consistency leads to some data are not backed up.

You talking as if this is some mom-and-pop shop that you run.

21h agoHN ↗

That isn't recovery from AWS's point of view. If the customer has data in another region, thats great for them but AWS isn't really a part of that, AWS doesn't know which data is fungible in every case. Sure they have some data is replicated, what they can't recover is the data THEY do not replicate.

20h agoHN ↗

Offsite to.. where? Sea? Data residency in Gulf states is very strict and basically nothing is leaving the countries

20h agoHN ↗

Typically 300 miles geographically but could be hard in some Gulf States

20h agoHN ↗

In a Gulf State 300 miles is still within ballistic missile range and any belligerent is going to target both places if at all.

Strictly speaking from a missile defense perspective there's an argument 2 sites are a waste of valuable interceptors.

19h agoHN ↗

They're likely going to target two datacenters, not the datacenters + your medium sized company's office NAS and the safe in the office manager's home.

(Encryption handles confidentiality concerns.)

18h agoHN ↗

(Encryption handles confidentiality concerns.)

Which is why data residency is such a stupid concept.

18h agoHN ↗

Yes and no. For example if you are doing Azure, technically Azure can see tenant traffic I believe and you need to use both a Platform Key and CMK for data rest. VMs need encryption at host turned on too.

There is nothing to say that a determined adversary may still get at your data so it needs to stay in country.

20h agoHN ↗

A datacenter not owned/run by a major US or Israeli company seems like it might be a good first step.

20h agoHN ↗

If a AWS customer chooses to store their data in a single AZ, that is a design choice. AWS is not taking a daily copy of a entire regions S3 cluster and driving it to some warehouse for a "just in case" situation. That is why Multi-AZ exists.

19h agoHN ↗

Isn't S3 claiming eleven nines of data durability?

https://aws.amazon.com/s3/storage-classes/

"Additionally, S3 stores data redundantly across a minimum of 3 Availability Zones by default, providing built-in resilience against widespread disaster."

I wonder if "can't restore some data" includes any S3 data?

I'd expect to lose EC2 instance EBS data in the event of a datacenter being destroyed, but I kinda assume I wouldn't lose S3 data? Now I'm wondering if RDS backups are more like EBS or S3...

20h agoHN ↗

If you had data at two facilities in different countries hundreds of miles apart (about 250 miles between Dubai and Bahrain), that would count as offsite backup most of the time.

Certainly, this event will inform people's disaster recovery plans, but when you're also looking at data residency requirements, small countries, and state level military action against your hosting provider, it can be hard to keep your data.

20h agoHN ↗

Not all data is allowed to leave all countries.

19h agoHN ↗

You apparently don't do business out here with the unwashed masses where "whadda mean with all that nonsense? It's cloud...it's by definition safe[1][2]!" is an all too common preconception.

[1] That's a quote, including the Boston accent. [2] The only one I had that was better was a C-level who said "why are you asking for all this money for security in Azure. It's Microsoft so it's already secure.". That, too, is a quote.

21h agoHN ↗

What a nightmare scenario to tabletop. How do you even begin to recover from something like this?

20h agoHN ↗

Works unless local laws specifically block you doing that which they do for some classes of data in some countries.

Multi-cloud in the same country (if that exists in the country and is far enough apart) maybe.

20h agoHN ↗

Do cloud providers even share data center locations so you can assess the "far enough" bit yourself?

20h agoHN ↗

No - and usually the reason is so they cannot be targeted.

19h agoHN ↗

You usually get city level location information. Depends on your definition for 'far enough' if that works for you.

me-south-1 is about 250 miles away from me-central-1, but that's not far enough in this instance. Given that, I think city level location information should be good enough.

250 miles is pretty good for weather or not specifically targeted destruction (wildfire / industrial explosions / arson), but it's clearly not enough if your data is in a building targeted in a regional war. Assuming datacenters remain targets in wartime, I think it's fair to assume if one datacenter in any particular country is attacked, all the rest of the datacenters in that country are likely to be attacked, too. In that case, offline storage (tapes and things) in inconspicuous locations might be the way.

9h agoHN ↗

Kinda they do? Atleast our company knows the exact location of ours

18h agoHN ↗

I wonder what exactly these laws prohibit. Like, does it apply to a fully encrypted cold-copy on Amazon Glacier? If you don't store the encryption key outside of UAE, I'd say this isn't even the same data that gets transferred to the third party, it's just some random blob. But I have no idea if the authorities of that country would agree and if it's even actually enforced for that matter, or if it's one of those laws that actually cause problems only if you follow them.

14h agoHN ↗

In this case, a local on-prem backup in your office-that-is-not-an-aws-datacenter would've worked and satisfy the law right?

(or an AWS Outpost thingy, assuming those things work even if the mothership is down)

19h agoHN ↗

For providers that just act as middlemen, I assume data that doesn't have a residency requirement is stored outside, so probably

a) letting customers in other areas know that their data is backed up to another continent

b) asking the AI model of your choice to translate the following into PR-speak: "Because of the boneheaded data residency requirements in this country, all your data is gone, and we weren't able to do anything about it - here's an empty copy of a re-setup version of whatever infrastructure we provide, glhf setting up everything from scratch, hope you had backups"

For customers who use such a provider or operate primarily in that area: Restore from local backups, or tell whoever depended on you that everything is gone and if you really didn't have backups, probably close up shop.

21h agoHN ↗

Backup both locally and everywhere no matter what.

21h agoHN ↗

50%+ of companies that lose all of their data go out of business in 6 months.

DR/BCP costs are readily justified by doing a Business Impact Analysis (BIA).. budget up to some fraction of risk cost * risk probability.

And a friendly reminder that replication isn't a tested data backup.

21h agoHN ↗

I think this is due to the data residency requirements in UAE. I'm working with a client in the health space and the government requirements requires me to store data only in UAE! Tried with AWS but they were not allowing any new instances and I had to go with Azure.

20h agoHN ↗

Hi, I'm from the future. You might want to consider storing the data somewhere besides an Azure datacenter in the UAE.

20h agoHN ↗

the government requirements requires me to store data only in UAE!

20h agoHN ↗

turns out reading comprehension skills have still not gotten better in the future

20h agoHN ↗

In the future, some companies begin to store their data on-premises away from big centralized datacenters. But many companies do not, due to costs and the general friction of changing how things are done.

If OP tells me the name of his company I can hop in my time machine and tell him how it plays out.

20h agoHN ↗

MENA is the on-prem capital of the world, don't worry.

9h agoHN ↗

off-topic but nice username. Fan of Charlie Kelly?

4h agoHN ↗

The 'at' in the middle tells me no. I also get Harvey Birdman, Attorney at Law vibes but that could just be me.

19h agoHN ↗

... but not only in an Azure datacenter!

(The obvious option here might be an encrypted backup on some hard drives in a safe in a local office.)

5h agoHN ↗

What other options I have?

I don't like Oracle. GCP has no servers in UAE. All other options were by local providers with not so friendly budget options

20h agoHN ↗

Are you me from the future? I'm not talking to future strangers. Tell me to deliver the bad news myself from the future, in the future.

11h agoHN ↗

It might be the you of Theseus. You really don't want to know what the future holds.

6h agoHN ↗

I'm waiting for a messenger from the future coming back to tell people that passkeys are stupid, and having passwords on post-it notes attached to the monitor was never a real problem, it's just the security industry that got threat modeling backwards for couple decades...

4h agoHN ↗

Don't worry. After the great captcha war of 2029 there was an awakening and things got better for a while (until the Flock smart glasses incident...)

7h agoHN ↗

Cool! I have a question about lottery numbers..

19h agoHN ↗

Hi, I'm from the past. When countries in the 2010s -- especially Western countries -- started seeing data residency requirements as an acceptable aspect of national policies, as opposed to a weird authoritarian thing that only China and Russia imposed on their citizens, we[1] spent a bunch of time explaining to their lawmakers that having geographical redundancy was a good thing, actually, and that you should stop insisting on where the data resided for jurisdictional purposes and start talking about where administrative access and encryption keys lived.

[1] OK, "we" here is probably just me -- it was one of those things where the chances of successfully convincing anyone was so small, and the commercial advantages of just nodding along, and then changing your product offering was so great, that really very few people raised it or had reason to. But somebody had to!

19h agoHN ↗

and start talking about where administrative access and encryption keys lived

Yeah, that was/is just another problem. Considering how that was actually handled in the real world before data residency laws came into force, I'm glad 'we' didn't convince those countries to put their citizens data at risk.

19h agoHN ↗

I'm not sure you were disagreeing with (past) me; but if you were, could you expand on your point?

19h agoHN ↗

I'm disagreeing with you. I in the before time, I had all sorts of conversations around this topic with any number of cloud providers that were like:

Us: We are concerned about our citizens (US) data, how are you managing the databases. Clout Provider (CP): They are only managed by fully background check employees. Us: Yeah, but where are they? What is their citizenship? CP: Um...mostly Eastern Europe. Lots in RU. (another CP proudly said "they're pretty much all in China...for cost containment"). Us: ...

Us: We are concerned about our citizens (EU) data, how are you managing encryption? CP: Everything is perfectly encrypted with hardware HSMs and all the FIPS and stuff. Us: So...where are the folks who run the HSMs? CP: Um...mostly SV. Some in the EU. Us: But can you assemble a quorum of US citizens for the HSM? CP: Of course! Us: ...

And on and on. Not to put too fine a point on it, many of us have no faith that vendors self policing international data protection in the face of government level pressure on companies and employees would work. Not that it can't, I don't think it would.

18h agoHN ↗

(I like the accidental pun of "Clout Provider" btw, which sadly conveys some of what they try to imply).

We may not be disagreeing that much. My argument was, and is, it's not about where the data is, it's about who has control over it. The counter-argument was "well if it's in another country, then we don't have jurisdiction, so it's going to be much harder". But what you need jurisdiction over is the people. Otherwise, you end up with multi-national corporate end-runs where you have shonky companies offering to store data locally, but who knows what department has control and access.

To be fair, the context I was having these conversations was countries arguing for data residency to combat the threat of mass surveillance (corporate and governmental) in the US, and the limited protections their users had relative to US nationals. But again, the problem is that it assumes that jurisdiction remains territorial: which is not how this was ever going to play out. The next wave after data residency requirements, beyond the usual extraterritorial intelligence community actions, was laws like the US CLOUD Act, the UK's Investigatory Powers Act, and Australia's TIA law, which effectively attempts to provide regular government departments and law enforcement with the legal ability to access data that would technically be on foreign soil.

My point was not that corporations should not self-police, but the concept of "it's stored here so we can oversee it" is not as clearcut as it seemed, and it risks introducing a new level of complexity to resiliently storing data. Which may be worth the price, but was never considered at the level this was discussed.

18h agoHN ↗

That's fair, and it sounds like we aren't that far apart. It is, in fact, about control. So I'll restate my central theme as "until the idea of enforceable data sovereignty requirements were enshrined in law, the cloud providers did not and would not delegate control of any body of data to 'controllers' that weren't in jurisdictions where they could be influenced/coerced to compromise that data". Was this a slippery slope/camel in the tent? Well...that's politics and it didn't have to be, but I see your point. But the reality is the push for data sovereignty wasn't done with the intention of enabling totalitarian follow-on legislation and it wasn't in and of itself a bad idea.

Best laid plans and all that.

19h agoHN ↗

Is there not more than one data center in your country? Is the power feed at your office too small to put a computer there?

19h agoHN ↗

I think both of these things are (very) often true, but it is also true that if I'm going to have backups, it is (all other things being equal) better to minimize correlated risk. The assumption in a lot of these conversations is that having the data "in one place" (ie inside a country) was "safer" than having it in "somewhere else". The tougher counterintuitive argument was that it can be safer to have data stored in multiple places, for some risk assessments -- and that for others, having data close by was less important, in the case of seizure or surveillance or illegal use, than who had legal or effective access to that data.

12h agoHN ↗

I think both of these things are (very) often true,

You think most countries have one data center or less, and most offices have less than two hundred watts of electricity supply?

19h agoHN ↗

This is a very engineer-centric view. I studied economics in school, so an analogy in that realm is ironically how all countries should specialize and raise the PPC curve. The reality of the situation was that in 2010 not many people understood how powerful big data actually was. Data sovereignty is actually quite logical when you consider the scale and power of not only the company, but the US as a whole. I can assure you that lawmakers were not thinking about efficient disaster recovery plans or back ups when they made the laws. You can also create reasonably diversified data silos within a country.

As an aside, it is quite crazy the world we live in. I am with the majority where I expected Amazon to be more redundant, but I still marvel at the assumption that a US dev can spin up multiple redundant and data sovereign servers in dozens of countries with efficient caching, failover and redundancy (enough to survive an earthquake or targeted missile attack) from their own home. Even a few hours of outage in a foreign country is considered unacceptable.

18h agoHN ↗

See my other answers, but briefly: no, they were not thinking about this, which is why we were raising it. I guess the counterintuitive point we were trying to get across is that with most things, the best way to keep it safe from being lost is to put it in a known place, and lock it away. But for data, the strategy -- for that scenario -- is to keep it in a lot of places, with heterogenous defense strategies. This is for data loss, of course, not access or surveillance or unlawful processing. But there is a cost as well as a benefit to deliberately limiting your options.

(I can feel someone saying "but surely having redundancy in one country is good enough, so I'll just say that I know relatively sane people who try to have hemispheric redundancy in their data, and also you never know when two different-in-every-quality-but one locations will suffer from the same disaster. Floods; heat-waves; national protests and strikes. It's surprising how often rare things happen!)

On your second point, it really is crazy. And also amazing that this is a capability that is -- or should be -- available to anyone in the world, not just in the US, and not just devs. Hopefully without also having to think about their data suddenly finding itself in a warzone.

17h agoHN ↗

This is for data loss, of course, not access or surveillance or unlawful processing.

This is why the minority of politicians who actually know about how this stuff works worry about where the data resides for jurisdictional purposes. If the government where the data resides can compel the folks who have physical and/or logical access to the physical machines that contain that data to give them access to that data, then that's game over for you.

«But you just don't permit that sort of breach to happen!» you might say. To which I reply "Yeah, right.".

Substantial physical separation of datacenters is very important, but the politics and policies of the location housing the data cannot be ignored.

12h agoHN ↗

I mean, in those rooms I was arguing over the best policies to prevent access and surveillance and unlawful processing, and what the potential cost-benefit analysis was. And what I was arguing against was an assumption that physically compelling all companies -- or worse, all citizens -- to keep their data within the borders of the host country, would protect you from these problems.

We'd have to explain that if the data was physically in Brazil, but hosted by a U.S. company, that would not stop that company from accessing that data remotely -- unless you specified that. We'd have to also explain that if you were intended to defend against US mass surveillance of non-US persons by the US intelligence services, intelligence services and SIGINT are univerally almost defined by their broad remit to target foreign nations on their own territory in violation of local law. And, finally, if you intended to use the prohibiting the movement of of data as a sanction against companies to punish them for violating data protection standards, as pre-GDPR law in the EU had as an ultimate last resort, and the GDPR often ends up relying on as a last resort, you would find that multinationals are more capable of putting up servers in your home territory and continuing to serve your citizens than they are of substantially changing their practices regarding data processing.

I don't want to sound nihilistic about this -- regulations can exist in these areas. But it's those politics and policies of the institutions with control over the data that are the most important part of this: not where the bits are kept. Especially when those bits are encrypted, and the keys and access controls are elsewhere.

5h agoHN ↗

But it's those politics and policies of the institutions with control over the data that are the most important part of this: not where the bits are kept. Especially when those bits are encrypted, and the keys and access controls are elsewhere.

Nah. Policies prohibit rule-followers from accessing data that you don't want accessed. Such policies are very important. But if you give your adversary effectively-unlimited physical access to the hardware where the bits are kept, that's game over. If you don't trust the governors of a region to honor the "don't tamper with this hardware" gentleman's agreement, and you very seriously care about preventing unauthorized access to the data that that hardware stores and processes, then you don't put that hardware in that region.

To point to a real-world example of this, there's not going to be an AWS Top Secret Cloud region in China, Russia, or -say- North Korea.

17h agoHN ↗

You can also create reasonably diversified data silos within a country.

I generally agree, although if a small country only had half a dozen or so redundant data centers then it would be relatively easy for a powerful adversary to wipe out all of the data centers and potentially have a significant economic impact on that country.

Having a backup data center in an ally country might make sense. Kind of like how I keep an encrypted backup hard drive at my parents house. Whenever I go to visit I pull it out and backup my laptop there too.

11h agoHN ↗

  > I can assure you that lawmakers were not thinking about efficient disaster recovery plans or back ups when they made the laws.

That's why the input of actual specialists in a field should be the one drafting the policies. I'm glad that professional lawmakers exist, I personally couldn't draw up a proper par if I had to, but they are not and can not be specialists in every field.

4h agoHN ↗

No, because "actual specialists" like to pretend power and politics don't exist. If you're a leader of a sovereign nation and let all your data be stored in US datacenters you're either stupid or corrupt

2h agoHN ↗

You are describing the US regulatory landscape, up until Trump's SCROTUS revoked the Chevron deference in 2024.

11h agoHN ↗

you know this is why china and russia were stealing data, so that they can provide you backups if you lost yours /s

11h agoHN ↗

I mean, it was clear and open that USA will spy on any data stored in there. Because foreigner do not get legal protections.

And second, the USA is in the middle of power grab that completely ensures any data stored there will be taken hostage wherever suitable for "negotiations".

9h agoHN ↗

start talking about where administrative access and encryption keys lived.

As soon as you start specify technologies, rather than "sovereignty" you end up needing to created specific legal tests to stop people getting around it.

"Data must be stored domestically" is a short hand for being held in the same legal jurisdiction. This means for somewhere like the UK, you get all that battle tested data protections law for free. (new laws require case history to be reliable. Ie, prosecuting under a new law is hard, because if its on the edge of being legal, it can create a precedent that undermines the entire law)

In civil code places, its different, but I don't know enough to offer even a half arsed opinion.

The reason why jurisdiction is important is because if you are storing data outside of your legal protection, when something goes wrong there is little you can do to discourage fuckery.

This is the problem with blinkered engineering thinking. Yes geographically distributed data storage is good. But as you also know, storing it in place with lots of other data, means that its a target. The more places its stored, the more physical security you need. This means that there is higher chance of people being bribed.

Its not a binary, its a multi-dimension graph, with no one answer. Every dimension has a tradeoff.

UAE's tradeoff was: not even trump would ignore all the wargaming that clearly shows kicking iran in the nuts would have inflation rising consequences

19h agoHN ↗

... in the UAE, so only less risky if your office is remote enough it doesn't also make a nice target (like, say, located within/near Dubai's Internet City).

17h agoHN ↗

and you don't upset whichever prince decides to compete

12h agoHN ↗

They're not firing nukes yet. Missiles don't blow up a whole building, let alone a city.

5h agoHN ↗

Not an option. Client is non-technical, a one person company

1h agoHN ↗

Managed on-prem perhaps? Setup a server at their office.

17h agoHN ↗

Check if the requirement is "only in the uae" or has to have the primary copy in the UAE

5h agoHN ↗

Yeah it's a law. Everything health related has to be in UAE.

7h agoHN ↗

Azure has outages every 6 hours. It's really pleasant.

5h agoHN ↗

data residency requirements in UAE

Is it a real law? I mean I would largely ignore those kinds of regulations on the basis of sheers stupidity. After all, lawmakers of the world tried to ban math for multiple times in the last few decades. Why would anyone consider conforming to laws like that? Especially since its not possible to enforce this law.

20h agoHN ↗

It was always a bad bet for billionaires like Bezos to become Trump enablers. You weren't buying a seat at the table, or the privilege of being left alone, you were just signing yourself up to be force-fed shit sandwiches over and over (And the shit-to-bread ratio gets worse as time goes on)

You should have used your considerable resources to fight. If only billionaires would oppose aspiring tyrants with the same zeal with which they oppose even minor tax increases.

20h agoHN ↗

He had little choice. The Trump tariffs could've been a massive, massive blow to Amazon, so I'm sure he felt he had to get out in front of them and buy some influence with the incoming administration.

See also Tim Cook. Doesn't make it right to suck up to Trump, but it was, and unfortunately still is, a rational move.

19h agoHN ↗

Bezos hasn't lost anything from this. He's only gotten richer.

6h agoHN ↗

Well, he did lose some data centers. Those aren't cheap.

The sort of global instability that is being created by all this is bad for everyone. It generally isn't good for business either.

We're so unstable and capricious now, countries are actually trying to decouple from American tech services like AWS and Amazon. That ain't great for Bezos.

Maybe inertia kicks in after Trump and things revert to the mean, but restoring confidence in the US again as a friendly, stable nation is going to be tough if we're always 4 years away from another Trump-type figure running the show.

20h agoHN ↗

I wonder if they'll start adding an underground bunker to new data centers so you can put an S3 replica there?

5h agoHN ↗

How long would it then take to be able to use the backed up data? Wait for a war to end and a replacement data centre to be built...? By that time most data probably no longer matters (e.g. business no longer exists).

It's more likely the entire data centre (not just backups) would need to be built underground (or cut and cover) at enormous expense. A price that perhaps for certain data sovereignty reasons the government of Bahrain (or companies in Bahrain requiring it) would be happy to pay?

Another way to do things on the cheap could be small-scale "covert hosting". Buy an apartment or house, maintain it to give an outside appearance of being an apartment or house, but inside it has a few racks of IT equipment. This has been done in the past for telephone exchanges in some countries, not for security reasons, but rather to hide an ugly bit of infrastructure that due to technology limitations of the time had to be located deep within a residential neighbourhood.

20h agoHN ↗

This interview with an AWS leader isn’t aging well, from CBS Sunday morning:

Pogue asked, "I don't mean to give anyone ideas, but let's say I figured out that one of these unmarked buildings was an AWS data center, and I blew it up. Are you saying that it's so backed up and redundant that you probably wouldn't notice?" Wood replied, "Yeah, you wouldn't notice. I mean, we might be a bit upset, but you wouldn't notice!"

https://www.cbsnews.com/news/cloud-computing-loudoun-county-...

20h agoHN ↗

multi-AZ doesn’t help against multi-AZ drones :)

19h agoHN ↗

black swan events.

this is one of one of those things - were in the current era either a cloud provider should provide automatic backups in another geographic zone.

if you're in us-east, then your back-ups should ideally be in eu-west + africa for redundancy.

18h agoHN ↗

It's not really a black swan event just an unlikely one. If it were one it would not have been something people talked about before.

18h agoHN ↗

That's absolutely something you can configure your S3 bucket to do if you want (I have one of mine replicating elsewhere).

The amount of data S3 stores "automatically replicating" to other geographical locations would make things prohibitively expensive, especially when you consider the daily delta, and how much of that is ephemeral or frequently mutated data that is stored. The bandwidth costs alone would be eye-watering, let alone the storage costs.

S3 cannot make any automated decision about whether data is, or isn't important, and if they did they'd only open themselves up to lawsuits if they guessed wrong. That's why it's made an option for the end user to enable replication if they want to, or choose to replicate their own data.

17h agoHN ↗

The amount of data S3 stores "automatically replicating" to other geographical locations would make things prohibitively expensive

It did work like this! And it was! My recollection is that the first S3 was out of SEA and had no user concept of region. Then “VDC” was added in virginia. That provided API endpoints in what became us-east-1. A bucket could be accessed from either location, and the original intent was for object store to replicate between them. By the time dub/eu-west-1 came along that was obviously not tenable; itd take 10s of gbs to replicate.

So S3 became regional. But the original sea/vdc deployments still had shared APIs and data in both regions. Your object would be stored in the region of the API you geolocated to via DNS, but read from either. _eventually_ all the data migrated to IAD, but those API endpoints were transparently proxying across the continent until 2013 or so.

And of course glacier had much more interesting takes on this with cross dc/az/region erasure encoding. But i dont think any of the wacky multi dimensional cross region stuff ever materialised in practice.

PS: hi!

13h agoHN ↗

Long time no see.

2013-2014 would be about right, they were still in SEA when I joined in 2013 but were actively working on decommissioning it.

18h agoHN ↗

The minute you cross borders you run into data sovereignty concerns. No idea if it applies at the subnational level (US states, Emirates in the UAE etc) but it wouldn't surprise me if it did in some cases.

16h agoHN ↗

It applies subnationally. Think about different gambling laws per US state, and if you’re running the infrastructure for that across state lines.

15h agoHN ↗

"Black Swan" would be a good name for an attack drone

6h agoHN ↗

Along with "Outside Context", "Tail Risk", "Impossible to Predict", "Force Majeure" and "Spanish Inquisition".

11h agoHN ↗

they should provide and option for multi-AZ with drone defenders now

19h agoHN ↗

"The damage to our infrastructure spanned multiple Availability Zones and exceeded what our regional and multi-AZ services are designed to withstand," AWS said in the status update

20h agoHN ↗

That is actually surprising to me. Claims like that are pretty common, they make sense and they should be true, so even though I don't really know AWS (/Backblaze/Azure/whatever) redundancy planning in enough detail, I used to trust them. It's really worrying when they outright say it will be ok, and then a week later it turns out to be not ok.

19h agoHN ↗

The caveat is always "if you're using the service correctly" which is not necessarily free. Meaning taking advantage of multiple geo zones, building in redundancy to your stack, etc. Like everything he said is possible if your technology stack living in AWS was designed to survive it. Everyone who has ever had the "we lost your data" email from AWS knows at the end of the day the cloud is just someone else's data center with neat provisioning tools and services.

19h agoHN ↗

The problem is that lots of people seem to be under the impression that they are doing it right because they are using AWS. They don't realize that AWS is a toolbox, not a 'ready made solution for redundancy against all catastrophes you are possibly exposed to'. They use that to their advantage by pricing such solutions at a level that people will either pay through the nose or will be left without recourse when AWS loses their data. It's stupid, but at the same time these beliefs are surprisingly wide spread.

17h agoHN ↗

If I were a bit more bloody-minded I would launch a service for vibe-coded apps that, under the hood, did everything "the right way" and just charged a flat fee + percent on the underlying.

I feel like if this was done correctly it would eat a bunch of the market, but I question how many people are actually willing to pay for "the right way". The last time I had that experience it was with Heroku which was quite a leaky abstraction.

17h agoHN ↗

I think a lot of people have built those services. But doing it right is more expensive, and people end up choosing the $5-$10/mo option over your $30/mo+ that does it right. Multiply those numbers by whatever multiple you want for higher end stuff.

12h agoHN ↗

And the LLM say: "You're absolutely right! We didn't need jaggederest's service. We can build it easily."

12h agoHN ↗

Well, one of the things I would be very, very focused on is LLM optimization, so that is something that could be mitigated, at least. Especially since I would be building with it, certainly, I would try to engineer it so that it was to Mr. Claude / Mr. Gippity's taste

12h agoHN ↗

1. There is no single "the right way", different app have different "right ways" 2. As for the flat fee, it only works initially when things are simple, as time goes on and your business and usecases you support grows, a flat fee won't work anymore.

11h agoHN ↗

I mean I would be limiting it to a very specific use case, and the flat fee would be "plus a percentage of the underlying", something like $50/mo + 20% of the AWS spend underlying.

And for that, with a simplified use case, well, they can scale up to $20k+ a month if they like, that would be ideal, and if they need a more complicated setup or enterpriseyness, migrate off with my blessing and available-not-required hands on support (again for a reasonable fee, maybe $10k if you want the white glove).

17h agoHN ↗

Not to be callous but even the 3-2-1 rule is pretty basic, the issue is that people don't apply it. But that's a hiring and strategy thing.

17h agoHN ↗

Yep, a lot of non-technical leaders believe that 'cloud' is synonymous with 'DR strategy' or even 'backup'. "We won't have to worry about being offline if our server goes down if we move to the cloud!" Some of these people fundamentally don't understand what the cloud is, their assumption is cloud means easy button that solves all your infrastructure and uptime problems.

16h agoHN ↗

You are pointing the fingers at the wrong people. Cloud providers pushed this idea onto executives via their marketing and conferencing channels.

Not disimilar to how AI naratives are pushed on executives these last two years.

14h agoHN ↗

But alongside the marketing blitz they also offer certifications for people who are supposed to execute these projects. Even the most foundational certification exam -- which simply tests your recall on what AWS service is for compute versus networking -- teaches and tests you about "Shared Responsibility Model" that draws the line at what the customer is still responsible for when they use cloud bases IaaS / PaaS / SaaS services.

If businesses are going to cloud but without engaging / listening to competent people who know these basics -- then the blame needs to be somewhat pointed back at those very business leaders I feel.

This is not obscure magical knowledge either that is tightly controlled. Any cloud vendor will freely teach you that. Or even a google search would.

12h agoHN ↗

You just have to take a look at the Rube Goldberg-esque abomination that is the AWS web interface for a few minutes to realize that it's not that simple. Granted, most executives won't do that...

11h agoHN ↗

Or encounter their CLI and it's lovecraftian combination of positional arguments, named arguments, piping, json documents, etc, distributed largely by random lot, as far as I can tell.

11h agoHN ↗

No. If you're the leader you should know what you're leading. It's literally the job.

But the problem is right there: "non-technical" leaders. Never work for one if you can avoid it.

1h agoHN ↗

All Tesla employees reading this are fervently agreeing, between their tears.

"We should build vacuum delivery tubes for passengers!"

16h agoHN ↗

And for a time, it even seems true as long as the natural disasters that impact various customer offices happen to not impact that one particular data center…

15h agoHN ↗

I mean, can't cloud kinda solve all of those if you just use the right things?

17h agoHN ↗

No, that quotation on the GP clearly states that AWS has enough redundancy within the same region that they will continue all services running on it if a datacenter is destroyed.

It's very clearly not about you being able to set-up redundancy for yourself.

17h agoHN ↗

That's actually true. AWS is designed to survive one datacenter being offline (which happened more than once, btw). When the first DC in ME was hit, AWS continued working normally, with only a few services experiencing issues.

But it's not designed to survive TWO datacenters going offline, and in a permanent fashion.

17h agoHN ↗

That's fair. The only reference was about one datacenter blowing up. You can't have unlimited redundancy.

12h agoHN ↗

Aren't they supposed to have 3 copies of everything? If so, two datacenters going offline at the same time should mean almost zero data loss.

11h agoHN ↗

No.

_You_ are supposed to have three copies of everything - and two of those should be off AWS.

1h agoHN ↗

Some systems in AWS do, but they typically need a quorum of at least two copies to work. Other systems only have 2 copies.

And in this case, it looks like all three DCs were damaged.

11h agoHN ↗

Get some training. You dont even understand the core concepts, and the difference between a data center and an availability zone...

17h agoHN ↗

Pretty much this, it’s your responsibility to use their tools to make sure your data is managed in such a way that any data destroyed is already elsewhere before the event.

12h agoHN ↗

which is not necessarily free

Not just in terms of service costs, but in time and complexity. In many cases building out that complexity is complicated and difficult. And sometimes the functionality you need isn't supported in the regions you use.

12h agoHN ↗

Amazon never said you had to mirror your data across multiple regions to prevent Amazon-caused data loss.

11h agoHN ↗

Yes it is, they didn't have adequate redundancy for this extremely foreseeable event (military site getting hit by missiles).

1h agoHN ↗

Negligence is not malicious action.

Both are bad, but having a missile hit your shop is never going to mean you are causing violence.

16h agoHN ↗

The devil is always in the details. Somehow I feel that when we offload the responsibility to some one else we get this feeling that the other person/entity would be doing full diligence and whatever else is required to carry out the job perfectly. However in reality most of the times they just do the bare minimum to pass your evaluation criteria to get the job.

14h agoHN ↗

This is almost always the case. It's one of the frustrating things about the software industry; because everything is much more complicated than the customer is able to comprehend, a software company can promise anything and the customer can't actually verify.

So any software company/project which actually took the time and effort to fully handle the enormous complexity, they can't sell themselves based on that fact because every other company (who didn't invest the effort) is also claiming it and the customer has no mechanism to verify the claims until some major rare event occurs.

And most of the effort is required precisely to handle those 1% of rare situations.

12h agoHN ↗

Especially with Amazon, who are well known for squeezing every last bit of profit from their employees, contractors etc., it doesn't really sound surprising. "Offsite backups?! Sure, you could have had that if you had found the right page in the AWS console and if you would have paid 50% extra!"

10h agoHN ↗

Perhaps, cheapness is always an factor but I'm potentially reading it as the customers perhaps not wanting data to move outside of the country and with AWS only have one datacenter in said country produced this result.

4h agoHN ↗

but the general assumption is your data isn't going to get lost if you use Amazon. so once the trust is gone it's gone.

8h agoHN ↗

Oh sorry, you picked "AWS Backup" but you actually needed to use "Backup AWS" to solve that problem. Perhaps hop on a call with our sales engineering and cost magnification teams to guide you to a better, more solutioned, tomorrow?

7h agoHN ↗

I sounded all legit until the comma after "solutioned". I suspect that might be a fake.

7h agoHN ↗

It all depends on how *off* off-site really is. You know, *off* isn’t a binary, it is a spectrum.

12h agoHN ↗

its also the fact that a lot of fundamental systems work in trade offs.

Do you want performance, or correctness.

Well, if you want performance you use write through caching and in the case of distributed storage: more nodes confirming the block before returning. Huge performance cost.

Outsourcing this just means someone else makes these tradeoffs, they will prioritise the general case- and they’re even more incentivised to move the needle towards things that are most visible to the end user.

In this case, performance.

You won’t notice that theres a third commit server off-site (unless that site is bombed), but you will notice slower writes- and the general case says that people will express comparative dissatisfaction with weaker performance and use it as a justification to use another provider.

7h agoHN ↗

if you want correctness* you use write through caching.

My mistake, if you want performance you choose write-back caching, and fewer nodes need to acknowledge the write. Sorry for clumsily typing the inverse when I was in a morning haze waking up :(

11h agoHN ↗

It's funny because this applies to children cleaning home bathrooms as a Saturday chore as well.

11h agoHN ↗

I think it applies to all humans of all ages in one way or another, these varying imperfections is what makes us what we are

10h agoHN ↗

its not true. If everyone actually did the bare minimum, everything would screech to a halt immediately. In fact people doing that imo is part of what causes the decline of empires. You can, in fact, trust people further than you can throw them.

Its a cousin of the mindset that the reason people don't steal is because they think they will be caught and rationally weigh up based on the value they gain and the chance of loss that its not a worthwhile action.

No, most of the time people steal because they think its wrong, and they dont want to do it.

Public trust is a real thing and varies massively by country. America is notably extremly low on this metric

7h agoHN ↗

One needs the state, the other needs intrinsic morals. One is a tax on all actions, the other forms working societies.

6h agoHN ↗

How are the behavior described imperfections ?

5h agoHN ↗

Huh, I was thinking US Congresspeople.

3h agoHN ↗

funny enough, there have been some huge problems involving congressman with both bathrooms and children in regards to moral diligence.

7h agoHN ↗

That is why contracts are more than one page in length. The details matter. I can remember receiving a contract class in Afghanistan about something as simple as moving gravel. Yeah, use your imagination with that and whatever absurd cartoon like fantasy you could dream up regarding "moving gravel" is still probably less strange than the real events that occurred.

I don't write contracts for a living, at least yet, but my learning so far is:

* clear goals: where is the end point and what does the product look like once it gets there in all required details

* defined test criteria: this is where you get to sue when they fuck shit up

* measures: there must be predefined measures. These can be wildly unrealistic at the start and require changes as the work occurs, which is ok, but there must be defined performance criteria that all parties are held to before work completion. In other worlds this is rewarded with bonus targets and penalties

4h agoHN ↗

but aren't amazon contracts pretty one sided?

5h agoHN ↗

I would love if someone could tell me the name for this phenomena. FWIW, using an LLM incites this same undue assurance as well. No matter how much you know that an LLM might just be hallucinating, it still happens anyway. It's a hard instinct to fight.

4h agoHN ↗

This is so true. After 20 years in the tech business I have so rarely seen perfect execution. It's mostly scrambling and chaos and a miracle anything works in the first place.

11h agoHN ↗

“It’s not the cloud. It’s just someone else’s computer.” - MM

11h agoHN ↗

It's probably a dedicated government type thing where the data is housed seperately from normal AWS

10h agoHN ↗

Well, listen, you do know that even when the largest most ambitious and most sophisticated things are shipped, each owner of each specific part (could be many many owners) basically , to the best of their ability, prayed that nothing particularly bad happens when it’s shipped off. Truly, that’s the best a mortal human can do, pray their part doesn’t break.

So then that big thing comes to you. It’s all kind of … held together by a prayer …

Trust me I’ve worked at these big places. You wouldn’t believe how much fucking luck and grace from God is allowing you to do anything with your digital life. It’s a mindfuck of a tangled mess out there, eternities worth of written code that only God ensures works together at this point, only to get more hidden with AI.

10h agoHN ↗

It really is important to understand the failure modes that the durability model accounts for and what it doesn't. It only accounts for "normal" failures, like an HDD reaching end of life.

For example, you mention Backblaze. Backblaze has public posts about their durability model. They claim to use 17:20 Reed-Solomon erasure encoding. That means there are 20 shards of a blob, and you can lose 3 of them and still reconstruct the blob.

Think about that for a second. If they store 4 shards in a datacenter, that means that a loss of that one datacenter is sufficient to lose the blob, forever. That entails that blobs are sharded across a minimum of 7 data centers, or the loss of one data center might mean permanent data loss. Which one do you think is true? (In fact it's pretty clear from Backblaze's public posts that they don't shard across data centers at all, only across racks within a data center.)

Now, AWS's availability guarantee — not their durability guarantee — entails that they use a less cost-effective erasure coding ratio. S3 is designed so that your blob is available even if a whole AZ goes down, and it's well known that most AWS regions have only 3 AZs. Therefore, if you tolerate the same number of shards lost to HDD failure as Backblaze in your durability model (3), then you might need 17:30 erasure coding to get the same durability and the required availability. That means S3 is storing way more physical bytes than Backblaze — 1.76x the logical size of the blob, instead of Backblaze's 1.18x. That's more expensive, but it also gives you better availability.

Which is also why One Zone S3 is cheaper — if you don't care about the availability guarantee, S3 can do what Backblaze does and save 33% on physical bytes, and they pass on 40–50% of those savings to the customer (this is fairer than it sounds — there's more overhead than physical storage bytes).

But here's the thing. AWS has more redundancy built in than Backblaze because they make availability guarantees in addition to durability guarantees. BUT the durability model is the same, which is why Backblaze can claim equivalent durability to S3. S3 in fact has better durability — they can survive the permanent loss of an AZ without necessarily losing blobs stored there (with the exception of One Zone blobs), and Backblaze cannot. But that's not actually a factor of the durability model, which is just taking into account normal events like HDD failure. Instead, S3 has durability that's more resilient to AZ loss because of their availability model. It's a side effect that isn't actually part of the durability promise!

6h agoHN ↗

As far as I'm aware Backblaze stores data only within a single datacenter (for a given region). This likely made sense in their original business model of being "offsite" copy of data.

But it very much breaks down for B2 where they're now storing original data. I hope they rethink this model. You do get what you pay for. There's a reason they're cheap.

4h agoHN ↗

My org isn't a Backblaze customer but a Wasabi one. Wasabi offers the ability to replicate buckets to other zones at double the cost, of course.

I would imagine Backblaze would offer something similar for original data storage for their customers to choose.

5h agoHN ↗

The question is about one AWS facility. The headline refers to facilities. The guy's answer may well be an honest and truthful one.

5h agoHN ↗

How much data is stored in the average aws data center? Assuming they do regularly back ups, it makes sense that recent data can’t be instantly backed up and so there will always be some data in transit or queued up no?

20h agoHN ↗

Isn’t the problem that multiple datacenters in one zone were blown up?

19h agoHN ↗

They had nine years and more money than god to build redundancy in an unstable region.

19h agoHN ↗

If it's an unstable region, then it might not result in a good return, especially since a blown up data center is 100% loss.

13h agoHN ↗

This is focusing in too close.

Bezos has been working with Trump. The instability has been hugely exacerbated by Trump. The issue lies at home.

9h agoHN ↗

Yeah, Iran blowing up half the Middle East? Totally Bezos' fault.

18h agoHN ↗

They did build redundancy but most of it was bombed.

17h agoHN ↗

If they're region-locking data appropriately (i.e. for people to comply with domestic storage / gdpr-style requirements) they really can't. Bahrain is only about 300 square miles / 80k hectares.

They could, of course, open other regional data centers in other countries, or say "data in this geo zone may be in any of X, Y or Z" countries, but for the latter that pretty starkly limits some of the major customers they'd have, I would guess, and for the former, well, they have other geo zones already, so if people weren't replicating to them, I'm not sure why adding me-east-1 me-west-1 me-central-1 would fix that issue, they just wouldn't replicate there either.

13h agoHN ↗

Yes, this, and the article seems pretty clear that only "some data stored exclusively in Bahrain" is affected. Add to that the option to store data with reduced redundancy (in S3, for example), and I don't really see what the drama is about.

Amazon only claims "99.999999999% durability" per year, even for the properly replicated stuff.[0]

[0] https://docs.aws.amazon.com/AmazonS3/latest/userguide/DataDu...

12h agoHN ↗

stored exclusively in

Now I'm thinking about legal/contractual rules that might force that kind of geographic risk.

I mean, logically you could have the Allowable Location send pre-encrypted backups to anywhere in the world, except (A) laws and regulations aren't always logical and (B) you still have the problem of keeping the decryption keys somewhere safe without leaving the key jurisdiction.

19h agoHN ↗

that quote is definitely making its way into a lawsuit

13h agoHN ↗

Yes. It's super hard to answer interviews or customer demos. Every sentence you say should include all preconditions that were said in previous answers, within the same context, because you may be quoted. You should think in live about all possible contexts in which your app might be used, and your speech should be as detailed as a contract.

Obviously here, he should have mentionned that they can recover a hit on a single data center, provided the customer chose multi-AZ hosting. That's probably why companies run their ads on "This watch is a legacy for your children" rather than any material claim.

16h agoHN ↗

There's something reassuring in this for me.

There's a lot of magic & handwaving from hyperscalers like AWS about redundancy. I always wondered about some of the engineering to make this absolutely (and literally) bullet proof. At the end of the day most of their answers when you push hard enough involved paying 2-3x to run everything across multiple zones/regions, and lots of awareness in your application to handle this.

In any case, I think it's good that when a data center blows up the data is lost. Noteworthy for future skynet situation, etc.

15h agoHN ↗

I think it's good that when a data center blows up the data is lost. Noteworthy for future skynet situation, etc.

Doesn’t really apply, because the only reason data was lost is because customers chose not to replicate it to other regions, either because of legal data residency requirements, cost, or just not bothering.

If Skynet wants to make sure it’s backed up, none of that prevents it from doing so. Although it would be amusing if Skynet was stopped by a billing alert when it tries to copy itself to another region.

15h agoHN ↗

aws egress fees: the real hero we didn't know we needed

12h agoHN ↗

AWS promises to keep data safe without requiring cross region redundancy. If you choose not to trust them that's valid but then why use them at all?

11h agoHN ↗

> AWS promises to keep data safe without requiring cross region redundancy.

A perfect example of a claim they never made.

9h agoHN ↗

S3 is advertised as having 11 9's. That means in the entire history of S3 they've only lost a handful of objects, and will continue to lose objects at this same rate.

5h agoHN ↗

You are confusing the concepts, and applying the 11 nines to the wrong problem.

Durability is about: "If I successfully store an object in S3, how unlikely is S3 to permanently lose that object because of storage failures?"

It does not answer: "Will I be able to access that object, after a rain of Shahab-1 or Shahab-2 burn all data centers in the 3 availability regions across which my S3 bucket is spread out..."

2h agoHN ↗

Eleven nines covers ordinary infrastructure failures, not acts of war that physically destroy the underlying facilities.

It's similar to any engineered artifact, like a skyscraper. They have a structural safety factor that defines their resilience to failures that the system is designed to tolerate, it doesn't make them immune to missile strikes.

8h agoHN ↗

You'd be surprised how many customers are lead to believe this by a combination of opaque marketing, aggressive sales, and willful ignorance.

I worked at a 2000 person shop where the CTO moved us from 1 on-prem dc (we were begging to add a secondary site for years due to outage risk) to 1 AWS region.

It's not exactly straightforward to run multi-region across an alphabet soup of AWS services without decent configuration / application awareness, paying at least double.

9h agoHN ↗

They should have asked AI to implement the redundancy.

7h agoHN ↗

How would have that changed anything?

7h agoHN ↗

First of all they could write an amazing slop blog PR post about it afterwards

7h agoHN ↗

If you still think AI produces slop then you aren't paying enough for your AI.

7h agoHN ↗

AWS has a history of making hand wavy explanations on the robustness of its infrastructure that were misleading at best. See comments about how regions are truly independent only for everyone to find out if us-east-1 goes down you could still be down even if you built in other regions. Those dependencies were not well documented and previously hand-waved away by AWS when it boasted about how its regions were truly independent.

Similar here, there’s a lot of detail that got hand-waved away by a sloppy “yeah we good” puff PR answer.

7h agoHN ↗

I suspect at this point, given the nature of the AWS org.. they actually don't even know themselves either.

12h agoHN ↗

Unless you have a synchronous like setup where you don’t acknowledge data writes unless the remote has aconowledged them first, you will lose data in case your datacenter is hit by a warhead.

Now, there are a few things to consider:

- AWS best practices recommend multiple AZs for workloads and cross-region backups for things like databases and other “stateful” data

- You have to read the fine-print on what AWS offers in terms of recovery: do they reffer to their own infrastructure when they say “you won”t notice” or your data

When Google’s Paris colocation facility was flooded and all AZs there went dark, they sent an email saying “restore from backup in another region and if we can restore your data, we will make it availbale to you”. They did not even issue credits for the downtime.

9h agoHN ↗

> and all AZs there went dark,

An extraordinary statement itself that shows the difference between a proper cloud where they AZs are at least 60 to 100 miles apart...and Google or Microsoft... pretend clouds...where those AZs are just firewalls across the same data center...

7h agoHN ↗

All the AZs flooded simultaneously?!

How big was this flood?

6h agoHN ↗

Well, this is how the world found out that Google AZs are just different rooms in the same building on the same floor.

5h agoHN ↗

Azure and GCP definitions of AZ allow a single datacenter to have multiple AZs. AWS has by far the strictest definition.

12h agoHN ↗

Clearly since the MBAs took over AWS standards are not anymore what they used to be. That marketing guy should not be talking to the press, as he does not have the skills, and if somebody happens to say...our data center we wont lose any data if we have an issue, without qualifying it will depend on what quality of service, and usage of our services you setup ...is the type of technical answer that should make a hiring interview stop at the moment.

He is also violating an enormous amount of compliance requirements, by disclosing the location of the data center, and having strange people inside making a tour. Did he vet the crew and their accompanying party? Did one of them accidentally left some kind of device within the insider perimeter? There at least one or two ISO certifications he is violating there. As customer I would be asking questions...

AWS always made very clear they wont copy your data to another region as only you know what your compliance and data residency requirements are. But at the same time they always said, its up to you to come your with your disaster recovery strategy based on your project requirements. And it has always been the case copying your critical data to another region is one of the first things on your check list.

And their Well Architected Framework and other docs make this plenty clear:

"It is a good practice to always make backups of your data, and copy these to another site (such as another AWS Region)."

Also...

"All DR strategies require that data sources are backed up within the AWS Region, and then those backups are copied to the recovery Region."

And also for single-Region / Multi-AZ architectures:

"Where possible, you should also copy data backups to another AWS Region as an additional layer of protection."

"AWS Architecture Blog — Disaster Recovery Architecture on AWS, Part II" has a whole section named "Backup to another AWS Region": "By copying your data to another Region, you can handle the largest scope of disasters."

https://aws.amazon.com/blogs/architecture/disaster-recovery-...

Or "Creating backup copies across AWS Regions" - https://docs.aws.amazon.com/aws-backup/latest/devguide/cross...

This whole thread of people literally saying , "on no I trusted them...I did not know they could lose my data", with no technical context...is the the kind of incompetence I would expect from a generation raised on vibe coding and llm prompt driven miseducation...

6h agoHN ↗

That marketing guy should not be talking to the press, as he does not have the skills

They have way more skills than the technical folk in doing that

10h agoHN ↗

articleAuthorId: 30dacdbc-6a8a-11e2-9d12-0018fe8a00b0

articleAuthorName: cbsnews (hidden byline)

articleSecondaryAuthors: n/a

articleEditors: n/a

Is the author a human or machine? Google shows 1 result for "30dacdbc-6a8a-11e2-9d12-0018fe8a00b0" and Brave Search shows 5 results.

7h agoHN ↗

1. This wasn’t a strike on a single data center, it was strikes on many data centers.

2. Since that interview, AWS has started selling versions of their storage that isn’t redundant. It is not surprising that when AWS sells non-redundant storage that some data is not recoverable.

7h agoHN ↗

This wasn’t a strike on a single data center, it was strikes on many data centers.

He didn't say "but if you hit many data centres then there is a problem". The premise was if you hit data centre, user won't notice, without caveat that there is a limit.

6h agoHN ↗

Amazon's Matt Wood isn't worried: "If something does happen or we have a power event or there's a flood in one specific location, that data is held redundantly in other locations as well."

Pogue asked, "I don't mean to give anyone ideas, but let's say I figured out that one of these unmarked buildings was an AWS data center, and I blew it up. Are you saying that it's so backed up and redundant that you probably wouldn't notice?"

Seems pretty silly to argue, but he certainly did say "one specific data center", and I don't think anyone even non-technical will conclude "it's safe if they all go down at once" from this statement.

3h agoHN ↗

The Iranians hit one specific data centre at a time, no?

5h agoHN ↗

Is there anything more loathsome than an interviewee who fails to do the journalist's job for them?

6h agoHN ↗

The whole "flawless victory" thing isn't aging well either.

20h agoHN ↗

They say "some" data, i wonder what percentage that really is. I haven't seen pictures but I find it hard to imagine all of me-south-1 was completely leveled to the point where's there's just nothing left. On the other hand, if you have 100 rows of racks and then randomly take out a contiguous 10% across both rows and columns it may be functionally equivalent to taking out everything.

19h agoHN ↗

I wouldn't be surprised if the engineers said "we can probably recover between 20-30% of the data but it will cost 200 hours of engineering and the data will be 7 months old by then" and the beancounters said "we'd rather have one news cycle rather than the news watching what we can and cannot recover + save those 200 hours, we'll just say it's all gone".

20h agoHN ↗

Uh.

Uh-oh.

It's not clear from their messaging if multiple availability zones were severely damaged, or if the damage to one availability zone was simply more than they planned for. If it's the latter, that's a big uh-oh.

The wording certainly seems very careful:

"The damage to our infrastructure spanned multiple availability zones and exceeded what our regional and multi-AZ services are designed to withstand"

19h agoHN ↗

Hey but that's exactly as per design. It is the customer's responsibility to store stuff elsewhere as DR backup, not AWS.

19h agoHN ↗

That's the excuse they'll say, yes, then we quote back to them "eleven nines" and they come up with an excuse for that too

18h agoHN ↗

Guess those multi-AZ promises have an asterisk when actual missiles are involved. Makes you double-check your own off-site backups.

18h agoHN ↗

"..even if something happens only once in a billion requests, that means it happens multiple times per day within S3."

but one in a trillion...

18h agoHN ↗

   me (Middle East)
   ├── me-south-1 (Bahrain)                                DOWN since 2026-04
   │   ├── mes1-az1                            me (Middle East)

├── me-south-1 (Bahrain) DOWN since 2026-04 │ ├── mes1-az1 DOWN │ │ └── mes1-mct1-az1 (Oman, Muscat) │ ├── mes1-az2 DOWN since 2026-03-01 │ └── mes1-az3 DOWN ├── me-central-1 (United Arab Emirates) │ ├── mec1-az1 │ ├── mec1-az2 DOWN since 2026-03-01 │ └── mec1-az3 DOWN since 2026-03-01 └── il-central-1 (Israel, Tel Aviv) ├── ilc1-az1 ├── ilc1-az2 └── ilc1-az3 DOWN │ │ └── mes1-mct1-az1 (Oman, Muscat) │ ├── mes1-az2 DOWN since 2026-03-01 │ └── mes1-az3 DOWN ├── me-central-1 (United Arab Emirates) │ ├── mec1-az1 │ ├── mec1-az2 DOWN since 2026-03-01 │ └── mec1-az3 DOWN since 2026-03-01 └── il-central-1 (Israel, Tel Aviv) ├── ilc1-az1 ├── ilc1-az2 └── ilc1-az3

13h agoHN ↗

Ok claude, now format that in a way humans can read

11h agoHN ↗

Sorry, I accidentally posted it twice while fixing it and on a slow connection.

18h agoHN ↗

The data center layout should look something like this:

    me (Middle East)
    ├── me-south-1 (Bahrain) DOWN since 2026-04
    │   ├── mes1-az1 DOWN
    │   │   └── mes1-mct1-az1 (Oman, Muscat) ???
    │   ├── mes1-az2 DOWN since 2026-03-01
    │   └── mes1-az3 DOWN
    ├── me-central-1 (United Arab Emirates)
    │   ├── mec1-az1
    │   ├── mec1-az2 DOWN since 2026-03-01
    │   └── mec1-az3         DOWN since 2026-03-01
    └── il-central-1 (Israel, Tel Aviv)
        ├── ilc1-az1
        ├── ilc1-az2
        └── ilc1-az3

Not sure about the Muscat local zone, whole me-south-1 region has been reported down despite Muscat still being operational.

If someone had told me a year ago that a whole AWS region could go down I'd called them crazy, but now me is close to exactly that happening.

See also previous discussion: https://news.ycombinator.com/item?id=49033240

17h agoHN ↗

us-east-1 is finally looking pretty stable for once.

17h agoHN ↗

11.3 Force Majeure. Except for payment obligations, neither party nor any of their affiliates will be liable for any delay or failure to perform any obligation under this Agreement where the delay or failure results from any cause beyond its reasonable control, including acts of God, labor disputes or other industrial disturbances, electrical or power outages, utilities or other telecommunications failures, earthquake, storms or other elements of nature, blockages, embargoes, riots, acts or orders of government, acts of terrorism, or war.

12h agoHN ↗

It's not terrorism because it's self-defence and it's not war because Donald Trump says so

12h agoHN ↗

It's beyond the reasonable control of Jeff Bezos. Even though he supports trump with billions of dollars, he didn't ask for this particular war.

12h agoHN ↗

I just gave the crazy man a lot of money knowing he wants to start a holy war. I did not want him to fund this particular holy war that turned out to be disadvantageous to me!

11h agoHN ↗

Kind of like building your house on a floodplain and claiming insurance for a flood. They'd still say it was an act of god.

6h agoHN ↗

Well, flood is an act of God. "Floodplains" is another lie invented by those "scientists" to get you pissed, probably on a break from their usual attempts to convince you magnets are not a genuine miracle.

3h agoHN ↗

Unless you have flood insurance in the US which is what those policies are specifically meant to cover and are backed by FEMA (for now)

2h agoHN ↗

Not true. Bezos could have purchase air defense munitions for his data centers. There is also the argument that this strike is the result of Bezos financially supporting Trump's second term and that without his support Trump's war would never have started

11h agoHN ↗

including acts of God

Serious question: how do we know something is or isn't an act of god?

5h agoHN ↗

In this case I suppose we can take the Ayatollah's word for it?

5h agoHN ↗

It's a legal term of art not directly tied to religion any more. Act of God is basically "natural disasters"; eg: hurricanes, tornados, earthquakes, etc.

17h agoHN ↗

Seems like their multi-AZ redundancy didn't account for missile strikes. Good reminder to always have your own backups.

17h agoHN ↗

There are many regulations over there which prevent data from leaving the country.

Sometimes those rules can have serious consequences.

13h agoHN ↗

We collectively gave this company trillions over the years and they still can’t get it right.

12h agoHN ↗

I wonder if this will cause a mini cyber insurance crisis. I don't think any of those data loss plans have been tested at scale.

12h agoHN ↗

Ahhh the downside of “data sovereignty” rules and the de-globalization meme strikes again. I’d bet a bajillion dollars the reason this happened is due to government thinking it’s a good idea to make it illegal to store data outside their country.

Hey EU, take note of this next time you create silly data residency requirements that don’t allow data to travel outside your region. Encryption is an easy solution to multi-region residency…as long as the European Commission doesn’t stupidly keep trying to make encryption illegal too!

Hint: Russia absolutely knows where your data centers are.

11h agoHN ↗

The EU is big enough to have several geographically independent data centers. without leaving the sovereign territory.

10h agoHN ↗

Yes there’s never been division or wars in Europe and the EU will always exist…

Germany and France in a debt spiral and turning inward/hyper-nationalist while massively re-militarizing means the EU is a safe place to structure your data with zero redundancy.

I wouldn’t bother worrying about key industries and functions, since, as history has shown, the EU is a bulletproof institution that no country has ever left.

And of course every data center in Europe has Israel-grade air defenses, it’s not like they are just sitting ducks for a fleet of drones to take out within 24 hours.

9h agoHN ↗

This is a valid point if you're not so confrontational about it.

2h agoHN ↗

The only valid tone to take with people who are on a crusade to ban encryption is disdain and mockery.

7h agoHN ↗

The EU is quite big; there are seven AWS regions there, each with three AZs.

11h agoHN ↗

They can recover most of it? I assumed that those AZs had been offline for so long because they were like gone gone.

11h agoHN ↗

My personal bet is there's an 80% chance this is caused by some internal bootstrapping problem that they've messed up. AIUI all the main cloud vendors are in trouble here. The automation project I work on is expressly designed to help folks solve this DR/bootstrapping problem. Soo many people get this wrong. Of course missiles don't help things, but I'd bet AWS is primarily to blame here. I'd love an actual technical report of why they can't recover things.

7h agoHN ↗

I work on the same kind of thing, and while we think hard about bootstrap problems, we always find new surprising ones. The problem is you never know until you do it, and creating a faithful test of restarting giant systems is economically impossible. Because if you say to the boss, "look, I need 1 million now to test against a maybe 100 million loss, maybe in 10 years" they don't give you the money (and rightly so).

5h agoHN ↗

Even if you did the $1MM test there is very low likelihood that the $100MM event would be fully mitigated 10 years down the line (after who knows how many changes - physical, logical, and even in the org chart).

The only way to approach readiness here is repeated investment - like one team doing the deep dive and another pulling cables and then constantly doing pre- and post-mortems.

10h agoHN ↗

So many AWS simps in this thread that don't actually understand how AWS scrimped out and fucked their customers. You can't all suck up to AWS at the same time, you need to serialize.

10h agoHN ↗

If all your backups are with AWS you don't have backups at all.

10h agoHN ↗

I suppose this is why Microsoft and Google want to build data centers in the Netherlands. Its not cheap but it is safe.

10h agoHN ↗

I wonder what kind of "some data" is: military installations in Arab states and Israel?

9h agoHN ↗

Stuff hosted in overpriced cloud has actually no backups ? Unbelievable!

7h agoHN ↗

I think people are overthinking this. I'm the first one to criticize AWS (I run a competitor, carolinacloud.io) but I don't think we should hold it against them when they lost data due to their hard drives being physically bombed.

7h agoHN ↗

Of course not.

The surprising thing is that multiple availability zones were bombed simultaneously.

And to my knowledge, they haven’t even yet said 2+ AZs were compromised.

6h agoHN ↗

Turns out that "availability zones" in cloud parlance seem to have nothing to do with geographical availability. Apparently they're logical splits and are more about billing than anything.

5h agoHN ↗

How often does an AZ go down while the other 2 still work?

5h agoHN ↗

An Availability Zone is one or more discrete data centers with separate and redundant power infrastructure, networking, and connectivity in an AWS Region. Availability Zones in a Region are meaningfully distant from each other, up to 60 miles (~100 km) to prevent correlated failures, but close enough to use synchronous replication with single-digit millisecond latency.

https://docs.aws.amazon.com/whitepapers/latest/aws-fault-iso...

They are far enough apart that tornados, floods, and fires cannot simultaneously impact multiple.

5h agoHN ↗

10s of miles is close enough that they are functionally in the same place in a military context though. The latest cheap massive wave attack drones fly hundreds of miles and hour which puts the DCs less than a minute apart by flight time. This means if you want protection against those risks you have to use multiple regions instead of, or in addition to multiple AZs.

4h agoHN ↗

A targeted, coordinated attack can compromise multiple AZs.

But AFAIK AWS has only reported data loss in mec1-az2.

Cross-AZ services like S3 should have no data loss, unless there is more damage than AWS has currently published.

4h agoHN ↗

Elsewhere in this threat comments mention AZ turning out in practice to be different areas of the same building, or just a matter of opinion and some config.

4h agoHN ↗

Why not? Why shouldn't we hold it against Amazon when they are making business deals with deeply unequal societies that rely on slave labor? This is what happens when you make deals with authoritarians. Expecting the blowback to not effect you is a child's mindset.

3h agoHN ↗

I think that's past the point OP was making. Amazon makes some pretty impressive guarantees of data durability but those guarantees only hold up if you take their advice and replicate data to other regions.

It's like saying "Oh this car manufacturer claimed that their cars were the safest in the industry but couldn't prevent this driver from flying through the windshield" and omitting the fact the person wasn't wearing a seatbelt.

1h agoHN ↗

No, it's more like willing to work with dictatorships across the world brings inherit risk.

No one was forcing Amazon to make deals with oligopolies and monarchies in the middle east.

No one was forcing Amazon to engage with US imperialism.

Like these things have extremely large threads connecting to one another. Now AWS, along with other American big tech companies, are legitimate military targets. This is what happens when you engage in these practices. You don't get to serve the interests of US imperialism and claim you're innocent. They deliberately signed up for this, there is a long history of this in US imperialism. It's nothing new.

Next time you don't want missiles raining down upon your data center, don't engage yourself in imperial politics. You won't be shocked next time when someone punches you in a face.

7h agoHN ↗

What AZs were affected?

I don’t think I’ve seen them say?

7h agoHN ↗

When doing Disaster Recovery (DR) planning as a SRE in the New York City area, I sometimes use the phrase "Hurricane Sandy 2" to describe an event so big that it knocks out the power/compute/etc for an entire region.

I may just start using AWS Bahrain 2 as an additional example.

5h agoHN ↗

Pretty soon we won’t need backups at all, just have AI regenerate all the data.

3h agoHN ↗

Ex AWS. AWS had one zone specific products. If those were in the affected AZ that data is lost. AWS has multiple multi-AZ products. S3 and EFS are good examples. If you loose an AZ for an extended periodic of time, the dataplane of those services will migrate shards around until you get back to your durability goals.