136 comments

[ 0.21 ms ] story [ 51.3 ms ] thread
Does this mean that even with 3 availability zones for Amazon S3 storage, that some data is lost?
Obviously what you understand is different from the reality after you actually follow all the footnotes.
Did they say its S3 data? Could also be single-az EBS or RDS.
All 3 data centres providing the redundancy were blown up by Iran.[1] The redundancy was localised to small geographic area and a single government--something customers of AWS were hopefully aware of when they entrusted AWS with their data.

It's always buyer beware for any claims of availability. Engineers completing a FMECA[2] will (or should) always state upfront what type of failure modes they've deliberately excluded (such as meteor strike) or else every FMECA would be full of failure modes that have never been measured, and are not worth anyone's time worrying about.

I do think however it'd be reasonable to include the prospect of war for calculating data centre / cloud service availability. Especially in a place such as Bahrain where the country is obviously concerned enough about the prospect of war to have built very permanent and expensive air/missile defence sites. New Zealand on the other hand--maybe not so important to consider.

[1] https://news.ycombinator.com/item?id=49033240

[2] https://en.wikipedia.org/wiki/Failure_Mode,_Effects,_and_Cri...

As far as I know, the attacks happened at different times. If Amazon knew that they had lost some data redundancy, shouldn’t they have been quickly mirroring that out of the region?
That would be a legal nightmare. They don't necessarily know what customers' data residency requirements are.
This. We have (well, had) customers running in me-south-1 and once the first AZ went down we wanted to proactively move their data to other regions even just as cold backups. But our legal department slapped that down pretty quickly.
Most likely, their own data residency terms prohibit this. It would be interesting to know if, when 2 out of 3 AZs got destroyed, customers got a heads up to move their data to a different region?
We received repeated, constant heads up to move our data by the first AZ much less second. The problem is that nobody is storing data in Bahrain unless there are data residency requirements for it.

nobody wakes up one morning and chooses to launch instances, CDN or S3 and would choose Bahrain as that without a requirement to, we were contractually and legally forbidden (in the middle as a vendor) to copy even encrypted data where we don't have the key out for redundancy, so the best we could do was tell our subcustomers to download all of their buckets to their office or some employee laptops at their office

AZs weren’t meant to be disaster resistant, eg, an earthquake or hurricane could take out a whole region.

Regions were always the scale of disaster isolation on AWS.

Regions are really the scale of disaster isolation only in extreme cases - such as global catastrophe (meteor strike taking out a city) or in this case, when actively targeted in war. I don't really see the same thing happening to a US or European region.
I think so.

Although the more paranoid AWS customers who turned on (and pay for) S3 cross region replication or similar cross region DR for other services would be fine.

They guaranteed 11 9's durability, didn't they?

e: Yep

https://aws.amazon.com/s3/storage-classes/

Amazon promises that level of durability for S3, which is a global service.

I assume this article is not specifically talking about S3, although right now archive dot today isn’t loading the article for me. I doubt AWS lost any S3 data.

Your EBS volume is just a virtual disk on a single computer.

I don't think this has any teeth. They don't compensate in the event of loss afaict.
Considering all the data they have globally, they might still be compliant.
Even if they have payable SLA on this, most SLAs have Acts of God and Acts of War exemption.
But do they have Act of Special Operation exemptions?
The SLA excludes force majeure.
[delayed]
Are you saying that most data loss happens because your data center gets blown up in a shooting war? Like, AWS is the first digital service provider to lose data in decades?
Force majeure carveouts are really common in every type of contract.

You should check your home insurance contract, for instance... It likely would not cover an ICBM strike.

Somehow I feel like the biggest post-apocalyptic problem will be the loss of home equity due to uninsured damage causing a collapse of financial markets.
I think it's what most people comparing provider SLAs would expect
Are you suggesting that their technical documents have separate availability numbers to predict geopolitical events and war?
Creative accounting works and is good because it works. If your customers give you more money because you lied to them, but it's legal, then it's good.
The footnote which says that is the design durability against equipment failure literally begins:

> In the unlikely case of the loss or damage to all or part of an AWS Availability Zone, data in a One Zone storage class may be lost. For example, events like fire and water damage could result in data loss

(comment deleted)
I'm sorry but your data is in another castle.
This is the flipside of data residency requirements that countries are now starting to require. If the EU wants to keep data in the EU, then great, but when the war comes and energy and infrastructure are hit, people would have wished for backups in North America, Asia, and the middle east.
data residency requirements in the EU don't categorically exclude data storage in other countries. The EDPB explicitly recognizes encrypted backups, for example, as valid as long as the keys remain in the EU and there's a secure transfer mechanism (p. 30) exactly for reasons such as disaster recovery.

https://www.edpb.europa.eu/system/files/documents/2021-06/ed...

The EU is big enough to house multiple regions for multiple cloud providers, all in the same jurisdiction (so they can actually be used within data sovereignty requirements). Not true for most of the other places in Asia/UK/South America/etc.
I'm comfortable enough with the Sydney and Melbourne AWS regions - about 700km (400 miles) apart and with (at least) 3 AZs in each. If something takes out enough AWS datacenters to lose some of work's or client data stored across all that, the uptime and resilience of the CRUD platforms I'm responsible for will not be very high on my personal priority list. (At least on AWS datacenter is within 10km of my home. I'm hoping that well before Australia gets involved in the sort of geopolitical conflict that might mean missile strikes against civilian infrastructure, I'll have headed bush to hang out with my off grid friends)
"Use the cloud" they say.

"It's the only way." the true believers say.

"You can't run your own computer systems.", they say.

"Trust us.", they say.

“No one ever got fired for buying IBM.”

What's you point? That if you had run your own DC in that region (because that was your business requirement) then you'd have better missile defense than AWS?

Or maybe AWS or DIY, you are always responsible for geographic diversity?

Anyone losing data over this lost it because they'd literally told AWS to only store it in one place.

HN shitposting is the point. If you want actual pragmatic commentary, use Lobsters.
You don't need better missile defense than AWS. You don't need missile defense at all because you won't be a target. 99.999999% of the land has no missile threat on it. You are actively increasing the threat to your business by running it on the same servers that military contractors run their software on.
> What's you point?

Not the OP, but the point is that decentralized infrastructure is a lot more resilient to attacks of any kind.

this is reductionist to the point of absurdism.

It can be true that using cloud storage, using managed services, and paying a premium is still worth it for a lot of people and organizations, and while not perfect, still a hell of a lot better compared to the fully in your control tape backups that you distribute to different physical locations every week.

I'm not a particular fan of relying on one provider or vendor lock in, but to pretend they don't provide a service with failure rates that are low enough to be very useful is a very short-sighted take.

Can you post the specs on your missile defense system for your home lab?

How are you dealing with the rise of low-cast swarm attacks from drones?

Is it land-based, sea-based or space-based and at what point of the trajectory do you target and do you use jamming?

I would but then I'd have to kill you.
No disaster recovery plan? No offsite backups? Someone failed to applied the most basic principles that have existed for decades.
The more dramatic contingency you have to plan for, the more expensive the plan gets.

Earlier this week I mentioned that if we lose enough data centres to bring our operation down, the first items in the to-do list becomes securing weapons, vehicles and fuel.

I had the same discussion with a manager about the backups of financial contracts for cleaning school facilities.

He just couldn't get past the notion that if the six copies in four buildings across two states were all simultaneously physically destroyed, then most likely there are also no more schools left standing, and hence the contracts to clean them are null and void. Also, payment is now in booze and ammunition, not dollars.

> to-do list becomes securing weapons, vehicles and fuel.

I toured a datacenter once back in the early 2000s and they showed me 30 days of generator fuel storage. When i asked them why 30 days and not 35 they replied "we're such a major customer of both electricity and fuel that if we don't get electricity or fuel for 30 days there's way bigger problems than your website not being online" hah.

That's probably already true for 7 days or less
7 days of unreliable electricity wouldn't be unheard of for a very large storm
Large storms regularly result in some customers with lack of utility power for more than 7 days for some customers. When storms take out major transmission lines and roads and bridges, you can end up with some pretty lengthy outages, and fuel deliveries will also be difficult.

Look at data center responses from Hurricanes Katrina and Sandy.

I would say, by 7 days in you'll probably have a good idea of if 30 days might not be enough.

You also need to understand whether the generator backup actually runs everything. Where I work it doesn't. Only "essential" systems get backup power.

And if the data center is more than about 5 years old it almost certainly was not planned with adequate backup power to run racks of GPUs.

(comment deleted)
That depends on the data. If this is EBS or single-AZ S3, then from Amazon's perspective this was correct. Backup responsibly (for any data that does need to be backed up) lives with the customer, and Amazon has no way of knowing about that. EBS data data is unrecoverable, and that's what's reported.

Now if this was multi-AZ S3 or whatever then this would be significant.

The article does not tell us what products were impacted.

I was unaware that Amazon even sold single AZ S3. 20% discount. Doesn't seem worth it. By the time I commit to purchasing S3 space, it has to be important data.

I get that S3 is convenient and reasonably performant, but it is not cheap at all.

That’s simply not true. I use S3 (well GCS mostly) for data that I wouldn’t be upset if it’s lost. And I pay the zonal discount for it.
You call it "discount" but it's a 20% discount on a 10x inflated price, so it's an 8x inflated price
Single AZ S3 has other benefits. The point isn't the price, it's that it's _highly performant_ since you can keep all of your reads in the same AZ
But that is something the customer needs to consider. AWS doesnt offer that as standard if your data is in one zone, and during a war even multiple zones in the same region may not be sufficient.
Nothing is ever real-time. Eventual consistency leads to some data are not backed up.

You talking as if this is some mom-and-pop shop that you run.

That isn't recovery from AWS's point of view. If the customer has data in another region, thats great for them but AWS isn't really a part of that, AWS doesn't know which data is fungible in every case. Sure they have some data is replicated, what they can't recover is the data THEY do not replicate.
Offsite to.. where? Sea? Data residency in Gulf states is very strict and basically nothing is leaving the countries
Typically 300 miles geographically but could be hard in some Gulf States
In a Gulf State 300 miles is still within ballistic missile range and any belligerent is going to target both places if at all.

Strictly speaking from a missile defense perspective there's an argument 2 sites are a waste of valuable interceptors.

They're likely going to target two datacenters, not the datacenters + your medium sized company's office NAS and the safe in the office manager's home.

(Encryption handles confidentiality concerns.)

A datacenter not owned/run by a major US or Israeli company seems like it might be a good first step.
Do you think they could ask for a backup
If a AWS customer chooses to store their data in a single AZ, that is a design choice. AWS is not taking a daily copy of a entire regions S3 cluster and driving it to some warehouse for a "just in case" situation. That is why Multi-AZ exists.
Isn't S3 claiming eleven nines of data durability?

https://aws.amazon.com/s3/storage-classes/

"Additionally, S3 stores data redundantly across a minimum of 3 Availability Zones by default, providing built-in resilience against widespread disaster."

I wonder if "can't restore some data" includes any S3 data?

I'd expect to lose EC2 instance EBS data in the event of a datacenter being destroyed, but I kinda assume I wouldn't lose S3 data? Now I'm wondering if RDS backups are more like EBS or S3...

1/f noise strikes again
If you had data at two facilities in different countries hundreds of miles apart (about 250 miles between Dubai and Bahrain), that would count as offsite backup most of the time.

Certainly, this event will inform people's disaster recovery plans, but when you're also looking at data residency requirements, small countries, and state level military action against your hosting provider, it can be hard to keep your data.

You apparently don't do business out here with the unwashed masses where "whadda mean with all that nonsense? It's cloud...it's by definition safe[1][2]!" is an all too common preconception.

[1] That's a quote, including the Boston accent. [2] The only one I had that was better was a C-level who said "why are you asking for all this money for security in Azure. It's Microsoft so it's already secure.". That, too, is a quote.

What a nightmare scenario to tabletop. How do you even begin to recover from something like this?
Backups in a different region?
Works unless local laws specifically block you doing that which they do for some classes of data in many countries.

Multi-cloud in the same country (if that exists in the country and is far enough apart) maybe.

Do cloud providers even share data center locations so you can assess the "far enough" bit yourself?
No - and usually the reason is so they cannot be targeted.
You usually get city level location information. Depends on your definition for 'far enough' if that works for you.

me-south-1 is about 250 miles away from me-central-1, but that's not far enough in this instance. Given that, I think city level location information should be good enough.

Not applicable in Gulf states
For providers that just act as middlemen, I assume data that doesn't have a residency requirement is stored outside, so probably

a) letting customers in other areas know that their data is backed up to another continent

b) asking the AI model of your choice to translate the following into PR-speak: "Because of the boneheaded data residency requirements in this country, all your data is gone, and we weren't able to do anything about it - here's an empty copy of a re-setup version of whatever infrastructure we provide, glhf setting up everything from scratch, hope you had backups"

For customers who use such a provider or operate primarily in that area: Restore from local backups, or tell whoever depended on you that everything is gone and if you really didn't have backups, probably close up shop.

Well… they have a good excuse.
How about ... stop bombing other countries? Trump is like a professional liar. From "no more forever wars" to "hey this is what must be done now" in a second. He is almost as good as Putin with regards to lies - the ultimate agent Krasnov. Minus the apparent dementia now.
It's a war that he started to benefit Israel, and is now causing suffering for so many others.
Israel had been trying to get every president to bomb Iran for decades. There's a reason they wouldn't do it themselves. They finally got a president stupid enough to listen.
You can't have a bunch of deranged mullahs threaten the entire region and world with missiles or nukes. Chanting "death to America" for half a century and killing thousands of Americans directly or indirectly. America should have blasted the hell out of the mullahs when they took hostages about 50 years ago. This is the first administration in a long time with the cohones to do something about it.
This is not exactly a nuanced view of the conflict, and in either case, the fact that you don't like that someone on the other side of the world is chanting death to America doesn't give you a bonus card for a free attack.

Seriously, it's like people, when deciding whether to launch a war or not, are not thinking "how will the other side react and will this conflict benefit me" but instead they are only thinking "does this nation deserve to get hit".

Well, news flash, your moral outrage does not translate into you not suffering more than your opponent during a conflict. It's a completely separate issue, and a personal issue between you and your priest or rabbi. When it comes to starting wars, you have to look at military capabilities and long term outcomes, not "does this nation deserve to be attacked".

You mean the same people who were told to keep the hostages until after Reagan got elected?

The same people who bought weapons sold by the Reagan administration to fund the Sandistas?

That’s some revisionist history you’ve got there.

> This is the first administration in a long time with the cohones to do something about it.

Containment worked WAY WAY better than "doing something about it" and surrendering the strait to them. They "solved it" by aerosolizing asbestos everywhere and still have no cleanup plan. Sometimes containment works much better, which is why intelligent foreign policy worked the problem from that angle.

Wasn't victory declared 1 year ago and also a number of months ago?

Do you know why Iranians were chanting "Death To America"? Because we toppled their democratically elected government in 1953 and installed a brutal dictator so that we could continue to pillage their oil.(1) When the mullahs finally toppled that dictator in 1979 they had reason to be angry with us. This is especially true because almost immediately afterwards, in the 1980s, we used our proxy Saddam Hussein to launch a war against Iran, in which over a million Iranians were killed. During this time we provided Hussein with the means to make both chemical weapons and biological weapons (2), which Iraq used extensively against Iran. I suspect if another country did the same to us, we'd be chanting "Death to XYZ" too.

(1)https://en.wikipedia.org/wiki/1953_Iranian_coup_d%27%C3%A9ta...

(2)https://irp.fas.org/congress/2002_cr/s092002.html

We're chanting "Death to Iran" - isn't turnabout fair play?
> You can't have a bunch of deranged mullahs threaten the entire region and world with missiles or nukes.

Can and do. It's called the United States of America.

Backup both locally and everywhere no matter what.
50%+ of companies that lose all of their data go out of business in 6 months.

DR/BCP costs are readily justified by doing a Business Impact Analysis (BIA).. budget up to some fraction of risk cost * risk probability.

I think this is due to the data residency requirements in UAE. I'm working with a client in the health space and the government requirements requires me to store data only in UAE! Tried with AWS but they were not allowing any new instances and I had to go with Azure.
Hi, I'm from the future. You might want to consider storing the data somewhere besides an Azure datacenter in the UAE.
Are you me from the future? I'm not talking to future strangers. Tell me to deliver the bad news myself from the future, in the future.
Hi, I'm from the past. When countries in the 2010s -- especially Western countries -- started seeing data residency requirements as an acceptable aspect of national policies, as opposed to a weird authoritarian thing that only China and Russia imposed on their citizens, we[1] spent a bunch of time explaining to their lawmakers that having geographical redundancy was a good thing, actually, and that you should stop insisting on where the data resided for jurisdictional purposes and start talking about where administrative access and encryption keys lived.

[1] OK, "we" here is probably just me -- it was one of those things where the chances of successfully convincing anyone was so small, and the commercial advantages of just nodding along, and then changing your product offering was so great, that really very few people raised it or had reason to. But somebody had to!

and start talking about where administrative access and encryption keys lived

Yeah, that was/is just another problem. Considering how that was actually handled in the real world before data residency laws came into force, I'm glad 'we' didn't convince those countries to put their citizens data at risk.

I'm not sure you were disagreeing with (past) me; but if you were, could you expand on your point?
I'm disagreeing with you. I in the before time, I had all sorts of conversations around this topic with any number of cloud providers that were like:

Us: We are concerned about our citizens (US) data, how are you managing the databases. Clout Provider (CP): They are only managed by fully background check employees. Us: Yeah, but where are they? What is their citizenship? CP: Um...mostly Eastern Europe. Lots in RU. (another CP proudly said "they're pretty much all in China...for cost containment"). Us: ...

Us: We are concerned about our citizens (EU) data, how are you managing encryption? CP: Everything is perfectly encrypted with hardware HSMs and all the FIPS and stuff. Us: So...where are the folks who run the HSMs? CP: Um...mostly SV. Some in the EU. Us: But can you assemble a quorum of US citizens for the HSM? CP: Of course! Us: ...

And on and on. Not to put too fine a point on it, many of us have no faith that vendors self policing international data protection in the face of government level pressure on companies and employees would work. Not that it can't, I don't think it would.

Is there not more than one data center in your country? Is the power feed at your office too small to put a computer there?
(comment deleted)
I think both of these things are (very) often true, but it is also true that if I'm going to have backups, it is (all other things being equal) better to minimize correlated risk. The assumption in a lot of these conversations is that having the data "in one place" (ie inside a country) was "safer" than having it in "somewhere else". The tougher counterintuitive argument was that it can be safer to have data stored in multiple places, for some risk assessments -- and that for others, having data close by was less important, in the case of seizure or surveillance or illegal use, than who had legal or effective access to that data.
This is a very engineer-centric view. I studied economics in school, so an analogy in that realm is ironically how all countries should specialize and raise the PPC curve. The reality of the situation was that in 2010 not many people understood how powerful big data actually was. Data sovereignty is actually quite logical when you consider the scale and power of not only the company, but the US as a whole. I can assure you that lawmakers were not thinking about efficient disaster recovery plans or back ups when they made the laws. You can also create reasonably diversified data silos within a country.

As an aside, it is quite crazy the world we live in. I am with the majority where I expected Amazon to be more redundant, but I still marvel at the assumption that a US dev can spin up multiple redundant and data sovereign servers in dozens of countries with efficient caching, failover and redundancy (enough to survive an earthquake or targeted missile attack) from their own home. Even a few hours of outage in a foreign country is considered unacceptable.

You can always selfhost
... in the UAE, so only less risky if your office is remote enough it doesn't also make a nice target (like, say, located within/near Dubai's Internet City).
It was always a bad bet for billionaires like Bezos to become Trump enablers. You weren't buying a seat at the table, or the privilege of being left alone, you were just signing yourself up to be force-fed shit sandwiches over and over (And the shit-to-bread ratio gets worse as time goes on)

You should have used your considerable resources to fight. If only billionaires had the drive to fight tyrants with the same vigor with which they'll fight a minor tax increase.

He had little choice. The Trump tariffs could've been a massive, massive blow to Amazon, so I'm sure he felt he had to get out in front of them and buy some influence with the incoming administration.

See also Tim Cook. Doesn't make it right to suck up to Trump, but it was, and unfortunately still is, a rational move.

Bezos hasn't lost anything from this. He's only gotten richer.
I wonder if they'll start adding an underground bunker to new data centers so you can put an S3 replica there?
This interview with an AWS leader isn’t aging well, from CBS Sunday morning:

Pogue asked, "I don't mean to give anyone ideas, but let's say I figured out that one of these unmarked buildings was an AWS data center, and I blew it up. Are you saying that it's so backed up and redundant that you probably wouldn't notice?" Wood replied, "Yeah, you wouldn't notice. I mean, we might be a bit upset, but you wouldn't notice!"

https://www.cbsnews.com/news/cloud-computing-loudoun-county-...

He forgot to add "if it's Multi-AZ" ;)
multi-AZ doesn’t help against multi-AZ drones :)
black swan events.

this is one of one of those things - were in the current era either a cloud provider should provide automatic backups in another geographic zone.

if you're in us-east, then your back-ups should ideally be in eu-west + africa for redundancy.

> "The damage to our infrastructure spanned multiple Availability Zones and exceeded what our regional and multi-AZ services are designed to withstand," AWS said in the status update
That is actually surprising to me. Claims like that are pretty common, they make sense and they should be true, so even though I don't really know AWS (/Backblaze/Azure/whatever) redundancy planning in enough detail, I used to trust them. It's really worrying when they outright say it will be ok, and then a week later it turns out to be not ok.
The caveat is always "if you're using the service correctly" which is not necessarily free. Meaning taking advantage of multiple geo zones, building in redundancy to your stack, etc. Like everything he said is possible if your technology stack living in AWS was designed to survive it. Everyone who has ever had the "we lost your data" email from AWS knows at the end of the day the cloud is just someone else's data center with neat provisioning tools and services.
Isn’t the problem that multiple datacenters in one zone were blown up?
They had nine years and more money than god to build redundancy in an unstable region.
If it's an unstable region, then it might not result in a good return, especially since a blown up data center is 100% loss.
that quote is definitely making its way into a lawsuit
They say "some" data, i wonder what percentage that really is. I haven't seen pictures but I find it hard to imagine all of me-south-1 was completely leveled to the point where's there's just nothing left. On the other hand, if you have 100 rows of racks and then randomly take out a contiguous 10% across both rows and columns it may be functionally equivalent to taking out everything.
I wouldn't be surprised if the engineers said "we can probably recover between 20-30% of the data but it will cost 200 hours of engineering and the data will be 7 months old by then" and the beancounters said "we'd rather have one news cycle rather than the news watching what we can and cannot recover + save those 200 hours, we'll just say it's all gone".
Hey but that's exactly as per design. It is the customer's responsibility to store stuff elsewhere as DR backup, not AWS.
That's the excuse they'll say, yes, then we quote back to them "eleven nines" and they come up with an excuse for that too