Ask HN: Is S3 down?

2589 points by iamdeedubs ↗ HN
I'm getting

{ "errorCode" : "InternalError" }

When I attempt to use the AWS Console to view s3

1,140 comments

[ 2.9 ms ] story [ 426 ms ] thread
Looks like it. Brief panic caused here.
We're getting errors indicative of an S3 outage too.
Not sure if it's related... but I'm having issues with Amazon Cloud-drive.
I would assume, I couldn't even share a screen shot of my evidence to my team on slack!
Lol... It's interesting how much depends on Amazon's infrastructure at the moment.
Also, interesting everyone is posting that various sites which are down in this thread. Feels like the internet is down!
(comment deleted)
any one get more info from AWS?
Having issues as well.. big issues..
Yup, same here. It has been a few minutes already. Wanna bet the green checkmark[1] will stay green until the incident is resolved?

[1] https://status.aws.amazon.com/

Still green now, 8 minutes in.
I've had a few non-Amazon providers tell me AWS things are not working in the last 5 minutes, no note from Amazon though.

Nice.

Just sent out a notice to our customers via our status page. I really wanted to be able to add a link back to AWS detailing the issue but that's a pipe dream I suppose.
Just went yellow

Edit: nevermind

Still green for me
(comment deleted)
Just went yellow

Increased Error Rates

We are investigating increased error rates for Amazon S3 requests in the US-EAST-1 Region.

https://status.aws.amazon.com/

Check individual services ...

Amazon Simple Storage Service (US Standard) Service is operating normally

Did it? Still fields of green for me.
While keeping the status green for s3, they have at least put up a notice at the top:

Increased Error Rates

We are investigating increased error rates for Amazon S3 requests in the US-EAST-1 Region.

Yeah I just now saw that. Probably regional cache clearing or something.
http://downdetector.com/status/aws-amazon-web-services looks like a reasonable alternative place to check/report downtime.
Which is coincidently down.
Maybe they are hosted on S3 facepalm or maybe they just got a surge in traffic
(comment deleted)
I just check Twitter, since Amazon's status is always a lie. My personal dashboard is still showing no problems. It's bad enough that the main public status is always green even when there's clearly a problem, but you'd think they could at least make the private status accurate.
Gah. It was up 3 minutes ago. Anyone have any suspicion this is another ddos episode? I saw that SO was down last night too: https://twitter.com/StackStatus/status/836450836322516992
Pretty confident that isn't it. S3 was returning InternalErrors for 22 seconds before it started timing out and/or returning 503s to all my requests.

I'd bet that something broke (causing InternalError responses) and then nodes started marking themselves as failed (causing the timeouts and 503s soon after).

I want to see the botnet capable DDoSing S3. That would be something.
Apparently, that's down too. Sigh.
I'm seeing green checkmarks across the board, but they just added a notice to the top of the page:

> Increased Error Rates

> We are investigating increased error rates for Amazon S3 requests in the US-EAST-1 Region.

I guess sub-1% to 100% failure rate is technically an "increase".
I guess file uploads and downloads are technically "API calls".
the worst thing is when your system cant handle these "increased error rates" as your control plane cascades failure due to something like this....

The worst "increased error rate" problem I had was when the API was failing and my autoscale system couldnt deal and launched thousands of instances because it couldnt tell when instances were launched (lack of API access) and the instances pummelled the fuck out of all other parts of the system and we basically had to reboot the entire platform....

Luckily, amazon is REALLY forgiving with respect to costs in these (and actually most) circumstance....

recalls numerous times

Yes. Yes they are. Thankfully.

These service health boards are more like advertisement page then actual status of the service.
I guess their bizarre thinking is something along the lines of: "unless we have proof that noone can access the service, we won't change the indicator from green to yellow.

Seriously: I don't understand why you guys stay with AWS.

> I don't understand why you guys stay with AWS.

Who do you recommend instead (assuming in-house or Hetzner-equiv is out of reach)? Google Cloud? Azure? Rackspace?

Google Cloud if you're looking for something similar. It's just so much better and cheaper. I think a lot of the resistance here towards that kind of move is just because people are inherently lazy and they aren't paying the bill themselves.

(I'm guessing a relatively large part is also selfish attachment to the market leader because of employment reasons. I hate wasting money, both for myself and for my employer, so I don't really understand this kind of thinking - but I do understand how it could flourish in a venture capital-rich time/locale.)

I also recommend reading:

https://thehftguy.com/2016/06/15/gce-vs-aws-in-2016-why-you-...

Google also doesn't have the best record for developer tools.
Let me guess. They cancelled your grandmother's RSS reader four years ago?
Google Cloud doesn't exactly have the greatest reliability/uptime either.
https://status.cloud.google.com/summary tells a different story or do you have other information?

I have used GCP for some time without being affected from any incident.

That is a pretty awesome page.. way better than a page full of green icons, during an obvious outage... I like that they have writeups a few days after the incidents....
I'm not sure what you mean. If anything that link underscores my point. GCE has absolutely had it's own catastrophic errors. Remember last April when ALL instances in ALL regions went down?

https://status.cloud.google.com/incident/compute/16007?post-...

The GCP services are usually within their SLA target so I don't see the problem with the incidents. So you know what you buy and can take actions if you need a higher SLA for your application.
> so I don't see the problem with the incidents

All instances going down in all regions is an order of magnitude worse than a single service going down in a single region. You're deluding yourself if you think GCE is any more reliable than any other reputable cloud hosting platform.

While I agree that was a horrific outage for us, there's a big difference between no external connectivity for a few minutes (note: internal IPs still worked fine, as did access to APIs through that mechanism) and "ALL instances in ALL regions went down".

Disclosure: I work on Google Cloud (and wouldn't want to be an incident responder at AWS today...)

At least Google has a post-mortem, at AWS everything would still show up green with some random note about 'increased error rates'.
GC's CDN doesn't cache files bigger than 4Mb. No Windows VMs. Bound to AWS for these 2 reasons.
As already mentioned, they do have Windows VM's but there are some caveats that indicate it's not fully baked yet. 1.) They require that each VM MUST have a public IP address so that Windows can talk to an activation server every 30 days. 2.) You cannot yet bring your own license.
OVH
Bad idea there, support is horrible.
What is your last datapoint on that?

The last year or two has seen a remarkable improvement according to those customers of mine that host there.

OVH doesn't even want to take my money to keep my server running. Their auto-billing process is busted and when it goes wrong they just delete your server.
That's not what I've seen. I misconfigured my auto-billing and got paged in the middle of the night by nodes mysteriously disappearing, but they released those machines minutes after my CC went through. Not that I'm a big fan of OVH but if you design your system to allow for failure you can't match their value for money.
What about something like B2 from https://www.backblaze.com/ ?
S3 in a single region is based out of multiple data centres / availability zone, with data distributed so that the loss of a single availability zone won't impact either data availability or durability, even to the point of being comfortable with complete physical destruction of an AZ. The same applies for Azure, GCP etc.

B2 is based out of a single DC (or at least, was at launch and I don't see anything that suggests that has changed?) You've got to decide what's most important to you. Data persistence or $$$.

> Seriously: I don't understand why you guys stay with AWS.

I tried them all and Amazon is still the best.

> Seriously: I don't understand why you guys stay with AWS.

Personally I've been using it for ages and I know most services inside and out. They do suffer downtime in some regions occasionally, but it'd be too expensive at this point to move.

And who doesn't suffer downtime? You can't avoid it; you just need a plan to deal with it. For example, having a backup replica bucket in another region and the ability to quickly switch your CDN over would probably be a good idea here; that's what I did.

If you want to go further you can replicate your data to another cloud provider entirely and use low TTLs to switch to a backup CDN if your system is that mission-critical (in the event of a worldwide AWS failure doomsday scenario).

All systems will fail you and it's our responsibility as IT professionals to have a plan to mitigate this.

Sunk cost fallacy.

I do agree that we should all plan for failures.

However, I also think it's a sign of failure in planning and architecture foresight if it's too expensive to move away from a particular cloud provider.

The sunk cost fallacy is when you (irrationally) decide to stick with what you're doing purely because you've already spent a lot of resources on it. It doesn't apply when you've done an economic analysis and found out it doesn't make sense to swap.

There are plenty of cases where it just wouldn't make sense to switch after looking at the costs, opportunity costs, etc. For example, if his site makes him $10 a month, outages cost him $1 a month that could be mitigated by moving, and it would cost $1000 of labor to swap providers. (Depends on interest rates.)

Perhaps it was originally a failure to not have a plan to easily move from a provider, but it doesn't seem unreasonable to me that right now it may cost too many hours of work to justify the move.

(You're right, I used that term incorrectly.)

Still stand behind the other two points I made in that post though.

It's not as though it would be impossible; our integration with AWS isn't that deep, it's not as though we use DynamoDB for our core data store or anything like that. But even migrating from one traditional datacenter to another isn't easy from an operational point of view.

There needs to be a clear financial win. Even taking into account the failures we've seen so far, I don't see a compelling reason to leave AWS.

Low TTL on DNS entries might do more harm than good: if your DNS provider gets seriously DDoS, being able to rely on caches can save the day.

Anyway, I agree with your conclusion.

Because you perceive public clouds only as virtual machine providers, that you can replace with other provider in two days. A detailed cloud migration consists of replacing some parts of your software to use managed services provided by a specific cloud provider, and AWS is still has the best service offerings IMHO. When you use these services carefully also you will see that AWS is very cheap and reliable enough. Outages like today's are happening in every platform and it is possible to mitigate them.

You can use Adwords as a self-service user. Without knowing so much of details you can run your ads but also you can bery easily ruin your budget. But many enterprise customers use it very differently than those users and they are extremely optimizing the cost. Cloud is the same. If you don't know how big customers use AWS, it is normal that you are surprised because AWS is still leading the market.

You say GCP is better than AWS. Which part is better? GCP does not have many services of AWS we benefit from. How can you compare totally different providers? You can only say AWS EC2 is worse than GCP. But you cannot compare whole platforms in one sentence.

Sorry for endless number of typos and mistakes. Obviously I was sleepy while I was writing this.
(Sorry, I'm late to reply, but since you addended your comment you might still be listening...)

After spending a year evaluating both AWS and GCP (with an emphasis on their managed database services; both SQL and no-SQL) my general feeling is this:

"Microsoft Windows is to Unix as AWS is to GCP".

(Or perhaps closer to the truth: "VMS is to Unix as AWS is to GCP".)

Baically AWS services seem like they are badly designed by buerocratic mediocre engineers following some bureocratic template for "a service".

GCP feels a lot saner (both API- and UI/console-wise). I often got the feeling it's designed by people who:

a) are smart and well-rounded in terms of experiences. It does take cleverness and experience to design something elegant that is also useful.

b) take pride in their work (it does show)

(And then, as a bonus: It's cheaper!)

You talk about SQL and No-SQL as managed services and it shows that your experience is limited to a classical application consisting of virtual machines and some data storage. However these are not the only services offered by both platforms and currently AWS has a richer feature set. For example Lambda and its deep integration with whole AWS platform is the biggest game changer from my point of view. If we are talking about virtual machines and databases, I can accept this comparison. However we are talking about 30+ services, some of them are even not available somewhere else and solving serious business problems in production and at scale. It is very wrong to put everything into basket and compare. Maybe GCP has better pub/sub service and AWS has better object storage. These should be compared seperately. Answering to your question, why do we still stay at AWS, because it is solving our problems in the most cost effective way and with reduced complexity, we are happy with it.
You're probably assuming too much again :)

I specifically spent a lot of time on Lambda and found it quite annoying compared to GCP AppEngine. So much bureaucracy. Just this thing that you have to specifically register every single Lambda API call and its parameters using an interface built by non-thinking people.. Sheesh.

For on-demand processing I just want a single HTTP-ish entry point, like AppEngine provides. (That way I can I move my service between different providers, if I wanted to move away from e.g. AWS.)

Anyway, I just updated my HN profile with more details about my experience. Please visit it to judge if I might know what I'm talking about.
Postgres on RDS
Come to NEXT in a week! :).
Any chance UDF iterators for Cloud Bigtable are in the works?

Being able to run distributed D4M/GraphBLAS queries in Cloud Bigtable would be killer.

"From NoSQL Accumulo to NewSQL Graphulo: Design and Utility of Graph Algorithms inside a BigTable Database" https://arxiv.org/pdf/1606.07085.pdf

I think it's more, "if the service can't do what people need it to do, that's a problem; if the service cluster gets wedged hard enough to stop responding to the requests of our monitoring system, that's a failure."

Which would make sense (and is sorta-kinda a best-practice) if Amazon wrote services such that they "crashed early"—but instead they're seemingly written so the backend lock up and be rendered completely useless at "doing its job" but will continue to run just fine.

Either of those two design decisions is potentially a good thing on its own, but they need to be considered in light of one-another if you want your status page to make any sense. If you want to report cluster failures, code your clusters to actually fail. If you want to keep your clusters up, write your monitoring checks as whole-stack acceptance tests.

> Seriously: I don't understand why you guys stay with AWS.

You don't seem to have enough experience to comment on the issue.

I always joke that if one of those statuses ever went to red, it means the zombie apocalypse has begun.
The good news is, if Amazon's services are marked as offline, you're allowed to use Amazon Lumberyard to control nuclear power plants.
The number of non-green marks is the number of ICBMs currently in flight towards an AWS data center.
I've heard (on the Fnord new show on the most recent CCC congress, so take it with a grain of salt and a bucket of humor) that Amazon's TOS are more or less void when a Zombie Apocalypse breaks out.

They had some convoluted but fairly specific wording in their TOS, whoever wrote must have had a lot of fun.

From https://aws.amazon.com/service-terms/

> 57.10 Acceptable Use; Safety-Critical Systems. Your use of the Lumberyard Materials must comply with the AWS Acceptable Use Policy. The Lumberyard Materials are not intended for use with life-critical or safety-critical systems, such as use in operation of medical equipment, automated transportation systems, autonomous vehicles, aircraft or air traffic control, nuclear facilities, manned spacecraft, or military use in connection with live combat. However, this restriction will not apply in the event of the occurrence (certified by the United States Centers for Disease Control or successor body) of a widespread viral infection transmitted via bites or contact with bodily fluids that causes human corpses to reanimate and seek to consume living human flesh, blood, brain or nerve tissue and is likely to result in the fall of organized civilization.

First the fall of human civilization has to be a real threat per the TOS so not sure they'll care.

Second, I know the lawyer and yes he had fun.

Then I guess it has begun, the page is now showing red. I'd put a picture on imgur but it's not loading.
if(true){ displayGreenCheck() }
In December 2015 I received an e-mail with the following subject line from AWS, around 4 am in the morning:

"Amazon EC2 Instance scheduled for retirement"

When I checked the logs it was clear the hardware failed 30 mins before they scheduled it for retirement. EC2 and root device data was gone. The e-mail also said "you may have already lost data".

So I know that Amazon schedules servers for retirement after they already failed, green check doesn't surprise me.

So just as a FYI the reason that probably happened to you is that the underlying host was failing. I am assuming they wanted to give you a window to deal with it but the host croaked before then. I've been dealing w/ AWS for a long long time and I've never seen a maintenance event go early unless the physical hardware actually died...
That's completely ridiculous, get some fucking RAID Amazon.

I order drives off newegg directly to my DC and I'm yet to lose data with the cheapest drives available in RAID10.

Yes, solving problems at your scale and AWS' are quite comparable.
Not saying my scale is the the same at all - but the fact they can't do something so simple that I can do it as a single individual is embarrassing at best.

Simple solutions to this do scale - Linode and DigitalOcean don't have such issues for example - and while they're not Amazon scale, they are quite large and I'd say they prove the concept.

EBS data is backed up in multiple redundant ways (using erasure encoding I think).

Local storage is not intended for permanent storage, and is more use at your own risk. That's also why most of the new EC2 instances don't even support local storage.

Availability =/= durability of course

EBS is incredibly expensive and slow, not really a good solution. It'd be nice if they offered a better local storage option.
I think most people rely on EBS and are happy with it. Sure it depends on the use case, but I think it works for most use cases.
Incredibly expensive and slow compared to what? A 500 GB SSD (gp2) costs $50/m, and has 1500 - 3000 IOPS. It's okay for most loads.

For higher performance, you can use

1. EBS Provisioned IOPS (kind of expensive)

2. Aurora (for DB use)

3. The new I3 instances (super fast local storage at a reasonable price.)

That's the cost of a new 500GB SSD per month! For the cost of three months' EBS storage and a couple of hours you could setup your own RAID array with a spare for backups and possibly get better uptimes than Amazon :P
Not to mention you'll get 3-5x better IOPS off a dedicated SSD.

I guess this just boils down again to Amazon not being cost effective enough for my use case in yet another way.

Actually, you can get way more than 3-5x better IOPS on your own SSD! Different types of storage have different types of trade-offs. EBS is great for some things, slow and overpriced for others...
Huh? What kind of 500 GB SSD costs $50? And again--500 GB on EBS is not stored on 500 GB of flash...they use erasure coding and distribute it over ~3x as much, roughly.

Oh, and good luck creating snapshots of your home RAID!

> Huh? What kind of 500 GB SSD costs $50?

Definitely not $50 to my knowledge but for ~$170 you can get a Samsung 850 EVO which is rated for 98k IOPS. They're fairly reliable drives and much, much faster than anything you'll get on EBS. You could be running that full 3x replication in less than a year of paying for EBS.

> Oh, and good luck creating snapshots of your home RAID!

LVM, ZFS and Btrfs all do snapshotting quite nicely. FreeNAS - commonly used for consumer grade NASes will automatically manage ZFS snapshots for you too. Amazon will sell you extra space to store snapshots, sure, but increasing the size of your devices usually solves that problem. And quite cost effectively as you can probably tell by now...

...and Dropbox can be replaced with SFTP and rsync. Can you roll your own? Sure. Will it work as reliably and effectively as EBS? Can I scale it easily? How many people are comfortable using snapshots on their own?
It's tried and true tech that any competent ops person can use quite easily. Been around much longer than EBS quite frankly.

Dropbox targets end users who don't have the knowledge required to use the alternative, if you're smart enough to use EBS you're probably smart enough to use ZFS snapshotting just as easily. Or could within a day or two. It's really not that hard.

Like I said, there are systems that pretty much manage the whole thing for you and just warn you when something is about to blow up like FreeNAS.

It could be replaced with Nextcloud

Shadow volume replication is entirely possible with several filesystems or Hot Copy kernel mod. Also LVM does snapshotting fairly easily

I may have got my prices a bit mixed up (I saw 120GB at Fry's for $60 last month) but my point stands.

Also why is discomfort such a big problem for folks? Learn stuff.

Not really that hard.

zfs snap tank/data@$(date '+Y%m%d')

zfs send tank/data@$(date '+Y%m%d') | zfs recv backup/data

advanced magic for off-system backup

zfs send tank/data@$(date '+Y%m%d') | ssh cheapdiskserver zfs recv tank/data

> 1500 - 3000 IOPS

So about as many as this SD card, and nothing compared to a real SSD.

In practice, it's actually fine for most purposes. It's equivalent to multiple striped 15K RPM magnetic disks, which used to be high-end enterprise storage a few years ago.

SD cards have much worse write IOPS.

> In practice, it's actually fine for most purposes.

It is, yes, but I wouldn't refer to it as comparable to an SSD.

> SD cards have much worse write IOPS.

Surprisingly not! Testing in ATTO I got read and write speeds that were almost identical, and a peak of 2000 IOps.

>It is, yes, but I wouldn't refer to it as comparable to an SSD.

EBS (gp2) is flash based, has far better performance than high end magnetic disks, with excellent latency and consistent performance. So, it's more comparable to SSD than anything else.

>Surprisingly not! Testing in ATTO I got read and write speeds that were almost identical, and a peak of 2000 IOps.

Really? Were you looking at 4K write? Typically that would be under 1 MB/s for an SD card.

(comment deleted)
but I never lost data off an usb stick how hard could it be!
(comment deleted)
Really?!?! Several times USB sticks (and USB HDs) failed on me and other people I work with.
Yes, the only way a server can die is from non-raided disks.
Otherwise they should at least be providing customers their data back.
I think you misunderstood the local storage. It is not intended to permanently store data. It's a volatile storage like RAM.
It's not just a RAID that can fail. And everyone who uses AWS should expect failures. You should build your infrastructure to handle such failures well.
They offer no RAID on local storage and only the expensive, IO restricted EBS as an alternative.
That what happens when cloud provider doesn't support live migration for VMs.
We have a slack emoji for it called greenish. It's the classic AWS green checkmark with an info icon in the bottom. Apparently it's NOT an outage if you don't acknowledge it. It's called alt-uptime.
I really liked it. But when trying to add it to my HipChat group it failed to upload. Why? S3 outage, what an irony.
AWS internal lingo calls this the "green-i"
So, global S3 outage for more than an hour now. Still green, still talking about "US East issue". I'm amazed.
It doesn't appear to be global; my app in eu-west-1 appears unaffected.

It's possible that the console won't work however as I believe that's served from us-east-1.

My site hosted on S3 is also running.
It's crazy how much better the communication (including updates and status pages) is of the companies that rely on AWS than AWS' communication itself.

https://status.heroku.com/incidents/1059

I feel for them. Imagine, 40 or 50 different engineering teams all responsible for updating their statuses. At this moment on the AWS status page I see random usage of the red, yellow, and green icons, even though all the status updates are "Increased error rates." What that tells me is that there's no unified communication protocol across the teams, or they're not following it. And just imagine what it's like being on the S3 team right now.

I notice even Cloudflare is starting to have problems serving up pages now.

Looks like they have fixed the issue with their health dashboard now.

From https://status.aws.amazon.com/ : Update at 11:35 AM PST: We have now repaired the ability to update the service health dashboard. The service updates are below. We continue to experience high error rates with S3 in US-EAST-1, which is impacting various AWS services. We are working hard at repairing S3, believe we understand root cause, and are working on implementing what we believe will remediate the issue.

FreshDesk makes extensive use of S3 and it's been unbearably slow to load for the past hour or so. All on S3 requests.
(comment deleted)
Uh-oh. Same here... and tried taking a screenshot of pinging s3.amazonaws.com and Slack upload hung.
Dropbox using this? Can't seem to sync
They used to, but i think they are off it now.
It is, of course the checkmark will stay green throughout this as Amazon doesn't care about actually letting its customers know they have a problem.
Down in US-East-1 as of 17:40 GMT. Amazon SES also down in US-East-1 as of a few minutes later.

Hearing reports of EBS down as well.

I confirm for SES being down in US-East-1 :-(
EBS is down for 30% of my servers as well
Down from the outside; The internal access (from within EC2) APIs still work.
Dead as a doornail for me