141 comments

[ 0.27 ms ] story [ 8.7 ms ] thread
(comment deleted)
Must be a day ending in Y
At this point, maybe it'd be more appropriate for people to post about github to HN when githubs actually working.
Oh thank god my pink unicorn site is back online, its had great uptime lately so thats nice.
> Update - We've identified an issue with a database primary and are failing over to a replica immediately

Seems like a weird thing to post on a status page. Shouldn't this have happened automatically and therefore precluded the need to inform users of it?

> Update - primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues
Can't even run self hosted github actions lol
Notice odd behavior on GitHub. Get gaslit by a green status page. Notice more odd behavior on GitHub. Think it must be me this time. See unusual action queuing. Ah, an incident on the status page. Go for a walk and check HN on my phone. The AI SDLC.
GitHub needs to completely bifurcate their enterprise/paid services from their free services at the infra level.
That's what I don't understand. They could mitigate their name so much if they just split free/paid/enterprise. It's already shown that enterprise is much more estable and is largely unaffected from service disruptions. Why don't they go one more layer? For sure it's worth the extra complexity.
There is no such thing as "just split" there is 20+ years of legacy decisions and even if the split is relatively clean it is still probably 1 years work for 200 people for maybe a marginal improvement.

The real money is going to go towards, "make this all more reliable".

According to their status pages (e.g. https://eu.githubstatus.com/, https://us.githubstatus.com/), their Enterprise Cloud uptime for Actions is significantly higher.
“GitHub Enterprise Cloud with data residency” is hosted on separate infrastructure and dedicated subdomains under *.ghe.com. It’s been around since November 2024z

It’s not the same thing as GitHub Enterprise Cloud hosted on the shared global network on github.com.

https://docs.github.com/en/enterprise-cloud@latest/admin/dat...

So confusing and so Microsoft. They love to have licensing so complicated their own sales people aren't up to date and have to rely on third party spreadsheets.

Edit:

Read that document - why do more people not do the self hosted option with GitHub Enterprise Server ?

At the point you are self hosting, you have many more options ranging from a simple ssh git server with pick your favorite cicd, to gitlab ce, to forgejo, and more.
Just to be clear, I am on Github Enterprise, and am also experiencing this disruption both privately and publicly on every org and project I have access to.
enterprise is mostly separate, is it not? uptimes are significantly more reasonable on the enterprise status pages
surely if they did that everybody would complain how github "lost its touch with open source since they now prioritize paid services"
They have that-ish as an option: https://docs.github.com/en/enterprise-cloud@latest/admin/dat...

I'm told that GitHub has asserted to us that moving to this model means we would not be exposed to github.com outages. It's not at feature parity with github.com though.

Thanks, this option is good to know.

We're currently on "GitHub Enterprise Cloud" on github.com and am affected by this outage (even though we use self-hosted runners!), but we're not on "GitHub Enterprise Cloud with data residency" on *.ghe.com, which I understand is/may not be affected by this outage?

This is what they've told us. It's represented as basically a separate deployment of the entire GHEC stack, so you're not exposed to the load/scaling issues they believe are the underlying cause of all the github.com outages today.
Do you have any meaningful level of faith in GitHub's ability to deliver on stability? At this point, I have none.
A lot of the stability problems come from trying to scale a free product on a still WIP cloud solution without costing too much; the same software running on separate, paid for infra has a lot better odds.
Meaningful is subjective, but yes I do. It was very stable for many years, and I do believe the recent issues are mostly or all because they were caught flat-footed by the rapid AI-driven load increases.

This has been a bad time, though. I'm ready to move back to self-hosting if they can't get it together or we move to GHDR and it's still bad.

The last large company I worked for switched from self hosted github enterprise to github.com in January 2025 or so.

I wonder how much egg is on that exec's face.

That's what Azure DevOps is supposed to do, but for some reason GitHub has a redundant enterprise division.
It would probably be better to run projects with extremely high commit/merge frequency on a separate "slop infrastructure", basically like MMOs move cheaters to their own servers ;)
I noticed earlier this week that my URL bar now pre-fills githubstatus.com instead of github.com when I type "gith"
For me this started happening when I type "stat" too.

Which surprised me because of the many hundreds of services I rely on regularly that also have status pages.

Good lord, hundreds?!

Uptime reliability is nkt additive, it's multiplicative, sometimes even logarithmic.

Downtime of 99% and 99% is 98.1%. I hope almost all those services are non-prod related.

A browser extension to indicate githubstatus being yellow/red:

https://chromewebstore.google.com/detail/is-github-down/lcfo...

A VSCode extension:

https://marketplace.visualstudio.com/items?itemName=RuslanRy...

Firefox:

https://github.com/matagus/github-status-checker

Caveat: https://news.ycombinator.com/item?id=49450924

(disclaimer: none of these are recommended by me to use - merely sharing to build on for this discussion)

Another useful disclaimer: all of these lets the developers push updates to your computer as they wish by default in most setups, these days you might want to decrease that kind of attack surface and just check the status in the official website when "git push" suddenly stop working.

I'm also not sure why you'd share links to software you don't even recommend yourself? Isn't it better to just not share that then? I think others could use search engine/LLM too if they need whatever, but most generally expect things shared in the comments here to actually at least have been looked at by the person sharing them.

Well, the alternatives are:

> It would be really cool if someone built an extension to show GitHub status live

"These already exist, why wouldn't you search before posting?"

> There exist quite a few extensions to show live GitHub status

"Why would you not post them if you know they exist?"

Or, what actually happened:

> Here are some extensions that might work for you to show live GitHub status

"Why would you recommend extensions to do this?"

Seems to me like, if your goal is to pick apart someone suggesting something, you'll find a way to.

> why you'd share links to software you don't even recommend yourself?

Overall, I disagree and think it's valid and responsible to share links to apps if I disclaimed that I don't recommend installing their linked applications in an environment that has access to private data.

One valid reason, to share links, is to illustrate that there are more than zero efforts to address this issue - the parent comment issue - through and even more convenient practice than firing up a web browser, following link, waiting for it to load, reviewing the material, to see if GitHub status is red or green today. So this reason attempts to build up the importance of the parent comment's idea, and suggests that legitimate verifiable work is advisable, to continue down that path of making it more convenient for developers have simpler, more human indicators about the reliability of their digital tools.

Pretty simple reason but that's my reason.

Can't believe this is my highest upvoted comment ever lol go check out my open source agent IDE if you are bored while you wait for GitHub to come back https://getness.dev
I can tell it’s getting bad because I went to close some Safari tabs and got confused because the tab next to this was [a dupe of] https://news.ycombinator.com/item?id=49330597.
You've a tab from 9 days ago open? Maybe you should be closing your tabs more often :p
I have tabs from like 6 months ago. Modern browsers have gotten pretty good at hibernating unused tabs and restore when needed. And this website is extremely light, it gets restored in milliseconds.
I hording my tab in case I need them later (I don't, 99% of the case)

but I always have sense of fear that I might need it someday

Every 3-6 months I purge all my tabs, but always dump all URLs to an html file in case I will need them in the future. You never know!
That is just called your browser history?
The browser history contains a lot of junk and I don't want that saved forever. Curated tab history is a bit different. It's stupid for sure, but also a bit comforting to keep the most important articles and such. For anything I really care about, there's always ArchiveBox.
I do that too but instead of html, I store them on discord and obsidian text file (double back up)

but I still keep it on my browser, maybe someday I would like to fetch LLM and describe my interest over time

I notice I sometimes wake up before my alarm because I get an SMS about GitHub being down.

It could be a new Microsoft feature !

i noticed it today as well - typing "g" is enough to go to the statuspage. and surprisingly often when i go there by accident, they have an incident…
As a joke I built an extension to the GitHub CLI called “omens” so you can run “gh omens” before you’re planning to use GitHub and it’ll tell you if it’s likely to work.

https://github.com/sandermvanvliet-stack/gh-omens

>go full bin chicken

LOL is this really something Aussie's say? That’s hilarious!

Bin chicken is local slang for these horrible birds ( https://en.wikipedia.org/wiki/Australian_white_ibis ), and “go full x” is also a common utterance, but I’ve never heard them said together like that.
Ibis’ are fantastic creatures, and don’t deserve the scorn you’ve heaped upon them!

Their bin-chicken status is testament to their adaptability, but it’s essentially our fault they’re that way.

They really are bin chickens for a reason. I remember in Lakemba in Sydney watching a young child trying to shoo an ibis from a bin where it was having its lunch. For a very long time I believed that they had migrated from Egypt because of their long beak - I think I asked my Mum once and she in tiredness just affirmed it.
So a normal Wednesday
> We've identified an issue with a database primary and are failing over to a replica immediately

This is why it's hard to take GitHub seriously. How can a single database cause an outage for everyone? This is amateur stuff. Have they no sharding or partitioning internally? Paying customers should not be impacted in the same way as free ones are.

I wonder what is this database, and why it is hard to fall-over automatically.
Possibly vitess from the latest update:

> primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues

Why shouldn't it? Most companies run on a single database server. If they can immediately fail over to a replica, that's doing it right.

Maybe you expect that part of GitHub to have a scale where a single database can't handle it, but evidently that isn't true.

We can criticise them for not splitting up free and paid customers but again, most companies don't do that.

Did you not read it? Just because there's a database primary doesn't mean there is 1 primary database. There's likely man redundancies and they have issue with how they're allocating traffic to them which is in turn causing an issue with how much traffic redundancies are receiving.
2.9B commits per month; 100M action runs per day; I think they probably have some sharding.
We can't keep living like this.
Except nobody moves to a privately hosted "gitweb"...
Obviously we can because we are choosing to. Because servers are scary.
[delayed]
Good for you. These kinds of migrations never feel productive at the time, but the right tools can make your life and work so much better.
I should make a business selling git hosting. Apparently it's really easy because it doesn't have to actually work.
Sure but first you gotta be Microsoft.
Microsoft and Oracle have the market cornered on selling stuff to corporations who are relaxed about whether it works.
>Update - primary failover briefly improved performance but did not fully mitigate, we've throttled inbound traffic and are investigating upstream Vitess issues

And now they're blaming their upstream vendor! Embarrassing stuff to be writing on a public page.

I read that as an upstream service they own, but I agree the wording a bit weird.
Why is that a problem if it _is_ an upstream vendor problem? (assuming it is)
Vitess is a distributed mysql database. Github could very well be managing it entirely on their own. I have only seen people managing their own vitess, its entirely open source afaik.
Things can go wrong, but really, its been a lot and we're normalizing that to an unhealthy degree...

I wonder if it was down that much, if users would get credits the way we pay when we use the services - its kind of ridiculous for a critical service to be down that much and all we do is "ah okay, its just github". Like, as if that was normal to be down that much...

I think the authors of SMTP had a healthy attitude to server uptimes:

   Retries continue until the message is transmitted or the sender gives
   up; the give-up time generally needs to be at least 4-5 days. 
https://datatracker.ietf.org/doc/html/rfc5321#section-4.5.4....
At the time, most email was either local to the host (big mainframe in the basement) or transferred once a day through scheduled dial-up connections during off-peak phone hours
Which is a good fit to a distributed, offline first version control system, if you use it like that.
I saw an application once about 25 years ago that was built on SMTP for communication. I think it was an inter library loan system, but can't remember for sure. I was a pretty smart way of not having to worry much about redundancy in the communications.
I don't think it's been "normalized". Github uptime is literally a joke in the tech community. They have first mover advantage and a behemoth behind them so they're not going away, but everyone knows how shit it's uptime is. It takes time for organizations to move away from services like this but I would bet anything that many are starting to try to move away, as well as new companies knowing they shouldn't use the service.
I agree. I think it is the kind of situation that goes “gradually and then suddenly”, to quote Hemingway.
must feel bad that at this point every dev checks github status before going to work like it was the weather app.
I have switched off from github to my own server besides two websites that depend on gitbub integration to deploy to clpudfare pages
Fun read about Azure and having 173 agents running a node: https://isolveproblems.substack.com/p/how-microsoft-vaporize...

Probably just a coincidence that Github started to have issues after beginning their move to Azure at the end of last year.

Azure's going to suffocate github. I'm curious to see what's next. Will self-hosting the code repository come back in vogue or will another social-coding platform take off?
My small org has definitely had internal discussions around self-hosting gitlab. We'll see what happens.
Curious what's driving the self-hosting discussion most, reliability, control, or compliance? Disclosure, I run a managed forgejo service
probably not. The people that have done it before or willing to do it now is probably a very % of the commit volume, they leaving wouldn't change much, probably not gonna even move the exponential growth needle.
We should worry less about what everyone else uses, and more about what we use. I'm self-hosting Gitea and thinking of upgrading to Forgejo. What are you using?
Yes

But GitHub, in its heyday, was a centre. Created a sense of "community"

It has taken a while, but since MS bought it it has been shirking that. I expected honest enshittification, but instead it has been technical collapse

What ever.

IMO we need a federation protocol for Git that can rebuild some sort of "community", but on solid foundations.

So we can find one another on our self hosted instances

There was never a "sense of community", there was only low friction on issue reporting and especially, commenting on an issue you had nothing else to do with. Every time an issue hit HN it would be flooded with bystander comments.
I actually just installed forgejo in My Home lab and loving it. Thinking about moving the companys code hosting to it next.
If you move the company too, what would make you pay someone to run forgejo instead of self-hosting it? Disclosure, I run Fjord, a managed dedicated forgejo provider.
I've moved from gitlab -> gitea+drone -> forgejo+woodpecker

Works great for small-medium scale

I used to work in a small independent team of 30 people within a large corp, half of which was dev. We used to run our own gitlab on-prem, our CI/CD was also on-prem. It worked perfectly, never had down time, devops guy could configure them on-demand to our needs. Me (and some other guys) also jumped in times to times to help (mostly just ssh into the servers for health check, disk partition, etc.). Then we grew (the biz team, dev was the same) and some new PMs with fancy Ivy League degrees came in and pushed for on-cloud Bitbucket. Things went to shit pretty fast after that ... Our codebase was only a few hundred thousands lines, there was only like hundreds of commits per day, the servers our git + CI/CD lived on never saturated ...
PMs deciding software infrastructure over the dev teams that the dev teams use, is wild.
they want that jira <-> bitbucket integration
From my experiences devs in large corp usually don't have a say in what kind of software infrastructure they can use ...
I've mostly worked for "engineering lead" corps and a PM has only ever affected outside software sold to users except for JIRA. There's no chance a PM would be able to affect the internal code storage/ci-cd systems at anywhere I've worked in the past unless they went on some proselytizing war path and convinced some senior lead engs to convince the rest of the org to accept it.

In the example, in my 20 year career, I've seen bitbucket in use once. I'm the one who usually manages that stuff since I'm an infra eng.

To be honest, I'm very much surprised Amazon hasn't eaten Microsoft's lunch here. Offer out-of-the-box AWS instances with Gitlab, SLAs, backups and the works, should be straightforward.
> Probably just a coincidence that Github started to have issues after beginning their move to Azure at the end of last year.

Right.

I'd like to think that Azure has improved in a meaningful way in the past half a decade or so... but it has not. Maybe less inexplicable 400 errors in random API calls, I guess.

And yet Microsoft keeps posting record growth. Unbelievable.

Keep in mind the seniority of the author. This was not written by a staff+ engineer.
"fun"

I thought I was cynical enough about goings on at Microsoft / Azure, but apparently not.

I've heard that companies that run enterprise Github instances on-premises have almost equally poor availability compared to github.com. There are probably plenty of reasons why they have so many outages, though I agree one of them is probably Azure. The outages started around the same time as the number of pushed commits skyrocketed, and the number of Dependabot vulnerabilities started growing as well, which in its turn increases the amount of PRs and Actions activity. Also, LLMs regularly import dependencies that are already out of date, meaning that when doing agentic development, one commit often results in several instant Dependabot PRs being created.
Ah so that's why I got a random "github-merge-queue Bot removed this pull request from the merge queue due to no response for status checks"
So what’s stopping you (or your org) from leaving GitHub?
I don't like any open source solution, Codeberg and Gitlab are more of the same and Cursor isn't ready yet.
I started the campaign to move us from Bitbucket to Github before the Microsoft acquisition. If I knew MS would be involved, I'd have campaigned for something else. I would never recommend any Microsoft product.

It took ages for us to get permission and licenses. We only recently finished moving the last repos over and shut down Bitbucket.

If AWS offered a product like GH, I could probably unilaterally start moving to that. Any other alternative would take 3-4 years of meetings, budgets, lawyers, and other such nonsense before we could declare we'd left Github.

Was wondering why all my Actions just stopped running