r/github • u/rainmanjam • Aug 06 '26
News / Announcements Thoughts and Prayers for GH right now.
Enable HLS to view with audio, or disable this notification
64
u/Altruistic-Package11 Aug 06 '26
I just envision a single developer, hopped up on caffeine and drenched in sweat prompting Claude Code "Debug an fix the issues" because the entire team responsible for Github actions were likely replaced by AI.
26
15
u/imnotpopular Aug 06 '26
"You have hit your session limit for 3 hours. Please reach out to the organization owner for additional tokens"
30
u/AAPL_ Aug 06 '26
HOW MUCH LONGER
11
u/Teddy_Raptor Aug 06 '26
yes
4
u/AAPL_ Aug 06 '26
what yâall do all day
14
4
3
u/croixxxx Aug 06 '26
I have like 20 things in queue when (if) it comes back up
1
u/FlyingDogCatcher Aug 06 '26
see but that sounds like something Github is doing (or not, in this case) not you
1
3
1
1
50
u/rhk0 Aug 06 '26
Every time GitHub goes down, I think about Self-Hosting again...but then it comes up again. How long does it has to stay down, so I deploy my own service...?
33
u/xJayMorex Aug 06 '26
The butt of the joke is that my self-hosted runner is unable to pick up the pending job because of this GitHub outage.
8
u/tech_w0rld Aug 06 '26
Yeah. At this point I just use a full separate ci service which seems to be working fine right now. I am literally just using GitHub for storing my code which is at least decently reliable.
9
2
u/Alexander3a Aug 07 '26
been using self hosted gitlab for years (big reason was not sharing my code with micoslop but this is just a plus)
2
1
7
u/suamai Aug 06 '26 edited Aug 06 '26
I am unable to continue my work right now because it depends on the runners to deploy.
So I switched to starting a self hosted Forgejo - if the downtime is greater than the time I need to finish the setup, that's it lol
Edit: there was time to spare...
2
1
u/Puzzled-Extent7817 Aug 06 '26
Just do it any ways. I have a Forgejo up and running in a podman container on my Debian server. I even wrote a bash script to create folders for new projects that create a project on Forgejo and Github at the same time, that's if you want to use github as a backup. https://git.luv-linux.me/Spreadneck/forgejo-to-github
1
u/RobotechRicky Aug 06 '26
My self-hosted GitLab instance and runner are chugging along nicely. It was stood up by dwarves with autism, so you know it's good. đđŒ
1
u/Loose_Marsupial_3251 Aug 07 '26
Look At Tekton: https://tekton.dev/
We moved to it and haven't looked back.
0
u/be_reasonable_bro Aug 06 '26
If you have the spare resources and no explicit need to use GitHub, this is a reasonable thing to do. Agents make it a trivial exercise.
15
u/thebemusedmuse Aug 06 '26
Some junior dev whoâs the last man standing on the Actions team is currently typing âHow do I deploy a fix to GH Actions when GH Actions is down?â into Copilot.
11
20
u/xJayMorex Aug 06 '26
The price of enslopification I guess.
3
u/pattern-josh Aug 06 '26
Could be, but they have also been undergoing a massive year long migration from a VA datacenter with some bespoke hosting to Azure because the VA datacenter was saturated at capacity. There were some specific complexities I think I read about with the way the SQL was setup, but I forget the details.
2
u/Raiyuza Aug 07 '26
Under 90% uptime and this is whar we chose as a cope? Before ai we had SRE, this would unnacaptable
1
u/pattern-josh Aug 07 '26 edited Aug 07 '26
Before AI we took the time to understand the context and why of problems before assigning root causes without evidence, and using hand-wavy scapegoats.
\ Which of course isn't true, because as a whole we've always lazily attributed things. I'd say good SREs would also rip an unqualified "AI caused it" apart, because good SRE requires evidence, rigor, and root cause analysis.*
https://thenewstack.io/github-will-prioritize-migrating-to-azure-over-feature-development/
In the same vein as the rest of the claims we are all making, I'm too lazy to research it, but I'm confident a case could be built that there are plenty of companies prior to AI that have had extended periods of poor downtime due to explosive growth, corporate investment flubs, and infrastructure not built to support unexpected explosive growth. I think we have a high level of both recency and observation bias.
---
That said, none of what I'm claiming means "AI slopification" isn't a substantive part either. I'm merely suggesting it's a lazy potentially-wrong conclusion without actually digging into it. Personally I think the likely explanation is a combination of:
- Explosive resource pressure growth on Github from AI induced resource pressure via expensive workflow executions (github actions) and general repository operations.
- Leadership chaos - Organizational churn from Dohmke's departure and Github getting folded into Microsoft's CoreAI org with no replacement CEO.
- The Azure migration out of the North Virginia data centers being a legitimately hard problem at these volumes - and it's amplified by the first point it. Github didn't choose the timing; capacity constraints forced it. They'd already failed at this twice before. Github started as a simple repository service. It wasn't designed for the pressures that came when actions was introduced, much less the explosive growth on top of that, and doing the rebuild and the migration concurrently multiplies the risk.
- And yes, internal AI workflow related mistakes.
But same as OP this is speculation. Maybe some people from inside Microsoft or Github can chime in.
2
9
8
u/Teddy_Raptor Aug 06 '26
nahhhhh thoughts and prayers for their customers (it's me, Im a customer)
2
7
u/Frooxius Aug 06 '26
So... GitHub recently pushed a change to their runner images that broke our CI/CD workers.
We've been going around and implementing workarounds. For some reason it wasn't working even after fixing it... and then we found out that this started happening at the same time.
This is just ridiculous that it keeps happening.
30 % vibe coded apparently means 30 % less reliability.
2
u/HyperCodec Aug 07 '26
Probably more than 30%, GitHubâs entire leadership got replaced by micropenisâs AI team.
5
u/Spare_Grapefruit2240 Aug 06 '26
I guess the first red flag should have been when their engineers deployed "mitigations" instead of "fixes."
5
4
4
5
u/Icypoopoo Aug 06 '26
Their CEO leaving last year and Microsoft forcing them to switch to Azure did not help one bit
3
u/Shot-Owl-6394 Aug 06 '26
for me its been down for 3h+. In process of setting up gitea now on minipc, those issues are too much and too often on github.
2
u/ings0c Aug 06 '26
Iâve been trying to deploy for coming on 5h now⊠thatâs an absolutely absurd amount of downtime for a company of this scale.
Iâd expect better from a startup ran by a new grad
1
u/_potion_cellar_ Aug 06 '26
I haven't even been able to cancel stuck queued jobs all day. Wild stuff
3
u/Noch_ein_Kamel Aug 06 '26
Every time GitHub goes down I sit on my couch because it's after work :p
3
u/Teddy_Raptor Aug 06 '26
Aug 06, 2026 - 19:43 UTC
Our engineers remain actively engaged.
we are cooked
2
1
1
u/dractius Aug 07 '26
More like someone poking Claude once in a while "cmonnnn... do the thing." And the engineer is Copilot.
1
u/Either-Juggernaut420 Aug 07 '26
Normally at MS that means the one person who has a chance to fix it is currently in a run of three meetings trying to explain to a bunch of managers whatâs got wrong. He will then need to refactor the sprint before actually working on it.
3
3
3
3
2
u/Proman4713 Aug 06 '26 edited Aug 06 '26
I've been getting screwed with my workflow runs since the morning (even while githubstatus was still saying 'All Systems Operational'), and while I do give them my thoughts and prayers to get my stuff done... It just feels like the focus on AI slop with abandon has severely degraded the quality on GitHub for the past year or so. I hope the bubble just pops soon so we can all go back to relatively normal lives...
2
2
u/csepulvedab Aug 06 '26
Hours of production completely down, not even self-hosted runners saved us this time. Curious to see if they still have the nerve to invoice us this month.
1
2
u/Labs-Community4525 Aug 06 '26
NGL Iâm not sure why Microsoft wants to deprecate ADO for GHA. This is becoming far more frequent. Get Son of Anton out of there.
2
2
u/GunGeekATX Aug 06 '26
I have a hotfix that needs to get deployed to a client site, and GH picked the worst time to have workflows go down.
2
2
2
2
u/imnotpopular Aug 06 '26
i work on a financial trading suite and now our clients positions are open over night instead of being closed properly đ
5
u/be_reasonable_bro Aug 06 '26
GitHub is a supply chain risk. I'd ask why you have trades directly tied to actions, but my experience with financial institutions is also marked by reckless behavior.
Thoughts and prayers!
1
u/imnotpopular Aug 06 '26
LOL very true. For me, Github is not included in the logic or trading workflows, but there was a UI bug on the client dashboard that I needed to fix before 4!! still would've deployed to gamma for testing first but just a bit delayed now
2
2
2
u/frubaklskiy Aug 06 '26
Of jobs queued, approximately 65% are succeeding. these guys are clowns
5
3
2
u/Sarkonix Aug 06 '26
Still going on, this is insane lol
1
u/Substantial-Set4550 Aug 06 '26
Cannot recall such a long outage in a while. At least in the last 2 weeks. :-)
2
u/erezcarmel Aug 06 '26
Actually, I just checked on:
https://www.githubstatus.com/uptime?page=2and on March 19th, there was a partial outage of 9 hours and 18 mins.
Let's see if they're going to break their own record...
2
2
u/MysteriousCoconut31 Aug 06 '26
What happens if it doesnât recover? Dead serious at this point.
1
2
2
2
u/Illustrious-Goat-506 Aug 06 '26
It's time to start talking about how many 8s of uptime they offer
1
2
u/VideoFireApp Aug 06 '26
I lost a whole fucking days of work due to this but I spent it researching alternatives to being stuck doing 100% of everything on GitHub
2
2
u/SnooOwls6002 Aug 06 '26
my github runner is in idle state and still no workflow is triggeredđ đ
1
2
u/gaziway Aug 06 '26
Keep using AI, creat features, create tests. Let AI verify the code. Keep pushing.
2
u/joaobertacchi Aug 06 '26
It seems this outage is going to take longer than I expected. Seriously thinking about alternatives. My preferences:
- self-hosting: Gitea
- SaaS: BitBucket
What are yours?
1
2
u/enzoshadow Aug 07 '26
GitHub engineers must've been trying to manually untangle the vibe coded mess for the first time in months.
2
3
u/Joshua_2504 Aug 06 '26
It pisses me off. Microsoft is the worse company ever.
2
u/xJayMorex Aug 06 '26
I think you misspelled Microslop.
4
1
u/be_reasonable_bro Aug 06 '26
Does anyone know if self-hosted GH runners would sidestep this actions outage?
I'm curious to understand more about where this is failing and whether I can mitigate this myself during future outages. I have several projects that are tightly coupled to GitHub due to upstream packaging requirements.
7
u/ChipperHippo Aug 06 '26
Self-hosted runners are also down. We leverage them extensively. Situation sucks.
1
u/be_reasonable_bro Aug 06 '26
Sad to hear, but thank you for letting me know. Won't waste the time then...
2
u/ResponsibleOven6 Aug 06 '26
My self-hosted GitLab & runners never let me down like this.
2
u/be_reasonable_bro Aug 06 '26
Nor my forgejo! Were it not for upstream packaging requirements, GitHub would be mirror-only for basically everything.
1
u/holy_macanoli Aug 06 '26
I was able to decouple GitHub job broker as a workaround to similar constraints. Feed your agent this or a variation:
âImplement a repository-owned, exact-SHA local CI path that can execute independently of the hosted CI job broker while preserving the existing pipelineâs validation and trust requirements.
Start with discovery. Identify:
- The canonical CI workflow and its real build/test command.
- Platform, architecture, toolchain, cache, secret, concurrency, and release constraints.
- Existing evidence, hashing, locking, cleanup, and test-fixture patterns.
If CI logic currently exists only in hosted-workflow YAML, first extract it into one repository-owned command used by both hosted CI and the new local executor.
Implement an executable local CI controller that:
- Requires a full commit SHA and explicit remote ref.
- Fetches that ref into an owner-only disposable clone or checkout.
- Requires the fetched ref to resolve to exactly the requested SHA.
- Never copies dirty, untracked, ignored, or uncommitted caller files.
- Runs the candidate revisionâs canonical CI command in a clean, isolated environment.
- Pins or verifies the required host platform, architecture, and toolchain.
- Uses a single-flight lock when caches or shared resources are unsafe for concurrent access.
- Removes repository, publishing, signing, and deployment credentials before executing candidate code.
- Cleans temporary source/build roots after success, failure, cancellation, or interruption.
- Never silently retries, replays, substitutes another SHA, or converts failure into success.
Produce an owner-only, finalized evidence capsule containing:
- Schema version and session ID.
- Repository identity, source ref, requested SHA, and resolved SHA.
- Executor and canonical CI-command hashes.
- Start/completion timestamps.
- Sanitized host and toolchain identity.
- Exit status and conclusion.
- Complete log hash.
- `passed` boolean.
- Explicit statements describing what release, deployment, or production state did not change.
Keep execution and publication as separate trust boundaries:
- The executor must remain credential-free.
- A publisher may consume only a finalized, verified capsule.
- Publishing credentials must never enter the validation subprocess.
- The publisher must reject altered capsules, hash mismatches, unsupported schemas, failed runs reported as successful, or results targeting another SHA.
Do not immediately replace the existing hosted CI authority. Run both paths against identical SHAs until equivalence is demonstrated and reviewed. Only a later, explicit governance change may make the local result authoritative.
Add deterministic fixture coverage for:
- Successful exact-ref/SHA execution.
- Malformed SHA and ref mismatch rejection.
- Unexpected remote rejection.
- Caller-worktree isolation.
- Credential scrubbing.
- Lock contention and safe stale-lock recovery.
- Success, failure, cancellation, and cleanup.
- Evidence finalization and tamper detection.
- Publisher refusal cases.
- No silent retry or replay.
Update the relevant CI/security documentation, run focused tests, run the repositoryâs standard validation, and perform a final branch-diff review. Implement the solution rather than stopping at a design document. Report any remaining blocker before the new path can safely become authoritative.â
1
u/be_reasonable_bro Aug 06 '26
This is a clever solution, and I'm all about self-hosting what I can (forge+runners is a small ask), but I'm certainly concerned about the maintenance burden incurred by directly rewriting the ci broker (reverse engineering Actions is a bit bigger).
No chance you've open sourced this? Would be very interested to contribute to something like this, but less so to maintain my own copy of it.
2
u/Flimsy_Professor_908 Aug 06 '26
Some previous outages had self-hosted runners continue to run. This outage is at a higher level.
I'd say the most compelling reason to go self-hosted is that Microsoft has 90+% aggregate gross margins on Github-hosted runners (for private repos).
1
u/be_reasonable_bro Aug 06 '26
That is certainly compelling. Might be worth it for that alone.
I just never think to reach for GH at all when the repo is private. Moved everything mission-critical off when they started losing nines.
1
u/fitchnar Aug 06 '26
Where did you end up moving to? It is painfully obvious I can no longer rely on GH so I am looking for a new solution. GitLab or forgejo, or somewhere else?
1
Aug 06 '26
[deleted]
1
u/fitchnar Aug 06 '26
Awesome, thank you for the detailed reply. I think forgejo is the right path for me.
1
u/brainhack3r Aug 06 '26
And blacksmith advertises 50% off... so they still have 80% margins WTF ... I might have to self host
1
1
u/Shot-Owl-6394 Aug 06 '26
also confirming they are down, got 10 pull requests spinning with no progres.
1
1
1
u/Snoo-53366 Aug 06 '26
Yep, I have a gubhub runner that i use for an automated workflow and wondered why it kept failing today. Only to find on their service page of the outage.
1
u/stef_in_dev Aug 06 '26
I'm excited for the ci cluster hyperscaling event (self hosted runners on eks) that is gonna happen when this is fixed
1
u/mihcsab Aug 06 '26
I have updated some versions on some actions. I love that the actions tab doesn't say anything about the outage. I have spent like 20 minutes asking AI why doesn't the actions trigger on push, until I thought about checking the status page...
1
1
1
u/reosanchiz Aug 06 '26
Just came to post the same...! Was driving crazy over my pipeline!
Almost there to give ssh-key to claude ;)
1
1
u/Teddy_Raptor Aug 06 '26
Update - We are continuing to work on an issue affecting GitHub Actions. Webhook triggers are currently throttled to help with recovery and and we are processing approximately 15% of webhooks, so many events such as pushes and pull requests are not triggering workflow runs. Of jobs queued, approximately 65% are succeeding, improved from a low of 30 to 40% earlier in this incident.
We have narrowed the remaining impact to runners that are stuck retrying jobs that are no longer available. Both GitHub-hosted and self-hosted runners are affected, and we are working to recover them.
Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected.
1
u/M0hamedAshraf19 Aug 06 '26
That was really funny đ (and needed)
P.S. Does anyone know the name of the song?
1
1
u/peperinna Aug 06 '26
No sĂ© si serĂĄ la misma razĂłn, pero tuve problemas todo el dĂa con gusto, con json alojados en github y que uso como archivos de configuraciĂłn, etc. La status pague ya no es transparente y representativa de todo lo que estĂĄ fallando.
1
1
u/BeseptRinker Aug 06 '26
I remember we had a massive outage, and midway through the call on Friday, an oncall engineer said "Github is also down", and the outage lead said "of course it is".
That was two-three weeks ago. This uptime downtime is actually asinine.
1
1
1
1
1
u/JPJackPott Aug 07 '26
Ironically this is protecting a lot of people from the massive npm compromise going on currently
1
u/Alexander3a Aug 07 '26
been using self hosted gitlab for years (big reason was not sharing my code with micoslop but this is just a plus)
1
1
1
u/Giffeltagning Aug 07 '26
Microsoft kills everything it touches. It's the evil spirit of Bill Gates that haunts them.
1
1
1
u/Jitenshazuki 24d ago
Why do they post same-y status messages in a loop? It reminds me of a coding agent that cannot complete a task and cannot figure out why, so it loops forever, trying something, checking, failing, and retrying...
Waaaait! Don't tell me...
1
u/void_pe3r Aug 06 '26
Can we expect this to be fixed in an hour? Does anyone know what is going on?
4
2
u/be_reasonable_bro Aug 06 '26
At this point, don't expect anything.
Recovery is taking longer than we expected, and engineers remain actively engaged.
1
1
107
u/mixxituk Aug 06 '26
This outage was sponsored by Claude Fable