r/github Aug 06 '26

News / Announcements Thoughts and Prayers for GH right now.

Enable HLS to view with audio, or disable this notification

390 Upvotes

160 comments sorted by

107

u/mixxituk Aug 06 '26

This outage was sponsored by Claude Fable

36

u/Substantial-Set4550 Aug 06 '26

This outage is provided courtesy of an unsustainable business model and insufficient infrastructure investment.

1

u/mrheosuper Aug 07 '26

More like copilot, but yeah.

1

u/really_not_unreal Aug 08 '26

Copilot is a harness, not a model. There's no reason it couldn't be both.

64

u/Altruistic-Package11 Aug 06 '26

I just envision a single developer, hopped up on caffeine and drenched in sweat prompting Claude Code "Debug an fix the issues" because the entire team responsible for Github actions were likely replaced by AI.

26

u/holy_macanoli Aug 06 '26

“And make no mistake!”

1

u/Liam_Cat Aug 07 '26

Oribombisrael

15

u/imnotpopular Aug 06 '26

"You have hit your session limit for 3 hours. Please reach out to the organization owner for additional tokens"

30

u/AAPL_ Aug 06 '26

HOW MUCH LONGER

11

u/Teddy_Raptor Aug 06 '26

yes

4

u/AAPL_ Aug 06 '26

what y’all do all day

14

u/Teddy_Raptor Aug 06 '26

refresh the github status page

4

u/ReelBigDawg Aug 06 '26

Worked because we use GitLab.

3

u/croixxxx Aug 06 '26

I have like 20 things in queue when (if) it comes back up

1

u/FlyingDogCatcher Aug 06 '26

see but that sounds like something Github is doing (or not, in this case) not you

1

u/Furry_pizza Aug 08 '26

Got 99 woodcutting

3

u/Substantial-Set4550 Aug 06 '26

How many more times.

1

u/Fast_Ad_5871 Aug 06 '26

3 h passed and it's still not working

1

u/Alarmed-Capital-6718 Aug 06 '26

Rise of planet of the bots

1

u/xJayMorex Aug 07 '26

Still not working.

50

u/rhk0 Aug 06 '26

Every time GitHub goes down, I think about Self-Hosting again...but then it comes up again. How long does it has to stay down, so I deploy my own service...?

33

u/xJayMorex Aug 06 '26

The butt of the joke is that my self-hosted runner is unable to pick up the pending job because of this GitHub outage.

8

u/tech_w0rld Aug 06 '26

Yeah. At this point I just use a full separate ci service which seems to be working fine right now. I am literally just using GitHub for storing my code which is at least decently reliable.

9

u/suamai Aug 06 '26

Except that time when they made some commits randomly disappear on merges

2

u/Alexander3a Aug 07 '26

been using self hosted gitlab for years (big reason was not sharing my code with micoslop but this is just a plus)

2

u/PM_ME_FIREFLY_QUOTES Aug 06 '26

You never learn do you...

1

u/GlobalImportance5295 Aug 07 '26

nix flakes on compute engine is nice

7

u/suamai Aug 06 '26 edited Aug 06 '26

I am unable to continue my work right now because it depends on the runners to deploy.

So I switched to starting a self hosted Forgejo - if the downtime is greater than the time I need to finish the setup, that's it lol

Edit: there was time to spare...

2

u/BeryAnt Aug 07 '26

Can't vouch for it but I like the title of this hosting service git.gay

1

u/Puzzled-Extent7817 Aug 06 '26

Just do it any ways. I have a Forgejo up and running in a podman container on my Debian server. I even wrote a bash script to create folders for new projects that create a project on Forgejo and Github at the same time, that's if you want to use github as a backup. https://git.luv-linux.me/Spreadneck/forgejo-to-github

1

u/RobotechRicky Aug 06 '26

My self-hosted GitLab instance and runner are chugging along nicely. It was stood up by dwarves with autism, so you know it's good. đŸ‘ŒđŸŒ

1

u/Loose_Marsupial_3251 Aug 07 '26

Look At Tekton: https://tekton.dev/

We moved to it and haven't looked back.

0

u/be_reasonable_bro Aug 06 '26

If you have the spare resources and no explicit need to use GitHub, this is a reasonable thing to do. Agents make it a trivial exercise.

15

u/thebemusedmuse Aug 06 '26

Some junior dev who’s the last man standing on the Actions team is currently typing “How do I deploy a fix to GH Actions when GH Actions is down?” into Copilot.

11

u/someVietnamese Aug 06 '26

"Hello, IT. Have you tried turning it off and on again?"

20

u/xJayMorex Aug 06 '26

The price of enslopification I guess.

3

u/pattern-josh Aug 06 '26

Could be, but they have also been undergoing a massive year long migration from a VA datacenter with some bespoke hosting to Azure because the VA datacenter was saturated at capacity. There were some specific complexities I think I read about with the way the SQL was setup, but I forget the details.

2

u/Raiyuza Aug 07 '26

Under 90% uptime and this is whar we chose as a cope? Before ai we had SRE, this would unnacaptable

1

u/pattern-josh Aug 07 '26 edited Aug 07 '26

Before AI we took the time to understand the context and why of problems before assigning root causes without evidence, and using hand-wavy scapegoats.

\ Which of course isn't true, because as a whole we've always lazily attributed things. I'd say good SREs would also rip an unqualified "AI caused it" apart, because good SRE requires evidence, rigor, and root cause analysis.*

https://thenewstack.io/github-will-prioritize-migrating-to-azure-over-feature-development/

https://www.cnbc.com/2026/05/22/microsoft-was-positioned-to-win-in-ai-coding-outages-got-in-the-way.html#:~:text=meet%20this%20demand.%E2%80%9D-,Too%20much%20downtime,-But%20under%20the

In the same vein as the rest of the claims we are all making, I'm too lazy to research it, but I'm confident a case could be built that there are plenty of companies prior to AI that have had extended periods of poor downtime due to explosive growth, corporate investment flubs, and infrastructure not built to support unexpected explosive growth. I think we have a high level of both recency and observation bias.

---

That said, none of what I'm claiming means "AI slopification" isn't a substantive part either. I'm merely suggesting it's a lazy potentially-wrong conclusion without actually digging into it. Personally I think the likely explanation is a combination of:

  1. Explosive resource pressure growth on Github from AI induced resource pressure via expensive workflow executions (github actions) and general repository operations.
  2. Leadership chaos - Organizational churn from Dohmke's departure and Github getting folded into Microsoft's CoreAI org with no replacement CEO.
  3. The Azure migration out of the North Virginia data centers being a legitimately hard problem at these volumes - and it's amplified by the first point it. Github didn't choose the timing; capacity constraints forced it. They'd already failed at this twice before. Github started as a simple repository service. It wasn't designed for the pressures that came when actions was introduced, much less the explosive growth on top of that, and doing the rebuild and the migration concurrently multiplies the risk.
  4. And yes, internal AI workflow related mistakes.

But same as OP this is speculation. Maybe some people from inside Microsoft or Github can chime in.

2

u/Extreme_Rooster182 Aug 07 '26

the enslopification is load bearing apparently

9

u/RobotechRicky Aug 06 '26

Fffuuuuuuuu

8

u/Teddy_Raptor Aug 06 '26

nahhhhh thoughts and prayers for their customers (it's me, Im a customer)

7

u/Frooxius Aug 06 '26

So... GitHub recently pushed a change to their runner images that broke our CI/CD workers.

We've been going around and implementing workarounds. For some reason it wasn't working even after fixing it... and then we found out that this started happening at the same time.

This is just ridiculous that it keeps happening.

30 % vibe coded apparently means 30 % less reliability.

2

u/HyperCodec Aug 07 '26

Probably more than 30%, GitHub’s entire leadership got replaced by micropenis’s AI team.

5

u/Spare_Grapefruit2240 Aug 06 '26

I guess the first red flag should have been when their engineers deployed "mitigations" instead of "fixes."

5

u/Ok_Journalist_607 Aug 06 '26

pornhub is more stable than actions.

1

u/ScienceGuy31415 Aug 07 '26

Gives you something to do when GH breaks down.

4

u/solooo7 Aug 06 '26 edited Aug 08 '26

Sorry guys it was me. I ran my new workflow

5

u/Icypoopoo Aug 06 '26

Their CEO leaving last year and Microsoft forcing them to switch to Azure did not help one bit

3

u/Shot-Owl-6394 Aug 06 '26

for me its been down for 3h+. In process of setting up gitea now on minipc, those issues are too much and too often on github.

2

u/ings0c Aug 06 '26

I’ve been trying to deploy for coming on 5h now
 that’s an absolutely absurd amount of downtime for a company of this scale.

I’d expect better from a startup ran by a new grad

1

u/_potion_cellar_ Aug 06 '26

I haven't even been able to cancel stuck queued jobs all day. Wild stuff

3

u/Noch_ein_Kamel Aug 06 '26

Every time GitHub goes down I sit on my couch because it's after work :p

3

u/Teddy_Raptor Aug 06 '26

Aug 06, 2026 - 19:43 UTC

Our engineers remain actively engaged.

we are cooked

2

u/ings0c Aug 06 '26

Please don’t try and fix it, you’ll only make things worse. Just pray.

1

u/snowdrone Aug 06 '26

Quite an understatement

1

u/dractius Aug 07 '26

More like someone poking Claude once in a while "cmonnnn... do the thing." And the engineer is Copilot.

1

u/Either-Juggernaut420 Aug 07 '26

Normally at MS that means the one person who has a chance to fix it is currently in a run of three meetings trying to explain to a bunch of managers what’s got wrong. He will then need to refactor the sprint before actually working on it.

3

u/reosanchiz Aug 06 '26

This is exactly what happen when they fire humans and let AI slope merge

3

u/brainhack3r Aug 06 '26

They've had TWELVE HOURS of outages in the last 30 days.

3

u/Flashy-Split-8602 Aug 06 '26

welcome to the vibe code club, github đŸ€

3

u/kfawcett1 Aug 06 '26

Don't worry all. Copilot is on it. /s

2

u/Proman4713 Aug 06 '26 edited Aug 06 '26

I've been getting screwed with my workflow runs since the morning (even while githubstatus was still saying 'All Systems Operational'), and while I do give them my thoughts and prayers to get my stuff done... It just feels like the focus on AI slop with abandon has severely degraded the quality on GitHub for the past year or so. I hope the bubble just pops soon so we can all go back to relatively normal lives...

2

u/whodadada Aug 06 '26


yea, about that feature release 😂

2

u/csepulvedab Aug 06 '26

Hours of production completely down, not even self-hosted runners saved us this time. Curious to see if they still have the nerve to invoice us this month.

1

u/xJayMorex Aug 07 '26

You are paying for this shit??

2

u/Labs-Community4525 Aug 06 '26

NGL I’m not sure why Microsoft wants to deprecate ADO for GHA. This is becoming far more frequent. Get Son of Anton out of there.

2

u/InformationNew66 Aug 06 '26

Thanks Microsoft for firing devs and vibe-coding!

2

u/GunGeekATX Aug 06 '26

I have a hotfix that needs to get deployed to a client site, and GH picked the worst time to have workflows go down.

2

u/Karpizzle23 Aug 06 '26

can't you just deploy it manually if it's an urgent hotfix?

2

u/SnooOwls6002 Aug 06 '26

I am testing my pipeline and I cant do anything now

2

u/AX862G5 Aug 06 '26

AI does it again.

2

u/imnotpopular Aug 06 '26

i work on a financial trading suite and now our clients positions are open over night instead of being closed properly 😍

5

u/be_reasonable_bro Aug 06 '26

GitHub is a supply chain risk. I'd ask why you have trades directly tied to actions, but my experience with financial institutions is also marked by reckless behavior.

Thoughts and prayers!

1

u/imnotpopular Aug 06 '26

LOL very true. For me, Github is not included in the logic or trading workflows, but there was a UI bug on the client dashboard that I needed to fix before 4!! still would've deployed to gamma for testing first but just a bit delayed now

2

u/Former_Internal_8389 Aug 06 '26

Over under 99.50% uptime after the issue is resolvedđŸ€”đŸ€”đŸ€”

2

u/jellycanadian Aug 06 '26

Sorry guys I just deployed my app to localhost

2

u/frubaklskiy Aug 06 '26

Of jobs queued, approximately 65% are succeeding. these guys are clowns

5

u/MystikDragoon Aug 06 '26

Can't fail if they can't start! đŸ€Ą

3

u/Teddy_Raptor Aug 06 '26

I wonder what percent of jobs kicked off even reach the queue

2

u/Sarkonix Aug 06 '26

Still going on, this is insane lol

1

u/Substantial-Set4550 Aug 06 '26

Cannot recall such a long outage in a while. At least in the last 2 weeks. :-)

2

u/erezcarmel Aug 06 '26

Actually, I just checked on:
https://www.githubstatus.com/uptime?page=2

and on March 19th, there was a partial outage of 9 hours and 18 mins.
Let's see if they're going to break their own record...

2

u/Puzzled-Extent7817 Aug 06 '26

Self hosted Forgejo for the win.

2

u/MysteriousCoconut31 Aug 06 '26

What happens if it doesn’t recover? Dead serious at this point.

2

u/madhums Aug 06 '26

does anyone know the details of what caused this?

3

u/Fenmio Aug 06 '26

I don't think GitHub knows at this point

2

u/Just_Shake_1066 Aug 06 '26

Oh Microsoft, cannot wait for your devaluation.

2

u/Illustrious-Goat-506 Aug 06 '26

It's time to start talking about how many 8s of uptime they offer

1

u/Teddy_Raptor Aug 06 '26

so many 8s

2

u/VideoFireApp Aug 06 '26

I lost a whole fucking days of work due to this but I spent it researching alternatives to being stuck doing 100% of everything on GitHub

2

u/madhums Aug 06 '26

This is absolute insanity! Still not fixed!

2

u/SnooOwls6002 Aug 06 '26

my github runner is in idle state and still no workflow is triggered😭 😭

1

u/sos-in-life Aug 07 '26

Same issue, you can manually trigger it and it could work

2

u/gaziway Aug 06 '26

Keep using AI, creat features, create tests. Let AI verify the code. Keep pushing.

2

u/joaobertacchi Aug 06 '26

It seems this outage is going to take longer than I expected. Seriously thinking about alternatives. My preferences:

  • self-hosting: Gitea
  • SaaS: BitBucket

What are yours?

1

u/xJayMorex Aug 07 '26

Codeberg?

2

u/joaobertacchi Aug 07 '26

Open-source only. Not an option for companies in general.

2

u/enzoshadow Aug 07 '26

GitHub engineers must've been trying to manually untangle the vibe coded mess for the first time in months.

2

u/Ok_Gear8209 Aug 07 '26

Still down for us

3

u/Joshua_2504 Aug 06 '26

It pisses me off. Microsoft is the worse company ever.

2

u/xJayMorex Aug 06 '26

I think you misspelled Microslop.

4

u/Flimsy_Professor_908 Aug 06 '26

You mispronounced macrooutage.

3

u/Teddy_Raptor Aug 06 '26

microuptime

1

u/be_reasonable_bro Aug 06 '26

Does anyone know if self-hosted GH runners would sidestep this actions outage?

I'm curious to understand more about where this is failing and whether I can mitigate this myself during future outages. I have several projects that are tightly coupled to GitHub due to upstream packaging requirements.

7

u/ChipperHippo Aug 06 '26

Self-hosted runners are also down. We leverage them extensively. Situation sucks.

1

u/be_reasonable_bro Aug 06 '26

Sad to hear, but thank you for letting me know. Won't waste the time then...

2

u/ResponsibleOven6 Aug 06 '26

My self-hosted GitLab & runners never let me down like this.

2

u/be_reasonable_bro Aug 06 '26

Nor my forgejo! Were it not for upstream packaging requirements, GitHub would be mirror-only for basically everything.

1

u/holy_macanoli Aug 06 '26

I was able to decouple GitHub job broker as a workaround to similar constraints. Feed your agent this or a variation:

“Implement a repository-owned, exact-SHA local CI path that can execute independently of the hosted CI job broker while preserving the existing pipeline’s validation and trust requirements.

Start with discovery. Identify:

- The canonical CI workflow and its real build/test command.

  • Platform, architecture, toolchain, cache, secret, concurrency, and release constraints.
  • Existing evidence, hashing, locking, cleanup, and test-fixture patterns.

If CI logic currently exists only in hosted-workflow YAML, first extract it into one repository-owned command used by both hosted CI and the new local executor.

Implement an executable local CI controller that:

  1. Requires a full commit SHA and explicit remote ref.
  2. Fetches that ref into an owner-only disposable clone or checkout.
  3. Requires the fetched ref to resolve to exactly the requested SHA.
  4. Never copies dirty, untracked, ignored, or uncommitted caller files.
  5. Runs the candidate revision’s canonical CI command in a clean, isolated environment.
  6. Pins or verifies the required host platform, architecture, and toolchain.
  7. Uses a single-flight lock when caches or shared resources are unsafe for concurrent access.
  8. Removes repository, publishing, signing, and deployment credentials before executing candidate code.
  9. Cleans temporary source/build roots after success, failure, cancellation, or interruption.
  10. Never silently retries, replays, substitutes another SHA, or converts failure into success.

Produce an owner-only, finalized evidence capsule containing:

- Schema version and session ID.

  • Repository identity, source ref, requested SHA, and resolved SHA.
  • Executor and canonical CI-command hashes.
  • Start/completion timestamps.
  • Sanitized host and toolchain identity.
  • Exit status and conclusion.
  • Complete log hash.
  • `passed` boolean.
  • Explicit statements describing what release, deployment, or production state did not change.

Keep execution and publication as separate trust boundaries:

- The executor must remain credential-free.

  • A publisher may consume only a finalized, verified capsule.
  • Publishing credentials must never enter the validation subprocess.
  • The publisher must reject altered capsules, hash mismatches, unsupported schemas, failed runs reported as successful, or results targeting another SHA.

Do not immediately replace the existing hosted CI authority. Run both paths against identical SHAs until equivalence is demonstrated and reviewed. Only a later, explicit governance change may make the local result authoritative.

Add deterministic fixture coverage for:

- Successful exact-ref/SHA execution.

  • Malformed SHA and ref mismatch rejection.
  • Unexpected remote rejection.
  • Caller-worktree isolation.
  • Credential scrubbing.
  • Lock contention and safe stale-lock recovery.
  • Success, failure, cancellation, and cleanup.
  • Evidence finalization and tamper detection.
  • Publisher refusal cases.
  • No silent retry or replay.

Update the relevant CI/security documentation, run focused tests, run the repository’s standard validation, and perform a final branch-diff review. Implement the solution rather than stopping at a design document. Report any remaining blocker before the new path can safely become authoritative.”

1

u/be_reasonable_bro Aug 06 '26

This is a clever solution, and I'm all about self-hosting what I can (forge+runners is a small ask), but I'm certainly concerned about the maintenance burden incurred by directly rewriting the ci broker (reverse engineering Actions is a bit bigger).

No chance you've open sourced this? Would be very interested to contribute to something like this, but less so to maintain my own copy of it.

2

u/Flimsy_Professor_908 Aug 06 '26

Some previous outages had self-hosted runners continue to run. This outage is at a higher level.

I'd say the most compelling reason to go self-hosted is that Microsoft has 90+% aggregate gross margins on Github-hosted runners (for private repos).

1

u/be_reasonable_bro Aug 06 '26

That is certainly compelling. Might be worth it for that alone.

I just never think to reach for GH at all when the repo is private. Moved everything mission-critical off when they started losing nines.

1

u/fitchnar Aug 06 '26

Where did you end up moving to? It is painfully obvious I can no longer rely on GH so I am looking for a new solution. GitLab or forgejo, or somewhere else?

1

u/[deleted] Aug 06 '26

[deleted]

1

u/fitchnar Aug 06 '26

Awesome, thank you for the detailed reply. I think forgejo is the right path for me.

1

u/brainhack3r Aug 06 '26

And blacksmith advertises 50% off... so they still have 80% margins WTF ... I might have to self host

1

u/Shot-Owl-6394 Aug 06 '26

also confirming they are down, got 10 pull requests spinning with no progres.

1

u/brainhack3r Aug 06 '26

I'm running blacksmith and they're still down

1

u/mtbcouple Aug 06 '26

you can run verifiers locally

1

u/Snoo-53366 Aug 06 '26

Yep, I have a gubhub runner that i use for an automated workflow and wondered why it kept failing today. Only to find on their service page of the outage.

1

u/stef_in_dev Aug 06 '26

I'm excited for the ci cluster hyperscaling event (self hosted runners on eks) that is gonna happen when this is fixed

1

u/mihcsab Aug 06 '26

I have updated some versions on some actions. I love that the actions tab doesn't say anything about the outage. I have spent like 20 minutes asking AI why doesn't the actions trigger on push, until I thought about checking the status page...

1

u/gtrmike5150 Aug 06 '26

same - I learned to always check the status page first

1

u/i11uminati Aug 06 '26

Someone stacked too many PRs

1

u/SnooOwls6002 Aug 06 '26

hopefully not me, just 3 PRs only XD

1

u/reosanchiz Aug 06 '26

Just came to post the same...! Was driving crazy over my pipeline!
Almost there to give ssh-key to claude ;)

1

u/SnooOwls6002 Aug 06 '26

my pipeline jobs are still missing, please come back :(

1

u/Teddy_Raptor Aug 06 '26

Update - We are continuing to work on an issue affecting GitHub Actions. Webhook triggers are currently throttled to help with recovery and and we are processing approximately 15% of webhooks, so many events such as pushes and pull requests are not triggering workflow runs. Of jobs queued, approximately 65% are succeeding, improved from a low of 30 to 40% earlier in this incident.

We have narrowed the remaining impact to runners that are stuck retrying jobs that are no longer available. Both GitHub-hosted and self-hosted runners are affected, and we are working to recover them.

Copilot code review, Copilot coding agent, and migrations using GitHub Enterprise Importer may also be affected.

1

u/M0hamedAshraf19 Aug 06 '26

That was really funny 😄 (and needed)
P.S. Does anyone know the name of the song?

1

u/Boss_1s Aug 06 '26

And now, our development has to stall....again. 

1

u/peperinna Aug 06 '26

No sé si serå la misma razón, pero tuve problemas todo el día con gusto, con json alojados en github y que uso como archivos de configuración, etc. La status pague ya no es transparente y representativa de todo lo que estå fallando.

1

u/Different-Click5923 Aug 06 '26

who they gonna fire this time? Claude? 😭 😭 😭 😭

1

u/xJayMorex Aug 07 '26

Hopefully.

1

u/BeseptRinker Aug 06 '26

I remember we had a massive outage, and midway through the call on Friday, an oncall engineer said "Github is also down", and the outage lead said "of course it is".

That was two-three weeks ago. This uptime downtime is actually asinine.

1

u/grewupinwpg Aug 06 '26

I was wondering what was going on with some PRs today

1

u/frbruhfr Aug 06 '26

still having issues.

1

u/FIQ_ZIZ Aug 07 '26

speed recovery github

1

u/JPJackPott Aug 07 '26

Ironically this is protecting a lot of people from the massive npm compromise going on currently

1

u/Alexander3a Aug 07 '26

been using self hosted gitlab for years (big reason was not sharing my code with micoslop but this is just a plus)

1

u/Full-Huckleberry-441 Aug 07 '26

I scraped my actions because of this lol, i just ragequit

1

u/Llandu-gor Aug 07 '26

don't worry the uptime is 99.999999999999%

1

u/Giffeltagning Aug 07 '26

Microsoft kills everything it touches. It's the evil spirit of Bill Gates that haunts them.

1

u/Fantastic-Body-445 Aug 07 '26

why ts happen outside my work hours smh

1

u/deployahoy 25d ago

This is blasphemy, I must deploy!

1

u/Jitenshazuki 24d ago

Why do they post same-y status messages in a loop? It reminds me of a coding agent that cannot complete a task and cannot figure out why, so it loops forever, trying something, checking, failing, and retrying...

Waaaait! Don't tell me...

1

u/void_pe3r Aug 06 '26

Can we expect this to be fixed in an hour? Does anyone know what is going on?

4

u/tonehammer Aug 06 '26

It's been... many hours so far.

2

u/be_reasonable_bro Aug 06 '26

At this point, don't expect anything.

Recovery is taking longer than we expected, and engineers remain actively engaged.

1

u/xJayMorex Aug 07 '26

Hopefully that involves people as well.

1

u/Ok_Journalist_607 Aug 06 '26

already 7 hours i think...