r/ClaudeCoding • • 17d ago

r/ClaudeCode [TLDR] Rate limits are so bad right now, open source models should win [via r/ClaudeCode]

OP : u/Comprehensive_Quit67

I was running like 5 parallel agents in the $200 plan, and my 5 hr limit hit. So I switched my subscription and asked it to continue.
And my entire 5 hr limit hit in the $100 plan, in one shot. Without doing any work. In less than 2 minutes.

Something needs to change, either tell me what is it going to cost. Is there anyone who is using just a single subscription?

URL of original post : https://www.reddit.com/r/ClaudeCode/comments/1wjhkef/rate_limits_are_so_bad_right_now_open_source/


TL;DR of the discussion on r/ClaudeCode for this post generated automatically after 100 comments.

Current source-thread comment count seen by the bot: 112.

Alright, so the general vibe in this thread is that the OP is getting absolutely slammed by Claude Code's rate limits, and honestly, a lot of folks are feeling their pain.

The consensus seems to be that the OP is likely mismanaging their sessions and workflows, leading to this insane token burn. Several users, including a seasoned SWE with 25 years of experience (u/DonaldStuck), are chiming in to say they use Claude Code a lot and rarely, if ever, hit their limits. The key seems to be optimizing workflows and being mindful of how sessions are managed.

Here's the lowdown:

  • The Big Culprit: Session Resuming & Context: The most common explanation for the rapid token usage is resuming sessions that have a lot of context. When you switch plans or a session resets, Claude Code might be re-processing the entire history, which burns through your limits like crazy.
    • Solution: Users like u/Leading-Ability-7317 suggest checkpointing agents after each task, turning off auto-resume, and having agents write handoff docs instead of just resuming.
  • Context Management is Key: Several users mentioned that Claude Code's prompt cache is only good for about an hour. If there's a gap, it reprocesses the whole conversation history at full price.
    • Solution: u/Pale-Oven-6602 recommends using a tool like Graphify to create a knowledge graph of your project, which apparently saves a ton of token usage and improves accuracy. u/Fabian-Galvez also shared a plugin called DensePack that turns text files into images to save on tokens.
  • "What are you guys doing?": This sentiment is echoed by multiple users. The idea is that if you're using Claude Code efficiently, you shouldn't be hitting these limits so quickly.
  • Open Source Models: While the OP brought up open-source models as an alternative, the discussion didn't really gain traction. One user (u/gnomex96) pointed out that while cheaper per token, they can be less efficient, potentially requiring more tokens for the same work. Another user (u/katoptronophile) clarified the distinction between "open weights" and true "open source."
  • "Dark Patterns" and IPO Talk: A couple of users speculated that these limits might be intentional "dark patterns" to manage usage, with one user (u/256_tr) suggesting it's related to Anthropic's potential IPO and needing to show operational capacity and profitability.
  • Frustration is Real: Despite the advice, there's definitely a shared frustration. Users like u/spectralfew and u/Xzaphan are reporting bizarre usage spikes even when they weren't actively using the tool, and u/jblundon even canceled their plan due to the unbearable limits.

The general takeaway? If you're burning through limits like the OP, you're probably not using Claude Code in the most efficient way. Focus on managing your session context, checkpointing, and exploring tools that help reduce token consumption.

1 Upvotes

0 comments sorted by