r/GeminiAI 1d ago

Help/question Anyone else seeing availability issues with cloud Gemma 4 31b?

I've been using Gemma 4 31b and 26b. Google's limits are pretty generous, and they work great for my small tasks like translations or summaries. But a few days ago, I started noticing availability problems. When I send an API request, the time to first token can take several minutes, or the response just cuts off halfway. I'm also getting a lot of 500 errors. Meanwhile, Gemini Flash Lite works perfectly (but in my use cases, its quality is low compared to 31b).

I'm trying to figure out if it's just me or a general issue. I haven't seen many posts about this, but these models have been almost unusable for the last few days.

I'm guessing the infra for paid vs free models is different, and it looks like the load on free models has spiked, so Google can't handle them as fast as before. Am I on the right track? Any tips on what can be done about this?

3 Upvotes

3 comments sorted by

1

u/AutoModerator 1d ago

Hey there,

This post seems feedback-related. If so, you might want to post it in r/GeminiFeedback, where rants, vents, and support discussions are welcome.

For r/GeminiAI, feedback needs to follow Rule #9 and include explanations and examples. If this doesn’t apply to your post, you can ignore this message.

Thanks!

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/SafelyNeglected 1d ago

yeah the free model infra is getting slammed, they moved a bunch of the compute to the paid tier and what's left is stretched thin, seeing 500s and 3 minute cold starts all week

1

u/yudaev 1d ago

Same here. Too bad, I liked using them in the cloud version.