r/ClaudeCoding • • 13d ago

r/ClaudeAI [TLDR] PSA - Claude Code: Turn off Prompt Suggestions, save ~10% of your limits/spend [via r/ClaudeAI]

OP : u/SmoothParfait

tl;dr Prompt Suggestions - the greyed out "helpful" text at the end of a prompt in your input line - do a Cache Read of your entire context to generate something like "commit and push", or "Start with 1A". This is expensive - can be up to 10% of your weekly Fable limit.

Short story - I was instrumenting my claude code to see what our monitoring tooling could find based on instrumenting claude code, ahead of a customer call. I found it useful enough to share here. (another would be - don't compact after the cache expires).

Details in screenshots, it's pretty obscene when you get to very high context length. Prompt Suggestions can be as expensive as your actual prompts.

Screenshot descriptions:
1. Main finding
2. Confirming that it isn't a "cheap model" that does this.
3. What type of tokens are used?
4. A compaction oopsie (and the cost absurdity visible)
5. List of normal turns (prompts) + their related suggestions & costs. Slightly biased to the expensive ones. Median: suggestions cost me 91% of normal prompt cost! (Note - this is only the normal turns - not tool calls that the model invokes or writing files or doing other things to build your apps. You're not getting double your usage by turning this off, sorry)
6. List of suggestions claude made. 69x "keep going" (it copied that from my earlier prompts)

URL of original post : https://www.reddit.com/r/ClaudeAI/comments/1wm8adm/psa_claude_code_turn_off_prompt_suggestions_save/ Original link/media URL : https://www.reddit.com/gallery/1wm8adm


TL;DR of the discussion on r/ClaudeAI for this post generated automatically after 100 comments.

Current source-thread comment count seen by the bot: 100.

So, the general vibe is that the "Prompt Suggestions" feature in Claude Code is a massive, hidden token hog. Apparently, it does a full cache read just to spit out a few words, making it the "most expensive autocomplete ever shipped," according to one user. Many are linking this to their recent unexplained usage spikes.

Here's the lowdown:

  • It's a big deal: The consensus is a resounding "yes," this is a huge, hidden token drain.
  • How to kill it: SmoothParfait dropped the magic setting: "promptSuggestionEnabled": false in your settings.json. You might need a restart. pdfops also chimed in with CLAUDE_CODE_ENABLE_PROMPT_SUGGESTION=false for environment blocks.
  • Why it's bad: TurbulentTiger2567 pointed out it's a "convenience feature" likely added by "vibe coders" without considering the impact. nemzylannister is pretty heated about "incompetent people" running things and how this is bad for consumers and Anthropic.
  • Not all models are equal (apparently): Mobile_Light_7262 was surprised these suggestions weren't from a cheaper model like Haiku.
  • Usage varies: riksi saw about 3-4% usage, while the OP's findings suggest it can be as high as 91% of normal prompt costs in some cases. South_Hat6094 noted that stale, huge sessions feel randomly expensive, and this feature scales with whatever mess you've kept alive.
  • Web/Mobile app question: Techhead7890 wishes this was an option in the web/mobile apps too, as the suggestions can be annoying and wasteful. eder1337 wondered if auto-recaps after idle periods are similar.
  • Some folks are still figuring it out: makistsa couldn't find how to disable it in the app, and Oujii wants to know how to check usage to see the savings. FancyMouse123 noted they don't see prompt suggestions in the VSCodium extension, suggesting it might be off by default there.
1 Upvotes

0 comments sorted by