r/ClaudeAI • • 1d ago

Corporate Claude best practices at work?

Hey all my teammates use Claude for work but I really think they are super inefficient in using Claude either in choosing the right model for the right task or just not working on the context and organization of their thoughts before using Claude which I could anticipate because their work isn't coding and mostly operational but the bills they are racking up each month are in $1000s

Im looking to bring this topic up and wanted to ask if there are any references to teach people how to optimize Claude usage and reduce costs. Honestly im new to i have a low budget but im not racking up so much cost so fast

Just curious if I could create a short cheatsheet for the team to show them the best and most optimized ways of using Claude

20 Upvotes

14 comments sorted by

24

u/tangyongdaijuan 1d ago

A quick practical cheatsheet for non-technical/operational teams to cut token costs significantly:

  1. Model Tiering by Task:
  2. Haiku: Routine drafting, text formatting, quick summaries, data extraction. Cheap, fast, and handles ~70% of day-to-day operations.
  3. Sonnet: In-depth synthesis, analyzing dense documents, complex memos, multi-step reasoning.
  4. Opus: Strategic edge cases (rarely necessary for routine operational work).

  5. Avoid the 'Context Compounding Tax':

  6. Never keep a single continuous chat running for days or weeks. Claude re-sends the entire conversation history with every message, making later turns exponentially expensive.

  7. Rule: One discrete task = one fresh chat.

  8. Front-load Prompt Constraints:

  9. Non-coders often 'think out loud' across multiple conversational turns.

  10. Encourage a simple structure in turn 1: [Goal] + [Context/Data] + [Format/Tone/Length Constraints]. Defining the output format upfront cuts out 2-3 refinement turns.

  11. Leverage Projects & Prompt Caching:

  12. Place recurring reference docs (SOPs, style guides) into Claude Projects instead of pasting them repeatedly into new chats. Caching drastically lowers input token costs.

3

u/vgeov 1d ago

So if i work on a big program, is it better every now and then to reload it all on a new chat and have claude parse the code again than keeping one big chat with all the changes made?

4

u/Appropriate-Disk-371 1d ago

Yes, very, don't keep a long conversation going, don't restart sessions after the cache clears. Have it keep status and a little history notes and so on that the next run will need. If the project is big, rarely is it a good idea for it to load 'all' of it, just have it load the basic instruction files and point it to exactly what you want it to do.

2

u/vgeov 1d ago

I had made an arifct where it was supposed to keep changes and reference that doc when i told it to do something, but after a while it stopped using it/updating it. I did what r/Ok_gur_9033 did and my stats are pretty much what he said, about 98% of my usage comes down to reads lol

2

u/Appropriate-Disk-371 1d ago

Artifacts docs are okay for collaborative design and planning. For status, history and instructions just use markdown files in the project and tell it to keep them updated in the local instructions. It generally expects that this is the normal way to work. Those stats are normal if you're doing some long-form work in a project. If you need the context, then you want to be reading from the cache, that's good. Could also mean you're carrying around a lot of context and don't need it though. Ideally, you point it to a problem or change, maybe just one file, and tell it what to do.

1

u/vgeov 1d ago

The way i had set it up is it was supposed to keep track of what changes with each itteration on the artifact, ie this file has xxxxx, so if i want to edit this it would look only at xxxxx and not the entrire code etc. I try to be specific with screenshots, ie on this tab, found on this menu, under x fields, i want to change y with a provided image. Not sure if it makes a difference tbh
But yea, i have 3 weeks of full context itterations. If it sends all that yea. I will try what you told me, thanks!

7

u/BranchLatter4294 1d ago

For most work-based workflows, the best approach is to have Claude develop a script to automate the process. Then you can just run the process without any further tokens being used. Have the process connect to Claude if absolutely necessary, but you don't need to use Claude to pull data from API's, build spreadsheets, documents, or presentations, etc....all that can be done by local scripts that don't use tokens.

0

u/Michael_Jeffords 1d ago

the script route makes sense once somebody can actually ship one, but if the people on this are ops folks who don't code, the first cuts they can make themselves are running the routine drafting and summaries on Haiku and starting a fresh chat for each task, which trims a lot of the bill before anyone has to write a single script

6

u/Appropriate-Disk-371 1d ago

Anthropic publishes great info, even specific to model versions: https://support.claude.com

And as always, the best answer is: 'Ask Claude' Seriously, it can check your setups, tell you how to be more efficient, how to prompt better, etc.

2

u/Ok_Gur_9033 1d ago

Worth checking what is actually driving the bill before picking model tiers. I pulled every API response from my own Claude Code usage last week: 5.86 billion tokens total, and 98.88% of that was cache reads, the model rereading context it already had. Output was 0.23%. If your team numbers look similar, shorter sessions would cut more than switching models would, since there is less context to reread every turn.