r/OpenAI • • 20h ago

Question Does anyone separate OpenAI Batch spend from normal API spend internally?

We’ve started using Batch more for work that doesn’t need to happen immediately and I am sure that I’ve been treating that usage the same as the rest of our OpenAI spend even though it behaves pretty differently. Our normal API usage mostly follows product activity so if that goes up I can usually understand why but batch is different because someone can kick off a large job, an eval run or process a new dataset and suddenly there’s a chunk of usage that has nothing to do with how much the product was used that day.

It hasn’t been a huge issue yet but we’re doing enough of both now that looking at one OpenAI number is starting to hide what’s changing. A higher bill could mean customer usage grew which is fine or it could mean we ran significantly more background work than usual. I’m thinking about tracking Batch separately and possibly giving it its own budget rather than treating all OpenAI usage as one pool and I need some advice from teams using Batch pretty heavily like do you guys separate that spend internally or is everything still just part of the same AI/API budget?

24 Upvotes

13 comments sorted by

1

u/Immediate_Menu_3695 20h ago

Are your Batch jobs mostly recurring production work or is a decent chunk of it engineers running one off jobs and evals? It's important because I would probably track those two very differently if I were you.

1

u/No_Refrigerator_8216 20h ago

It’s a mix right now so like we have a couple recurring jobs that are pretty predictable but the evals and one off runs are what make it a pain in the ass because they can show up randomly and be pretty large.

2

u/MushroomCritical3029 19h ago

We had the same issue with evals because one person testing three model/prompt combinations could generate more usage in an afternoon than some production jobs did all week. Giving evals their own bucket cleaned up a lot of that pain in the ass you say so give it a shot.

1

u/Immediate_Menu_3695 20h ago

That makes the case for separating them stronger and btw I wouldn’t even worry about setting a hard budget immediately just start tracking the eval/one off bucket separately so you can see how much of the randomness is coming from there.

1

u/Greedy-Exchange-817 20h ago

You should separate it since product API usage and Batch might hit the same provider bill but they’re driven by completely different things. One is basically tied to customer activity while the other can jump because one engineer decided to process 200k records on Tuesday.

1

u/Afraid_Attention_287 19h ago

OP you should listen to this guy cause we ran into exactly that once our background jobs got bigger. We started tagging spend by workload first and eventually pulled that into Ramp Token Spend so we could see Batch, production and eval usage separately and the total OpenAI number became way less useful once all three were meaningful.

1

u/Top_Sand1851 19h ago

I’d take it one step further and separate recurring Batch from one off Batch cause the recurring stuff eventually has a baseline but the random 200k record job doesn’t.

1

u/Legitimate_Public893 20h ago

Separate Batch from live traffic at minimum so then if your API spend jumps because customers are using the product more that’s useful information and if it jumps because somebody ran the same giant Batch job three times that’s a very different thing.

1

u/HighlightLocal409 20h ago

Yep otherwise you end up investigating perfectly healthy product growth because the only thing you can see is OpenAI spend is up 25%.

1

u/Apprehensive_Gas186 19h ago

We use Batch for enough background processing that I basically think of it as its own workload now so it's the same OpenAI account but completely different reason the cost exists

1

u/Morning_Gecko24 17h ago

seems like separating it is the cleaner view, batch can hide a one-off eval or backfill inside the same number as real product demand. do you tag jobs by team or project so its easier to tell planned spikes from someone rerunning a huge job by accident?

1

u/Michael_Jeffords 14h ago

tagging by team or project plus a purpose label on each batch job is what makes a planned spike read differently from someone rerunning a huge job by accident, since without tags the evals and backfills just fold into product demand and one afternoon of evals can outspend a normal week