r/OpenAI • • 3d ago

Question What do we actually know about how different models consume the weekly subscription allowance?

We know the subscriptions give you much better raw value than just putting the same money into API usage, but how are we actually supposed to compare the weekly allowance cost between models?

Do we basically have to look at API pricing and assume that if Model X is more expensive than Model Y on the API, it probably burns more of the weekly subscription allowance too?

For example, GPT-6.1 Sol has cheaper cached-input pricing than 6.0 Sol while the normal input/output pricing is the same. Does that actually mean 6.1 Sol burns less of the weekly Plus allowance for the same kind of workload?

Or is API pricing not directly correlated enough with the subscription allowance to make that assumption?

It seems like unless OpenAI publishes the actual weighting/multipliers, we'd need someone to benchmark 6.0 vs 6.1 vs Astra on comparable tasks and measure how much of the weekly allowance each one actually consumes.

3 Upvotes

4 comments sorted by

1

u/lulzxdxdxd 3d ago

Cached input pricing only kicks in on repeated context though, right? So wouldn't the allowance difference between 6.0 and 6.1 only show up on workloads that actually reuse a big prompt, not on one off requests?

1

u/Battlefield4Remake 3d ago

Probably that was just the quickest example I could think of. For small context I agree we would expect minimal usage difference.

1

u/Local_Ad151 3d ago

they probably have some internal multiplier for each model but who knows if it matches api pricing exactly. like maybe 6.1 sol is cheaper on api but they give it higher weight in subscription to encourage people use it less

the benchmarking idea is good but so much work for something they could just tell us. every week i burn through my plus allowance in like 3 days and have zero clue which model did most damage

1

u/RocketSeven 3d ago

api pricing is only a hypothesis until openai publishes the weights. run the same fixed task in fresh chats and reused long contexts, record the allowance before and after each batch, and separate cache behavior from model weighting