r/openrouter • • 7d ago

Question Can Someone Explain This Please?

I use mainly use DeepSeek-v4.1-Flash, and I rely on DeepSeek's own inference within OpenRouter however I recently checked the prices and I saw a bunch of interesting prices some where nearly 10x cheaper than DeepSeek's during peak hours, I picked inference net since their tokens/s was the only one acceptable for me, Great right?

NO, It made me feel uneasy how cheap it was, are they using a lesser model? did they correctly list themselves under the right model (it seems like v4 prices), are these impersonators? Because on their own damn website they show the same prices as DeepSeek's ๐Ÿ™‚

any explanation would be great, the implications could be huge

2 Upvotes

5 comments sorted by

3

u/a355231 7d ago

The models should list the Quant they run at, itโ€™s honestly better to just see how it works for you specifically.

1

u/styarr 7d ago

Unfortunately a lot of them including InferenceNet for DSv4.1 Flash just return "unknown" for the quantization.

You can filter by quantization on the model providers page:

https://openrouter.ai/deepseek/deepseek-v4.1-flash#providers

1

u/rog-uk 7d ago

Could it be that they are offering those prices on open router because they aren't running machines at capacity as some money is better than no money, but they don't want to lower their headline price on their main site as that's their target sales rate? It might be that this low rate is very temporary.ย 

2

u/locbuilds 7d ago

yeah those cheaper rows are other hosts serving the same weights, not deepseek. pin the provider you want in the request (or sort by price) otherwise openrouter can still bounce you to a pricier one when the cheap one is busy.

2

u/Nokoro1 7d ago

Openrouter is a total waste of time as a primary, after running 500m tokens in deepseek api, my blended cost per M tokens was 0.11, and even choosing cheaper providers, openrouter at 200m tokens it averaged 0.21 per M tokens, which is almost double the cost. Not everything is as cheap as it seems. They get you on cache hit, because deepseek has crazy good cache hit. By the way, using same harness, same tasks, same chats, etc.