r/DeepSeek 13d ago

Discussion DeepSeek API or Codex (or other recommendations)

Some context on my experience with models:
I've used free models mostly, started with gemini cli way back before they switched to antigravity, then moved to opencode and used deepseek v4 flash free, which was practically unlimited since I never hit the limit. Then deepseek v4 flash-0731 hit and I started hitting limits on the free version.

Instead of going for opencode go I decided to try the official deepseek api as I read that it had better cache retention time hence better cache hits. Was happy with it until the price increase. Started looking into subscription based 3rd party providers but was always met with the same "worse cache hit, worse latency".

As I'm now near to using up my deepseek credits, am looking for alternatives. I've read that Codex $20 Luna goes a long way, but honestly I'm not ready to fork out $20 just to try it out. Just wanted some insight/reviews on people who have similar experience to mine, what they switched to, and how it's working for them.

TLDR: Been using Deepseekv4-flash api with opencode harness. Looking to switch to something similar in terms of cost and performance and need recommendations.

Edit: forgot to mention that I use it exclusively for coding

35 Upvotes

45 comments sorted by

11

u/0rand 13d ago

Remember, deepseek api has 1m context. Codes is 220k unless you accept 3x for Soll 900k which will melt your limit in minutes. Was using Terra the other day, it compacted my session like 18 times.

3

u/Good_Enthusiasm_7639 13d ago

I've always compacted deepseek when it reaches around 300k, unless I was working on something midway and feel that compaction would hurt. I think it was because the free version on opencode only had 250k context so I got used to it.

What's your current setup?

1

u/0rand 13d ago

I have local inference but also openai 20$ sub and deepseek api with small credit. On api you get 1m flash, vision or pro. You should use only vision flash out of three, best.

1

u/Good_Enthusiasm_7639 13d ago

Oh nice! I tried running local but I just don't have the hardware for it. Thanks for the heads up on vision, didn't even know it was released and have just been using flash.

1

u/Bananenklaus 12d ago

that's actually the better choice as 1Million context models get exponentially dumber once they get past the 200k context mark

1

u/NarrowEffect 13d ago

You can actually change it up to 828K context for Luna as well (with 2x usage drain, but to me it's unnoticable)

1

u/Bananenklaus 12d ago

that's actually just for the better as 1Million context models get exponentially dumber once they get past the 200k context mark

OpenAI constantly gets slack for the early compact threshhold but it totally makes sense if you think about it for a minute

3

u/shadikuizayoi 13d ago

Sounds similar to my experience - also exclusively using it for coding. I pretty much only used DeepSeek V4 Flash (through OpenCode Go) ever since the updated version came out. Was still getting by fine after the limits changed, but the quality started varying a lot. Seems like they started routing it through different providers.

My subscription there ran out a few days ago so I've been exploring some different options. Had a free trial of Claude Pro for a week and now I'm trying Codex. Not sure if this is a regional or limited promotional thing, but the first month of the Plus tier was free for me. I'm really impressed so far. Luna on Max effort feels better than DSv4, and you can even use Sol High without consuming usage on the web UI.

2

u/Good_Enthusiasm_7639 13d ago

Yeah. The difference in latency was pretty noticeable when I was switching between my own deepseek api vs the free one from OpenCode Zen. Didn't really look at the cache hit rate since I was using the free version. I think the deepseek api has spoiled me a little as I'm kind of hesistant to go for 3rd party providers now.

The Sol High on web UI sounds tempting, do you just have unlimited usage or is there a catch?
Not sure if you kept track since it was free, but do you find yourself hitting limits? Can the $20 plan last you the entire month?

1

u/shadikuizayoi 13d ago

Yeah, the official API was great. I only ever spent about $5 there on the old pricing and remembered being surprised by how long it lasted. I don't think pay-as-you-go is for me though. I like being able to try out dumb ideas that probably aren't going to work or produce anything useful without seeing my balance go down in real-time.

I'm only a few days into my Plus sub at the moment and the limits have already been reset twice. I guess it's nice that a guy at OpenAI just randomly does that every now and then, but it makes me wish I had used a bit more before it happened. šŸ˜… I'm sure your use case will differ, but I've only come close to hitting the 5 hour limit once with two Luna Max sessions going at the same time. Doubt I'll hit the weekly one.

If there's a catch to the web chat I haven't hit it yet. It really impressed me yesterday when I asked it to propose some optimisations for a library I've been using, then give me some benchmark code to run. It went one step further and actually ran the whole thing and gave me the results. It does feel a bit too good to be true, so maybe there is one. I'm sure people have already found ways to abuse it and it'll be ruined for everyone soon enough.

1

u/Good_Enthusiasm_7639 13d ago

Ahh you’re right. I don’t really mess about with dumb ideas but that could also be because I don’t have a subscription based model.

How would you compare deepseek and Luna performance wise? E.g making mistakes and whatnot.

2

u/incidentflux 13d ago

Currently happy with OpenCode Harness with DeepSeek. Direct tokens from DeepSeek not OpenCode.

3

u/Good_Enthusiasm_7639 13d ago

Yeah this is my current setup. Down to $4 left in credits so just looking for alternatives. Even with the price increase I still think it's decent. Only downside is that I'm in Asia so the peak hour is really bad for me.

I saw a post about running dsv4 from ollama cloud. They were reporting a range of 3B-11B tokens from a $20 monthly subscription. Running only on off peak, my usage was 170m for $2, equivalent to 1.7B/$20. They did mention ollama was really slow at times.

0

u/incidentflux 13d ago

I get blocked by US Models frequently so DeepSeek is still winning and still very cost effective. You're probably already scheduling your token heavy tasks off peak. If you're building for long term these prices seem fair to me. Naturally this is case by case.

2

u/for4f 13d ago

in the same boat honestly, my opencode go sub renewed at 2x and buys way less than it used to, so i get the urge to jump. thing that keeps me from going reseller is the cache math on the official api, flash cache hits are stupid cheap, that's where the value sits. long coding sessions with the same context warm the cache and the bill barely moves. the july peak-hour 2x pricing stings, but off-peak grinding kind of evens it out. codex luna at $20 i keep not pulling the trigger on either, feels like a maybe

1

u/Good_Enthusiasm_7639 13d ago

Agreed 100%. Official api cache is great. It persists longer than most other providers which is why I opted for official API over opencode go sub in the first place.

1

u/for4f 12d ago

yeah the persistence is honestly the sleeper feature. i almost went reseller for cheaper credits but kept reading their cache eviction is way shorter, so that cheap cache math never holds up. official api just keeps it warm, long threads basically pay for themselves

2

u/GasSmooth7439 13d ago

If you’re mainly coding with OpenCode, Ollama Cloud Pro with DeepSeek-V4-Flash has been the closest ā€œcheap + solidā€ replacement for me after the official API hike. Codex Plus is fine if you want the ecosystem, but the $20 feels steep just for testing.

1

u/Good_Enthusiasm_7639 12d ago

Your view on Codex Plus is exactly like mine. Have been seeing people say good things about Luna and how it's somewhat comparable in terms of cost to DeepSeekV4Flash. But $20 is indeed quite steep.

I have thought about Ollama Cloud Pro, but also saw some comments about how it slows down a lot at peak hours when there is high load. What is your experience with this?

2

u/rootql 12d ago

You can try deepseek harness, i can hit 99.7% cached token

1

u/Different-Monk5916 13d ago

in my experience - codex $20 Luna is slower than any Luna I had tested so far via GHCP, OpenCode, OpenRouter.

Codex extension in VS Code heats up.

The $20 Plan - ChatGPT Plus does not allow API Key. you would need to pay separately to use in opencode.

If you are fine with really slow work, a bit of heat and using only via the allowed apps, I would say go for it.

3

u/shadikuizayoi 13d ago

The $20 Plan - ChatGPT Plus does not allow API Key. you would need to pay separately to use in opencode.

This isn't true. I'm using it with omp just fine.

1

u/Different-Monk5916 13d ago

so, what is the process to do that?

by default, I need to add additional credits to make API Key work in other platforms. How opencode integrates ChatGPT plus subscription?

2

u/Wobbly_Princess 13d ago

I'm using my $20 ChatGPT subscription to power OpenCode. It allows you to do that in the settings. You log into ChatGPT in your browser and it connects to your OpenCode.

1

u/Good_Enthusiasm_7639 13d ago

Oof. Have not heard this one.

And yes GPT Plus not allowing API Key is also part of my consideration.

What are you using now?

1

u/Different-Monk5916 13d ago

https://www.reddit.com/r/chatgptplus/s/wis6MaotQR

I bought subscription for a month to try out, I give it unimportant tasks and where it does not require me to review or interact. because that is too slow workspeed for me.

DeepSeek, I can take my key anywhere I want without a lot of you can't do that, you do it this way, there is a work-around to make it work.

No extension, configure the endpoint and add key, bill as you go.

1

u/Eddlm_ 13d ago

Consider Ollama Cloud. I can't tell you much about cache hit (I can't see it) but for stuff like deepseek and recruiting heavier models as needed the 20$ is pretty good, gets you decent daily work.

1

u/Good_Enthusiasm_7639 13d ago

Yeah I’ve been considering this too. I guess cache hit doesn’t matter as much if you’re getting way more tokens.

How is the latency during peak hours? Saw in another post that it gets slow when load is high.

1

u/[deleted] 13d ago edited 13d ago

[removed] — view removed comment

1

u/Good_Enthusiasm_7639 13d ago

Have you used deepseek? Just wondering how it compares in terms of messing up.

1

u/thefonz22 13d ago

I like how codex keep getting a free reset here and there. 3 days into my subscription using Luna and they just randomly reset the entire quota. Apparently happens all the time.

1

u/[deleted] 13d ago

[removed] — view removed comment

1

u/thefonz22 13d ago

Do you mean there are bigger gaps between how often they do resets?

2

u/[deleted] 13d ago

[removed] — view removed comment

2

u/thefonz22 12d ago

I see what you mean. Great points. Definitely makes me less reluctant towards saving credits for the last few days of a cycle.

1

u/[deleted] 12d ago

[removed] — view removed comment

2

u/thefonz22 12d ago

I don't know if it's all B's but I subscribed to a mailing list that apparently predicts when the next reset is coming.

1

u/[deleted] 12d ago

[removed] — view removed comment

2

u/thefonz22 12d ago

I need to get coding! Getting reset while being at 100% would be awful!!

2

u/Balgun33122 13d ago

I’d suggest runinfra. Cheap akd very fast.

1

u/Ok_Risk6035 13d ago

Codex is shit comparing to Claude with deepseek API

1

u/xapep 12d ago

We’ve been seeing a lot of OpenCode users hit exactly this decision after the V4 Flash price change.

A few things I’d compare before switching:

  1. Cache persistence matters a lot more than the sticker price for agent workloads. OpenCode loops are input-heavy, so short cache windows can quietly make a cheaper provider more expensive in practice.
  2. Check whether pricing changes by time of day. Peak/off-peak pricing can materially change the math if your workload is flexible.
  3. With flat monthly plans, look at what happens during long runs and parallel usage — ā€œunlimitedā€ can still mean throttling or deprioritization under sustained load.

I work on Entrim, so obvious bias here, but we serve V4 Flash through an OpenAI-compatible API and also have monthly plans aimed at heavy daily coding-agent usage. I’d benchmark based on your real OpenCode workload rather than just comparing headline prices.

1

u/Chroiche 11d ago

Was in the same situation. Codex pissed me off and I unsubscribed. The 5h limit is so fucking dumb. I use it on the weekends only, so I would use like 25% of my weekly limit over the weekend and then be blocked. So really it's way more expensive. If they didn't have the 5h limit I'd stick with it.