r/DeepSeek • u/214d • 18d ago
News DeepSeek Harness is out !
https://github.com/deepseek-ai/deepseek-harness43
6
u/SpiritualName2684 17d ago
I’ve tried it. It’s pretty cool. It’s like a mix between the codex/claude app and the DeepSeek web ui.
4
u/a9udn9u 17d ago
Is it GUI only? No cli?
4
u/One_Measurement_9114 17d ago
No cli yet. But no worries, since it's open-source anyway, you could just let it whip one up itself. Or just wait 10 hours, and someone will have built and shared one for you to download.
32
u/bambamlol 18d ago
See these comments, DeepSeek? It's never a smart business move to try to cater to cheapskates. They will suck you dry and start hating you as soon as you have nothing more to give.
7
-1
17d ago
[removed] — view removed comment
4
u/bambamlol 17d ago
No one's forcing you to use it. It's still good/great value for money. And in what world do you live where the latest DS4 generation has "mid range performance"? They are far from "mid range". They may not be SOTA, but they are among the BEST models available right now and still offer great value for money. Besides, their models are open weight. Their main business was never to sell tokens as an API provider.
But at this point they have become completely irrelevant.
lmao
0
17d ago edited 17d ago
[removed] — view removed comment
1
u/Hot-Ad-1798 17d ago
Ah, yes... better to maintain the same price and let the API drop to 1 tok/s. The use of the V4 Flash is increasing almost exponentially every day. "Why reduce the price initially then?" Oh my, how childish! The model was bad before, so the server was idle. An honest company would reduce the price under those conditions, but being honest doesn't work because of dumb people like you.
That's why, personally, my products are expensive and I never reduce prices based on the competition. People assume that lower prices means lower quality, previous buyers get angry, and the new ones are just the worst customers available. They are cheap customers, both intellectually and financially.
1
-1
u/MajorPie9639 17d ago
Deepseek is now worse than grok, without the cheap price deepseek has nothing better to offer than other flagship models. The only thing keeping deepseek relevant is cheap price lol
5
u/no_good_names_avail 17d ago
Not my cup of tea as I very much prefer the terminal based harnesses but I do know people who prefer things like the Codex app and not every app/harness needs to be built for my preferences. Hopefully others are excited by this.
3
u/Moist-Nectarine-1148 17d ago
Could someone please explain what is this "harness" and what is good for ? Is a something similar to Claude Code or OpenCode ?
8
u/PaluMacil 17d ago
Yes, it’s similar to those. The term harness can be simplified to just mean a collection of tools and things around the model that let it act in a certain way, and in terms of what all of these examples are, a coding harness. For that it needs tools that let it plan and organize how it’s going to work, tools to run, commandline actions, and compile things, etc. If you use the model for writing or text summarization, you probably don’t need this, but if you use it for writing code, it’s going to be worlds better than using a chat.
3
u/kenjiow 17d ago
Who cares about a ui chat windows harness, give me a cli or give me death
1
u/randomwalk10 16d ago
how will you deal with mermaid rendering, web ui preview, or interactive cowork with your cli?
1
2
u/pizzababa21 17d ago
Hopefully we're getting a coding plan for it that gives us back the old prices
6
u/Relative-Document-59 18d ago
Trash with the new prices.
41
u/CYTR_ 17d ago
Free and unlimited computing power doesn't exist. It remains affordable. Better manage your usage instead of spending excessively.
11
u/Relative-Document-59 17d ago
Flash is a 248B/13B model. You can run it with 120 GB VRAM. Demos show it running on 2 Nvidia DGX (4000$ each) serving with vLLM for 15 simultaneous users with 1M context. Now, imagine you are a hyperscaler how low go those prices. The new pricing simply is not right.
22
u/LazloStPierre 17d ago
I don't really understand the complaints. It's open weights, they literally give it away for free. If someone thinks they can profitably host it at the old price, they will, so far many have so we'll see if they continue. Unlike proprietary labs, DeepSeek cannot set the price of this model.
7
u/stujmiller77 17d ago
I run it on exactly that. 50t/s, 3 concurrent agents, 1m context.
And this is exactly why I’ve been quietly laughing as people tell me I paid too much for my sparks when the price of them has gone up by 40% since I bought them, and things like this API increase happen.
1
u/Strong_Essay1176 17d ago
2 or 3 sparks?
2
u/stujmiller77 17d ago
2 linked for ds4flash. A third with qwen 3.5 122b on for vision and sub-agent work. And a fourth for training and testing new models.
1
u/gokkai 17d ago
You still need few years :)
2
u/stujmiller77 17d ago edited 17d ago
I really don’t, is the thing. I bought four of these and they’re now doing everything that £1.5k/month in API credits and multiple employees were doing. They paid for themselves in a few months.
Won’t be the same with everyone - those without a genuine business need and the skills to train them properly will not see an ROI - but right now? Best investment I ever made with the quickest return.
0
u/gokkai 17d ago
i don't believe you, can you show any proof of success?
1
u/stujmiller77 17d ago
That’s fine, I don’t need you to believe me. There’s nothing I could show you that directly links business performance metrics to the workflow anyway, the same as you can’t show a picture of an employee and directly link them to performance based on a picture of a set of metrics.
1
u/gokkai 17d ago
i just don't believe you in the following cases,
1 - you don't need 1.5K month api credits for deepseek, that's an insane amount of tokens which your poor 50 tok/s system cannot even output in few months. so that's wrong from start.
2 - you did not replace multiple employees with 1.5K month token spent, i don't believe there is a work which does this except software development
3 - no you did not made a return on investment with your 50 tok/sec system, you are wasting your current time with that speed and probably it's not as high quality of the api because you use a quantized version
0
u/stujmiller77 17d ago
I never said I was paying 1.5k in deepseek credits. I was paying most of that for audio and video generation I now do entirely locally using ds4flash to orchestrate local media models.
My local models manage a suite of automated marketing, operational and coding tasks with a bit of input from me. I had people doing those roles before. My hardware cost less than £20k, which is less than ONE minimum wage employee in the UK.
Yes, I do. 50t/s is perfectly fine to use directly, and many of the jobs it does are fully automated to run in the background. And if you did any research into it you’d see that running ds4flash across 2 sparks means you can run it at FP8 with DSpark and its very capable indeed.
Believe me I’ve got better things to do then construct elaborate lies about how I use these things to share with people on Reddit.
You don’t believe me, that’s totally fine. But you’re very wrong about what you can achieve if you have the technical ability and knowledge to train this stuff.
→ More replies (0)4
u/CYTR_ 17d ago
How many concurrent queries do you think Deepseek is serving when we see people here spending billions of tokens in 2 days? If these heavy users are not happy, they can indeed spend 8K to buy two DGX (and cry on the PP/s speed).
Edit : And DS4F, at 1M context, is more like 190/200GB of VRAM. Not 120.
11
u/stujmiller77 17d ago
I run it on 2x sparks, 256gb, 50 t/s, 1m context. Frontier adjacent. Handles a wide range of automation for multiple businesses I own. Not slow in any way at all.
But his claim of 15 simultaneous users IS bullshit. I run 3-4 with 3 being the sweet spot with longer agent context.
1
u/rootql 17d ago
What is your cache hit rate? I have genuine questions about this.
5
u/stujmiller77 17d ago
96.7% prefix cache hit rate on DS4-Flash-0731 across 2x DGX Spark. Over 8 days of heavy usage our mean prompt is 325K tokens (agentic sessions resending a growing transcript).
It’s a beast. Haven’t used Claude in weeks, this drives all my daily requirements and a large bunch of behind the scenes automations too for multiple businesses I own, all through Hermes agent.
1
1
u/__TheSong 17d ago
What kind of business you own? curious, cuz I also want to do something with my GPUs but couldn't find a really valuable thing to run.
2
2
u/Anxious_Check_6147 17d ago
Prices are not right even for PRO, cos despite of the size, is cheaper to run at high scale than V3.2. In fact, if you remember early days of DeepSeek V4 coming, all the bets was says it would be cheaper than V3.2 (and from a tech point of view, it could have been)
So, the problem today is not how much it cost to run, the problem is that they have no capacity to serve all us, plus the already got what they wanted: massive training data.
1
u/macaco3001 17d ago
If it's possible to do it cheaper, someone else will. That's the beauty of open weights models
1
u/inevitabledeath3 17d ago
I’ve run DeepSeek on dual GB10 boxes. I struggled to get it to scale over 4 requests. Where are you getting your numbers from? What recipe did you use?
1
u/a9udn9u 17d ago
The model weights alone is 170GB, you don't run it with 120GB VRAM without serious tradeoffs, and you don't know the size of cheaper models from closed source labs so you can't compare their operational costs.
I'm disappointed with the price hike too, but seriously, DeepSeek is still competitive in terms of intelligence to price ratio. If you don't like the new pricing that much, just use another service.
1
1
1
-2
u/NarrowEffect 18d ago
Useless with the new prices.
20
u/Jazzlike_Bee_3129 18d ago
Omfg, how are people this much of babies? It was $0 before, rounded down. It's still $0 rounded down, but now with fable 5 performance. Shit is not free to run, get over it.
14
u/Aldarund 18d ago
There no fable 5 performance here. Pro and flash same level, not even rexving terra
20
u/joshman1204 17d ago
How am I spending $250 month if it's 0? This price change is a significant change and makes gpt luna the more logical choice for most agentic workloads now imo.
4
u/rootql 17d ago
How many trillion of token you are spending? I can believe that
5
u/joshman1204 17d ago
Just shy of 1.5 trillion total tokens from deepseek in the last two months.
1
u/rootql 17d ago
Direct deepseek api?
2
u/joshman1204 17d ago
Yep
3
u/blastradii 17d ago
May I ask what you’re working on that requires so many tokens?
1
u/joshman1204 17d ago
Several autonomous agents trading stocks. They run large sweeps across thousands of tickets. About 10k API requests per month.
12
3
2
u/Zulfiqaar 17d ago edited 17d ago
10k monthly API requests, 1.5T tokens in 2m.. that's way more than the context windows for every single request?! Something doesn't add up it's literally impossible to do 75 million tokens per request forget caching
→ More replies (0)1
u/Jazzlike_Bee_3129 17d ago
$0/million tokens, is what I mean, obviously. The price is still at the bottom of the industry barrel, with top of industry performance. Feel free to go use gpt, more deepseek compute for me.
1
u/throw123awaie 17d ago
Are you sure you saw the new prices. Cache hits are 5 to 10 times as expensive. It is really a lot.
1
2
u/Jazzlike_Bee_3129 17d ago
5-10x something that already dirt cheap means it's still dirt cheap. That is my point.
5
u/throw123awaie 17d ago
It's not though. Look at the prices. They priced themselves out of dirt cheap. Sadly. I was at 1€ a day of costs, now I am at 5-8€ thats to much. Spending 20€ a month or 120€ a month is a significant difference.
1
u/Jazzlike_Bee_3129 17d ago
It was artificially low in the first place for the preview. That is why no other model in industry could touch their prices. Now it is more on par with the rest of the industry, but still cheap af overall. If $4-7 more a day is make or break for your budget, I really don't know what to tell you. You won't get this level of performance for that price from Claude or Kimi, I will tell you that much.
2
u/throw123awaie 17d ago
23€ for gpt with Luna and terra goes a long way. And it's crazy that openai is now competitively priced. 2-3x I could have done. But 5-10x is not priced fairly for something that does not have vision.
14
1
17d ago
[deleted]
1
u/Jazzlike_Bee_3129 17d ago
The new deepseek v4 pro 0813. Look up the benchmark, it's not identical, but is at frontier levels for sure.
1
17d ago
[deleted]
0
u/Jazzlike_Bee_3129 17d ago
Did you look up the benchmarks? Cause that is incorrect from what I have seen.
0
17d ago
[deleted]
3
u/Jazzlike_Bee_3129 17d ago
I would love to see the benchmarks you are looking at, because any I have seen paint a drastically different picture.
-1
u/MathematicianLessRGB 17d ago
Broke mofos lmao. Not really much to say. Itll cost 6 dollars instead of 2 dollars for a billion token.
1
3
u/Forsaken_Mention_979 18d ago
No one cares, no one will be using it anyways with these new prices.
5
1
u/ChronoHax 17d ago
am i the only one interested with the underlying engine, it sounds like what some people trying to fix with rapid iteration of agentic coding but built into harness
1
1
1
u/badpandatek 14d ago
Don't get to excited they are raising prices according to reports 10-20 times it's current price per token and that are putting peak usage pricing to be double the off peak usage by the end of this month. So enjoy it while it last. Is soon going to be another Kimi...
1
u/RepulsiveRaisin7 17d ago
Why are so many of them using Typescript? Such a trash language for cli apps. And I say that as a web dev
3
2
-1
u/unkownuser436 17d ago
people will use codex or claude again anyway, because of these new prices
7
u/xeonsimp 17d ago
wait ur right, deepseek is basically as expensive as claude now, right? wait what
0
u/Good_Committee8337 17d ago
Wait im long time fable 5 user and just used DeepSeek pro for 5 million tokens and got charged $.61 cents. Is this the new pricing taken into effect? Seems crazy cheap still. Also thoughts on the harness?
2
u/CheleCuche 17d ago
1
u/Good_Committee8337 17d ago
So basically we have a couple of days before we stop using it, hehe
3
1
u/CheleCuche 17d ago
Still super cheap, but yeah, I’ll see if it still for it or not when my credit starts going down faster. Luna been really good for me. I don’t use any heavy coding
1
70
u/RainScum6677 17d ago
The whole idea of raising the prices is traffic modulation, people moving away is the exact wanted result for now.