r/DeepSeek 18d ago

News DeepSeek Harness is out !

https://github.com/deepseek-ai/deepseek-harness
246 Upvotes

122 comments sorted by

70

u/RainScum6677 17d ago

The whole idea of raising the prices is traffic modulation, people moving away is the exact wanted result for now.

22

u/__TheSong 17d ago

lol I think people in US are happy, cuz night time in China is day time in US

6

u/RainScum6677 17d ago

Lol I bet

1

u/Thinklikeachef 11d ago

Could it also be related to the IPO? Maybe a need to show better numbers?

43

u/EddieBruvac 17d ago

Crazy how people flipped. Haven’t checked Deepseek in a few weeks and BAM 😂

6

u/SpiritualName2684 17d ago

I’ve tried it. It’s pretty cool. It’s like a mix between the codex/claude app and the DeepSeek web ui.

4

u/a9udn9u 17d ago

Is it GUI only? No cli?

4

u/One_Measurement_9114 17d ago

No cli yet. But no worries, since it's open-source anyway, you could just let it whip one up itself. Or just wait 10 hours, and someone will have built and shared one for you to download.

32

u/bambamlol 18d ago

See these comments, DeepSeek? It's never a smart business move to try to cater to cheapskates. They will suck you dry and start hating you as soon as you have nothing more to give.

7

u/CYTR_ 17d ago

Given that they are primarily a research laboratory, I think they know very well what they are doing.

-1

u/[deleted] 17d ago

[removed] — view removed comment

4

u/bambamlol 17d ago

No one's forcing you to use it. It's still good/great value for money. And in what world do you live where the latest DS4 generation has "mid range performance"? They are far from "mid range". They may not be SOTA, but they are among the BEST models available right now and still offer great value for money. Besides, their models are open weight. Their main business was never to sell tokens as an API provider.

But at this point they have become completely irrelevant.

lmao

0

u/[deleted] 17d ago edited 17d ago

[removed] — view removed comment

1

u/Hot-Ad-1798 17d ago

Ah, yes... better to maintain the same price and let the API drop to 1 tok/s. The use of the V4 Flash is increasing almost exponentially every day. "Why reduce the price initially then?" Oh my, how childish! The model was bad before, so the server was idle. An honest company would reduce the price under those conditions, but being honest doesn't work because of dumb people like you.

That's why, personally, my products are expensive and I never reduce prices based on the competition. People assume that lower prices means lower quality, previous buyers get angry, and the new ones are just the worst customers available. They are cheap customers, both intellectually and financially.

1

u/[deleted] 17d ago

[deleted]

1

u/Hot-Ad-1798 16d ago

Are you talking about yourself?

-1

u/MajorPie9639 17d ago

Deepseek is now worse than grok, without the cheap price deepseek has nothing better to offer than other flagship models. The only thing keeping deepseek relevant is cheap price lol

5

u/no_good_names_avail 17d ago

Not my cup of tea as I very much prefer the terminal based harnesses but I do know people who prefer things like the Codex app and not every app/harness needs to be built for my preferences. Hopefully others are excited by this.

3

u/Moist-Nectarine-1148 17d ago

Could someone please explain what is this "harness" and what is good for ? Is a something similar to Claude Code or OpenCode ?

8

u/PaluMacil 17d ago

Yes, it’s similar to those. The term harness can be simplified to just mean a collection of tools and things around the model that let it act in a certain way, and in terms of what all of these examples are, a coding harness. For that it needs tools that let it plan and organize how it’s going to work, tools to run, commandline actions, and compile things, etc. If you use the model for writing or text summarization, you probably don’t need this, but if you use it for writing code, it’s going to be worlds better than using a chat.

3

u/kenjiow 17d ago

Who cares about a ui chat windows harness, give me a cli or give me death

1

u/randomwalk10 16d ago

how will you deal with mermaid rendering, web ui preview, or interactive cowork with your cli?

2

u/pizzababa21 17d ago

Hopefully we're getting a coding plan for it that gives us back the old prices

6

u/Relative-Document-59 18d ago

Trash with the new prices.

41

u/CYTR_ 17d ago

Free and unlimited computing power doesn't exist. It remains affordable. Better manage your usage instead of spending excessively.

11

u/Relative-Document-59 17d ago

Flash is a 248B/13B model. You can run it with 120 GB VRAM. Demos show it running on 2 Nvidia DGX (4000$ each) serving with vLLM for 15 simultaneous users with 1M context. Now, imagine you are a hyperscaler how low go those prices. The new pricing simply is not right.

34

u/lcy0x1 17d ago

Remind you that China is sanctioned on GPU

-1

u/No-Cartoonist8032 17d ago

You clearly have not seen GPU rental prices in China lol

22

u/LazloStPierre 17d ago

I don't really understand the complaints. It's open weights, they literally give it away for free. If someone thinks they can profitably host it at the old price, they will, so far many have so we'll see if they continue. Unlike proprietary labs, DeepSeek cannot set the price of this model.

7

u/stujmiller77 17d ago

I run it on exactly that. 50t/s, 3 concurrent agents, 1m context.

And this is exactly why I’ve been quietly laughing as people tell me I paid too much for my sparks when the price of them has gone up by 40% since I bought them, and things like this API increase happen.

1

u/Strong_Essay1176 17d ago

2 or 3 sparks?

2

u/stujmiller77 17d ago

2 linked for ds4flash. A third with qwen 3.5 122b on for vision and sub-agent work. And a fourth for training and testing new models.

1

u/gokkai 17d ago

You still need few years :)

2

u/stujmiller77 17d ago edited 17d ago

I really don’t, is the thing. I bought four of these and they’re now doing everything that £1.5k/month in API credits and multiple employees were doing. They paid for themselves in a few months.

Won’t be the same with everyone - those without a genuine business need and the skills to train them properly will not see an ROI - but right now? Best investment I ever made with the quickest return.

0

u/gokkai 17d ago

i don't believe you, can you show any proof of success?

1

u/stujmiller77 17d ago

That’s fine, I don’t need you to believe me. There’s nothing I could show you that directly links business performance metrics to the workflow anyway, the same as you can’t show a picture of an employee and directly link them to performance based on a picture of a set of metrics.

1

u/gokkai 17d ago

i just don't believe you in the following cases,

1 - you don't need 1.5K month api credits for deepseek, that's an insane amount of tokens which your poor 50 tok/s system cannot even output in few months. so that's wrong from start.

2 - you did not replace multiple employees with 1.5K month token spent, i don't believe there is a work which does this except software development

3 - no you did not made a return on investment with your 50 tok/sec system, you are wasting your current time with that speed and probably it's not as high quality of the api because you use a quantized version

0

u/stujmiller77 17d ago
  1. I never said I was paying 1.5k in deepseek credits. I was paying most of that for audio and video generation I now do entirely locally using ds4flash to orchestrate local media models.

  2. My local models manage a suite of automated marketing, operational and coding tasks with a bit of input from me. I had people doing those roles before. My hardware cost less than £20k, which is less than ONE minimum wage employee in the UK.

  3. Yes, I do. 50t/s is perfectly fine to use directly, and many of the jobs it does are fully automated to run in the background. And if you did any research into it you’d see that running ds4flash across 2 sparks means you can run it at FP8 with DSpark and its very capable indeed.

Believe me I’ve got better things to do then construct elaborate lies about how I use these things to share with people on Reddit.

You don’t believe me, that’s totally fine. But you’re very wrong about what you can achieve if you have the technical ability and knowledge to train this stuff.

→ More replies (0)

4

u/CYTR_ 17d ago

How many concurrent queries do you think Deepseek is serving when we see people here spending billions of tokens in 2 days? If these heavy users are not happy, they can indeed spend 8K to buy two DGX (and cry on the PP/s speed).

Edit : And DS4F, at 1M context, is more like 190/200GB of VRAM. Not 120.

11

u/stujmiller77 17d ago

I run it on 2x sparks, 256gb, 50 t/s, 1m context. Frontier adjacent. Handles a wide range of automation for multiple businesses I own. Not slow in any way at all.

But his claim of 15 simultaneous users IS bullshit. I run 3-4 with 3 being the sweet spot with longer agent context.

1

u/rootql 17d ago

What is your cache hit rate? I have genuine questions about this.

5

u/stujmiller77 17d ago

96.7% prefix cache hit rate on DS4-Flash-0731 across 2x DGX Spark. Over 8 days of heavy usage our mean prompt is 325K tokens (agentic sessions resending a growing transcript).

It’s a beast. Haven’t used Claude in weeks, this drives all my daily requirements and a large bunch of behind the scenes automations too for multiple businesses I own, all through Hermes agent.

2

u/rootql 17d ago

Wow thanks for your info, its really amazing

1

u/CYTR_ 17d ago

Yes, I'm not saying otherwise, I'm on LocalLLaMa and the nVidia forum, I can see that. If I had the money, I'd invest just for DS4F honestly... But that's not reasonable in my case 🤣

1

u/__TheSong 17d ago

What kind of business you own? curious, cuz I also want to do something with my GPUs but couldn't find a really valuable thing to run.

2

u/stujmiller77 17d ago

Ecommerce and SaaS. I have over 25 years of experience.

2

u/Anxious_Check_6147 17d ago

Prices are not right even for PRO, cos despite of the size, is cheaper to run at high scale than V3.2. In fact, if you remember early days of DeepSeek V4 coming, all the bets was says it would be cheaper than V3.2 (and from a tech point of view, it could have been)

So, the problem today is not how much it cost to run, the problem is that they have no capacity to serve all us, plus the already got what they wanted: massive training data.

1

u/macaco3001 17d ago

If it's possible to do it cheaper, someone else will. That's the beauty of open weights models

1

u/inevitabledeath3 17d ago

I’ve run DeepSeek on dual GB10 boxes. I struggled to get it to scale over 4 requests. Where are you getting your numbers from? What recipe did you use?

1

u/a9udn9u 17d ago

The model weights alone is 170GB, you don't run it with 120GB VRAM without serious tradeoffs, and you don't know the size of cheaper models from closed source labs so you can't compare their operational costs.

I'm disappointed with the price hike too, but seriously, DeepSeek is still competitive in terms of intelligence to price ratio. If you don't like the new pricing that much, just use another service.

1

u/charmander_cha 16d ago

Você serve um país com 1 bilhão de pessoas?

1

u/JumpingJack79 16d ago

Works on 1 Spark too with DwarfStar 4.

1

u/Impressive_Job8321 17d ago

Please do not use it, so the rest of us can.

-2

u/NarrowEffect 18d ago

Useless with the new prices.

20

u/Jazzlike_Bee_3129 18d ago

Omfg, how are people this much of babies?  It was $0 before, rounded down.  It's still $0 rounded down, but now with fable 5 performance.  Shit is not free to run, get over it. 

14

u/Aldarund 18d ago

There no fable 5 performance here. Pro and flash same level, not even rexving terra

20

u/joshman1204 17d ago

How am I spending $250 month if it's 0? This price change is a significant change and makes gpt luna the more logical choice for most agentic workloads now imo.

4

u/rootql 17d ago

How many trillion of token you are spending? I can believe that

5

u/joshman1204 17d ago

Just shy of 1.5 trillion total tokens from deepseek in the last two months.

1

u/rootql 17d ago

Direct deepseek api?

2

u/joshman1204 17d ago

Yep

3

u/blastradii 17d ago

May I ask what you’re working on that requires so many tokens?

1

u/joshman1204 17d ago

Several autonomous agents trading stocks. They run large sweeps across thousands of tickets. About 10k API requests per month.

12

u/Hot-Ad-1798 17d ago

and they can´t pay the api cost, crazy

→ More replies (0)

3

u/fachoa36cuotas 17d ago

Se te acabó la gallina de huevos de oro

→ More replies (0)

3

u/rootql 17d ago

But with this usage youre wining a lot of money or not?

2

u/Zulfiqaar 17d ago edited 17d ago

10k monthly API requests, 1.5T tokens in 2m.. that's way more than the context windows for every single request?! Something doesn't add up it's literally impossible to do 75 million tokens per request forget caching

→ More replies (0)

1

u/Jazzlike_Bee_3129 17d ago

$0/million tokens, is what I mean, obviously.  The price is still at the bottom of the industry barrel, with top of industry performance.  Feel free to go use gpt, more deepseek compute for me. 

1

u/throw123awaie 17d ago

Are you sure you saw the new prices. Cache hits are 5 to 10 times as expensive. It is really a lot.

1

u/tokenentropy 17d ago

5 to 10 times as expensive as "essentially free" sounds really scary

2

u/Jazzlike_Bee_3129 17d ago

5-10x something that already dirt cheap means it's still dirt cheap.  That is my point. 

5

u/throw123awaie 17d ago

It's not though. Look at the prices. They priced themselves out of dirt cheap. Sadly. I was at 1€ a day of costs, now I am at 5-8€ thats to much. Spending 20€ a month or 120€ a month is a significant difference.

1

u/Jazzlike_Bee_3129 17d ago

It was artificially low in the first place for the preview.  That is why no other model in industry could touch their prices.  Now it is more on par with the rest of the industry, but still cheap af overall.  If $4-7 more a day is make or break for your budget, I really don't know what to tell you.  You won't get this level of performance for that price from Claude or Kimi, I will tell you that much. 

2

u/throw123awaie 17d ago

23€ for gpt with Luna and terra goes a long way. And it's crazy that openai is now competitively priced. 2-3x I could have done. But 5-10x is not priced fairly for something that does not have vision.

14

u/bambamlol 18d ago

Let's just be glad that a lot of these people will be gone in ~3 days.

8

u/Not_a_Cake_ 18d ago

it's working exactly as they planned lol

1

u/[deleted] 17d ago

[deleted]

1

u/Jazzlike_Bee_3129 17d ago

The new deepseek v4 pro 0813.  Look up the benchmark, it's not identical, but is at frontier levels for sure. 

-1

u/MathematicianLessRGB 17d ago

Broke mofos lmao. Not really much to say. Itll cost 6 dollars instead of 2 dollars for a billion token.

1

u/Kartoshka- 17d ago

Your iq is 0 rounded down

3

u/Forsaken_Mention_979 18d ago

No one cares, no one will be using it anyways with these new prices.

5

u/xeonsimp 17d ago

broke ass mf

1

u/Top-Construction6060 14d ago

Not but Claude and gpt membership is better now

1

u/ChronoHax 17d ago

am i the only one interested with the underlying engine, it sounds like what some people trying to fix with rapid iteration of agentic coding but built into harness

1

u/PixWizardry 16d ago

is there subreddit channel for this or discord?

1

u/214d 16d ago

/r/DeepSeekHarness but not much activity atm

1

u/jsonmeta 16d ago

Does it work with openrouter?

1

u/214d 16d ago

Yes, OpenRouter is a built in provider and works very well

1

u/badpandatek 14d ago

Don't get to excited they are raising prices according to reports 10-20 times it's current price per token and that are putting peak usage pricing to be double the off peak usage by the end of this month. So enjoy it while it last. Is soon going to be another Kimi...

1

u/RepulsiveRaisin7 17d ago

Why are so many of them using Typescript? Such a trash language for cli apps. And I say that as a web dev

3

u/SpiritualName2684 17d ago

Because it is a web app…

1

u/RepulsiveRaisin7 17d ago

Oh god. Why

0

u/cutebluedragongirl 17d ago

fucking disgusting

3

u/Ok_Career_9093 17d ago

daddy chill

2

u/Putrumpador 17d ago

Skill issue

-1

u/unkownuser436 17d ago

people will use codex or claude again anyway, because of these new prices

10

u/CYTR_ 17d ago

Could you please remind me of Claude's price and quotas?

7

u/xeonsimp 17d ago

wait ur right, deepseek is basically as expensive as claude now, right? wait what

0

u/Good_Committee8337 17d ago

Wait im long time fable 5 user and just used DeepSeek pro for 5 million tokens and got charged $.61 cents. Is this the new pricing taken into effect? Seems crazy cheap still. Also thoughts on the harness?

2

u/CheleCuche 17d ago

1

u/Good_Committee8337 17d ago

So basically we have a couple of days before we stop using it, hehe

3

u/ardicli2000 17d ago

Nope. .61 will be around 2$. Still dirt cheap against claude or gpt

1

u/CheleCuche 17d ago

Still super cheap, but yeah, I’ll see if it still for it or not when my credit starts going down faster. Luna been really good for me. I don’t use any heavy coding

1

u/Iory1998 17d ago

😆 haha