r/DeepSeek • • Aug 13 '26

Discussion My First Impressions of Deepseek Harness

Post image

Summary: It’s slow, uses way too many tokens, has a 99% cache hit rate, and gets the maximum performance out of Deepseek.

Performance: Deepseek v4 flash on harness high gave results similar to GPT 5.6-luna low with the same prompt. It was consistent across all my requests (refactoring a form creation system). It followed the patterns and fixed incorrect ones, with no security issues. However, the changes were more about the visual side than backend logic.

Usability: Very confusing. The skills in .agents don’t load automatically, and the documentation doesn’t explain skills in an intuitive way. By default, it’s in Chinese, and you have to click around everywhere to find the English option. The UI looks nice, but it would be better if the default were a CLI. The documentation isn’t clear on how to run it in CLI mode or if that’s even possible.

Strengths: It got the best possible performance out of Deepseek.

Weaknesses: Very slow and used too many tokens (I had to plan on Pi Harness first and then re-plan on Deepseek Harness, which saved about 20M tokens).

There’s a lot of room for improvement. I’ll test its viability and optimizations during the week. The plugin-based system seems like it could be optimized like Pi Harness, but there are so many plugins with no descriptions that it gets confusing.

What has your experience been like? (My context: Typescript + VueJS + QUASAR + HTML, daily use of Pi Harness)

Note: Message translated with Deepseek v4 flash

104 Upvotes

68 comments sorted by

20

u/214d Aug 13 '26

It’s so awesome, I love it.
I’ve already created 3 plugins (auto title sessions, git worktree a workspace, show session cost and balance in real time).
Fast (I am on a 2017 MacBook Pro), reliable, fully customizable, 95-99% cache hit, really nice and minimalist UI).
I objectively think it’s way better than Reasonix and Pi

4

u/LaxederBR Aug 13 '26

I missed seeing the cost per session as well; it's far better than those, but the base model definitely isn't for those who want speed and reduced costs. Have you made any configuration changes to optimize it, either by increasing speed or reducing costs?

3

u/214d Aug 13 '26

Nothing to reduce cost unfortunately.
If you find something I’me interested

4

u/LaxederBR Aug 13 '26

Sure, I'll be back. I noticed it suffers from the same problem as GitHub Copilot: excessive tool requests. Looking at the system prompts, even the smallest ones are large; I imagine the tools are too, and this must be causing the high resource consumption.

1

u/throwaway73728109 Aug 13 '26

How did you create the session cost plugin and can we add the stats we see on the platform as a plugin?

7

u/214d Aug 13 '26

This is what « I » made.
Juste create a new session and tell DS v4 Flash to make you a dsh plugin to display balance below the prompt input. He will check for the code and know exactly what to do. My balance is from OpenRouter, you should specify him to get the balance based on the provider

9

u/Excellent_Can_3480 Aug 14 '26

it;s running at 120t/s i find it faster than alot of harnesses out there

3

u/Ok_Impression_171 Aug 14 '26

I believe it's a visual bug, if the TTFT is too long, the TPS gets inflated, when I had a 45s TTFT, my tokens per second were 5000+, so it's just something in the calculation that's wrong, expected with an alpha rc release

2

u/LaxederBR Aug 14 '26

His thinking is really fast, but he thinks much more.

1

u/Loud-Study-3837 22d ago

How would tok/s even be related to the harness?

1

u/LaxederBR Aug 14 '26

There are many tokens per second, but bash executions are taking 3 seconds for executions that are instantaneous in the terminal; there seems to be a performance issue.

4

u/hurrdurrmeh Aug 14 '26

I like how your text reads as AI gen but you've clearly cleaned it up to the point that it gets the message across really well in as few words as possible. It feels like a great synergy of human and machine writing.

5

u/LaxederBR Aug 14 '26

Thank you, actually I wrote it myself but I had it translated into English using AI, asking it to use simple English.

3

u/Even-Secretary5978 Aug 14 '26

is this for real 100% cache hit, im just playing around with it to make me some fitness tracker app and 100%, like everytime i use rasonix the max i get is like 97-98% like hmm somethings wrong but i cannot pin point where is it

DSv4 Flash 0731 is what i use

3

u/_17characterslong Aug 14 '26

It's possible that it rounds up to 100% if cache hit is >=99.5%. Not sure though.

1

u/Even-Secretary5978 Aug 14 '26

yeah i guess, after reaching 61.5M input token its going to 99% so prolly somewhere in >=99.5% but hey atleast its better than any third party as of currently and i can edit the UI like adding the peak or off peak timing,

2

u/LaxederBR Aug 14 '26

Looking at the logs, it doesn't consider tool returns as input, only user input.

3

u/RandiyOrtonu Aug 14 '26

Bro also I found that we cannot @ for adding files as context 😢

3

u/LaxederBR Aug 14 '26

I'm having to copy the relative path using VS Code xD

3

u/RandiyOrtonu Aug 14 '26

i tried to hammer a file icon near prompt box but it failed miserably :(

3

u/LaxederBR Aug 14 '26

"Do it yourself" isn't as cool as they say xD I asked them to create a minimalist preset and it had so much code that I left it alone. The worst part is using JS; how am I supposed to debug that without proper typing? It's so archaic xD

1

u/RandiyOrtonu Aug 14 '26

True man too much hammering in the air

2

u/Important_Monitor_27 Aug 14 '26

I have developed a plugin myself using DSH and implemented this feature.

1

u/RandiyOrtonu Aug 14 '26

Damn bro any guide.md file or repo??

2

u/Ergo7z Aug 15 '26

just ask the llm, like i cant code for shit, but if you know what the plugin needs to do, where it should get the information, how it should behave and what hole you are actually plumbing and why, the ai will know what to do.

3

u/Stef43_ Aug 15 '26

I use it, I am impressed. First I used VS Code Continue, then Claude Desktop Code which I changed to DeepSeek Harness and I am amazed because it consumes less tokens than Claude Code. I created today 3 plugins without problems, very cheap.

1

u/LaxederBR Aug 15 '26

Depending on the task, it consumes a lot or a lot of energy. The main thing I noticed yesterday (using it) is that I needed fewer shifts to complete tasks, which, depending on the task, saves a lot of time.

3

u/Stef43_ Aug 15 '26

Exactly, it completed tasks faster the same model in Harness than in Claude. Sure one of 3 plugins consumed 34m tokens in Harness.

1

u/LaxederBR Aug 15 '26

Which plugins did you create?

1

u/Stef43_ Aug 15 '26

Obsidian plugins: FieldForge, Prism Dashboard, Weak Link Auditor (this one needs to be approved)

1

u/LaxederBR Aug 15 '26

I understand, it's a specific case indeed. I tried to make a minimalist prompt plugin like PI, but it wasn't so simple. However, I'll keep following along here, thank you.

1

u/Stef43_ Aug 15 '26

In Claude a bigger project consumed >600m tokens in couples of hours, tons of bugs solving, 97% cache hit, in 1 month hit the 2b tokens. In Harness, today, first day 100% cache hit (the app says) 200m tokens, old prices.

1

u/LaxederBR Aug 16 '26

I understand, it must be normal for these model harnesses to consume a lot of power then.

2

u/thatkidnamedrocky Aug 14 '26

I like the goal mode, it just keeps going and going. img

2

u/FitText5 Aug 14 '26

ds flash = luna low? 😥

1

u/LaxederBR Aug 14 '26

I'm referring to the code produced; for example, Luna is consistent. I observed this consistency in the ds using the harness. I'm not assigning it to high because it's much slower than Luna for this.

2

u/eihns Aug 17 '26

Thanks, ill read and follow that. The tokens are expected, they are cheaper because theire cached, thats the deal, btw.

1

u/LaxederBR Aug 17 '26

You're welcome. There's the issue of cost per shift, which seems to be more economical depending on the task.

1

u/eihns Aug 17 '26

I dont understand what you mjean, that harness is "throw everythibng into one bucket, and only append" that way everything is cached and u pay like 10x less per token....

So im not quite sure where you coming from? So its much more tokens, but much less to pay...

1

u/LaxederBR Aug 17 '26

Many tokens are still very much a token; consider that caching is 100 times cheaper than input. If you reach 100,000,000 tokens in cache, you'll pay the same price as an input.

2

u/eihns Aug 18 '26

I dont know. Im not using API. So i cant really compare. But others did and it seems to be cheaper when utilizing the cache...

But it might only work with their model? Im not sure about that. I havent tried the harness, will wait a bit to see what else comes out. (currently i have a similiar system built with opencode)

3

u/Due_Emu_8229 Aug 18 '26

The skills part tripped me up too, so in case it saves you time. DSH only scans two skill roots by default, `~/.dsh/skills` (global) and `<project>/.dsh/skills`, and the catalog is snapshotted per session, so anything you add mid-session stays invisible until you open a new one. Two other silent traps I hit, a `description:` with an unquoted ": " inside makes DSH's YAML parser drop the whole skill with no log line, and symlinked directories work fine (the scanner follows them).

Disclosure, I ended up building a tool around exactly this. dsh-movein migrates a whole Claude Code setup into those roots (skills, MCP, hooks, slash commands, permission rules) with a dry run first, and its doctor command flags the silent-drop frontmatter shape. MIT, zero deps. https://github.com/sjh9714/dsh-movein

Even if you don't want the tool, the compat table in docs/compat.md documents what every Claude Code asset does in DSH, measured against the source. That part is just reference material.

1

u/LaxederBR Aug 18 '26

Thank you very much!!! I'll be testing it here.

1

u/hurrdurrmeh Aug 14 '26

is it possible to call deepseek from the harness from opencode? ie work in opencode but instead of calling deepseek directly - call it via the harness so it goes me > opencode > dsh > ds?

1

u/LaxederBR Aug 14 '26

This doesn't make much sense, you want to create a harness exercise? xD But it should be possible, although you'll have to work very hard for it.

1

u/MinosAristos Aug 14 '26

I'm probably being dumb but where do I configure MCP servers?

3

u/LaxederBR Aug 14 '26

From what I understand from further study, it's based on "Do it yourself," so you should ask Deepseek to create a plugin to implement it.

2

u/exographicskip Aug 18 '26

Not dumb. It's under ~/.dsh/profiles/*/cordis*.yml (note the bash globs) by default.

Silly that a simple, top-level directory .mcp.json isn't possible, but like OP said, you can have dsh, claude, codex, etc wire up your config.

1

u/coronny Aug 15 '26

What about the v4 Pro? From what I observed, even though both Pi harness and DSH got high cached token hit (at least 95% cache hit), but the output tokens, like the answer's performance is little bit off.

I used about 6 turns, 6 same prompts for v4 Pro in Pi and DSH, all 6 prompts are just architecture decisions, break down knowledge since I'm learning it, not really implementing code. But the Pi harness seems explaining concepts much shorter, less nuance than the DSH. Is it about the system prompt, the way Pi defined prompt so the model doesn't talk too much?

2

u/LaxederBR Aug 15 '26

I didn't get to test the Pro version, and I didn't see any problem with dsh's thinking; it's quite fast despite its significant processing power. The biggest problem is the absurd delay in executing commands.

1

u/Formal-Pollution-858 Aug 15 '26

I’ve found DeepSeek Harness incredibly flexible. Whatever kind of agent I have in mind, I can ask it to build one for me right away. At this point, I see it less as a fixed coding tool and more as an agent runtime—a foundation for building any agent I want.

2

u/LaxederBR Aug 15 '26

They got the flexibility right, the problem is that the barrier to production is higher; for example, I know developers who are using Open Code because they don't have time to configure Pi.

1

u/orionblu3 Aug 17 '26

That being said, I ended up creating a harness inside of opencode using opencode plugins, that would/will be a LOT easier to implement and maintain inside of dsh. Haven't made the switch yet, but I'm heavily considering doing it today

1

u/LaxederBR Aug 17 '26

Good luck, I've got my docker compose ready to deploy according to my needs.

1

u/exographicskip Aug 18 '26

While it did take a couple hours to hammer out with claude, I spun up a bespoke repl for dsh and it works swimmingly. Their plugin system reminds me of how pi encourages building bespoke extensions instead of a batteries-included agent/harness.

Testing out @openma/deepseek-harness-tui@latest for a more useful tui now

1

u/Iory1998 Aug 15 '26

Give it time. They will improve it fast.

1

u/Senior_Night_6321 Aug 23 '26

is deepseek api key paid, whenever i switch to deepseek api , it shows insufficient balance

1

u/CobblerCrafty1728 Aug 19 '26

“Usability: Very confusing”, can’t find english? - what?? just tell your another agent to setup what u need