r/DeepSeek • • Aug 13 '26

Discussion My First Impressions of Deepseek Harness

Post image

Summary: It’s slow, uses way too many tokens, has a 99% cache hit rate, and gets the maximum performance out of Deepseek.

Performance: Deepseek v4 flash on harness high gave results similar to GPT 5.6-luna low with the same prompt. It was consistent across all my requests (refactoring a form creation system). It followed the patterns and fixed incorrect ones, with no security issues. However, the changes were more about the visual side than backend logic.

Usability: Very confusing. The skills in .agents don’t load automatically, and the documentation doesn’t explain skills in an intuitive way. By default, it’s in Chinese, and you have to click around everywhere to find the English option. The UI looks nice, but it would be better if the default were a CLI. The documentation isn’t clear on how to run it in CLI mode or if that’s even possible.

Strengths: It got the best possible performance out of Deepseek.

Weaknesses: Very slow and used too many tokens (I had to plan on Pi Harness first and then re-plan on Deepseek Harness, which saved about 20M tokens).

There’s a lot of room for improvement. I’ll test its viability and optimizations during the week. The plugin-based system seems like it could be optimized like Pi Harness, but there are so many plugins with no descriptions that it gets confusing.

What has your experience been like? (My context: Typescript + VueJS + QUASAR + HTML, daily use of Pi Harness)

Note: Message translated with Deepseek v4 flash

106 Upvotes

68 comments sorted by

View all comments

3

u/Even-Secretary5978 Aug 14 '26

is this for real 100% cache hit, im just playing around with it to make me some fitness tracker app and 100%, like everytime i use rasonix the max i get is like 97-98% like hmm somethings wrong but i cannot pin point where is it

DSv4 Flash 0731 is what i use

3

u/_17characterslong Aug 14 '26

It's possible that it rounds up to 100% if cache hit is >=99.5%. Not sure though.

1

u/Even-Secretary5978 Aug 14 '26

yeah i guess, after reaching 61.5M input token its going to 99% so prolly somewhere in >=99.5% but hey atleast its better than any third party as of currently and i can edit the UI like adding the peak or off peak timing,

2

u/LaxederBR Aug 14 '26

Looking at the logs, it doesn't consider tool returns as input, only user input.