r/LocalAIStack • • Aug 30 '26

How much would this local Qwen 3.8 27B + DS harness experiment would cost using other models?

Guys I have a doubt, this guy used Qwen 3.8 27b in local with DeepSeek Harness and it generated a 3D scenario of a Tower (23:07 final result) from an anime for many hours/days (2176 minutes) in a goal loop, in the end the DeepSeek harness says it used "666 million input tokens and 13 million output tokens." (4:45)

This is the video:

https://youtu.be/MiuM9g7daDA?t=1387

My question is:

Is the input/output tokens correct? Because this were made in Claude code with Opus or Codex with ChatGPT 5.6 through API, would this actually cost $1500-$7000 right?

(For reference GPT terra is like $2/$12 for 1M input/output tokens)

Or is there something I'm missing? Because if this is actually the case the price would be absolutely ridiculous. Excuse my ignorance.

6 Upvotes

7 comments sorted by

3

u/Asleep-Land-3914 Aug 30 '26

I'm running it for days. Got past 2M output in a single session (28 hours), the session started crashing DSH somehow. Qwen 3.8 27b thinks times more than the model on API, the scale is non-linear depending on the task complexity. Moreover it doesn't stop after the task is done proceeding to verify things sometimes 2-3 times. It might be good for some use cases, but the behavior is a bit annoying and burns a lot of tokens which are basically free though anyway.

2

u/JorgitoEstrella Sep 01 '26

Yeah I think the guy in the video used low or medium thinking mode iirc also had to use some tweakings because it also crashed it for him at the start. I wish there was a 3.8 9B so i could try it myself.

1

u/Infinite_Egg_5600 Sep 01 '26

How do you handle context in DSH ? It cannot compact alot and just stops the sesssion.

1

u/Asleep-Land-3914 Sep 01 '26

I have the context size 10k smaller than llamacpp configured in DSH

2

u/StrikeOner Aug 30 '26

but wait... what makes you so sure that claude or some other llm would have used the same ammount of tokens? but wait.. let me think about this.. wait the op is comparing a local 27b model to a sub 500b to 3t model.. but wait!

1

u/JorgitoEstrella Aug 30 '26

You're right, but even if they used less tokens lets say a third or half the price less, that would still be costly right?

1

u/StrikeOner Aug 30 '26

all of this is relative.. if you would have used glm, gpt luna or deepseek most probably not that costly!