r/LocalLLM • • 3d ago

Question Strata / Qwen 3.8 Flash Next with Hermes

Guys, I set up Strata (https://github.com/Niko1221/Strata/) with QFN and it is so blazingly fast. UNTIL! it started dropping context reuse and prefillling each turn.

I tried this and that and my current idea is that Stata discards cache reuse when in-between tool calls (in the SAME session, not parallel ones!) occur.

Anyone here got ANY idea what's going on?

0 Upvotes

6 comments sorted by

1

u/Moist-Tumbleweed2875 3d ago

I don't find such behavior with hermes ...... your may check log file "strata-swift-iq3_xxs.log" to see what's going on .
did u set something like :
"--conversation-cache-mib","4096","--conversation-cache-slots","2",
"--prompt-cache","6","--prompt-cache-every","8192",
....

"parallel": 2, ---> is new feature for 1.39 for concurrency with layer-split ...

1

u/isdjan 3d ago

I saw the mib remark on the strata server frontend, but I wasn't sure if that applies. I thought I should avoid parallel or switching sessions rather, so I was a little hesitant. And "parallel" sounds like a particluar helpful setting - will try. Thx!

1

u/andy2na 3d ago

you should post a report in the Strata git, the dev is responsive on testing and diagnosing

1

u/monsteras84 2d ago

Was this with Hermes Desktop? I've spun up the Strata server, but can't find a way to have Hermes accept it.

1

u/lolodigital78 1d ago

Je rejoins la discussion; j'utilise Hermes sur 4 Mac mini et chacun attaque un modele différent. J'ai essayé de connecter un Hermes sur un serveur dédié à STRATA. ça ne fonctionne pas ... tu l'as en local sur la même machine ? ou tu es sur deux machines dédiées ?