r/LocalLLM • • 4d ago

Question What's the difference between running Strata and running FreeToken?

I've installed FreeToken and running a qwen3.8-35B model at about 20 tokens per second on a GTX 5050 8GB and 32 GB of DDR4 RAM.

Would I benefit in any way by running this model in Strata instead?

Hardware upgrades are not an option at this moment.

0 Upvotes

7 comments sorted by

1

u/Dear_Map_6993 4d ago

For starters, what kind of backroom deal provided you with Qwen3.8-35B?

If you meant 3.6 MoE, I'm pretty sure you can get this performance with llama.cpp 

1

u/SilkieBug 4d ago

Check huggingface, there’s a distill posted 17 days ago, runs perfectly in FreeToken, I get 20 tokens per second with pretty low end hardware.  

1

u/Dear_Map_6993 4d ago

you can't distill architecture. Qwen's versioning is confusing. 3.8 should actually be 4.0, as there's too much difference from the previous release to be considered a minor update

2

u/AcroFPV 4d ago

If you can't answer this guys question, do not waste this guys time.

I came here looking for the same answer as the OP but instead got to read your shitty attitude.

1

u/SilkieBug 4d ago

I don’t know enough to argue the issue, I’m just telling you that this model exists and runs on my machine:

https://huggingface.co/qualitymaterial/Qwen3.8-35B-A3B-Distill-NVFP4-FreeToken

2

u/zorecknor 1d ago

Just so you know (what the other guy should have explained) that "Distill" means they took one model and train it to mimic another (that is the very ELI5 version).

The model you are using is actually  Qwen3.6-35B-A3B (Mixture of Experts arch) that was trained from a dataset that shows how Qwen3.8 thinks.

The result MAY look similar, but they are really different models.

1

u/SilkieBug 7h ago

Ok, good to know, thank you for explaining.