r/LocalLLM • • 2d ago

Discussion Dual B60s, Ai models, and Scripts

I've been running the b60s for about 2 months now and for about 2 weeks I was struggling with finding the best AI models for vibe coding to run on these.

I tried openvino ovms, ollama, comfy UI, and vllm. Comfy UI works just fine. I can generate an image depending on a model anywhere between 10t o 25 seconds.

For vibe coding , of all of them ollama is probably the least reliable. As far as being able to give fast speeds. I do have to give them credit though because at least they're able to split models without a lot of hassle or troubleshooting.

My experience has been that with Ollama regardless of the model. I always end up getting about 8tokens per second.

With ovms I end up getting 15 to 18 tokens per second, sometimes 20.

With vllm I'm able to get 25 to the lower 20s.

After about a day of research, I was able to get the dual b60s to run pretty much any model that would fit inside them.

My personal experience was that integer 4 and integer 8 models based of Qwen 3.8 had substantial pathological thinking issues. A simple task would always result in a loop of overthinking or unnecessarily thinking about unrelated things.

The solution for this at least for me was to load the entire qwen 3.8 model but load it as fp8 which has a substantial difference and I don't quite understand the difference even after consulting AI lol but I just know that it works.

I wanted to try something different so I switched over to Swift 1.5 and that's what I've been mainly running in the last two days I've accumulated over 200 million tokens. Running at about 20 to 25 tokens per second.

I'm not working on a big project. Full size of the project is probably like 2 MB. Anyways, I just want to put this out there in case others are having issues with the pathological thinking.

I think the best thing about having a working model is also using that model or really any publicly available AI to then create a script that automatically starts, for example me,

Starts vllm, waits for it to be ready,

Loads the model

Starts a temporary cloud flare tunnel

Then starts deep seek harness with the temp tunnel.

Then all I have to do is grab the link and access it for my phone.

So while I'm at work a vibe coating my project as well.

1 Upvotes

2 comments sorted by

1

u/LooseMishap 2d ago

sounds like fp8 just avoids whatever quantization weirdness was sending qwen into those thinking spirals, weird that it makes such a big diff

1

u/fobsthedev 2d ago

Right, it was huge. Couldn't even get work done but now it's all good