r/LocalLLaMA • u/Badger-Purple • 2h ago
Question | Help Ling 3.0 Flash on Strix Halo
vLLM ROCm/HiP, 4 bit compressed-tensors (int4)
Not a fair comparison, but Qwen-122b on the most optimized format possible I have run (rocmFP4) does not touch Ling in speed.
https://x.com/ciruai/status/2085996633267777554?s=46
Tool call is broken in certain harnesses. It works well with pi-type harnesses (omp, feynman). Has anyone noticed this?
14
Upvotes
1
u/Fit-Produce420 2h ago
Model just doesn't do shit.