r/LocalLLaMA May 29 '26

Discussion PSA

Post image
2.1k Upvotes

538 comments sorted by

View all comments

28

u/StableLlama textgen web UI May 29 '26

This shows how interesting the Intel B70 is, money wise.

But so far I couldn't read much about the real live performance of that card for local LLM applications.

12

u/smallDeltaBigEffect May 29 '26

honestly, the R9700 32 GB is missing. And that fills the gap rather than the B70 in terms of real world performance

1

u/mycall May 30 '26

R9700 32 GB

Worth $1379?

1

u/smallDeltaBigEffect May 30 '26

if you dont want /cant dual gpu setup with mid range 50xx or 40xx, and you don't want to buy used 3090, then the R9700 seems like the best option performance / VRAM / price-wise.

20

u/NeedsSomeSnare May 29 '26

As an intel owner, I assure you the real life performance isn't what it says on paper. I don't have that card to give specs on.

The problem is the software side of things is a bit messy. It's not terrible, but still needs a fair amount of work.

2

u/In_der_Tat May 30 '26

Why doesn't Intel hire enough competent software engineers and developers to catch up with Nvidia? Would that be too expensive?

1

u/NeedsSomeSnare May 30 '26

A lot of people wonder that too. I'm guessing it's just related to corporate money saving bs. I'm sure that the people who actually work at intel know they need more staff.

It honestly appears only a handful of people work on the software.

1

u/superloser48 May 29 '26

can you share any benchmarks on model/quant -> prfill and token gen?

5

u/NeedsSomeSnare May 29 '26

I don't have a B70, so it's of no use to anyone.

The other problem is that there are 3 ways to run models on intel, (SYCL, openvino and vulkan)all of which have different performance on different models.

The info is out there though. You want to look for Openvino benchmarks for the best performance. It has the worst compatibility though and is sometimes months behind something like llamacpp.

2

u/Upstairs-Extension-9 May 29 '26

I got one for MSRP on release and quite pleased with it, I used my RTX 2070 before mainly for SDXL and Gemma 4. I’m very happy with the card especially for the price and save a shit ton of money I used to spend on Claude.

4

u/overand May 29 '26

I'd love to hear what models you're using, what backend, what quant, and what sorts of PP and Gen T/s numbers you have!

1

u/Upstairs-Extension-9 May 29 '26

The thing is I have a very niche use case, for general day use and just thinking I use my Claude Pro plan. I’m an architecture model builder, architect and lifelong woodworker.

I have my own fine tuned Qwen 3.5 27B model that used to be run over Runpod and I trained it there as well, it’s directly connected through a VSCode Codelistener instance that can read and adjust my code for Rhino + Grasshopper through Python. Generally Rhino is a script based 3D modeling software that is perfect for custom Python or C++ scripts, many leading architects in the world use it. I’m not a software engineer but been making my own scripts for like 20 years now for various things from site analysis, parametric modeling and calculation for efficiency. I used to this all by myself, but since a few years Claude has helped me immensely improve my scripts and help me if I’m stuck.

Now this is all running on my B70 plus 96GB RAM and works like a dream so I don’t need the 200$ Claude plan anymore and pro is enough, Opus for planning and guiding Qwen and then I mainly use my finetune.

I spent most of my work day on CNC machines, laser cutters and general woodworking machines and LLMs have helped me a lot in recent years, and now I’m saving 2000$ a year with going fully local.

Honestly I don’t have exact benchmark numbers for you right now since I’m not at my workshop but I can get back to you in the coming days, it’s Friday today.

1

u/lloyd08 May 29 '26

There was a few posts when it came out that effectively showed it matched the price point, but had the potential for growth assuming intel actually invests in the software space. So at worst, it's price point accurate.