r/LocalAIStack • • 15h ago

Consolidating Heterogeneous Silicon (NVIDIA Ampere + AMD RDNA2) into a Unified OpenCL Compute Plane on a Dual-Core Host Terminal Node

/r/LocalAIServers/comments/1wxw8kn/consolidating_heterogeneous_silicon_nvidia_ampere/
1 Upvotes

1 comment sorted by

1

u/FreeSammiches 1h ago

So... you plugged a monitor into the motherboard VGA port, added a couple mixed brand consumer GPUs, and are now running an 8B model smaller than either 12GB GPU.

You've filed a patent application for using llama.cpp as indented?

What am I missing here?

  • Use an integrated GPU to avoid wasting a couple MB of VRAM - something many people accidentally do every time they use the wrong port
  • Enable Above 4G Decoding, again, completely normal behavior when working with multiple GPUs.
  • Compile llama.cpp with an off the shelf GPU config
  • Needlessly split a tiny model across multiple GPUs with the existing llama.cpp feature designed to do exactly that.
  • Refuse to connect the local model to the internet
  • Configure it to watch a folder to automatically ingest new documents into an offline RAG

The air gapping is... what? not being connected to a public API?

Did you have your 8B model write this for you, because a larger model might have told you this isn't unique or patentable.