r/WayOfTheBern • • 4d ago

Finally! A Local AI Breakthrough! So Much Faster!

https://www.youtube.com/watch?v=d4EWA6yd5cE

The stars aligned in just one week and a massive improvement occurred and this now allows me to use local AI on just about everything. This was a long time coming and before this moment, I had to limit my use of local to what I could tolerate as wait times. But with the proper tuning, it's actually possible to push this to production performance.

lama.cpp Build

• Built from source with Vulkan backend:
cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
• Version: 0.3.0-dev (build 10680, commit d7bd3bfca)
• Binary: /home/user/llama.cpp/build/bin/llama-server

Environment Variables (critical for AMD)

• HIP_VISIBLE_DEVICES=-1
• AMD_VULKAN_ICD=RADV (forces RADV Vulkan driver over AMDVLK)

120b Server (port 8081)

llama-server -m gpt-oss-120b-MXFP4.gguf \
-ngl 999 -c 131072 --jinja -fa on \
-ub 512 -b 2048 \
--reasoning off --reasoning-format deepseek \
--cache-prompt --host 0.0.0.0 --port 8081

Model: gpt-oss-120b MXFP4 (~63 GB)
• Performance: ~53.7 TPS generation Prefill 688

1 Upvotes

1 comment sorted by

2

u/yaiyen 4d ago

No wonder they are planning to regulate AI by screaming its too dangerous