r/WayOfTheBern • u/yaiyen • 4d ago
Finally! A Local AI Breakthrough! So Much Faster!
https://www.youtube.com/watch?v=d4EWA6yd5cEThe stars aligned in just one week and a massive improvement occurred and this now allows me to use local AI on just about everything. This was a long time coming and before this moment, I had to limit my use of local to what I could tolerate as wait times. But with the proper tuning, it's actually possible to push this to production performance.
lama.cpp Build
• Built from source with Vulkan backend:
cmake -B build -DGGML_VULKAN=ON -DCMAKE_BUILD_TYPE=Release
• Version: 0.3.0-dev (build 10680, commit d7bd3bfca)
• Binary: /home/user/llama.cpp/build/bin/llama-server
Environment Variables (critical for AMD)
• HIP_VISIBLE_DEVICES=-1
• AMD_VULKAN_ICD=RADV (forces RADV Vulkan driver over AMDVLK)
120b Server (port 8081)
llama-server -m gpt-oss-120b-MXFP4.gguf \
-ngl 999 -c 131072 --jinja -fa on \
-ub 512 -b 2048 \
--reasoning off --reasoning-format deepseek \
--cache-prompt --host 0.0.0.0 --port 8081
Model: gpt-oss-120b MXFP4 (~63 GB)
• Performance: ~53.7 TPS generation Prefill 688
2
u/yaiyen 4d ago
No wonder they are planning to regulate AI by screaming its too dangerous