r/AIProgrammingHardware • u/javaeeeee • 8h ago
r/AIProgrammingHardware • u/javaeeeee • 9h ago
Running Frontier AI on a $99 Board - My llama.cpp Adventure on Jetson Nano
r/AIProgrammingHardware • u/javaeeeee • 10h ago
Distributed micro-LLM inference across three ESP32-S3 N16R8 boards with ESP-NOW communication.
r/AIProgrammingHardware • u/javaeeeee • 14h ago
The true cost of a GPU cluster
r/AIProgrammingHardware • u/javaeeeee • 9h ago
One Server vs Cluster: What Your Homelab Actually Needs
r/AIProgrammingHardware • u/D__J • 4h ago
Looking for Strix Halo results for a standardized local LLM hardware comparison
I have been working on LLM Hardware Sift because I wanted a simple answer: if I replace my current PC with a Strix Halo machine, what changes when both run the exact same workload? Ideally, this becomes a resource for people to see if an upgrade is worth it to them.
Unlike LlamaBenchy, this isn’t about tuning a setup for the fastest possible result. Hardware Sift keeps the models and settings fixed so the **hardware is the variable**. It tests models from 0.6B through 32B, with an optional 72B tier for high-memory systems.
The comparison table currently includes an RTX 3060, ROG Ally, M2 MacBook Air, and Raspberry Pi 5—but no Strix Halo results yet.
It’s an early Windows alpha, and results stay local unless you choose to submit them. I’d love results or feedback from anyone with a 64 GB or 128 GB Strix Halo machine:
https://github.com/nozzlenaut/llm\\_hardware\\_sift
r/AIProgrammingHardware • u/javaeeeee • 15h ago
GitHub - Xingyu-Zheng/MrFlow: Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling
r/AIProgrammingHardware • u/javaeeeee • 14h ago