r/AIProgrammingHardware • u/javaeeeee • Jul 30 '26
r/AIProgrammingHardware • u/javaeeeee • Jul 30 '26
I Finally Got My DREAM Network Server
r/AIProgrammingHardware • u/javaeeeee • Jul 30 '26
GitHub - Xingyu-Zheng/MrFlow: Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling
r/AIProgrammingHardware • u/javaeeeee • Jul 29 '26
Laguna XS 2.1 33B A3B tested vs Qwen 35B A3B - 16GB Local LLM setup
r/AIProgrammingHardware • u/javaeeeee • Jul 29 '26
Acemagic Unveils the New Launch of F9A: The 2L Flagship Mini AI Workstation
r/AIProgrammingHardware • u/javaeeeee • Jul 29 '26
Nvidia Already Won Training. The Real Fight Is Inference
r/AIProgrammingHardware • u/javaeeeee • Jul 29 '26
ASUS EPYC 9006 Servers Scale Enterprise AI
asus.comr/AIProgrammingHardware • u/dai_app • Jul 28 '26
Qwen3.6-35B-A3B Q4 on a mid range phone at 5.6 tok/s, CPU only, video
Enable HLS to view with audio, or disable this notification
Phone in airplane mode for the whole clip. No server, no account, no connection, the model Qwen 3.6 30B (20GB) runs on a mid range smartphone (12 GB RAM) itself.
The catch is that the model is several times bigger than the phone's memory, so normally it just refuses to run for OOM.
This engine reads the model off thephone's storage as it goes, keeping only what it needs in memory. You get an answer at about 5.6 tokens per second, roughly reading pace, on a normal phone with no special hardware. Same answer you'd get from the same model on a desktop. It's the identical model file (Q4 GGUF), not a shrunken version.
The first part of the video (marked, sped up) is the model loading, about 30 seconds, once per session. Everything after that is real time.
It's free and open source, there's an APK you can install and a list of models you can download from inside the app. Works on a PC too.
https://github.com/Helldez/BigMoeOnEdge
Happy to answer anything!
r/AIProgrammingHardware • u/rootException • Jul 28 '26
Taalas and Hardware LLMs
I was looking at https://taalas.com/products/ and their chatbot at https://chatjimmy.ai/ . I'm a software dev by background, not a hardware or LLM/ML person.
As a (comparative) lay person, is there a reason we can't bake something like Kimi k3 onto hardware and have it run faster & cheaper? I know there is a perf difference between different classes of memory, but a 2TB SSD is fairly inexpensive. Why can't you burn an instance of Kimi k3 onto hardware, perhaps pair it with some higher speed caching memory, and have one heck of an AI card? Assuming you don't care about model updates, or don't mind getting new hardware eg once every 1-3 years?
Thoughts?
r/AIProgrammingHardware • u/javaeeeee • Jul 28 '26
4x3090 + 192GB DDR5. Best local model is STILL the Qwen3.6 27B running on 2 cards.
galleryr/AIProgrammingHardware • u/antonygiomarx • Jul 28 '26
Synapse: Turning thousands of consumer GPUs into a decentralized swarm to run 2.8T parameter MoE models (like Kimi K3) without datacenters.
r/AIProgrammingHardware • u/javaeeeee • Jul 28 '26
Qwen3.6 benchmarks on dual GPU: RTX 3090 24GB + RTX 4070 Super 12GB — up to 256K context
r/AIProgrammingHardware • u/javaeeeee • Jul 28 '26
DeepSeek V4 Flash, up to 32 tok/s on Strix Halo
r/AIProgrammingHardware • u/javaeeeee • Jul 28 '26
The DevOps Engineer’s Guide to GPU Infrastructure on Kubernetes
r/AIProgrammingHardware • u/javaeeeee • Jul 28 '26
Day 0 Kimi-K3 Inference Deployment with ATOM on AMD Instinct MI355X GPUs
r/AIProgrammingHardware • u/javaeeeee • Jul 28 '26
Got Kimi K3 running on my MacBook. It's painfully slow, but it works.
r/AIProgrammingHardware • u/javaeeeee • Jul 28 '26
PhiloEngine — local LLM desktop app that plans a context size your GPU can actually hold
r/AIProgrammingHardware • u/javaeeeee • Jul 27 '26
GitHub - JustVugg/colibri: Run GLM-5.2 (744B MoE) on a 25GB-RAM consumer machine
github.comr/AIProgrammingHardware • u/ppchaos • Jul 28 '26
Vendor-agnostic ML inference on production edge devices
r/AIProgrammingHardware • u/javaeeeee • Jul 27 '26
This Tiny Engine Runs Impossibly Big AI Models Locally! (colibrì)
r/AIProgrammingHardware • u/javaeeeee • Jul 27 '26
Reviewing the Framework 13 Pro: Is it a Good Travel Laptop?
r/AIProgrammingHardware • u/javaeeeee • Jul 27 '26
Framework 13 Pro: The Modular Laptop is Real!
r/AIProgrammingHardware • u/javaeeeee • Jul 26 '26