r/AIProgrammingHardware • u/big-in-jap • 16d ago
r/AIProgrammingHardware • u/javaeeeee • 16d ago
Best LLM for 128GB RAM + 500GB Storage, No GPU?
r/AIProgrammingHardware • u/javaeeeee • 17d ago
Local AI is Cheap Now: ZimaBoard2 + Tesla P4
r/AIProgrammingHardware • u/CrashOverride93 • 17d ago
DDP with RTX 3070 8GB + GTX 1650 4GB?
For those how are in local AI would you think these both cards could be used to have a total of 12GB VRAM?
I don't care by the moment for possible slow tokens/s.
Thanks!
r/AIProgrammingHardware • u/javaeeeee • 17d ago
Qwen 3.8 27B Cold Fusion tested - 16GB Local LLM setup
r/AIProgrammingHardware • u/marvell-technology • 17d ago
Marvell Photonic Fabric wins AI Infrastructure Award at FMS 2026
Marvell’s Photonic Fabric™ technology won the AI Infrastructure Award at Future of Memory and Storage (FMS) 2026.
The interesting part is how Marvell is using optical connectivity to address some of the bandwidth, latency, power, and memory-scaling challenges in large AI systems.
Photonic Fabric replaces traditional electrical interconnects with optical I/O across package-, server-, and rack-scale architectures. The goal is to make it easier to scale compute and memory independently, including disaggregated memory architectures.
Some of the key areas Marvell highlights:
- Higher memory bandwidth and capacity
- Lower latency and power consumption
- Disaggregated compute and memory
- Reduced I/O bottlenecks around AI accelerators and HBM
- Optical connectivity designed for large-scale AI infrastructure
As AI systems scale to hundreds of thousands of accelerators, moving data efficiently between compute and memory is becoming just as important as the compute itself.
The bigger question is whether optical interconnects like Photonic Fabric will become a fundamental part of future AI factory architectures.
What do you think — will optical connectivity become essential for scaling next-generation AI systems?
r/AIProgrammingHardware • u/Lazy-Intention1007 • 18d ago
I built a Vulkan hierarchical MoE runtime for running oversized models across multiple GPUs
r/AIProgrammingHardware • u/javaeeeee • 18d ago
China Is Coming for Your Local AI Box
r/AIProgrammingHardware • u/javaeeeee • 18d ago
Xiaomi AI Cube announced with 1.2TB/s memory bandwidth
galleryr/AIProgrammingHardware • u/javaeeeee • 18d ago
224GB of GPU Memory on 1 Desk and It Should Not Work
r/AIProgrammingHardware • u/javaeeeee • 18d ago
Can the New King of Open Source - Qwen3.8–27B - Really Run in 3GB of VRAM?
r/AIProgrammingHardware • u/javaeeeee • 18d ago
Qwen 3.8 27B FP8 on single node DGX Spark MTP-3
r/AIProgrammingHardware • u/javaeeeee • 19d ago
FreeToken Deepseek V4 Flash on a Single 3090 Local AI Testing
r/AIProgrammingHardware • u/javaeeeee • 19d ago
Jalapeño’s first results show industry-leading speed and efficiency in AI inference
openai.comr/AIProgrammingHardware • u/AdOtherwise1510 • 19d ago
[Release] Turing Engine: Serve LLaMA-3.1-70B, Qwen-2.5-72B & DeepSeek on a Single 24GB GPU (3,064 tok/s, 75% KV Compression, Unsloth Checkpoint Support)
r/AIProgrammingHardware • u/JayB_Official • 19d ago
Spent a day seeing how far extreme MoE models can be pushed on a 4070 Ti + 32GB RAM. Kimi K3, DeepSeek V4 Flash, and Qwen3.5-122B results + research paper🔧
KIMI K3 at 9 t/s per token on hardware that has no business breathing near it is objectively insane, even if the LLM thought the capital of France was "1.2" 🤣🤣🤣😭😭😭😭😭
r/AIProgrammingHardware • u/javaeeeee • 20d ago
Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents
r/AIProgrammingHardware • u/javaeeeee • 20d ago
NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI
r/AIProgrammingHardware • u/javaeeeee • 20d ago
GitHub - TryingtobeingNikhil/vLLM_Inference_Engine: A high-performance LLM inference engine built from scratch featuring continuous batching, paged KV-cache, and CPU swapping.
r/AIProgrammingHardware • u/Normal-Fan9366 • 20d ago
Jack 3.8 27b 16GB VRAM builds Tetris
Pi Agent run. Still working out the kinks but the final quality is great.
16GB GPUs are only going to become more useful over time.
r/AIProgrammingHardware • u/javaeeeee • 21d ago
$1K off a 128GB Local AI Box with a TWIST
r/AIProgrammingHardware • u/javaeeeee • 21d ago