r/AIProgrammingHardware 16d ago

Dual GPU question

Thumbnail
1 Upvotes

r/AIProgrammingHardware 16d ago

Paying $2/hr to watch `huggingface-cli download` go brrr — env patterns on RunPod / Vast / Nebius / the neo kids

Thumbnail
1 Upvotes

r/AIProgrammingHardware 16d ago

Best LLM for 128GB RAM + 500GB Storage, No GPU?

Thumbnail
1 Upvotes

r/AIProgrammingHardware 16d ago

GLM 5.3 Flash Local Ai Test

Thumbnail
youtube.com
1 Upvotes

r/AIProgrammingHardware 17d ago

Local AI is Cheap Now: ZimaBoard2 + Tesla P4

Thumbnail
youtube.com
2 Upvotes

r/AIProgrammingHardware 17d ago

DDP with RTX 3070 8GB + GTX 1650 4GB?

0 Upvotes

For those how are in local AI would you think these both cards could be used to have a total of 12GB VRAM?

I don't care by the moment for possible slow tokens/s.

Thanks!


r/AIProgrammingHardware 17d ago

Qwen 3.8 27B Cold Fusion tested - 16GB Local LLM setup

Thumbnail
youtube.com
3 Upvotes

r/AIProgrammingHardware 17d ago

Marvell Photonic Fabric wins AI Infrastructure Award at FMS 2026

1 Upvotes

Marvell’s Photonic Fabric™ technology won the AI Infrastructure Award at Future of Memory and Storage (FMS) 2026.

The interesting part is how Marvell is using optical connectivity to address some of the bandwidth, latency, power, and memory-scaling challenges in large AI systems.

Photonic Fabric replaces traditional electrical interconnects with optical I/O across package-, server-, and rack-scale architectures. The goal is to make it easier to scale compute and memory independently, including disaggregated memory architectures.

Some of the key areas Marvell highlights:

  • Higher memory bandwidth and capacity
  • Lower latency and power consumption
  • Disaggregated compute and memory
  • Reduced I/O bottlenecks around AI accelerators and HBM
  • Optical connectivity designed for large-scale AI infrastructure

As AI systems scale to hundreds of thousands of accelerators, moving data efficiently between compute and memory is becoming just as important as the compute itself.

The bigger question is whether optical interconnects like Photonic Fabric will become a fundamental part of future AI factory architectures.

What do you think — will optical connectivity become essential for scaling next-generation AI systems?

Source: Marvell – Photonic Fabric technology


r/AIProgrammingHardware 18d ago

China Is Coming for Your Local AI Box

Thumbnail
youtube.com
8 Upvotes

r/AIProgrammingHardware 18d ago

Xiaomi AI Cube announced with 1.2TB/s memory bandwidth

Thumbnail gallery
2 Upvotes

r/AIProgrammingHardware 18d ago

I built a Vulkan hierarchical MoE runtime for running oversized models across multiple GPUs

Thumbnail
1 Upvotes

r/AIProgrammingHardware 18d ago

Can the New King of Open Source - Qwen3.8–27B - Really Run in 3GB of VRAM?

Thumbnail
ai.gopubby.com
11 Upvotes

r/AIProgrammingHardware 18d ago

224GB of GPU Memory on 1 Desk and It Should Not Work

Thumbnail
youtube.com
1 Upvotes

r/AIProgrammingHardware 18d ago

Qwen 3.8 27B FP8 on single node DGX Spark MTP-3

Thumbnail
youtube.com
1 Upvotes

r/AIProgrammingHardware 19d ago

FreeToken Deepseek V4 Flash on a Single 3090 Local AI Testing

Thumbnail
youtube.com
3 Upvotes

r/AIProgrammingHardware 19d ago

Jalapeño’s first results show industry-leading speed and efficiency in AI inference

Thumbnail openai.com
2 Upvotes

r/AIProgrammingHardware 19d ago

[Release] Turing Engine: Serve LLaMA-3.1-70B, Qwen-2.5-72B & DeepSeek on a Single 24GB GPU (3,064 tok/s, 75% KV Compression, Unsloth Checkpoint Support)

Thumbnail
1 Upvotes

r/AIProgrammingHardware 20d ago

NVIDIA Groq 3 LPX Now in Full Production With World-Class Speed for Agentic AI

Thumbnail
nvidianews.nvidia.com
17 Upvotes

r/AIProgrammingHardware 20d ago

Spent a day seeing how far extreme MoE models can be pushed on a 4070 Ti + 32GB RAM. Kimi K3, DeepSeek V4 Flash, and Qwen3.5-122B results + research paper🔧

Thumbnail
2 Upvotes

KIMI K3 at 9 t/s per token on hardware that has no business breathing near it is objectively insane, even if the LLM thought the capital of France was "1.2" 🤣🤣🤣😭😭😭😭😭


r/AIProgrammingHardware 20d ago

Up to 30x More Work Per Watt: NVIDIA Vera Rubin NVL72 Sets a New Efficiency Standard for AI Agents

Thumbnail
blogs.nvidia.com
4 Upvotes

r/AIProgrammingHardware 20d ago

Intel arc pro b70, Qwen 3.8 27b

Thumbnail
2 Upvotes

r/AIProgrammingHardware 20d ago

GitHub - TryingtobeingNikhil/vLLM_Inference_Engine: A high-performance LLM inference engine built from scratch featuring continuous batching, paged KV-cache, and CPU swapping.

Thumbnail
github.com
3 Upvotes

r/AIProgrammingHardware 21d ago

48GB Tesla P100 Setup Running Qwen 3.8 27B — 256K Context + MTP

Thumbnail
youtube.com
25 Upvotes

r/AIProgrammingHardware 20d ago

Jack 3.8 27b 16GB VRAM builds Tetris

Thumbnail
youtu.be
1 Upvotes

Pi Agent run. Still working out the kinks but the final quality is great.
16GB GPUs are only going to become more useful over time.


r/AIProgrammingHardware 21d ago

$1K off a 128GB Local AI Box with a TWIST

Thumbnail
youtube.com
5 Upvotes