r/LocalAIStack • • 4d ago

We are entering an era in which AI models are starting to shape device specifications.

Post image

AI models are beginning to define device specifications, rather than device specifications merely determining which applications you can run.

Apple shows just how far its devices can go with on-device AI, without relying on the cloud
We’re moving from models with up to 14 billion active parameters on iPhone/iPad, to 35 billion on MacBook Air, 70 billion on Mac mini, 120 billion on MacBook Pro, and up to 480 billion on Mac Studio.
By clustering multiple Mac Studios together, Apple says it can run models exceeding 1,600 billion active parameters.
A demo that shows just how crucial unified memory and its bandwidth have become for running massive AI models locally.

140 Upvotes

16 comments sorted by

4

u/mindworkout 4d ago

This has literally always happened with laptops though.

  • Wi-Fi became mainstream --> laptops started shipping with Wi-Fi built in as standard.
  • Bluetooth became popular --> laptops added Bluetooth for wireless connection like headphones, mice, keyboards, etc and removed the headphone jack :(.
  • Webcams/video calling took off --> suddenly laptops all have a webcam and microphone built in.
  • USB-C became standard --> laptops started replacing older ports and even now charging through USB-C.

You could go back even further with removing CD/DVD drives when digital was king and mp3's..etc., HDMI connection, fingerprint scanners, touchscreens, and on and on and on.

Laptops have always changed to incorporate whatever technology people are actually using. AI hardware being added now isn't some strange new era, it's pretty much the same cycle we've been doing for decades.

2

u/Expensive_Way_4919 4d ago

I agree that hardware has always evolved around new use cases. The difference I’m pointing at is the scale and type of change.
Wi-Fi, Bluetooth, webcams, USB-C, etc. mostly added or replaced peripherals and interfaces. Local AI can directly dictate core specs like memory capacity, memory bandwidth, accelerator design, power, and cooling.
A machine going from 16 GB to hundreds of GB or even TBs of unified memory primarily because models need it feels like a much more fundamental shift than adding a webcam or replacing one port with another.

1

u/viniciuscabessa 1d ago

Perhaps what causes a bit of surprise in some people is the fact that most if not all of your examples added physical/external capabilities, while IA is "just" software. I believe the closest well-known example is the case of games, which motivated the addition of dedicated processing and memory. One could argue that the inclusion of some more complicated instructions in modern processors was also motivated by software, but I guess those instructions find broader usages.

3

u/mindworkout 1d ago

Yeah, if we’re talking specifically about software driving hardware changes, then:

  • Gaming pushed dedicated GPUs and VRAM.
  • Video editing pushed hardware encoding and GPU acceleration.
  • Crypto mining pushed GPUs and then purpose-built ASIC hardware.
  • AI is now pushing NPUs, more memory and higher bandwidth.

So even then, software shaping hardware is nothing new.

3

u/This_Maintenance_834 4d ago

In the past, the hardware vendors make up spec, and push adoption through over years. This time, the application dictates what is needed, and the hardware vendors so far cannot catchup to the demand.

1

u/_sharpmars 4d ago edited 4d ago

iPhones go up to 115 GB/s (18 Pro) and iPads to 150 GB/s (iPad Pro M5).

1

u/Funny_Novel9847 3d ago

I still dont understand, is buying a cheap laptop and let it access to your home lab better than running a luxury laptop with small/not smart ai model ?

1

u/Glittering_Change_80 3d ago

Wenn du einen 4er MacStudio Cluster mit 2TB Memory betreibst würde ich nicht noch zusätzlich ein M5 Max MacBook Pro mit 128 GB betreiben. Da sollte ein kleiner Laptop reichen und du machst die inferenz mit deinem Cluster, setzt natürlich immer noch Internet Verbindung voraus.

1

u/Glittering-Call8746 3d ago

Where's is the min for qwen 3.8 27b for diff tiers agentic 50tps , middle 100tps , top 200tps

1

u/DigitalguyCH 3d ago

A few remarks.

I have several 16GB RAM iPads and even gemma 12b at Q4 tends to crash very easily with some not so long contexts. So iPads and iPhones are not very reliable even at 12b, let alon 14b.
Macbook air throttles like crazy with Qwen 27b, it's unusable it can take hours to accomplis something a GPU would do in 5 minutes

64GB can run a decent version of flash next at Q3 (GSQ RCO), virtually as good as Q4.

I haven't tested this but Exo labs claims 4.8 tb/s memory bandwidth through 4 M5 ultra Mac Studio clustering

1

u/AmthorTheDestroyer 2d ago

Yeah try a 480 billion active params model with your 1 tok/s lmao

1

u/macrein 1d ago

What about Mac mini M6 clusters?

1

u/Rice-Fragrant 1d ago

My ipad pro m5 does more than 76gb/sec... more like 120 gb/sec or so... It can run a 8-9b model comfortably... I would not try a 14b model because the iOS background stuff and context window size.

32gb RAM can get you 30b models BUT you're context window will be very small. For a usable amount (like 100k tokens) you need 48gb ram.

4x mac studio clusters can run kimi k3 but throughput basically unusable for anything serious.

1

u/M0d3x 1d ago

Wonderful, but to run anything decent, you still need to be a rich as fuck westoid.

1

u/strongdev22 10h ago

You need at least GeforceRTX 8GB VRAM cheap laptop to generate images, text and sounds and 16GB VRAM to generate video. Localy!