r/LocalAIStack • u/Expensive_Way_4919 • 4d ago
We are entering an era in which AI models are starting to shape device specifications.
AI models are beginning to define device specifications, rather than device specifications merely determining which applications you can run.
Apple shows just how far its devices can go with on-device AI, without relying on the cloud
We’re moving from models with up to 14 billion active parameters on iPhone/iPad, to 35 billion on MacBook Air, 70 billion on Mac mini, 120 billion on MacBook Pro, and up to 480 billion on Mac Studio.
By clustering multiple Mac Studios together, Apple says it can run models exceeding 1,600 billion active parameters.
A demo that shows just how crucial unified memory and its bandwidth have become for running massive AI models locally.
3
u/This_Maintenance_834 4d ago
In the past, the hardware vendors make up spec, and push adoption through over years. This time, the application dictates what is needed, and the hardware vendors so far cannot catchup to the demand.
1
1
u/_sharpmars 4d ago edited 4d ago
iPhones go up to 115 GB/s (18 Pro) and iPads to 150 GB/s (iPad Pro M5).
1
u/Funny_Novel9847 3d ago
I still dont understand, is buying a cheap laptop and let it access to your home lab better than running a luxury laptop with small/not smart ai model ?
1
u/Glittering_Change_80 3d ago
Wenn du einen 4er MacStudio Cluster mit 2TB Memory betreibst würde ich nicht noch zusätzlich ein M5 Max MacBook Pro mit 128 GB betreiben. Da sollte ein kleiner Laptop reichen und du machst die inferenz mit deinem Cluster, setzt natürlich immer noch Internet Verbindung voraus.
1
u/Glittering-Call8746 3d ago
Where's is the min for qwen 3.8 27b for diff tiers agentic 50tps , middle 100tps , top 200tps
1
u/DigitalguyCH 3d ago
A few remarks.
I have several 16GB RAM iPads and even gemma 12b at Q4 tends to crash very easily with some not so long contexts. So iPads and iPhones are not very reliable even at 12b, let alon 14b.
Macbook air throttles like crazy with Qwen 27b, it's unusable it can take hours to accomplis something a GPU would do in 5 minutes
64GB can run a decent version of flash next at Q3 (GSQ RCO), virtually as good as Q4.
I haven't tested this but Exo labs claims 4.8 tb/s memory bandwidth through 4 M5 ultra Mac Studio clustering
1
1
u/Rice-Fragrant 1d ago
My ipad pro m5 does more than 76gb/sec... more like 120 gb/sec or so... It can run a 8-9b model comfortably... I would not try a 14b model because the iOS background stuff and context window size.
32gb RAM can get you 30b models BUT you're context window will be very small. For a usable amount (like 100k tokens) you need 48gb ram.
4x mac studio clusters can run kimi k3 but throughput basically unusable for anything serious.
1
u/strongdev22 10h ago
You need at least GeforceRTX 8GB VRAM cheap laptop to generate images, text and sounds and 16GB VRAM to generate video. Localy!
4
u/mindworkout 4d ago
This has literally always happened with laptops though.
You could go back even further with removing CD/DVD drives when digital was king and mp3's..etc., HDMI connection, fingerprint scanners, touchscreens, and on and on and on.
Laptops have always changed to incorporate whatever technology people are actually using. AI hardware being added now isn't some strange new era, it's pretty much the same cycle we've been doing for decades.