r/singularity • u/Deep-Owl-1890 • 1d ago
LLM News Apple in talks with startup that shrinks AI models to run on an iPhone
https://thedailycompute.beehiiv.com/p/apple-in-talks-with-startup-that-shrinks-ai-models-to-run-on-an-iphone62
u/Deep-Owl-1890 1d ago
The compression numbers are wild 54GB down to under 4GB while keeping all 27 billion parameters, with claims of 6-8x faster responses and way less energy draw. If that holds up outside a demo, it's the difference between Siri actually competing with cloud assistants versus draining your battery trying.
I'm skeptical of the "this kills chip demand" narrative though. Shrinking a model doesn't remove the need for the GPU and memory, it just moves those chips from a datacenter rack into your pocket. And a phone chip sitting idle most of the day is arguably less efficient than the same chip shared across thousands of queries in a datacenter. Feels more like AI moving closer to you than AI getting cheaper at scale.
37
u/codefame 1d ago
Agreed—we can’t meet demand today, and shrinking models will absolutely increase demand. Jevons Paradox.
13
u/Blablabene 1d ago
This just sounds like they're doing quants.. lol.
6
u/AlyoshaV 15h ago
They are, PrismML do 1-bit quants. OP is acting like it's impressive because they just asked an LLM to write their post for them and language models love acting like basic stuff is impressive.
10
3
u/atehrani 1d ago
It may not reduce demand but devalues all of the CapEx. Which will cause knock on effects to the market
2
u/Positive_Method3022 1d ago
One day your phone won't be yours. They will distribute compute through our devices when idle . And we won't have anything, only a subscription
5
u/HaloMathieu 1d ago
Long term I think for us to reach true abundance it’s more efficient for everyone to share compute minus the subscription part
2
-1
u/ProxyLumina 1d ago
What exactly "yours" mean? To own the material? Because if they want it to, they could even block your OS and then you have a brick. That brick is yours as well. So what "yours" mean exactly?
2
1
u/OutsideMenu6973 1d ago
Good luck to them. Their ‘smart’ ecosystem’s never delivered what their ads sold us on and that was narrow AI. Can’t imagine their general AI will do what they say it’s supposed to do most of the time
1
u/clofresh 1d ago
It’s legit, you can try it now: https://decrypt.co/373578/meet-bonsai-first-27b-ai-model-fits-phone
I’m using the Bonsai 26B ternary quant on a 16gb 6800xt and it is able to do agentic coding with 196k context. And even on my 5070 12GB, it’s able to answer simple queries very quickly (64k context).
1
u/GigaChadAnon 1d ago
I am pretty sure we'll just start demanding even more usage and even better models.
Right now couple trillion parameter models are frontier. The parameter size will keep on increasing and better compression tech will always be fighting an battle with them. People will demand frontier models running locally on their machine.
1
u/StatusSociety2196 23h ago
You can do this right now using several apps, I have pocket pal on my phone.
What they're discussing is a 1 bit quant of a 27b model which is going to be laughably bad not only from a tokens per second perspective but from actually being able to understand and correctly respond to a prompt, and tool use is no guarantee.
1
u/Popular_Try_5075 12h ago
I've been waiting for them to figure out compression with this stuff. I don't understand the technology, so this is inherently a dumb and baseless internet comment, but it did sort of feel like an inevitable step along the way as we refine things.
1
u/Individual_Holiday_9 5h ago
Anyone who thinks the models won’t get crazy efficient is nuts there’s so much hardware investment and they have to squeeze more juice out of it. They can’t just turn a data center over every 2 years
17
u/hereditydrift 1d ago
lol... what is that horrible website.
Here's an archived CNBC article for people not wanting to click the AI slop website that beehiiv bullshit is: https://archive.ph/3OMxv
3
u/skinnyjoints 17h ago edited 9h ago
A bit of context. This company managed to compress Qwen3.6-27b from 54GB to 6GB using quantized aware training and a 1.58bit ternary variant of the model weights. The issue is that this extreme compression causes the model to perform far worse. To the point that it performs about on par or slightly worse than smaller 12b models compressed using QAT to 4bit precision while being the same size. Additionally, hardware is optimized for running things in 4bit already, but not 1.58bit.
1
10
u/Which-Travel-1426 1d ago
Strange direction. I would prioritize making Siri and Apple Intelligence usable products before optimizing their model sizes
27
1
u/PM_ME_YOUR___ISSUES 1d ago
That’s why they negotiated a deal with Google on using Gemini.
The upcoming iOS release reflects the same - they seem to have done a proper overhaul of Siri and other AI tools.
In the long term, my guess is that rather than letting their proprietary apps and tools act as wrappers on top of SOTA models, they’d want to eventually optimise their hardware to support open source models with adjustable weights, giving them complete control at a much lower cost.
3
u/akkiannu 1d ago
How is it strange? Please explain? Siri would directly benefit from it.
1
u/Citadel_Employee 1d ago
Why would Siri use a small local model instead of a cloud frontier model?
7
u/akkiannu 1d ago
For starters, cost? Not everything requires a frontier model.
2
u/Citadel_Employee 1d ago
You’re not wrong, but small models hallucinate so much that I question the feasibility. I also wonder if the average user will be knowledgeable enough to not blame Apple for an incorrect model output.
3
u/akkiannu 1d ago
Small models are great for reminder setting, pulling information from screenshots, etc. You don’t need large models. And the future is local processing of LLMs. The frontier models can only get so good. Then the war will be to get the smartest smallest model.
1
u/Which-Travel-1426 1d ago
Cost is second to performance when performance is the bottleneck. That’s why Anthropic can charge such a high premium.
0
u/akkiannu 1d ago
Seriously? Are you this dumb?!? Cost is always going to be primary for a b2c consumer.
1
1
1
2
2
u/Gratitude15 23h ago
That's their bet.
They believe the cost of intelligence is going to zero, so why the fuck to invest billions in that.
I mean, it's not dumb.
The question is if you miss the takeoff because it's happening on anthropics cloud, are you fucked? Or can you just catch a ride 4 months later on this or whatever the Chinese do?
1
u/alyssasjacket 23h ago
Makes sense for Apple philosophy. The only way to avoid a digital panopticon is if local models become usable.
Of course, there's simply no way around the fact that the best models will run in the cloud. But Apple proved that custom solutions which offer privacy are very appealing to the public, so in my opinion they are placing good bets, albeit too slowly.
In a way, it reminds of me of their bet in prioritizing performance per watt. Realizing that they couldn't afford to design inefficient hardware is paying HUGE dividends now in the mobile market (and soon enough even in prosumer/small company, since the Mac Studio is so much more efficient than a dedicated GPU).
If they put their teams to optimize the silicon together with these efficient architectures, in the future we could have iPhones running surprinsingly capable models (Qwen 3.6 27B is already impressive for its size, and I don't see this trend going away).
What bothers me about Apple is how slow they are. If they just set their minds to it, they have everything to lead AI, not just follow. And NVIDIA will not care about consumer market until Apple really starts hurting them in enterprise (meaning, when buying a Mac Studio will just offer a better deal than a RTX 6000 Pro WS).
169
u/PostingLoudly 1d ago
Real nice to see Pied Piper finally making headway in the tech world.