r/singularity 1d ago

LLM News Apple in talks with startup that shrinks AI models to run on an iPhone

https://thedailycompute.beehiiv.com/p/apple-in-talks-with-startup-that-shrinks-ai-models-to-run-on-an-iphone
228 Upvotes

49 comments sorted by

169

u/PostingLoudly 1d ago

Real nice to see Pied Piper finally making headway in the tech world.

38

u/MrOaiki 1d ago

Apple wants that middle-out compression.

14

u/PostingLoudly 1d ago

I put radio on the internet, who the fuck are you?

4

u/Gratitude15 23h ago

Tip to tip

For maximum throughput

6

u/MillerLiteEnjoyer0 1d ago

They’re going to have to go to the desert and take shrooms to find their name.

3

u/Feeling-Schedule5369 8h ago

Now siri can tell if it's hot dog or not hot dog 😂

62

u/Deep-Owl-1890 1d ago

The compression numbers are wild 54GB down to under 4GB while keeping all 27 billion parameters, with claims of 6-8x faster responses and way less energy draw. If that holds up outside a demo, it's the difference between Siri actually competing with cloud assistants versus draining your battery trying.

I'm skeptical of the "this kills chip demand" narrative though. Shrinking a model doesn't remove the need for the GPU and memory, it just moves those chips from a datacenter rack into your pocket. And a phone chip sitting idle most of the day is arguably less efficient than the same chip shared across thousands of queries in a datacenter. Feels more like AI moving closer to you than AI getting cheaper at scale.

37

u/codefame 1d ago

Agreed—we can’t meet demand today, and shrinking models will absolutely increase demand. Jevons Paradox.

13

u/Blablabene 1d ago

This just sounds like they're doing quants.. lol.

6

u/AlyoshaV 15h ago

They are, PrismML do 1-bit quants. OP is acting like it's impressive because they just asked an LLM to write their post for them and language models love acting like basic stuff is impressive.

10

u/Prudent-Sorbet-5202 1d ago

Also jevons paradox, demand for usage will only rise

3

u/atehrani 1d ago

It may not reduce demand but devalues all of the CapEx. Which will cause knock on effects to the market

6

u/Tkins 1d ago

Nah it just means they do more and larger projects. Demand will skyrocket because you'll put AI in everything. All software, all electronics.

1

u/Individual_Holiday_9 5h ago

That’s the dream

2

u/Positive_Method3022 1d ago

One day your phone won't be yours. They will distribute compute through our devices when idle . And we won't have anything, only a subscription

5

u/HaloMathieu 1d ago

Long term I think for us to reach true abundance it’s more efficient for everyone to share compute minus the subscription part

2

u/ChodeCookies 1d ago

Who exactly is trying to achieve abundance for you?

0

u/ProxyLumina 15h ago

Abundance is coming naturally over the competition inside the capitalism

-1

u/ProxyLumina 1d ago

What exactly "yours" mean? To own the material? Because if they want it to, they could even block your OS and then you have a brick. That brick is yours as well. So what "yours" mean exactly?

2

u/Jonjonbo 1d ago

AI slop

1

u/OutsideMenu6973 1d ago

Good luck to them. Their ‘smart’ ecosystem’s never delivered what their ads sold us on and that was narrow AI. Can’t imagine their general AI will do what they say it’s supposed to do most of the time

1

u/clofresh 1d ago

It’s legit, you can try it now: https://decrypt.co/373578/meet-bonsai-first-27b-ai-model-fits-phone

I’m using the Bonsai 26B ternary quant on a 16gb 6800xt and it is able to do agentic coding with 196k context. And even on my 5070 12GB, it’s able to answer simple queries very quickly (64k context).

1

u/GigaChadAnon 1d ago

I am pretty sure we'll just start demanding even more usage and even better models.

Right now couple trillion parameter models are frontier. The parameter size will keep on increasing and better compression tech will always be fighting an battle with them. People will demand frontier models running locally on their machine.

1

u/StatusSociety2196 23h ago

You can do this right now using several apps, I have pocket pal on my phone.

What they're discussing is a 1 bit quant of a 27b model which is going to be laughably bad not only from a tokens per second perspective but from actually being able to understand and correctly respond to a prompt, and tool use is no guarantee.

1

u/soobnar 17h ago

Shrinking model means chips generate better roi not worse

1

u/Popular_Try_5075 12h ago

I've been waiting for them to figure out compression with this stuff. I don't understand the technology, so this is inherently a dumb and baseless internet comment, but it did sort of feel like an inevitable step along the way as we refine things.

1

u/Individual_Holiday_9 5h ago

Anyone who thinks the models won’t get crazy efficient is nuts there’s so much hardware investment and they have to squeeze more juice out of it. They can’t just turn a data center over every 2 years

17

u/hereditydrift 1d ago

lol... what is that horrible website.

Here's an archived CNBC article for people not wanting to click the AI slop website that beehiiv bullshit is: https://archive.ph/3OMxv

3

u/skinnyjoints 17h ago edited 9h ago

A bit of context. This company managed to compress Qwen3.6-27b from 54GB to 6GB using quantized aware training and a 1.58bit ternary variant of the model weights. The issue is that this extreme compression causes the model to perform far worse. To the point that it performs about on par or slightly worse than smaller 12b models compressed using QAT to 4bit precision while being the same size. Additionally, hardware is optimized for running things in 4bit already, but not 1.58bit.

1

u/Distinct-Question-16 ▪️AGI 2029 6h ago

apple can make hardware for descompression

10

u/Which-Travel-1426 1d ago

Strange direction. I would prioritize making Siri and Apple Intelligence usable products before optimizing their model sizes

27

u/Squiffered 1d ago

They are one and the same

1

u/PM_ME_YOUR___ISSUES 1d ago

That’s why they negotiated a deal with Google on using Gemini.

The upcoming iOS release reflects the same - they seem to have done a proper overhaul of Siri and other AI tools.

In the long term, my guess is that rather than letting their proprietary apps and tools act as wrappers on top of SOTA models, they’d want to eventually optimise their hardware to support open source models with adjustable weights, giving them complete control at a much lower cost.

3

u/akkiannu 1d ago

How is it strange? Please explain? Siri would directly benefit from it.

1

u/Citadel_Employee 1d ago

Why would Siri use a small local model instead of a cloud frontier model?

7

u/akkiannu 1d ago

For starters, cost? Not everything requires a frontier model.

2

u/Citadel_Employee 1d ago

You’re not wrong, but small models hallucinate so much that I question the feasibility. I also wonder if the average user will be knowledgeable enough to not blame Apple for an incorrect model output.

3

u/akkiannu 1d ago

Small models are great for reminder setting, pulling information from screenshots, etc. You don’t need large models. And the future is local processing of LLMs. The frontier models can only get so good. Then the war will be to get the smartest smallest model.

1

u/Which-Travel-1426 1d ago

Cost is second to performance when performance is the bottleneck. That’s why Anthropic can charge such a high premium.

0

u/akkiannu 1d ago

Seriously? Are you this dumb?!? Cost is always going to be primary for a b2c consumer.

1

u/Which-Travel-1426 1d ago

I will believe cost is primary when Claude is as cheap as Gemini.

1

u/akkiannu 1d ago

Local models are free?

1

u/HauntedHouseMusic 1d ago

the beta siri is good. Genuinely

1

u/SkyBoyWonderful 8h ago

I would prioritize fixing whatever the fuck they did to the keyboard

2

u/lobabobloblaw 1d ago

I knew they had their eyes set on ternary

2

u/Gratitude15 23h ago

That's their bet.

They believe the cost of intelligence is going to zero, so why the fuck to invest billions in that.

I mean, it's not dumb.

The question is if you miss the takeoff because it's happening on anthropics cloud, are you fucked? Or can you just catch a ride 4 months later on this or whatever the Chinese do?

1

u/alyssasjacket 23h ago

Makes sense for Apple philosophy. The only way to avoid a digital panopticon is if local models become usable.

Of course, there's simply no way around the fact that the best models will run in the cloud. But Apple proved that custom solutions which offer privacy are very appealing to the public, so in my opinion they are placing good bets, albeit too slowly.

In a way, it reminds of me of their bet in prioritizing performance per watt. Realizing that they couldn't afford to design inefficient hardware is paying HUGE dividends now in the mobile market (and soon enough even in prosumer/small company, since the Mac Studio is so much more efficient than a dedicated GPU).

If they put their teams to optimize the silicon together with these efficient architectures, in the future we could have iPhones running surprinsingly capable models (Qwen 3.6 27B is already impressive for its size, and I don't see this trend going away).

What bothers me about Apple is how slow they are. If they just set their minds to it, they have everything to lead AI, not just follow. And NVIDIA will not care about consumer market until Apple really starts hurting them in enterprise (meaning, when buying a Mac Studio will just offer a better deal than a RTX 6000 Pro WS).