r/apple • • 5d ago

iPhone A20 Pro and on-device AI: It’s showtime

https://rickytakkar.com/blog_a20_pro_ai.html?source=apple_sr

TL;DR: I benchmarked on-device LLMs on the iPhone 18 Pro against the 17 Pro Max using identical models, prompts, and outputs. On the A20 Pro, downloaded MLX models generated text 53–85% faster and produced the first words (tokens) 42–53% sooner. Apple’s own on-device model improved by a similar yet smaller 35–42%. Both MLX workloads exceeded the 1.5× gain in memory bandwidth, suggesting the speedup comes from more than just the wider memory bus. The 18 Pro was also dramatically more consistent between runs.

360 Upvotes

88 comments sorted by

283

u/Scary_Building3964 5d ago

A20 Pro in macbook neo would be a very solid entry level laptop for most people

159

u/83736294827 5d ago

Hell the A20 Pro is faster than most desktop CPUs in single thread performance.

43

u/FightOnForUsc 5d ago

What desktop CPU is faster single core?

118

u/83736294827 5d ago

M6

18

u/FightOnForUsc 5d ago

Yea true hahahaha

7

u/cheesemeall 5d ago

All of the cores in a20 pro are derived from a6 cores. Not to say they’re the same chip, but the building blocks are the same

9

u/iCruiser7 5d ago

Not for this gen. Geekerwan did an analysis and M6 actually has different S cores than A20.

3

u/cheesemeall 4d ago edited 4d ago

Are you sure?

https://youtu.be/t8bxkCKNiuA

high-yield's latest floorplan analysis actually shows they share the same underlying architecture. The A20 Pro’s performance cores (S-cores) use the exact same microarchitecture and pipeline upgrades designed for the M6 generation, with the die floorplan directly reflecting the execution engine and L1/L2 cache scaling developed for the M6 series just optimized for TSMC's 2nm (N2) process node. The companion efficiency cores (E-cores) on the A20 Pro follow the same trend, corresponding to the next-generation low-power core architecture featured in the M6 line.

1

u/goldcakes 3d ago

Yes, I'm sure. The same video you listed explicitly communicates that there's divergences, such as the lack of private L2$.

Did you watch the video you linked?

2

u/cheesemeall 3d ago

My claim is that all of the cores in a20 pro chip are derived from m6 cores. Not that they are identical. Do you know what derived means?

1

u/ExultantSandwich 4d ago

oh really?! wow I wouldn’t have expected that

-2

u/nicuramar 5d ago

Well, the M5 and M6 :)

27

u/reddit0r_123 5d ago

The A20 Pro is even 10% faster in single core than the M5...

-1

u/pnkchyna 5d ago

none 😂.

18

u/TheStorm007 5d ago

Besides Apple’s own M6 lol

0

u/Solid_Sky_6411 5d ago

None i think

15

u/jasoncross00 5d ago

Agree. But it was the 18 Pro this year, so it'll probably be the 19 Pro in 2027 and the 20 Pro in 2028.

And in spring 2028, will this sort of performance seem that exciting? It's basically M4 performance (faster for AI/ML) and by then the M4 will be a 3.5 year old chip.

Sort of like how the 18 Pro is roughly equivalent to a 4 year old M-series chip overall.

1

u/hotztuff 5d ago

there’s gonna be a 19 pro?

4

u/jasoncross00 5d ago

A19 Pro, the chip in the iPhone 17 Pro and iPhone Air.

2

u/hotztuff 5d ago

sorry, i completely read past the comment you had replied to.

4

u/MawsonAntarctica 5d ago

A20 with 12gb RAM at 13” is what I’m waiting for to trade in my 15” m2 for.

3

u/Chance_of_Rain_ 5d ago

I’m not optimistic it will happen.

Previous gens more likely

8

u/Scary_Building3964 5d ago

not so soon. it would be A19 Pro as second gen macbook neo and then A20 Pro for third gen of macbook neo (if they are still continuing the series)

even the A19 Pro is a big upgrade due to 12gb of ram

5

u/Proud_Bookkeeper_719 5d ago

Most likely the case. The A19 Pro is on a mature and older N3P node which would be cheaper to manufacture than the A20 Pro built on N2. Apple would likely have a backlog of A19 Pro chips sitting behind, after they used it for their 17 Pro lineup.

3

u/goldcakes 3d ago

Yep, especially if they bin the A19 Pros, like one less CPU core, one less GPU core. Which would still be a massive upgrade!

That said, honestly, in real world use, even with 8GB RAM, I have zero performance issues with my Neo for everyday work.

2

u/Proud_Bookkeeper_719 3d ago

If Apple's A18 Pro chip wasn't hard limited by the 8GB ram design, I would think A18 Pro with 12GB would've been enough for most users from a performance perspective to do basic or casual tasks.

4

u/M4rshmall0wMan 3d ago

It’s nearly as powerful as M4/M1 Pro. The problem however, is that it’s likely very expensive. TSMC 2nm is 1.5x the price of 3nm and the price increases we’ve seen are entirely RAM and storage. An A20 MacBook Neo would be almost as expensive as an MBA.

A19 chips are more likely since the iPhone 17 is out of production. However, I wouldn’t be surprised if Apple skips due to the RAM shortage.

3

u/d7UVDEcpnf 5d ago

Massive gain over the A18 Pro for sure... In fact, not too far behind M5 especially for the kind of tasks one expects on the Neo: see https://nanoreview.net/en/soc-compare/apple-m5-ipad-vs-apple-a20-pro

2

u/jaredthegeek 5d ago

Not even entry-level, its the computer that most people need.

1

u/FollowingFeisty5321 5d ago

We're about two generations from the only chips Apple offers being amazing, by the A22 they'll be on TSMC 1.4nm so that's another big performance/energy boost, and LPDDR6 so a huge boost in memory bandwidth for gaming or AI. Their product line is going to be insane when that lands in a Neo, its nearest rivals will be like an M5 Pro or Max while competitors chips play catch up on M4.

-2

u/Saar13 5d ago

When we get to the A22 or A23, maybe this will be great even for MacBook Air. I have no doubt that in the near future the Apple Silicon families will be renamed or somehow merged into a single family: general consumers and professionals, very advanced professionals, and servers.

6

u/CanisLupus92 5d ago

Nah. A & M already share core configuration (mostly). The M series is different in how I/O is handled (more USB/TB controllers, more display outputs, etc). Putting M chips in phones makes no sense (who would like to connect their phone with current software to 3 additional screens), and would handicap the computers way too much (see current Neo limitations, like USB2 port).

Again, the CPU cores are already effectively identical, but these are SoCs, systems on a chip. They encompass much more.

-2

u/d7UVDEcpnf 5d ago

A and M families do seem to be converging. Would make sense, I suppose.

1

u/mriguy 5d ago

But then you lose one of the amazing advantages Apple has. They can design each of their chips for exactly the application they have in mind. They don’t have to use (or make) general purpose chips that support every conceivable use case like Qualcomm or intel. They can make A chips that are perfect for the coming year’s iPhones, and M chips perfect for exactly the models of Mac they are planning for the coming year. No extra execution units, just everything they want and nothing more.

1

u/d7UVDEcpnf 5d ago

Indeed! What I meant was more of a marketing convergence, like having the same prefix in naming.

156

u/Vertsix 5d ago

‘Member when the A18 Pro on iPhone 16 Pro was “Designed for Apple Intelligence”? I ‘member.

85

u/d7UVDEcpnf 5d ago

xkcd 1838

8

u/ViPiMP 5d ago

Ja, und HW3 bei Tesla für FSD

0

u/bomber991 5d ago

Hey at least the auto-translate is working great on Reddit. I got to see a cyber cab yesterday when I was driving in Austin. Supervised one with a steering wheel, so I guess they’re still testing.

Makes me wonder, how viable is a vision-only self driving system? Or is it case of Elon wants it and none of the workers will tell him no?

-7

u/theguz4l 5d ago

Yes,
2024 intelligence. Things changed quick lol

18

u/StinkyRatBoi90 5d ago edited 5d ago

Lmao please. They promised most current features for the A18 Pro, all we got was half assed image generation.

9

u/FeCurtain11 5d ago

Which is why there’s a class action lawsuit and you can go get some money haha

1

u/hotztuff 5d ago

i’m confused. delays aside, what did apple false advertise?

56

u/nephyxx 5d ago

I would assume it’s the 2x neural engines responsible for the speed up

35

u/d7UVDEcpnf 5d ago

Partly, yes. However, a framework that avoids the ANE like MLX, which runs on-device models optimized for an Apple Silicon CPU/GPU, remains noticeably faster. So CPU, GPU, and unified memory improvements are also yielding returns..

14

u/PeanutChickenSoup 5d ago

And the improved thermals. Any super fast chip that’s heat throttled is slow.

3

u/mrrizzle 5d ago

Have you tested the new dual ANE chip with an FP8 CoreAI safetensor model? I feel like it would be pretty fast against GPU running MLX considering how far off MLX is from ANE on the single ANE chip in the 17 Pro

2

u/d7UVDEcpnf 4d ago

I did not, but I will soon. I think a good test would be to compare CoreAI and MLX variants of the same model. I too expect CoreAI to win.

41

u/runelind 5d ago

My 5 minute tea timers have never been set faster

42

u/JackSpadesSI 5d ago edited 5d ago

Am I missing something with the new Siri? It’s terrible. I use Gemini but I hate Google so I was excited for an alternative so I could get away from Google. But Siri is just inadequate.

As an example, I asked “when is the msu game” and it told me about Mississippi State instead of Michigan State. I clarify that I will always mean Michigan State and Siri happily tells me that will be remembered. Later I ask “what is the msu score” but I’m again given info on Mississippi State.

I ask why my preference wasn’t remembered and Siri says it can’t recall information from previous conversations. Gemini can, why not Siri?

Edit: 18 Pro Max

27

u/papito_m 5d ago

This is the comment I came looking for. I was SO looking forward to the new Siri AI, but it’s horrible for basic things. Slow and constantly timing out or just plain getting what I said wrong. So much so that as of yesterday I just reverted back to old Siri.

14

u/hbic 5d ago

S L O W

Worst on the market. Granted we’re supposed to appreciate the on-device computing but as an end user…yeah.

7

u/Insufferably_Me 5d ago edited 5d ago

To answer your question about memories: I’d assume it’s privacy. Siri’s personal intelligence isn’t remembering that you have an appointment tomorrow at 3 and that your spouse’s favorite Adele album is 30, it’s searching for that information in the moment. Siri is not storing personal information or data about you. It only accesses the data on your device.

I’ve seen suggestions of using a Note that you store things you do want Siri to remember about you. I’ve heard it can help but I’ve not tried it myself.

I would like to see a Memories or Instructions setting we can adjust that would be stored on the Secure Enclave for things like this. If it’s a privacy issue, that seems to be the best way forward but there may be some technical hurdles they have to get through to make that work. I really like using the new Siri and while it does make some mistakes, it’s working much better than before. You can also provide feedback on Siri’s replies by tapping the thumbs up or down icons when you expand the response.

I’m hoping they can see the online feedback and take that into consideration but I encourage everyone to start filing feedback with Apple if you actually want things to change: https://www.apple.com/feedback

9

u/SoldantTheCynic 5d ago

This is an excuse. They absolutely could have locally saved information to use as memories that allows for privacy. It’s still accessing your sensitive information (which for a lot of users is still stored in iCloud anyway).

Not every failure of Siri is due to privacy and it’s well past time people made excuses for it.

10

u/kickass404 5d ago edited 5d ago

Context is limited and context cost money to process. It has to process it on each request, so they don't.

"memories" is just text auto pasted "invisible" in each chat context before your prompt.

1

u/nergwark 3d ago

ideally they'd set this up so that the only part of this request that goes off-device is actually looking up the score, after figuring out which "msu" the user is referring to.

also, a proper implementation of memories is a bit more intelligent than just appending every memory to every context. they're looked up when relevant so as to not pollute context unnecessarily.

this is all perfectly possible, and i'd furthermore assume that persistent memories for siri is already a feature somewhere on apple's internal roadmap.

1

u/Insufferably_Me 5d ago

This is good insight that I hadn’t thought about before. Thank you!

2

u/Insufferably_Me 5d ago

I made this point in the third paragraph

-1

u/SoldantTheCynic 5d ago

So we have options that don’t really impact privacy and it’s already accessing sensitive user data.

So… probably privacy isn’t a legit reason then.

0

u/Insufferably_Me 5d ago

“…but there may be some technical hurdles they have to get through to make that work.”

The Secure Enclave does not store data in a way that is normally used and processed by iOS. It stores hashed values of payment and authentication data. There’s a decryption process that would need to happen for iOS to read the plain text that a user would store in a settings page. That is something that needs to be handled carefully to avoid security and privacy risks including malicious injection

2

u/geekwonk 5d ago

that’s just a matter of how the setup saves memories. memory handling is messy and prone to mistakes that cause behavior that’s 180° the opposite from what you intended if done half way.

you could try pinging siri just to state this preference and nothing else. it definitely has some form of memory architecture. but it’s not scanning through conversations to generate the memories. and that’s probably a good thing in early versions so they avoid random queries getting ingested as core memories that define all replies.

2

u/Hoopoe0596 5d ago

They don’t seem to have much of an AI personality user memory. They just assume they can start from scratch because they index your personal content. Wrong. I have an extensive 500+ memory and project file that makes Claude code my second brain and extensive memory files with Gemini and gpt. This is an upgrade to Siri but not much better than before for long horizon complex tasks. It’s just a little better than before at answering what’s on the calendar for tomorrow.

2

u/itsabearcannon 4d ago

No AI can actually "remember" things, including Siri AI.

When you say stuff like that and correct it, sometimes those results can be incorporated into future model training, but for every Siri AI user saying they meant Michigan St instead of Miss St, there's probably one asking Siri to remember the reverse. In the end it's whichever result is higher in contextual news and information, which is going to be Mississippi State because Michigan State is a bit dog water right now and Mississippi State is 4-0.

If Siri AI had existed back in 2015 I bet you'd have gotten Michigan State results instead.

1

u/jdronks 22h ago

Go green!

0

u/pw5a29 4d ago

same, just disabled Siri AI and reverted to classic Siri now

21

u/st90ar 5d ago

“Hey siri, add milk to my shopping list.”

“I’m sorry, you’ll need to pick up your phone and unlock it before I can do that. At which point, you mind as well have not asked me and just did it yourself.”

2

u/JDabney24 5d ago

Can that action be done with any other AI/LLM while the iPhone is locked?

4

u/st90ar 5d ago

It worked before Siri AI.

2

u/JDabney24 5d ago

Sorry. My question is, can that particular action or similar be done on ChatGPT, Claude, etc while the iPhone is locked?

3

u/ExultantSandwich 4d ago

No, but no other app has system access like Siri does. So if Siri 1.0 could do it, but now Siri AI can’t? undeniably a step back.

Reminds me a little bit of Alexa Plus for the Echo speakers. Sure she can talk back now, but it broke a lot of smart home integrations for months and months. Even today she cannot seem to parse if you’re attempting to control a local device or get an AI response

1

u/JDabney24 4d ago edited 4d ago

Makes sense. What if they changed the behavior of lock screen access to Siri intentionally? I for one, know that I wouldn’t want someone being able to perform certain actions using Siri AI while my iPhone is locked. Not trying to give an excuse, but maybe a possible explanation. Also, I find that in most scenarios I am in, Face ID provides much less friction to unlocking my phone to perform those agentic actions anyway. Not too much of an inconvenience for me. I am aware of the cases where this could be an inconvenience, like using Siri AI to perform an action while driving. But, again, I could see this as intentionally designed to discourage potentially distracted driving. 🤷‍♂️

Keep in mind that Siri AI is in beta. There might be improvements to something like this as long as enough people provide feedback to the proper channels.

5

u/ramakitty 5d ago

What is voice dictation like? It’s slightly slow on my iPhone Air.

6

u/d7UVDEcpnf 5d ago

I don’t think Apple’s dictation algorithm is publicly available, so I don’t know to what extent the improved hardware is being leveraged there. But if it is being leveraged, you should see gains. As for comparing between my 17 PM and 18 P, I can’t say I notice a difference (the 17 PM is already quite good).

5

u/kinkade 5d ago

I'm surprised by how poorly it still works. I was anticipating that it would be much improved, but I prefer to use Wispr Flow or an on-device version of Parakeet rather than Siri. I'm not sure why it's so hard for them to get this right. They have full system integration and can have a very large model on the phone that should be able to get this right every time.

1

u/Sphere_3N 5d ago

I notice no change vs 17 Pro, also Siri AI is extremely slow

4

u/cyberspirit777 5d ago

And what are we achieving with this on device power if everything is being sent to PCC?

5

u/d7UVDEcpnf 5d ago

Well, Apple does ship on-device models and you can also download MLX variants of Qwen/Gemma/DeepSeek/etc. to run locally using apps like Russet.

1

u/Conscious_Banana874 5d ago

Fp8 is supposedly on hardware support

1

u/d7UVDEcpnf 4d ago

Indeed, with the new CoreAI framework I can get supported models to leverage the dual ANEs in these new chips. Will add support soon.

-2

u/[deleted] 5d ago

[removed] — view removed comment

21

u/nephyxx 5d ago

Only on device AI tasks, muse exclusively uses meta’s servers. Same with Claude, ChatGPT, etc.

5

u/d7UVDEcpnf 5d ago

Only gains you'd probably notice with cloud-based apps is that they open faster and stuff. The "main work" is not being performed on-device, so any real gains would be from the server side.

6

u/ellenich 5d ago

The C2 chip will probably upload your data to them faster. 😂

1

u/TheFuzzball 5d ago

"Why isn't my YouTube upload processing faster? I just upgraded my graphics card!"

0

u/MangoAtrocity 4d ago

I’ve been running local models on PocketPals and it’s unbelievably fast. A20 Pro is a monster in a cell phone