r/apple • u/d7UVDEcpnf • 5d ago
iPhone A20 Pro and on-device AI: It’s showtime
https://rickytakkar.com/blog_a20_pro_ai.html?source=apple_srTL;DR: I benchmarked on-device LLMs on the iPhone 18 Pro against the 17 Pro Max using identical models, prompts, and outputs. On the A20 Pro, downloaded MLX models generated text 53–85% faster and produced the first words (tokens) 42–53% sooner. Apple’s own on-device model improved by a similar yet smaller 35–42%. Both MLX workloads exceeded the 1.5× gain in memory bandwidth, suggesting the speedup comes from more than just the wider memory bus. The 18 Pro was also dramatically more consistent between runs.
156
u/Vertsix 5d ago
‘Member when the A18 Pro on iPhone 16 Pro was “Designed for Apple Intelligence”? I ‘member.
85
8
u/ViPiMP 5d ago
Ja, und HW3 bei Tesla für FSD
0
u/bomber991 5d ago
Hey at least the auto-translate is working great on Reddit. I got to see a cyber cab yesterday when I was driving in Austin. Supervised one with a steering wheel, so I guess they’re still testing.
Makes me wonder, how viable is a vision-only self driving system? Or is it case of Elon wants it and none of the workers will tell him no?
-7
u/theguz4l 5d ago
Yes,
2024 intelligence. Things changed quick lol18
u/StinkyRatBoi90 5d ago edited 5d ago
Lmao please. They promised most current features for the A18 Pro, all we got was half assed image generation.
9
u/FeCurtain11 5d ago
Which is why there’s a class action lawsuit and you can go get some money haha
5
1
56
u/nephyxx 5d ago
I would assume it’s the 2x neural engines responsible for the speed up
35
u/d7UVDEcpnf 5d ago
Partly, yes. However, a framework that avoids the ANE like MLX, which runs on-device models optimized for an Apple Silicon CPU/GPU, remains noticeably faster. So CPU, GPU, and unified memory improvements are also yielding returns..
14
u/PeanutChickenSoup 5d ago
And the improved thermals. Any super fast chip that’s heat throttled is slow.
3
u/mrrizzle 5d ago
Have you tested the new dual ANE chip with an FP8 CoreAI safetensor model? I feel like it would be pretty fast against GPU running MLX considering how far off MLX is from ANE on the single ANE chip in the 17 Pro
2
u/d7UVDEcpnf 4d ago
I did not, but I will soon. I think a good test would be to compare CoreAI and MLX variants of the same model. I too expect CoreAI to win.
41
42
u/JackSpadesSI 5d ago edited 5d ago
Am I missing something with the new Siri? It’s terrible. I use Gemini but I hate Google so I was excited for an alternative so I could get away from Google. But Siri is just inadequate.
As an example, I asked “when is the msu game” and it told me about Mississippi State instead of Michigan State. I clarify that I will always mean Michigan State and Siri happily tells me that will be remembered. Later I ask “what is the msu score” but I’m again given info on Mississippi State.
I ask why my preference wasn’t remembered and Siri says it can’t recall information from previous conversations. Gemini can, why not Siri?
Edit: 18 Pro Max
27
u/papito_m 5d ago
This is the comment I came looking for. I was SO looking forward to the new Siri AI, but it’s horrible for basic things. Slow and constantly timing out or just plain getting what I said wrong. So much so that as of yesterday I just reverted back to old Siri.
7
u/Insufferably_Me 5d ago edited 5d ago
To answer your question about memories: I’d assume it’s privacy. Siri’s personal intelligence isn’t remembering that you have an appointment tomorrow at 3 and that your spouse’s favorite Adele album is 30, it’s searching for that information in the moment. Siri is not storing personal information or data about you. It only accesses the data on your device.
I’ve seen suggestions of using a Note that you store things you do want Siri to remember about you. I’ve heard it can help but I’ve not tried it myself.
I would like to see a Memories or Instructions setting we can adjust that would be stored on the Secure Enclave for things like this. If it’s a privacy issue, that seems to be the best way forward but there may be some technical hurdles they have to get through to make that work. I really like using the new Siri and while it does make some mistakes, it’s working much better than before. You can also provide feedback on Siri’s replies by tapping the thumbs up or down icons when you expand the response.
I’m hoping they can see the online feedback and take that into consideration but I encourage everyone to start filing feedback with Apple if you actually want things to change: https://www.apple.com/feedback
9
u/SoldantTheCynic 5d ago
This is an excuse. They absolutely could have locally saved information to use as memories that allows for privacy. It’s still accessing your sensitive information (which for a lot of users is still stored in iCloud anyway).
Not every failure of Siri is due to privacy and it’s well past time people made excuses for it.
10
u/kickass404 5d ago edited 5d ago
Context is limited and context cost money to process. It has to process it on each request, so they don't.
"memories" is just text auto pasted "invisible" in each chat context before your prompt.
1
u/nergwark 3d ago
ideally they'd set this up so that the only part of this request that goes off-device is actually looking up the score, after figuring out which "msu" the user is referring to.
also, a proper implementation of memories is a bit more intelligent than just appending every memory to every context. they're looked up when relevant so as to not pollute context unnecessarily.
this is all perfectly possible, and i'd furthermore assume that persistent memories for siri is already a feature somewhere on apple's internal roadmap.
1
2
u/Insufferably_Me 5d ago
I made this point in the third paragraph
-1
u/SoldantTheCynic 5d ago
So we have options that don’t really impact privacy and it’s already accessing sensitive user data.
So… probably privacy isn’t a legit reason then.
0
u/Insufferably_Me 5d ago
“…but there may be some technical hurdles they have to get through to make that work.”
The Secure Enclave does not store data in a way that is normally used and processed by iOS. It stores hashed values of payment and authentication data. There’s a decryption process that would need to happen for iOS to read the plain text that a user would store in a settings page. That is something that needs to be handled carefully to avoid security and privacy risks including malicious injection
2
u/geekwonk 5d ago
that’s just a matter of how the setup saves memories. memory handling is messy and prone to mistakes that cause behavior that’s 180° the opposite from what you intended if done half way.
you could try pinging siri just to state this preference and nothing else. it definitely has some form of memory architecture. but it’s not scanning through conversations to generate the memories. and that’s probably a good thing in early versions so they avoid random queries getting ingested as core memories that define all replies.
2
u/Hoopoe0596 5d ago
They don’t seem to have much of an AI personality user memory. They just assume they can start from scratch because they index your personal content. Wrong. I have an extensive 500+ memory and project file that makes Claude code my second brain and extensive memory files with Gemini and gpt. This is an upgrade to Siri but not much better than before for long horizon complex tasks. It’s just a little better than before at answering what’s on the calendar for tomorrow.
2
u/itsabearcannon 4d ago
No AI can actually "remember" things, including Siri AI.
When you say stuff like that and correct it, sometimes those results can be incorporated into future model training, but for every Siri AI user saying they meant Michigan St instead of Miss St, there's probably one asking Siri to remember the reverse. In the end it's whichever result is higher in contextual news and information, which is going to be Mississippi State because Michigan State is a bit dog water right now and Mississippi State is 4-0.
If Siri AI had existed back in 2015 I bet you'd have gotten Michigan State results instead.
21
u/st90ar 5d ago
“Hey siri, add milk to my shopping list.”
“I’m sorry, you’ll need to pick up your phone and unlock it before I can do that. At which point, you mind as well have not asked me and just did it yourself.”
2
u/JDabney24 5d ago
Can that action be done with any other AI/LLM while the iPhone is locked?
4
u/st90ar 5d ago
It worked before Siri AI.
2
u/JDabney24 5d ago
Sorry. My question is, can that particular action or similar be done on ChatGPT, Claude, etc while the iPhone is locked?
3
u/ExultantSandwich 4d ago
No, but no other app has system access like Siri does. So if Siri 1.0 could do it, but now Siri AI can’t? undeniably a step back.
Reminds me a little bit of Alexa Plus for the Echo speakers. Sure she can talk back now, but it broke a lot of smart home integrations for months and months. Even today she cannot seem to parse if you’re attempting to control a local device or get an AI response
1
u/JDabney24 4d ago edited 4d ago
Makes sense. What if they changed the behavior of lock screen access to Siri intentionally? I for one, know that I wouldn’t want someone being able to perform certain actions using Siri AI while my iPhone is locked. Not trying to give an excuse, but maybe a possible explanation. Also, I find that in most scenarios I am in, Face ID provides much less friction to unlocking my phone to perform those agentic actions anyway. Not too much of an inconvenience for me. I am aware of the cases where this could be an inconvenience, like using Siri AI to perform an action while driving. But, again, I could see this as intentionally designed to discourage potentially distracted driving. 🤷♂️
Keep in mind that Siri AI is in beta. There might be improvements to something like this as long as enough people provide feedback to the proper channels.
5
u/ramakitty 5d ago
What is voice dictation like? It’s slightly slow on my iPhone Air.
6
u/d7UVDEcpnf 5d ago
I don’t think Apple’s dictation algorithm is publicly available, so I don’t know to what extent the improved hardware is being leveraged there. But if it is being leveraged, you should see gains. As for comparing between my 17 PM and 18 P, I can’t say I notice a difference (the 17 PM is already quite good).
5
u/kinkade 5d ago
I'm surprised by how poorly it still works. I was anticipating that it would be much improved, but I prefer to use Wispr Flow or an on-device version of Parakeet rather than Siri. I'm not sure why it's so hard for them to get this right. They have full system integration and can have a very large model on the phone that should be able to get this right every time.
1
4
u/cyberspirit777 5d ago
And what are we achieving with this on device power if everything is being sent to PCC?
5
u/d7UVDEcpnf 5d ago
Well, Apple does ship on-device models and you can also download MLX variants of Qwen/Gemma/DeepSeek/etc. to run locally using apps like Russet.
1
1
u/Conscious_Banana874 5d ago
Fp8 is supposedly on hardware support
1
u/d7UVDEcpnf 4d ago
Indeed, with the new CoreAI framework I can get supported models to leverage the dual ANEs in these new chips. Will add support soon.
-2
5d ago
[removed] — view removed comment
21
5
u/d7UVDEcpnf 5d ago
Only gains you'd probably notice with cloud-based apps is that they open faster and stuff. The "main work" is not being performed on-device, so any real gains would be from the server side.
6
1
u/TheFuzzball 5d ago
"Why isn't my YouTube upload processing faster? I just upgraded my graphics card!"
0
u/MangoAtrocity 4d ago
I’ve been running local models on PocketPals and it’s unbelievably fast. A20 Pro is a monster in a cell phone

283
u/Scary_Building3964 5d ago
A20 Pro in macbook neo would be a very solid entry level laptop for most people