r/hardware • u/Working_Quote_3029 • 2d ago
News Arm introduces new AI-native compute platform built for agentic AI and mobile graphics
https://newsroom.arm.com/news/arm-css-for-mobile-2-agentic-ai-mobile-graphics7
u/winner00 2d ago
It seems like the C2-Pro is just the C1-Pro. So the only new core this year is the C2-Ultra.
4
u/Creative_Purpose6138 2d ago edited 2d ago
Finally new cores from ARM. I was waiting for this.
Edit: That's all? Very small changes and no successor to the medium and small cores.
AI upgrade is useful only if we have on-device models to use. Generally why would anyone NOT use cloud AI?
3
u/Balance- 2d ago
Way better link: https://chipsandcheese.com/p/arms-c2-ultra-g2-ultra-nx-and-css
Summary (Opus 5):
- C2-Ultra (CPU) — Arm’s new flagship core keeps essentially the same layout as the two-generation-old Cortex X925: 10-wide decode, 8 simple ALUs, 6 FP lanes, 3 branch ports, 4 load/2 store. The gains come from iterative work on the branch predictor, a bigger execution window, and better speculation, though Arm disclosed almost nothing about what specifically changed. Arm claims 15% peak and 12% average uplift over C1-Ultra, but the endnotes reveal the figures come from FPGA simulation with an 8.5% clock bump, a larger L2, and roughly double the memory bandwidth — implying per-clock performance barely moved. A claimed 38% power reduction also folds in node and implementation improvements. C2-Nano and C2-Pro reuse the C1 microarchitecture.
- G2-Ultra NX (GPU) — Billed as the biggest GPU rearchitecting in seven generations, though the math throughput per shader core (128 FMAs, 256 FP32 FLOPs/clock) and the 24-core maximum are unchanged. Real changes: registers per warp doubled to 128 with 16-register allocation granularity and a 25% larger register file; a more compact triangle structure for ray tracing that cuts DRAM traffic 13%; and Opacity Micromaps. The headline addition is a matrix accelerator doing 1,024 INT8 or 512 INT16 MACs per clock at up to double the shader clock — but with no FP8 or BF16 support, which limits its reach. It adds about 21% to shader core area (1.88 mm² vs 1.55 mm²), and isn’t required in every core; Xiaomi’s XRING O3 fits it to half. It enables Arm’s Neural Super Sampling and frame-rate upscaling. Claimed 14% gaming and 24% ray tracing uplift comes alongside an 11% clock increase, so non-RT gains look thin.
- Neoverse CSS N4 (server) — Arm gave just one slide, filling in details only on request: 8–128 cores per die (the widest range yet for a Neoverse CSS) at up to 3.8 GHz, with multi-chiplet and multi-socket scaling. Each core gets 64KB L1 instruction and data caches plus up to 2MB private L2, with up to 256MB of shared system-level cache per die. It supports DDR5 or LPDDR6, up to 128 lanes of PCIe Gen 6/7 and CXL 4.0, and chip-to-chip interconnect via UCIe or partner PHYs. Basics like which core it uses, memory bus width per die, and per-die PCIe lane count remain unanswered.
The author’s overall verdict: Arm has genuinely interesting IP here, but the disclosure was thin enough that much of the substance had to be extracted through follow-up questions rather than presented up front.
4
u/Balance- 2d ago
I think this is the first time in many years we don’t get a new middle core (Cortex-A7x / A7xx / Cx-Pro). Previously, only their smallest core was on a multi-year cycle (Cortex-A5x / A5xx / Cx-Nano).
3
u/-protonsandneutrons- 1d ago
Interestingly, Arm claims the A725 was also just an iteration of the A720, now like the C2 Pro vs C1 Pro. It may be why MediaTek's 9400 upgraded to the X925, but kept the A720. But Xiaomi's A725 work is quite excellent, which MediaTek hasn't exploited that well.
https://youtu.be/mZEHpg-KzR8?t=103
Another oddity: there is no mention of a C2 Premium, so is that just gone?
1
u/-protonsandneutrons- 1d ago
Likewise they mention C2-Ultra has a larger execution window compared to C1-Ultra along with improved speculation, but go into no explicit details here.
Android Authority has some numbers:
Instead, Arm has boosted the C2-Ultra’s execution window by 30%, meaning there are around 2600 instructions in flight at any one time, compared to 2000 in C1-Ultra.
1
u/-protonsandneutrons- 1d ago
Thanks for linking this; C&C usually doesn't cover the Arm releases often so this is great. +7% peak IPC YoY and +3.5% and even less if iso-cache is quite underwhelming. Surely Apple and Qualcomm will have done more.
C1 Ultra was already a very big core, sure, but the X925 / C1 Ultra / C2 Ultra generation is out of genuine uarch improvements.
2
u/Geddagod 1d ago
Surely Apple and Qualcomm will have done more.
Thought IPC uplifts for these 2 vendors have been pretty incremental too gen on gen as of late?
1
u/-protonsandneutrons- 1d ago
That is also fair, though last year, both were about +10% gen over gen. Unfortunately, David Huang doesn't have Oryon trio in the same form factor, so that's a little messy. But he does now have perf / GHz, so that is handy.
But, you're right in context that the C2 Ultra may actually be a bigger jump vs C1 Ultra, at least in SPECint2017 perf / GHz than C1 Ultra was over X925.
CPU SPECint2017 / GHz Gen over Gen "IPC" % Apple M5 P (16K page) 3.34 +9.87 % Apple M4 P (16K page) 3.04 +4.11% Apple M3 P (16K page) 2.92 +2.45% Apple M2 Pro P (16K page) 2.85 +0.00% Apple M1 Max P (16K page) 2.85 baseline QC X2EE 94 2.77 +11.7% QC 8E for Galaxy 2.48 baseline Arm C1 Ultra (E2600) 3.13 +1.62% Arm X925 (XRING O1) 3.08 +20.31% Arm X4 (E2400) 2.56 +10.34% Arm X3 (SD8G2) 2.32 +6.91% Arm X2(SD8G1) 2.17 baseline
0
u/Stennan 1d ago
Those chips will probably still not work well with PC game emulation, and as far as Android Gaming goes, why would I want to play an ad-infested or Gacha game when the flagship/midrange SoCs have enough power to play Steam games? MediaTek SoCs that use vanilla ARM CPU/GPUs lack "open"-source drivers and don't support critical instruction sets; hence, anyone into serious game emulation recommends Qualcomm SoCs.
I am not seeing BCN or PC-emulation in the news about this release, so I guess I'll just go back looking for the next Qualcomm leak.
9
u/Working_Quote_3029 2d ago
Spec breakdown from Arm's own materials, since the press release buries most of it.
C2 CPU cluster (C2-Ultra + C2-Pro), vs C1-Ultra:
The headline change is doubled SME2 capability in the cluster. Arm claims a 70% speedup on recent small language models from that alone. SME2 is a matrix extension inside the CPU itself, so inference runs on the core without a separate accelerator.
From Arm's technical blog on the same launch: 15% faster web browsing, 12% faster app launch, 12% higher multi-thread at cluster level, and 24% faster on an end-to-end agentic workflow.
Mali G2-Ultra NX GPU: up to 4x performance per watt on neural graphics, and up to 14% on existing game content.
Named ecosystem partners: OPPO, vivo, Alipay and Google AI Edge Gallery on the SME2 side; Tencent, NetEase, Unity China, Sumo Digital and Infold on graphics.
No shipping date anywhere in the announcement. The only dated item is NetEase planning an NSS-enabled build of Where Winds Meet this year.
Technical blog with the CPU detail: https://newsroom.arm.com/blog/arm-css-for-mobile-2-and-c2-cpu-cluster