It’s not just Apple’s new A20 Pro mobile chip. Arm, the company designing the instruction set used by mobile SoCs as well as full core designs licensed by companies such as MediaTek, Nvidia, and others, is now introducing a new generation of CPU architectures as well. Arm replaced the Cortex cores with the Lumex C1 series last year, and now follows up with a new C2 generation that pushes the single-threaded performance further.
The new generation of cores forms a platform called Arm CSS for Mobile 2, consisting of complete blocks for processor design, specifically CPU cores forming the Arm C2 series and the Arm Mali-G2 GPU (CSS stands for Computer Subsystem; the first generation was last year’s platform with Lumex C1 cores).
It appears that after just one generation (the C1), Arm has dropped the Lumex branding it introduced last year as a replacement for the Cortex designation that had been used for its cores for years prior to that. The company apparently decided that it was not exactly a marketing success or an iconic-sounding brand; it looks like the new core series is now simply called Arm C2 rather than Arm Lumex C2.
C2-Ultra: The Flagship Core
The new flagship of Arm’s mobile processors and the company’s most powerful core will be the C2-Ultra, which is intended to serve as the “prime” core in mobile chips, with one or two of these cores used to achieve high single-threaded performance (as was previously the case with the Cortex-X1 through X925 core lines). However, it’s possible that SoCs intended for notebooks and PCs could integrate a larger number of them. These cores prioritize single-threaded performance, potentially at the expense of efficiency, and therefore occupy a relatively large area of the chip.
The core is designed to decode 10 Arm instructions per cycle. The fetch stage should also be able to read 40 bytes per cycle, corresponding to 10 instructions, while the rename and dispatch stages also send on up to 10 instructions per cycle for further processing.
The processor’s execution backend contains 8 ALUs (six of them are simple units for less demanding instructions and two are complex units capable of executing all instructions), plus three separate units for processing branch instructions. The vector/FPU section has six FMA units with SVE2 support, apparently still 128 bits wide. There are four Load/Store units in total, supporting up to four Load instructions per cycle, but only two Store instructions (and two Load instructions alongside them) per cycle.
These numbers of execution units are the same as in the Cortex-X925 and C1 Ultra cores—it appears that Arm increased core performance through other changes. The branch predictor has reportedly been improved in particular, with better accuracy and faster recovery from mispredictions.
The core has also received deeper queues and buffers for out-of-order processing. The processor’s Reorder Buffer should apparently have a depth of 768 entries, and the processor has 320 physical general-purpose registers and 448 physical registers for FPU and SIMD instructions for out-of-order code processing. The queue for Load instructions (memory reads) has 28 entries, while the queue for Store instructions (memory writes) has 60 entries.
The prefetcher (i.e. speculative loading of data from memory before the core needs it) is also supposed to be improved. Along with branch prediction, this is one of the two critical areas where every new CPU generation tries to improve.
The core has a 64 KB L1 cache for instructions and a 128 KB L1 cache for data (both are 4-way associative). The L2 cache is private and can be 2 MB (with 8-way associativity) or 3 MB (with 12-way associativity), depending on the manufacturer’s choice. The 3 MB L2 cache version will presumably offer higher performance at the expense of more chip area, and the benchmarks cited below most likely use this variant. The L3 cache is then shared at the level of the entire CPU cluster within the interconnect logic.
12% Increase in Single-Threaded Performance
According to Arm, the C2-Ultra core can achieve up to +15 % higher single-threaded performance than last year’s Lumex C1-Ultra core. Web browsing performance is also supposed to be 15% higher, application startup is 12% faster, and multi-threaded performance is supposed to improve by the same percentage (this figure applies to processors that would have 2 C2-Ultra cores with 6 C2-Pro mid cores and two C1-SME2 matrix units, versus a processor with 2× C1-Ultra, 6× C1-Pro, and only one C1-SME2 unit). Arm shows a slide according to which the 15% single-threaded performance improvement is achieved in some tests such as Geekbench 6.3 or Speedometer 3.1, while the increase is lower in others. On average, the new core is supposed to deliver +12 % higher single-threaded performance.
If the core were configured at a clock speed and voltage at which it delivered the same performance as the previous C1-Ultra core at its maximum clock speed, the C2-Ultra would consume 38% less power according to Arm. However, similar efficiency figures are not particularly meaningful in practice, since instead the new core will be used at a higher clock speed, where it delivers higher performance but has lower efficiency (so its improvement over the previous generation will be less pronounced).
IPC Improved by Up to 6 or 7%
This performance increase is a combination of architectural improvements that increase performance per 1 MHz of clock speed (so-called IPC), as well as increased clock speeds achieved either through architectural improvements or a new manufacturing process. Chips with these cores are already expected to typically be manufactured using TSMC’s 2nm process. Arm made these comparisons between a processor with C1-Ultra cores (with 2 MB L2 cache) at 4.1 GHz and C1-Pro at 3.5 GHz (with a C1-SME2 unit at 2.0 GHz), while for the new generation it models the C2-Ultra core at 4.45 GHz (with 3 MB L3 cache), C2-Pro at 3.6 GHz, and the C1-SME2 units increased to 3.0 GHz.
If the effect of clock speed is removed, the new C2-Ultra core comes out with “up to” +6 % higher IPC than the C2 Ultra core, if we base this on the figure of up to 15% higher single-threaded performance. Arm itself reportedly officially states that the core should have up to +7 % better IPC, presumably based on a somewhat different set of tests or a different averaging method. If we use the lower figure of a 12% average increase in single-threaded performance, however, the average IPC improvement would work out to only about +3.2 % (approximately).
Catching-up with the Single-Threaded Performance Leaders?
The Lumex C1-Ultra core ran at up to 4.2 GHz in the MediaTek Dimensity 9500 chip and achieves roughly 3300 to 3500 points in the single-threaded Geekbench test (however, there can be considerable variability here depending on the conditions and phone). If its performance could be improved by 15% over this result, the single-threaded score of chips with C2-Ultra could reach roughly 3800–4200 points.
This would not put C2-Ultra ahead of the Apple A20 Pro or M6 (which are reportedly capable of reaching up to 4700 points), but it could directly compete with AMD’s upcoming Zen 6 cores, although we have yet to see how those turn out. If Zen 6 performance ends up being relatively lower or, for example, the core does not reach particularly high clock speeds, Arm could have higher single-threaded performance. Or at least when we just look at Geekbench 6.x for rough orientation (and it is also worth keeping in mind that Geekbench scores in Windows, which we usually look at for AMD processors, tend to be somewhat lower than scores given by the iOS and Android/Linux versions of the benchmark).
C2-Pro and C2-Nano: The Mid and Small Cores Are a Refresh
Mobile hybrid SoCs will continue to combine different types of cores. In high-end chips for more expensive phones, Arm expects two “prime” C2-Ultra cores and six C2-Pro middle cores will typically be used. Cheaper processors may have just one C2-Ultra core, three C2-Pro cores, and four little C2-Nano cores—or four C2-Pro cores and four C2-Nano cores without a prime core. Even cheaper processors could then have a 2+6 or 2+4 combination. The very cheapest processors may be based solely on two or four little C2-Nano cores, without any middle cores.
Arm has not provided details about the C2-Pro and C2-Nano cores, however. According to Chips and Cheese, which analyzed Arm’s information and communicated with the company, these cores should be refreshes of the Lumex C1-Pro and Lumex C1-Nano cores from the previous generation. There will apparently only be smaller changes and modifications, so they can be expected to have similar IPC to their predecessors. Their performance will therefore be increased by higher clock speeds made possible by the chips being manufactured using better and newer manufacturing technologies (especially in the case of TSMC’s 2nm process).
SME2 Units Have Not Received a New Generation This Time
Alongside these cores, CPU clusters in processors of the CSS for Mobile 2 generation should also be paired with SME2 units, which are shared accelerators for matrix calculations located outside the cores and are already used by both Arm and Apple’s architectures. However, these are not receiving a new architecture at this time; the new processors will still integrate units from the previous generation (C1-SME2), although the reference processor design uses two units per cluster instead of one, providing a significant increase in the total performance. The improvements in AI performance using SME2 that Arm advertises in its materials are therefore due to this doubling of the accelerators (and also apparently the aforementioned 50% increase in their clock speed).
The New GPU Also Focuses on AI Acceleration
Arm does not provide many details about the new Arm Mali-G2 GPU architecture either. However, it revealed that it has paid particular attention to AI application performance. The most powerful version of the Mali G2-Ultra NX architecture adds units for AI acceleration directly at the GPU level, similar to what is found in AMD and Nvidia architectures (the tensor cores in GPUs) and, for the last two generations, in Apple’s GPUs as well. This acceleration is supposed to improve the energy efficiency of neural network execution on the GPU (performance per 1 W) by up to 4×.
This should not be the only new feature, however, as Arm has reportedly also revamped the RTU units for ray-tracing calculations and the general-purpose compute units. Gaming graphics capabilities should therefore hopefully improve as well. In traditional graphics applications, the Mali G2-Ultra NX is reportedly set to deliver a 14% improvement.
First Actual Silicon in the Coming Months
These architectures will appear in new generations of mobile SoCs from manufacturers of phone chips—except for those that develop their own core architectures (Apple and Qualcomm, currently). The C2 series cores will therefore primarily be used by MediaTek and Samsung. Their new chips (and phones) will presumably arrive toward the end of 2026 and in 2027.
It is not clear whether Nvidia will use them—its RTX Spark processors currently use Cortex-X925 and Cortex-A725 cores licensed from Arm, which are two generations older than the newly introduced C2 series. Nvidia also has a team developing in-house CPU cores, which it is now introducing in Vera server processors, so future generations of RTX Spark processors may no longer need to use any Lumex C1 or C2 cores from Arm.
- Read More: 50% better than x86? Nvidia introduces Vera processor with Olympus cores
- Read More: Nvidia’s plans with Spark CPUs: Next generation, lower-cost chips
- Read More: Nvidia confirms next-gen processors with new Rigel cores
Sources: Arm, Chips and Cheese, ComputerBase
English translation and edit by Jozef Dudáš
⠀
















