Arm has announced its next-generation Neoverse CSS N4 platform, bringing a massive upgrade in semi-custom compute designs to cloud and data center markets.

Built on TSMC’s advanced N3P process, the new platform delivers up to 128 cores per die running at speeds up to 3.8 GHz, effectively doubling the core density of its predecessor.

The Compute Subsystem (CSS) serves as a pre-configured, validated semi-custom program that enables cloud providers and chipmakers to rapidly design tailor-made silicon. Customers can fine-tune core counts, cache sizes, memory protocols, and I/O connectivity for specialized applications. CSS IP already powers high-profile enterprise hardware, ranging from cloud CPUs at Microsoft Azure and Google Cloud to data processing units (DPUs) manufactured by NVIDIA Corp. and Intel Corp.

Engineered for efficiency, the Neoverse CSS N4 platform scales beyond single-die limits through multi-chiplet and multi-socket designs using standard UCIe interconnects. Each die supports up to 256 MB of L3 cache alongside 2 MB of L2 cache per core. On the memory and I/O front, N4 offers support for high-speed DDR5 or next-generation LPDDR6, backed by up to 128 lanes of PCIe 7/6 and CXL 4.0 connectivity.

Compared directly to the previous Neoverse CSS N2 platform — which capped at 64 cores and PCIe 5.0 — the N4 architecture represents a leap in cloud infrastructure performance. In a standard 128-core configuration running at 3.0 GHz, Arm estimates the N4 achieves twice the socket performance of Neoverse N3, a 1.25-fold boost in performance-per-watt efficiency, and a 1.75-fold increase in memory bandwidth.

Historically, Arm’s energy-efficient N-series cores have powered specialized accelerators or cloud workloads like Microsoft’s Azure Cobalt 100, while the maximum-performance V-series cores anchor heavy-duty server processors like AWS Graviton and Google Axion. Codenamed Dionysus, the N4 IP is aimed at companies building next-generation infrastructure, though silicon in consumer or enterprise server hardware is still months away from commercial deployment.

Simultaneously, Arm expanded disclosures regarding its high-performance AGI CPU platform. Built on Neoverse V3 cores using a 3-nanometer manufacturing process, the dual-die enterprise CPU scales up to 136 cores and 272 MB of L3 cache at clock rates reaching 3.7 GHz.

Unlike traditional x86 server chips from AMD Inc. and Intel, the AGI CPU integrates memory and I/O directly onto the compute die, dropping memory latency below 100 nanoseconds while supporting up to 6 TB of DDR5-8800 memory capacity.

Arm revealed that Oracle Corp. and ByteDance have signed on to deploy its AGI processor designs, joining a customer roster that already includes Meta Platforms Inc., OpenAI, Cloudflare Inc., Lenovo, and SAP.

While real-world benchmarking data remains sparse, Arm projects its flagship silicon will deliver more than double the rack-level performance of rival current-generation x86 server systems.

Even as Arm ventures into fully realization-ready silicon designs, company leadership emphasizes that licensing customizable IP to hyperscale cloud partners remains the core engine of its data center business.