NVIDIA has shared architectural and performance details of its custom Armv9.2 'Vera' CPU, an 88-core, 176-thread monolithic chip built around its new 'Olympus' core design. NVIDIA positions Vera for data analysis, agentic AI, and large-scale software workloads, claiming significant per-core gains over AMD's EPYC 'Turin' 9755 in select benchmarks.
NVIDIA has released detailed architectural and performance information on its custom Armv9.2 'Vera' CPU, a chip the company describes as designed to operate either independently or within its NVL rackscale systems. According to TechPowerUp, the CPU is built around NVIDIA's new 'Olympus' core, which the company divides into four main areas: front end, mid-core, execution engine, and cache subsystem. The design targets data analysis, high-performance single-threaded tasks, agentic AI, and large-scale software operations.
Key architectural features include a 10-wide decode engine, a neural branch predictor capable of handling up to two taken branches per cycle, and a deep out-of-order execution pipeline meant to keep resources busy during control-flow-heavy workloads. NVIDIA also implemented a form of simultaneous multithreading that partitions core resources between threads to avoid the 'noisy neighbor' contention seen in some traditional x86 SMT implementations, according to TechPowerUp. The chip connects via a Scalable Coherency Fabric offering 3.4 TB/s of bandwidth and a 164 MB unified L3 cache, with external memory support for up to 1.5 TB of LPDDR5X running at 1.2 TB/s aggregate bandwidth.
NVIDIA compared Vera against AMD's EPYC 'Turin' 9755, a chiplet-based x86 chip, across memory bandwidth/latency, application workload performance, core IPC, and agentic AI/RL scenarios. Per Wccftech, NVIDIA claims up to a 1.9x IPC uplift over Zen 5 in workloads with dense control flow and long dependency chains, up to 2.3x faster branch prediction, and up to 4.3x faster backend ops per cycle in select tests. NVIDIA also reported Vera delivering roughly 4x higher per-core memory bandwidth than EPYC Turin (about 12.7 GB/s per core versus 3.1 GB/s), lower and more consistent core-to-core latency due to its monolithic design versus EPYC's chiplet architecture, and workload-specific gains such as up to 1.8x in Python-heavy SPEC CPU 2006 tests, 2.6x in graph traversal, and a 20% uplift in ClickHouse analytics workloads.
NVIDIA frames Vera's monolithic design as a direct contrast to the chiplet approach used in modern x86 server CPUs, arguing that eliminating die-to-die hops reduces latency and bandwidth bottlenecks that become more pronounced under heavy, irregular memory access patterns typical of agentic AI. These are NVIDIA's own preliminary performance disclosures against AMD EPYC Turin, and independent third-party benchmarking of production Vera silicon was not part of the provided source material.
What We Know
| Spec | Detail |
|---|---|
| Architecture | Custom Armv9.2 'Olympus' core |
| Cores / Threads | 88 cores / 176 threads |
| Decode Width | 10-wide decode engine |
| Branch Prediction | Neural branch predictor, up to 2 taken branches per cycle |
| Memory Support | Up to 1.5 TB of LPDDR5X (SOCAMM2) per CPU |
| Memory Bandwidth | 1.2 TB/s aggregate memory bandwidth |
| Fabric | Scalable Coherency Fabric with 3.4 TB/s bandwidth |
| L3 Cache | 164 MB unified L3 cache across the die |
Frequently Asked Questions
How many cores and threads does the NVIDIA Vera CPU have?
The Vera CPU has 88 cores and 176 threads, according to TechPowerUp.
How much memory can the NVIDIA Vera CPU support?
NVIDIA states the Vera CPU can support up to 1.5 TB of LPDDR5X memory via SOCAMM2 modules, delivering 1.2 TB/s of aggregate memory bandwidth.
What CPU core architecture does Vera use?
Vera is built on NVIDIA's custom Armv9.2 'Olympus' core, featuring a 10-wide decode engine, a neural branch predictor, and a deep cache hierarchy with multiple prefetch engines.
How does Vera compare to AMD's EPYC Turin in NVIDIA's benchmarks?
NVIDIA claims Vera delivers up to a 1.9x IPC uplift, up to 2.3x faster branch prediction, and roughly 4x higher per-core memory bandwidth compared to AMD's EPYC Turin 9755, along with lower and more consistent core-to-core latency due to its monolithic design, per Wccftech.
