El Capitan is an HPE Cray EX255a installed at the Lawrence Livermore National Laboratory (LLNL), in California. It was designed for the National Nuclear Security Administration (NNSA) and combines exascale computing, artificial intelligence, multiphysics simulation, parallel storage, and liquid cooling in a single industrial-scale platform.
In November 2024, the system debuted on the TOP500 with 1.742 exaFLOP/s. A new measurement raised the result to 1.809 exaFLOP/s in November 2025. On the June 2026 list, El Capitan moved to second place globally, maintaining 1.809 exaFLOP/s and remaining one of the most important exascale architecture references on the planet.
| Indicator | Value |
|---|---|
| HPL Rmax (Jun/2026) | 1.809 EF/s |
| Rpeak TOP500 | 2.821 EF/s |
| Measured HPL power | 29.685 MW |
| HPL efficiency | 60.94 GF/W |
Engineering point: the value of 29.685 MW from TOP500 is the power measured during the HPL benchmark. LLNL publishes ~35-36 MW as the order of magnitude of the system's peak, while the ECFM modernization raised the electrical capacity of the complex to 85 MW. These three numbers represent different layers of the infrastructure.
Public scope: El Capitan supports classified missions. This material exclusively uses public technical information. Power and cooling diagrams represent the disclosed functional architecture; they do not reproduce security, redundancy, or protection details that are not public.
What does it mean to operate at exascale?
Exascale is the class of systems capable of performing at least 10^18 floating-point operations per second on reference workloads. This frontier changes the scale of scientific problems: more spatial resolution, more variables, more scenarios, better uncertainty quantification, and greater use of artificial intelligence techniques integrated into simulation.
HPL measures dense linear algebra performance and remains the main TOP500 metric. HPCG seeks to reflect memory access and communication patterns closer to scientific applications. In June 2026, El Capitan recorded 17.406 PF/s on HPCG. On HPL-MxP, which explores mixed precision and approximates AI workloads, the system reached 16.7 exaFLOP/s in a measurement published by TOP500.
| Date | HPL Rmax (EF/s) |
|---|---|
| Nov/2024 | 1.742 |
| Jun/2025 | 1.742 |
| Nov/2025 | 1.809 |
| Jun/2026 | 1.809 |
As of June 2026, El Capitan holds 2nd place on the TOP500.
General architecture: HPE Cray EX255a + AMD MI300A
The architecture was designed as an integrated system. Compute, fabric, storage, and cooling are not independent subsystems; final performance depends on the balance between all of them. The HPE Cray EX was conceived for high-density workloads, with direct liquid cooling and an integrated Slingshot network.
| Physical scale | Value |
|---|---|
| Compute cabinets | 90 |
| Total nodes | 11,520 |
| MI300A APUs | 46,080 |
| Rabbit modules | 720 |
AMD Instinct MI300A: CPU and GPU in the same package
The MI300A is an APU for HPC: it integrates a Zen 4 CPU, a CDNA 3 GPU, and shared HBM3 memory in the same package. This unified memory reduces the need to copy data between physically separate CPU and GPU memories — an operation that, in discrete architectures, adds latency, power consumption, and programming complexity.
| Component | Per MI300A |
|---|---|
| CPU | 24 Zen 4 x86 cores |
| GPU | 6 CDNA 3 XCDs |
| Memory | 128 GiB HBM3 |
| HBM3 bandwidth | ~5.3 TB/s theoretical |
| Vector FP64 | 61.3 TFLOP/s |
| Coherency | CPU/GPU with shared memory and Infinity Fabric |
Anatomy of the compute node
Each compute node has four MI300A units. That means 96 Zen 4 CPU cores, four accelerated complexes with HBM3 memory, and 512 GiB of global memory per node. LLNL itself publishes 250.8 TFLOP/s of peak FP64 per node.
The node's external connection uses four HPE Slingshot 200 Gb/s interfaces. In aggregate terms, network injection reaches 100 GB/s per node. In a machine with more than eleven thousand nodes, the fabric stops being an auxiliary network and becomes part of the computer itself.
Why this matters for AI: distributed training and accelerated simulation are limited by data movement time. A unified architecture reduces internal CPU-GPU traffic; a high-bandwidth fabric reduces wait time between nodes.
Scale: 11,520 nodes and 46,080 APUs
The total configuration published by LLNL includes 32 login nodes, 11,424 batch nodes, and 64 debug nodes, totaling 11,520 nodes. Each compute node uses four MI300A units. The platform reports 1,105,920 CPU cores and 5,898,240 GiB of total memory.
| Metric | Value |
|---|---|
| Compute cabinets | 90 |
| Total nodes | 11,520 |
| MI300A APUs | 46,080 |
| CPU cores | 1,105,920 |
| HBM3 in the system | 5,760 TiB |
| Rabbit modules | 720 |
Benchmark vs. installed system: the June 2026 TOP500 lists 11.34 million "cores" for the measured configuration. This number follows the benchmark's counting methodology and should not be directly compared to the 1,105,920 x86 cores published by LLNL. For capacity engineering, use each source's definition and document the scope.
HPE Slingshot 11: the cluster's nervous system
The Slingshot 11 uses a Dragonfly topology. The goal is to reduce the number of hops between endpoints while, at the same time, avoiding an explosion in optical fiber usage. High-radix switches form groups with local copper connections; optical links connect distant groups.
| Parameter | Public value |
|---|---|
| Ports per switch | 64 |
| Bidirectional bandwidth per switch | 25.6 Tb/s |
| Port | up to 200 Gb/s unidirectional |
| Interfaces per compute node | 4 x 200 Gb/s |
| Aggregate injection per node | 100 GB/s |
| Topology | Dragonfly |
Storage: moving data at the speed of computation
El Capitan introduced Rabbit, a near-node storage layer designed to reduce I/O latency, absorb bursts, and decrease interference on the main fabric. Each Rabbit module combines SSDs and a dedicated AMD EPYC processor. The current configuration on the hardware portal lists 720 modules.
Global persistence uses Lustre. The /p/lustre4 filesystem, associated with El Capitan, is currently published with 359 PiB of capacity — approximately 400 PB in decimal units. LLNL architecture material documents throughput exceeding 2.6 TB/s at the global tier and a Rabbit layer on the order of 21 PB.
Architectural pattern: for AI clusters, the lesson learned is clear: storage should not be sized only in TB or PB. IOPS, aggregate throughput, metadata, checkpoints, burst buffers, and the distance from the data to the GPU determine time to solution.
Software stack: TOSS 4, Flux, ROCm, and MPI
Exascale hardware requires a software stack prepared for scale. El Capitan uses TOSS 4 (Tri-Lab Operating System Stack), based on Red Hat Enterprise Linux and adapted to integrate the MI300A, Slingshot, Rabbit storage, and large-scale operation.
| Layer | Technology / function |
|---|---|
| Operating system | TOSS 4 |
| Scheduler | Flux |
| GPU runtime | AMD ROCm / HIP |
| MPI | HPE Cray MPI |
| Math libraries | AMD rocBLAS and HPC ecosystem |
| Parallel filesystem | Lustre |
| Programming | C/C++, Fortran, Python, and heterogeneous models |
Flux is particularly relevant because it was created within the LLNL ecosystem to orchestrate complex, hierarchical workloads. In corporate AI clusters, the equivalent is to think of scheduler, containers, observability, quotas, and data orchestration as first-class components — not as items added later.
AI + HPC: why El Capitan is more than a benchmark
El Capitan was designed for high-fidelity simulations and new workflows that combine physics, data analysis, and artificial intelligence. LLNL cites applications in materials discovery, design optimization, advanced manufacturing, digital twins, and AI assistants trained on classified data.
HPC + AI convergence: in traditional HPC, the goal is to accelerate numerical simulation. In AI, the goal is to feed accelerators with matrices and data at scale. In El Capitan, the same infrastructure serves both worlds: HBM3, accelerators, low-latency fabric, parallel storage, and mixed precision.
Power consumption: correctly reading the numbers
On the June 2026 TOP500 list, the measured power associated with HPL is 29,684.62 kW. Dividing the HPL performance of 1.809 exaFLOP/s by this power, we get the published efficiency of approximately 60.94 GFLOP/s per watt.
LLNL's hardware overview publishes 36.0 MW of peak power for the total configuration and 90 compute cabinets. The ASC siting article uses the order of magnitude of ~35 MW. These values are consistent with the density of about 400 kW per rack published by LLNL in 2026.
| Scenario | Power | Energy if 24x7 for 1 year |
|---|---|---|
| HPL TOP500 | 29.685 MW | ~260.0 GWh |
| Peak LLNL | 36.0 MW | ~315.4 GWh |
The annualized values above are engineering calculations for order-of-magnitude sizing. Actual operation depends on load, maintenance, availability, job profile, and data center efficiency.
From the grid to the rack: the electrical infrastructure
To host exascale systems, LLNL carried out the Exascale Computing Facility Modernization (ECFM). The project brought a 115 kV transmission line to a new switchyard, installed two 40 MW substation transformers, 13.8 kV switchgear, relays, feeders, and secondary substations within the facility.
The computing floor's electrical capacity was raised from 45 MW to 85 MW. The project's goal was not just to power El Capitan, but to prepare the complex to operate two exascale machines simultaneously in future cycles.
What is not being inferred: open sources do not publish all details of UPS, generator sets, autonomy, selectivity, redundancy classification, or security controls. This material does not fill these gaps with assumptions.
The data center: Building 453 and the ECFM modernization
Building 453 was designed for multiple generations of supercomputers. LLNL reports 48,000 ft² of computer room floor — just over 4,450 m² — plus mechanical and electrical infrastructure at utility scale. The building received LEED Gold certification in 2010.
| Metric | Value |
|---|---|
| Machine floor | 48,000 ft² |
| ECFM electrical capacity | 85 MW |
| Cooling capacity | 28,000 TR |
| High-voltage supply | 115 kV |
In AI projects, this is the main paradigm shift: the data center stops being a room that houses servers and becomes an electromechanical plant whose main load is high-density computing.
Direct Liquid Cooling: 400 kW per rack
In June 2026, LLNL published that El Capitan's racks reach densities of approximately 400 kW. The facilities team cites ~25 kW/rack as the point beyond which air stops being a practical solution for this class of architecture.
The HPE Cray EX uses cold plates connected directly to the highest-dissipation components. The system is described by HPE as 100% fanless and direct liquid-cooled. Eliminating fans in the compute cabinets reduces parasitic energy and allows much more compute to be concentrated per square meter.
Thermal equation: nearly all the electrical power consumed by the hardware ends up as heat. A 36 MW IT system can generate a thermal load of the same order of magnitude. Converting 36 MW to tons of refrigeration, we get approximately 10,236 TR.
Two loops, heat exchangers, and towers
LLNL describes two circuits: one with treated water and another with glycol-based fluid. Sensors track temperature, flow, and pressure, and controls respond to rapid changes in computational load. The facility uses more than 2,000 ft of piping in this cooling chain.
Water circulates at a relatively high temperature, around 85 °F (~29.4 °C), reducing the need for intensive mechanical refrigeration. Final heat rejection occurs mainly through evaporation in the ECFM towers.
The industrial-scale thermal plant
The ECFM expanded the complex's water cooling capacity from 10,000 to 28,000 tons of refrigeration. This is equivalent to approximately 98.5 MW of nominal thermal capacity. The project added six 3,000 TR towers, plus pumps, piping, and heat exchangers.
Cooling capacity is not IT consumption: the 28,000 TR belong to the complex and were sized to support the facility's evolution, including the coexistence of exascale machines. It is not correct to interpret this number as "El Capitan consumes 98.5 MW of cooling."
Power per rack and design implications
The ratio of 90 compute cabinets x ~400 kW/rack produces 36 MW, exactly the order of magnitude published as the platform's peak power. In a conventional data center, 400 kW in a single rack changes practically everything: busbar, cabling, protection, distribution, weight, piping, CDU, layout, and operational safety.
| Item | Traditional data center | El Capitan / AI Factory class |
|---|---|---|
| Density per rack | 5-20 kW typical | ~400 kW/rack at El Capitan |
| Cooling | Air / containment | Direct liquid cooling |
| Distribution | Conventional PDUs | High current + specific architecture |
| Network | 10/25/100G | 200G per port, multi-rail |
| Storage | Predominantly capacity | Capacity + TB/s + near-node |
| Operation | Relatively stable load | Load and heat jumps from jobs |
For AI engineers, the conclusion is direct: accelerator clusters must be designed jointly with electrical and mechanical systems. Trying to "fit" a cluster of hundreds of kW per rack into a room originally designed for 10 kW/rack usually creates distribution and cooling bottlenecks.
Global context: TOP500 in June 2026
El Capitan led the TOP500 from November 2024 to November 2025. On the 67th edition, in June 2026, LineShine debuted in first place with 2.1984 exaFLOP/s. El Capitan moved to second position while maintaining 1.8090 exaFLOP/s.
| # | System | HPL Rmax (EF/s) | Power (MW) |
|---|---|---|---|
| 1 | LineShine | 2.1984 | 42.220 |
| 2 | El Capitan | 1.8090 | 29.685 |
| 3 | Frontier | 1.3530 | 24.607 |
| 4 | Aurora | 1.0120 | 38.698 |
| 5 | JUPITER Booster | 1.0000 | -- |
For the purpose of this material, the change in ranking does not reduce the technical value of the case. El Capitan remains one of the most well-documented architectures for studying the convergence of HPC, AI, high electrical density, parallel storage, and liquid cooling.
What El Capitan teaches AI data centers
- Compute cannot be sized in isolation: GPU, memory, fabric, and storage form a system.
- Electrical density becomes the main layout parameter: 400 kW/rack changes the room's engineering.
- Liquid cooling must enter the conceptual design, not as a last-minute retrofit.
- Energy must be contracted and delivered at utility scale; 30-40 MW of IT already requires a substation and long-term planning.
- Near-node storage and filesystems of hundreds of PB show that I/O is part of computational performance.
- The fabric needs to be calculated by injection, bisection, latency, and topology — not just by the nominal port speed.
- Power, temperature, flow, and pressure telemetry needs to keep pace with workload dynamics.
- Building capacity should include growth margin and coexistence of hardware generations.
EnQ Digital's view: an AI Factory is the combination of data center, GPU, network, storage, software, and energy. Business value emerges when these layers are treated as a single architecture.
Reference model for a corporate GPU project
El Capitan operates at an exceptional scale, but the principles can be translated into corporate clusters. A 1 to 10 MW GPU project should start from the workload and derive power, cooling, network, and storage.
| Stage | Engineering question |
|---|---|
| Workload | Training, inference, HPC, RAG, video, or simulation? |
| GPU | How many GPUs, TDP, memory, and intra-node interconnect? |
| Rack | What is the maximum density in kW/rack and weight per rack? |
| Cooling | DLC, in-row/in-rack CDU, supply temperature, and heat rejection? |
| Power | MW of IT, headroom, substation, distribution, and expansion? |
| Fabric | Ethernet/RoCE or InfiniBand; 200/400/800G; oversubscription? |
| Storage | Throughput, metadata, checkpoint, object, and archive? |
| Operation | Scheduler, observability, tenancy, security, and automation? |
The first calculation should be made in kW per rack and MW per room. The second, in GB/s per GPU and per cluster. The third, in TB/s of storage. Only afterward does it make sense to discuss the total number of racks and floor area.
Technical checklist for an AI Factory
- Define the IT power envelope per rack and per room.
- Define temperature, flow, and fluid quality in the liquid cooling circuit.
- Size CDUs, heat exchangers, and heat rejection capacity.
- Validate energy availability at the point of connection and energization timeline.
- Model GPU step loads and the electrical/mechanical system's response.
- Define east-west network with adequate latency and bisection bandwidth.
- Size storage by throughput and metadata, not just by capacity.
- Plan observability: power, temperature, flow, pressure, GPU health, and fabric.
- Plan phased growth without blocking maintenance or expansion.
- Validate compatibility between server, rack, CDU, manifold, fluid, and facility water.
Practical rule: the AI data center should be thought of as a thermal and electrical machine that hosts computing, not as a traditional IT room.
Essential glossary
| Term | Definition |
|---|---|
| exaFLOP/s | 10^18 floating-point operations per second. |
| Rmax | Performance effectively measured in HPL. |
| Rpeak | Theoretical peak calculated for the platform. |
| HPCG | Benchmark that emphasizes memory and communication. |
| HPL-MxP | Mixed-precision benchmark, relevant for HPC/AI convergence. |
| APU | Package that integrates CPU and GPU/accelerator with shared memory. |
| HBM3 | Third-generation High Bandwidth Memory. |
| DLC | Direct Liquid Cooling, with liquid cooling close to the component. |
| CDU | Cooling Distribution Unit. |
| Dragonfly | Interconnection topology for large-scale clusters. |
| Lustre | Parallel filesystem used in HPC. |
| Rabbit | Near-node local storage developed for the El Capitan ecosystem. |
| TOSS | Tri-Lab Operating System Stack. |
| Flux | Resource manager and scheduler developed at LLNL. |
Sources
- TOP500 — El Capitan system page: top500.org/system/180307/
- TOP500 — June 2026: top500.org/lists/top500/2026/06/
- HPC @ LLNL — El Capitan: hpc.llnl.gov/hardware/compute-platforms/el-capitan
- HPC @ LLNL — Hardware Overview.
- LLNL — Introducing El Capitan.
- LLNL — El Capitan verified world's fastest supercomputer.
- LLNL — Powering up / ECFM.
- LLNL — How El Capitan keeps its cool.
- ASC — Facilities and El Capitan.
- HPE — El Capitan press release and HPE Cray Supercomputing EX QuickSpecs.
- HPC @ LLNL — Parallel File Systems.
Verification note: data consulted and updated as of August 22, 2026. Supercomputer rankings and specifications may change with new measurements, upgrades, and new TOP500 editions.