El Capitan Supercomputer Shifts Focus to Generative AI Efficiency and Real-World Performance
The El Capitan supercomputer, boasting a peak performance of 1.809 exaflops, is now primarily utilized as a 'gigafactory' for generative AI, moving beyond its initial role in physics simulations. This shift emphasizes real-world efficiency under sustained Large Language Model (LLM) workloads, cooling scalability, and software stack compatibility. The supercomputer demonstrates significant performance, achieving approximately 12.4 tokens per second per watt on Llama 3 70B, which is nearly double the performance of the Frontier supercomputer's 6.8 tokens per second per watt. The operational focus for El Capitan has evolved from theoretical speed to practical metrics such as cost per trained token and tokens per second per watt, making it a preferred choice for national-scale AI infrastructure. This strategic pivot reflects a broader market trend where supercomputers are increasingly evaluated on their ability to deliver efficient and scalable AI training capabilities rather than just raw computational power.