English 箭头
Podcast Cover

[The Engineering and Vision Behind NVIDIA’s Cambridge-1 Supercomputer]-[NVIDIA’s Marc Hamilton on Building the Cambridge-1 Supercomputer During a Pandemic - Ep. 137]

NVIDIA AI Podcast · B2 · 2021-03-02

Technology
Or study on the web version

📋 Summary

The Architecture and Ambition of Cambridge-1

NVIDIA's Cambridge-1 supercomputer represents a milestone in high-performance computing (HPC) and artificial intelligence. Announced in October 2020, the project was designed to be one of the fastest AI supercomputers globally, specifically tailored to accelerate medical research and address healthcare challenges, including those exacerbated by the COVID-19 pandemic. Mark Hamilton, who leads NVIDIA’s Worldwide Solutions Architecture and Engineering team, provides insight into the rigorous process of site selection, architectural design, and the strategic decision to open this resource to external partners.

Site Selection and Sustainability

The construction of a supercomputer begins long before the hardware is installed. NVIDIA conducted a survey of nine co-location data centers in the Cambridge, UK area to find a site that met specific criteria. A primary requirement was that the facility must utilize 100% renewable energy. NVIDIA ultimately partnered with Kao Data, a company named after Sir Charles Kao—a fitting choice given that Kao discovered the fiber optic cable technology that is now essential to networking the thousands of components within the Cambridge-1 system.

The Scalable Unit: The DGX SuperPOD

At the core of Cambridge-1 is a modular design philosophy. NVIDIA utilizes its "DGX SuperPOD" architecture, which relies on a "scalable unit" consisting of 20 DGX A100 systems. This modular, "cookie cutter" approach allows organizations to scale computing power effectively without building bespoke, one-off systems.

Cambridge-1 utilizes 80 of these DGX A100 nodes. Hamilton notes that this specific configuration is remarkably efficient: "with only 20 servers, you can build one of the top 500 supercomputers in the world." This efficiency is a dramatic shift from traditional CPU-only servers, which would require hundreds or thousands of units to achieve comparable performance. Furthermore, this design consistently ranks on the "Green 500" list, highlighting its status as one of the world's most energy-efficient supercomputers.

Overcoming Construction Challenges

The deployment of Cambridge-1 was executed during a global health crisis, necessitating innovative construction methods. To maintain safety protocols while ensuring the system was built correctly, the engineering team utilized "telepresence robots." These mobile units, equipped with tablets and cameras, allowed remote engineers to inspect cable connections and system lights within the data center, effectively overcoming the limitations of on-site personnel availability.

Furthermore, the physical design of the data center was optimized through "computational fluid dynamics" (CFD). This allowed the team to model the airflow and cooling of the server rooms—referred to as the "apartment" within the data center—to ensure optimal performance. To further streamline the installation, cabling was pre-bundled in a warehouse environment before being shipped to the site, allowing for a rapid, modular assembly process.

Democratizing Access to AI Supercomputing

Historically, supercomputers were often closed systems. However, Cambridge-1 is the first NVIDIA supercomputer explicitly designed to be "open to our partners and to our customers." By utilizing cloud-native software like Kubernetes, NVIDIA has enabled a secure environment where multiple entities—including pharmaceutical leaders like AstraZeneca and GSK, and research institutions like King's College London—can run experiments simultaneously.

This shift is significant because it allows organizations that are experts in genomics or vaccine development, such as Oxford Nanopore, to leverage world-class AI infrastructure without needing to master the complexities of building a supercomputer from scratch. By providing an industry-standard, open-technology platform, NVIDIA is accelerating the timeline for critical medical breakthroughs, potentially reducing the time required for complex research from years to weeks.

🎯Key Sentences

1
Mark's here right now, so let's welcome him on.
2
Well, I'm excited to be on the call today
3
So just to kind of level set for people
4
Sure. Let me go back to October when we first announced the system.
5
And at the time, it was nothing more than that, an announcement.
Expand All

📝Key Phrases

1
level set
2
from the ground floor
3
how fitting
4
state of the art
5
building block
Expand All

📖 Transcript

Hello, and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. In October of 2020 at GTC, NVIDIA announced plans to build Cambridge One, Cambridge One is expected to be one of the fastest supercomputers in the UK and one of the most powerful AI supercomputers in the world.
One of its first applications will be healthcare.
Cambridge One will be used by researchers to solve medical challenges, including those brought about by COVID-19.
Joining us today to talk about Cambridge One is Mark Hamilton.
Mark leads the Worldwide Solutions Architecture and Engineering team at NVIDIA.

ListenLeap Brings You Into Real Context Learning

🎨 Interesting Content
🌍 Real Materials
📱 Listen Anytime
Or study on the web version