Hello, and welcome to the NVIDIA AI Podcast.
I'm your host, Noah Kravitz. In October of 2020 at GTC, NVIDIA announced plans to build Cambridge One, Cambridge One is expected to be one of the fastest supercomputers in the UK and one of the most powerful AI supercomputers in the world.
One of its first applications will be healthcare.
Cambridge One will be used by researchers to solve medical challenges, including those brought about by COVID-19.
Joining us today to talk about Cambridge One is Mark Hamilton.
Mark leads the Worldwide Solutions Architecture and Engineering team at NVIDIA.
He has an extensive background in high-performance computing and data centers, having worked across aerospace, defense, and other industries.
Mark's here right now, so let's welcome him on.
Mark, thank you so much for taking the time to join the NVIDIA AI podcast.
Well, I'm excited to be on the call today and to get a chance to talk a little bit more about Cambridge One.
Absolutely. So just to kind of level set for people, we're recording this in mid-February and So that's roughly four months or so after this announcement in October of 2020.
So what can you tell us about Cambridge One broadly and also kind of where we're at as we record today?
Sure. Let me go back to October when we first announced the system.
And at the time, it was nothing more than that, an announcement.
And often people, the process of building a supercomputer can be a little bit mysterious to people, right?
Systems often hear about them taking years or sometimes multiple years to build a large system like this.
But we really needed to start from the ground floor.
So NVIDIA didn't have any data center space in the UK where we wanted to build the system.
And so we started by going off and doing a survey of different data centers that are in the whole sort of Cambridge UK tech That's where it made sense for us to build it.
Now the good news is, is there were nine different data centers in that space, so-called co-location centers that met our overall requirements.
We started talking to each one of those and gathered a lot of information before we actually finally selected our final site.
And what were some of the criteria that you were looking for?
Well, you know, given that this system was going to be in the UK where we have a lot of employees now. in where we have many customers, we really like the idea of it being a UK-owned data center.
Okay. is of course, as everyone knows, computers and supercomputers use power.
And so we wanted a data center that had 100% renewable energy.
We were lucky to find that in our final choice, which was a company named Cowdata.
Very interesting that The company, Cow Data, is actually named after Sir Charles Cow, who in 1966, in the same town where the data center is located, was credited with the discovery of the fiber optic cable carrying communication signals.
How fitting. It's quite ironic now in that a supercomputer is, of course, a collection of many, many different computers.
Those computers need to be networked together.
In the Cambridge One system, it has 80 of our DGX A100 systems, It'll be connected with several thousand fiber optic cables that connect the servers to our NVIDIA Mellanox high-speed networking switches.
In addition, once you go outside of the data center, to go off and get data sources to connect to the system, fiber optic cables are also being used for that.
So I think Sir Charles Kao would be Very excited to see the Cambridge One system being built where he invented the fiber optic cable.
No, absolutely. That's a great little detail that kind of underscores the history behind this.
So let's talk a little bit about, you mentioned that to a lot of folks, the notion of a supercomputer, let alone how one comes to be designed and built. is kind of mysterious.
Can you demystify it a little bit? Talk about, I don't know, maybe start with... why 80 nodes and what kind of computing power might that result in?
Maybe some of the challenges in putting a a system like this together, let alone doing it during a global health crisis.
And so NVIDIA back in 2016, launched our DGX family of AI supercomputers.
DGX was designed to really demonstrate What was the state of the art of what could be built for high performance computing and AI in a single server?
And after we built DGX, we realized that, again, it wasn't enough to have one DGX, that most state of the art researchers would want to use a collection of them.
So at the time, back in 2016, we went to build Nvidia's first AI supercomputer. which was the predecessor to the super pot.
And again, we learned a lot of things when we were doing that.
And people have built supercomputers for high performance computing for decades.
But what was new was this supercomputer would be used not only for high-performance computing, but would be used for state of the art AI software development.
And so The DGX SuperPod design that's going into Cambridge One represents the fifth generation of our supercomputers.
There were a few requirements that we had.
One, just like DGX is designed to be the fastest single system, for AI research in computing, we wanted the DGX SuperPOD to be able to scale to build the largest AI supercomputers in the world.
And so in fact, we have another version that Cambridge One is based on that uses 560 DGX1s.
That system is currently the fifth fastest supercomputer in the world on the November 2020 top 500 list.
And it also is the world record holder in all of the MLPerth benchmarks.
Along the way, we realized that not everyone wants, not everyone needs or can afford to start off with 560 DGX systems, right?
And so we wanted to have a reasonable building block In HPC parlance, this building block is often called a scalable unit.
Think about it, it's a cookie cutter design and think about you want these sort of Lego blocks that are more than one server, but a collection of servers that can be connected together and then can scale up.
And so we settled on this building block size of 20 DGX A100 systems.
That building block size was was driven by a couple of things.
One, there was some implications in the network design.
It was a good building block for the network.
Second, it was a reasonable size that many different commercial companies would use as a starting point to build bigger systems.
So we started and we defined that building block.
20 DGX A100s, that's the scalable unit. It actually, only 20 servers is fast enough is an HPC system to get you on the top 500 list of the world's 500 fastest supercomputers.
This is amazing in that if you use a CPU only server without GPUs to build a similar sized system, you would need hundreds, or maybe even a thousand or more CPU-only servers.
And now with only 20 servers, you can build one of the top 500 supercomputers in the world.
It turns out that particular config with 20 DGX A100s is also the world's most energy efficient supercomputer based on a companion list called the Green 500, published twice a year.
And those numbers are based, again, on the November top 500 and November Green 500 list.
So that was our building block. And you can go up from there to literally thousands of different DGX systems put together in different super pod sizes.
As I stated, the largest we've built has been 560.
We settled on the 80 DGX A100s, As a reasonable-sized supercomputer that would meet many of the day-to-day needs of NVIDIA researchers in healthcare in other fields.
And if you remember, This system will be used by a number of our healthcare customers, It's actually unique.
It's the first time we're taking an NVIDIA supercomputer used by our engineers in opening it up to our partners and to our customers to use.
But we have across NVIDIA several thousand DGX systems today. that are organized in a number of different supercomputers.
And so 80 nodes is a common building block and is what we decided to deploy in the UK.
So I wanted to ask you, you brought this up about Cambridge One being the first NVIDIA supercomputer. if I've got that right, to be open to partners to use as well for research.
How and why did that decision come to be made?
Well, it was a rather easy one to make. Every time we talk to our customers about DGX and the DGX SuperPOD, one of the great value propositions to the customer when they buy a DGX SuperPOD is that they are buying exactly the same supercomputer, the same building block that Nvidia is.
So if they ever run into some challenge some problem with any of our software, they're running the same exact configuration that we have thousands of NVIDIA engineers using back at headquarters.
And so it's just natural for customers to want to use the same sort of configuration.
The second it is, is again, it's a cookie cutter process.
It significantly accelerates the time from when all the systems actually shipped to your data center to have the computer up and running.
Typically that's about a one month period of time from having the computer show up at the door of your data center to having them put together and ready to put the first users on.
And the way you build a supercomputer is the way you build any sort of large system on the public cloud.
So of course, you build it on top of so-called cloud native software.
So our super computers use a technology called Kubernetes, that's in widespread use across all of the public clouds as well as across industry around the world.
It's a set of open source software that lets you set up, administer and manage your cluster.
And again, we were running an internal version of that, and we simply had to make a few different security enhancements to the software in order to in order to let customers share the same system in a secure way with our researchers, the same way that you might have two different companies using Amazon or Google or Microsoft's public cloud.
And so again, that was the reason we decided to open this up.
Specifically in the UK, there's a huge concentration of healthcare companies.
You have very large global companies like AstraZeneca and GSK, You have some great research universities working on healthcare, and you have great companies in the genomics space.
Oxford Nanopore, which is one of the leading companies making a genome sequencer.
And so think about it just in the current context of the COVID pandemic.
Think of if you're trying to develop a new vaccine and you have an idea of saying, well, I would like to take this AI accelerated approach to building this vaccine.
And I think that if I had access to a supercomputer like Cambridge One, I might be able to do it in a few weeks instead of a few years.
Clearly, AstraZeneca already has their vaccine done today without using, they did use DGXs in the development, but not a big supercomputer like this.
But of course, there will be many more diseases in the future where we need medicines, we need vaccines, we need the likes.
So we've always had companies say, hey, what if we could use your system for just a few days or just a few weeks? work together to solve these grand challenge types of problems. recently had the chance to speak with Julie Bernauer from NVIDIA about the construction of Selene, another NVIDIA supercomputer that It was built in Santa Clara in California.
I wanted to ask you what... makes Cambridge One and Selene similar or different, as it were, but then also specifically about the use of telepresence robots in the construction and maintenance of the system?
Sure. So let's step back to answer that and sort of talk through, what have we been doing since that announcement in As I said, it took a month or two to go off and select the site.
Once CalData was selected, Think about it, moving into a data center can be a little bit like building a house.
The data center itself, what CowData builds, is almost like the empty shelf of a house or the empty shelf of a building.
And you have... air conditioning and you have power, a lot of both coming to the edge of the building or coming right inside the building, but then you just have empty floor space inside.
What a co-location center is and what NVIDIA is doing is we're not using the entire cow data center.
There's a number of other customers in there, we'll talk about later, but building our own apartment inside this apartment building that hasn't been completely built inside.
Now, the difference is unlike an apartment building where you just have to decide how big you want the rooms to be and maybe where to put the plumbing.
There's a lot of more detailed analysis that needs to go into place or on the power and the cooling to build out a supercomputer.
And so we actually use another supercomputer technology called CFD or computational fluid dynamics.
And we actually model the. space of the apartment and decide where we want to put the servers and the computer racks.
How many do we put next to each other? How do we build the rooms within our apartment building and the likes?
After doing all that analysis in conjunction with support from CalData, we finally came up on the design for our part of the data center. our so-called apartment where Cambridge One will be built.
And it's an apartment that has three equal sized rooms, in each room inside has two rows of data center racks.
These are these big sort of refrigerator size racks that you might see in pictures of a data center.
And again, we have 24 of those racks, 12 on one side, 12 on the other side. inside each room or enclosure.
And that room has its own power and air conditioning coming into the room.
And it is like a room. It has doors on it. sort of sliding glass doors going into the room, and that's to control the airflow.
So you keep the cold, air coming in separate from the hot air going out of the room, just like you would when you're cooling your house.
We came up with this design. We simulated it all to show that it would work, and then we went to work.
So the enclosures were built out on site.
It's just like putting up the walls of a room inside an apartment building.
The computer racks We wanted to use the same exact rack as Celine, just so that once we went and started installing the systems, we would know exactly know that they had the right space, the right weight and support for it.
Those racks were only available in the US and had to be manufactured in the US.
Remember, each DGX-1 weighs 123 kilograms.
It's a nice One, two, three kilograms. Easy number for the Europeans to remember.
And US pounds, I never remember the exact US pounds.
So again, we require special racks as well as to manage all the cables.
So again, think about it. You build the apartment. inside this building, three different rooms.
You then put in all the racks. You then have these thousands of fiber optic cables So you lay in effect almost like ladders being laid horizontally instead of vertically on top of the racks.
Three different sets of these ladders. This is where all the cables go.
So up to now, all of this design is exactly or very, very similar to what we use for Selene. all the way down to the fact that since it's a standard building block, when we take all of these cables, rather than try to connect all the cables on site and think about that.
If you've ever hooked up an internet cable to your computer, think of hooking up thousands of cables, right?
And how do you keep these straight? And so what we do is we work off site and basically a warehouse And we lay out all of the cables in bundles. of dozens or sometimes hundreds of cables together and ship these pre-bundled bundles of cables to the data center where they just plug in on one end into the servers, another end into the network switches.
And so not only does it Does it make the installation faster?
It also... fairly complex process. Yeah.
Mark, as I'm listening to you and I'm visualizing this, I'm also looking at my my desk that I'm sitting at, which has two laptops, one keyboard, a couple of hard drives, and it's an absolute disaster zone of cables.
Now, you mentioned something about a telepresence robot.
Yes. Now, when we were building out the latest iteration of Selene, in Santa Clara, California, which remember was finished just in time for the November 2020 top 500 list.
We were right in the middle of the COVID shutdown.
And so again, construction and some essential workers were allowed in the data center, but for safety protocols, but a very limited number of employees that we let in the data center at any one time.
And not only during the operation of the system, but especially as It's being set up.
There are some things right before all the networking is done that you really need a pair of eyes to go in and see.
And so we bought just one of these little telepresence robots, right?
It looks like a little stick on wheels with a, tablet on top and it can be remotely controlled and driven around the data center.
And it can go and roll right up to one of the systems and And then the engineer working remotely can see what cables are plugged in.
Once it's powered on, they can see what lights are or red or green on the system.
And it's been a very useful tool with Selene in keeping the system running.
Now, one of the things was with Celine, when we were building the system, Now before the power and the air conditioning was turned on, the doors to the enclosures, think of the door into the room in your apartment building,
They were left open, right? But once you start operation of the system, the doors have to be closed because that keeps the cold air in or the hot air out.
Just like on a hot day, you need to keep the door to your apartment building closed.
And so We were only able to look at the outside of the enclosures with Celine without sending inside the data center.
Our design team came up with a very simple, but what has been a hugely useful addition in modification to Cambridge One.
It's our first data center that the doors now are motion activated.
And so where is the construction process now?
You mentioned in the coming months when the system's online.
How far out are we from Cambridge One being ready to get up and running?
Well, you know, we have our GTC GPU Technology Conference coming up in April.
And so your listeners can rest assured that We are on track to make sure that we have some of the first results from Cambridge One and some of our partners to be ready to go to talk about it.
GTC. So that's in mid-April. A registration is opening soon.
And so make sure to Mark those dates down on your calendar to check back then. for our first results, but we're fairly far along and we're just going through the remaining steps, installing the final set of servers in networking gear, and then going through and doing all of the software checkout and working with our initial users, test users of the system.
I'm a bit of an outsider here, obviously, and I'm easily excited.
I'll be the first to admit. But I mean, this sounds remarkable to me, the pace at which from the announcement to having some results by mid-March GTC, I'm sorry, mid-April, It just sounds remarkable that it all came together so quickly.
Is it in fact, or is it kind of this is the state of the art these days?
Normally it takes much longer, right? And again, there's a couple of things that helped.
One is there are literally thousands of different co-location data centers around the world.
Some of them are just running simple, very low power web servers or database servers. they're not really suitable to build a supercomputer in.
And so we've gone through as part of our our DGX Ready data center program, we've pre-selected and pre-qualified a number of different data centers around the world that are ready to go to deploy DGX systems in.
And so that's a first step that helps. A second is having the exact same building block, this 20 node DGX scalable unit.
We've installed dozens of different supercomputers around the world based on that building block.
And so again, everything from the initial setup to managing and monitoring the system as we go, we've gotten better and better with each one.
Finally, again, because every... one of these is basically a replica of the other.
I mentioned that because this is our first one that will be opened up to external non-NVIDIA users.
Our software stack is going to be a little bit separate.
What we've already done and had running for a while now is we have basically a miniature version of Cambridge One. just 20 systems instead of 80. that's set up in one of our Santa Clara data centers.
And we're already working with our partners like AstraZeneca, GSK, and Oxford Nanopharm, for in running some experiments on that mini setup.
And I really think in the future, more and more people will build supercomputers like this.
Of course, if you're a national lab or if you're a university that you want to teach your grad students how to do it from scratch, they may still build bespoke systems and that's fine.
We're happy to work with those customers to do that.
But as more and more enterprises adopt AI, enterprises don't know how to build supercomputers, right?
Oxford Nanopore has never built a supercomputer before.
They built one of the world's most amazing genome sequencers, but not a supercomputer.
And so again, I think that there's a need there to have these sort of industry standard supercomputers based on open technology that an enterprise can just buy and quickly deploy is they want to ramp up their AI capabilities.
So as great as AstraZeneca, GSK and Oxford Nanopore have been to work with his early partners on Cambridge One, Of course, not all the smart people in the world, not all the great medical researchers work at companies.
Many of them work at universities. So we're so proud and excited that King's College London, KCL, has been one of our first partners that we've been working with.
We're already running some KCL code on the mini Cambridge One system.
We're really looking forward to partnering with other universities across the UK as well as the companies that I mentioned and others to come on the Cambridge One system.
Well, again, Mark, thank you so much for taking the time to come on and tell everybody about the progress.
It's really cool from a geeky standpoint, obviously, but then the applications that are going to be run out of the box couldn't be more important right now. to the world at large.
So it's really great, important stuff. Thank you.
Thank you for having me on the show today.
Thank you.